Optical remote sensing image weak supervision saliency target detection method

By building a conditional guidance module and a denoising module, combining the diffusion model and knowledge distillation mechanism, multi-scale feature extraction and noise recovery are used using graffiti annotation information, the problems of low accuracy and high labeling cost in complex backgrounds in the significance target detection of optical remote sensing images are solved, and high-precision significant object detection is achieved.

CN120431319AActive Publication Date: 2025-08-05JIANGNAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510562905.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-05
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing optical remote sensing image significance object detection method based on convolutional neural networks has poor effect in complex backgrounds, low accuracy, and relies on accurate pixel-by-pixel labeling, resulting in high labeling costs.

Method used

The weakly supervised significance object detection method of optical remote sensing images is adopted to build a conditional guide module and a denoising module, combined with the diffusion model and knowledge distillation mechanism, and multi-scale feature extraction and noise recovery are used to optimize the detection results through the iterative denoising process.

Benefits of technology

Under weak supervision conditions, high-precision significance target detection is achieved, reducing dependence on pixel-by-pixel labeling, improving the accuracy of detection results and boundary refinement effect, and reducing labeling costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431319A_ABST
    Figure CN120431319A_ABST
Patent Text Reader

Abstract

The invention discloses an optical remote sensing image weak supervision saliency target detection method, and relates to the technical field of optical remote sensing image processing, and the method comprises the steps: obtaining an optical remote sensing image training sample set, and each optical remote sensing image training sample comprises an optical remote sensing image and a graffiti labeling mask; constructing a network architecture of a saliency target detection model, wherein the saliency target detection model comprises a condition guidance module, a denoising module and a knowledge distillation module; training by using the optical remote sensing image training sample set to obtain a saliency target detection model; and detecting a to-be-detected optical remote sensing image by using the trained saliency target detection model to obtain a saliency target detection result. According to the method, the step-by-step generation capability of the diffusion model and the knowledge transmission mechanism of knowledge distillation are combined, high-precision saliency target detection of the optical remote sensing image is realized under a weak supervision condition, and the accuracy of a saliency target detection result is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of optical remote sensing image processing, and in particular to a method for weakly supervised salient target detection in optical remote sensing images. Background Art

[0002] Salient object detection, a fundamental task in computer vision, aims to rapidly locate and segment the most visually appealing target regions within complex scenes by simulating human visual attention mechanisms. With the advancement of deep learning technology, salient object detection methods based on convolutional neural networks have become mainstream and are currently widely used in fields such as agriculture, forestry, and environmental science.

[0003] With advances in satellite and aerial photography technology, the application of salient object detection in optical remote sensing images has become increasingly widespread. However, optical remote sensing images contain complex background interference from a variety of elements, such as buildings, vegetation, and water bodies. Furthermore, the wide variety of remote sensing targets exhibit significant differences in imaging scale and morphological characteristics. Consequently, existing convolutional neural network-based salient object detection methods are less effective and accurate in detecting salient objects in optical remote sensing images. Summary of the Invention

[0004] In response to the above problems and technical requirements, this application proposes a weakly supervised salient object detection method for optical remote sensing images. The technical solution of this application is as follows:

[0005] A weakly supervised salient object detection method for optical remote sensing images comprises the following steps:

[0006] Acquire an optical remote sensing image training sample set, where each optical remote sensing image training sample includes an optical remote sensing image and a graffiti annotation mask thereof, where the graffiti annotation mask indicates a salient target in the optical remote sensing image;

[0007] A network architecture for a salient object detection model is constructed, and the salient object detection model is trained using an optical remote sensing image training sample set. The salient object detection model includes a conditional guidance module, a denoising module, and a knowledge distillation module. The conditional guidance module is used to extract multi-scale features of the optical remote sensing image and generate conditional guidance information and a conditional prediction mask. The denoising module generates a noise mask and recovers salient objects from the noise mask based on the conditional guidance information to obtain a denoised prediction mask. The knowledge distillation module is used to use the conditional prediction mask generated by the conditional guidance module as a soft label to guide the denoising module to recover salient objects and guide feature learning within the conditional guidance module.

[0008] The trained salient object detection model is used to detect the optical remote sensing image to obtain the salient object detection results.

[0009] Its further technical solution is that the conditional guidance module includes a cascaded feature extraction encoder and a convolutional decoder; the feature extraction encoder is designed based on the pyramid vision Transformer, the feature extraction encoder includes multiple stacked encoding layers, different encoding layers extract guidance features of optical remote sensing images at different scales and then transmit them to the convolutional decoder, the convolutional decoder includes decoding layers corresponding to the encoding layers of the feature extraction encoder; the optical remote sensing image of the optical remote sensing image training sample is input into the feature extraction encoder, and the guidance features of multiple scales are extracted through multiple encoding layers and then transmitted to the convolutional decoder, the guidance features of multiple scales extracted by the feature extraction encoder are decoded through multiple decoding layers to obtain conditional guidance information and conditional prediction masks of multiple scales; as the number of encoding layers increases, the scale of the extracted guidance features decreases, and as the number of decoding layers increases, the scale of the decoded conditional guidance information and conditional prediction masks increases;

[0010] Each decoding layer consists of two convolution blocks and a prediction output layer cascaded in sequence. Each convolution block includes a 1×1 convolution layer, a batch normalization layer, and a ReLU activation function layer connected in sequence from input to output. The convolution decoder adopts a top-down feature transfer mechanism. The first decoding layer inputs the guided features extracted by the last encoding layer into the two convolution blocks of the decoding layer to decode and obtain conditional guided information with the same scale as the guided features extracted by the last encoding layer. Starting from the second decoding layer, the conditional guided information output by the previous decoding layer is upsampled and fused with the guided features extracted by the corresponding encoding layer. The two convolution blocks of the input decoding layer are decoded to obtain conditional guided information with the same scale as the guided features extracted by the corresponding encoding layer. The prediction output layer consists of a cascaded 1×1 convolution layer and a Sigmoid activation function layer. The prediction output layer maps the conditional guided information and outputs it as a conditional prediction mask.

[0011] A further technical solution is that the denoising module includes a noise mask generation module, a denoising encoder and a denoising decoder; the noise mask generation module gradually adds Gaussian noise to the graffiti annotation mask of the optical remote sensing image training sample according to the forward process of the diffusion model according to the iterative time step to obtain multiple noise masks, and each noise mask corresponds to an iterative time step; the denoising encoder is designed based on the Transformer encoder, and the denoising encoder includes multiple stacked denoising encoding layers, and the denoising encoding layers correspond one-to-one to the encoding layers of the feature extraction encoder of the conditional guidance module; the denoising decoder includes a denoising decoding layer that corresponds one-to-one to the denoising encoding layer of the denoising encoder;

[0012] The optical remote sensing image and noise mask of the optical remote sensing image training sample are spliced and input into the denoising encoder. After being extracted through multiple denoising coding layers, noise coding features of multiple scales are obtained and then transmitted to the denoising decoder. The noise coding features of multiple scales extracted by the denoising encoder are decoded through multiple denoising decoding layers to obtain denoising prediction masks of multiple scales. As the number of denoising coding layers increases, the scale of the extracted noise coding features decreases, and as the number of denoising decoding layers increases, the scale of the decoded denoising prediction mask increases.

[0013] Its further technical solution is that each denoising coding layer includes multiple denoising coding modules, and each denoising coding module includes a Transformer feature extraction module, a conditional enhancement module, a multi-head attention mechanism module, a layer normalization layer and a feedforward network processing module connected in sequence from input to output;

[0014] The conditional enhancement module integrates the iterative time step information into the noise features extracted by the Transformer feature extraction module, and fuses it with the guided features extracted by the conditional guidance module to obtain the conditional enhancement features. The conditional enhancement features are processed by the multi-head attention mechanism module, the layer normalization layer and the feedforward network processing module to obtain denoising coding features with time and space perception characteristics; the iterative time step information indicates the iterative time step corresponding to the noise mask.

[0015] Its further technical solution is that each denoising decoding layer includes two 3×3 convolutional layers and a denoising output layer connected in sequence from input to output;

[0016] The first denoising decoding layer fuses the denoising coding features extracted by the last denoising coding layer and the conditional guidance information extracted by the last decoding layer of the convolution decoder of the conditional guidance module, and then decodes the two 3×3 convolutional layers input to the denoising decoding layer to obtain denoising decoding features with the same scale as the denoising coding features extracted by the last denoising coding layer. Starting from the second decoding layer, the denoising decoding features output by the previous denoising decoding layer are upsampled and fused with the denoising coding features extracted by the corresponding denoising coding layer and the conditional guidance information extracted by the corresponding decoding layer of the convolution decoder of the conditional guidance module. Then, the two 3×3 convolutional layers input to the denoising decoding layer are decoded to obtain denoising decoding features with the same scale as the denoising coding features extracted by the corresponding denoising coding layer. The denoising output layer includes a cascaded convolutional prediction head and a Sigmoid activation function layer. The denoising output layer outputs the denoising decoding feature map as a denoising prediction mask.

[0017] Its further technical solution is that the knowledge distillation module includes a conditional denoising distillation module and a conditional guided self-distillation module;

[0018] The conditional denoising distillation module distills the conditional prediction mask generated by the last decoding layer of the conditional guidance module As soft labels, guiding the denoising module to recover the fine structure of the salient object based on the difference between the soft labels and the denoising prediction masks generated by each decoding layer except the first decoding layer in the denoising module;

[0019] The conditional guided self-distillation module generates the conditional prediction mask generated by the last decoding layer of the conditional guided module. As a soft label, according to the difference between the soft label and the conditional prediction mask generated by each decoding layer except the first decoding layer and the last decoding layer in the conditional guidance module, the features of each decoding layer in the conditional guidance module are guided to be consistent.

[0020] A further technical solution is to train a salient object detection model including:

[0021] The optical remote sensing image training sample is input into the network architecture of the established salient target detection model. The joint loss function L is calculated based on the conditional prediction mask and denoising prediction mask output by the salient target detection model and the graffiti annotation mask of the optical remote sensing image training sample. total =L c +L d +L mocd , and performing model training using the optical remote sensing image training samples according to a joint loss function;

[0022] Among them, the conditional guidance loss L c Used to measure the difference between the conditional prediction mask generated by the conditional guidance module and the graffiti annotation mask of the optical remote sensing image training sample; denoising loss L d Used to measure the difference between the denoising prediction mask generated by the denoising module and the graffiti annotation mask of the optical remote sensing image training sample; knowledge distillation loss L mocd It is used to measure the difference between the conditional prediction mask generated by the conditional guidance module and the denoised prediction mask generated by the denoising module, as well as the difference between the internal features of the conditional guidance module.

[0023] A further technical solution is that the feature extraction encoder of the conditional guidance module includes M stacked encoding layers, and the denoising encoder of the denoising module includes M stacked denoising encoding layers;

[0024] The knowledge distillation loss in, is the conditional prediction mask generated by the last decoding layer of the conditional guidance module The denoising prediction mask generated by the M+1-i denoising decoding layer of the denoising module Conditional denoising loss, is the conditional prediction mask generated by the last decoding layer of the conditional guidance module The conditional prediction mask generated by the M+1-jth decoding layer of the conditional guidance module The self-distillation loss is , i is an integer parameter and 1≤i≤M-1, j is an integer parameter and 2≤j≤M-1.

[0025] Its further technical solution is that the conditional prediction mask generated by the last decoding layer of the conditional guidance module The denoising prediction mask generated by the M+1-i denoising decoding layer of the denoising module Conditional denoising loss in, is the loss weight of the M+1-i denoising decoding layer of the denoising module.

[0026] Its further technical solution is that the conditional prediction mask generated by the last decoding layer of the conditional guidance module The conditional prediction mask generated by the M+1-jth decoding layer of the conditional guidance module Self-distillation loss in, is the loss weight of the M+1-jth decoding layer of the conditional guidance module.

[0027] The beneficial technical effects of this application are:

[0028] The present application proposes a method for weakly supervised salient target detection in optical remote sensing images, which is based on a diffusion model architecture and constructs two branches: a conditional guidance module and a denoising module. The conditional guidance module is used to accurately extract the image features of the optical remote sensing image and incorporate them into the denoising process as prior knowledge, providing effective guidance information on the position, boundary and morphology of the salient target for the denoising process. At the same time, the denoising module based on the Transformer-CNN hybrid architecture fully combines the global relationship modeling advantages of the Transformer to capture long-distance dependencies and the ability of CNN to efficiently extract local features, thereby significantly improving the detection accuracy. The multi-step iterative recovery of salient targets based on the diffusion model ensures that fine target boundaries and target areas are obtained under weak supervision conditions when the annotation information is incomplete, effectively reducing the dependence of traditional salient target detection methods based on convolutional neural networks on pixel-by-pixel precise annotation data, greatly reducing the annotation cost while improving detection performance.

[0029] By introducing a knowledge distillation mechanism into the diffusion model architecture, the knowledge distillation transfer mechanism is effectively combined with the gradual generation capability of the diffusion model, further improving model performance. The conditional denoising module is used to transfer the knowledge of the conditional guidance module to the denoising module, accelerating the convergence of the denoising module and effectively solving the problems of unstable training and slow convergence caused by the multiple iterative denoising mechanism of the diffusion model. The conditional guided self-distillation module is used to transfer the high-level semantic information generated by the conditional guidance module downward to the intermediate feature representation, making the soft labels generated by the conditional guidance module more discriminative, thereby providing more accurate supervision for the denoising module.

[0030] Furthermore, compared to traditional knowledge distillation methods, the conditional guidance module of this application's method is not a pre-trained static teacher network, but a dynamic knowledge source optimized simultaneously with the denoising module. This method compensates for the information loss of weak supervision by generating high-quality soft labels. This online distillation strategy enables the conditional guidance module and the denoising module to promote each other and make progress together, forming a virtuous cycle. This effectively overcomes the problem of inaccurate boundary positioning under weak supervision conditions, achieves high-precision salient object detection in optical remote sensing images under weak supervision conditions, and significantly improves the accuracy of salient object positioning results and the effect of boundary refinement. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 FIG. 4 is a structural diagram of a salient object detection model in an embodiment.

[0032] Figure 2 It is a structural diagram of a conditional guidance module in an embodiment.

[0033] Figure 3 It is the structural diagram of the denoising encoding module of the denoising model.

[0034] Figure 4 It is a schematic diagram of the PR curve of the comparative experiment.

[0035] Figure 5 This is a comparison chart of the salient target detection results of the comparative experiment. DETAILED DESCRIPTION

[0036] The specific implementation of this application will be further described below with reference to the accompanying drawings.

[0037] This application proposes a weakly supervised salient object detection method for optical remote sensing images. The specific steps of the salient object detection method are as follows:

[0038] Step 1: Obtain an optical remote sensing image training sample set, where each optical remote sensing image training sample includes an optical remote sensing image and a graffiti annotation mask thereof, where the graffiti annotation mask indicates a salient target in the optical remote sensing image.

[0039] To reduce the high annotation costs associated with traditional salient object detection methods, which rely heavily on pixel-by-pixel full-supervision, this application uses a weakly supervised salient object detection method based on graffiti annotation. Graffiti annotation is a sparse annotation method that only uses simple lines or colors to mark key areas of the target object in the image, rather than drawing a complete mask pixel by pixel. This method can effectively reduce annotation costs.

[0040] Download optical remote sensing image data from public optical remote sensing image datasets (such as ORSSD and EORSSD) and cloud platforms (such as GEE) and annotate graffiti using manual annotation or automated algorithmic generation. For example, use drawing software to mark key areas such as salient objects and background in the acquired optical remote sensing image with simple lines or colors. These lines and colors serve as the graffiti annotation results. Then, write a script to batch convert the graffiti annotation results into binary images containing salient objects and background. This binary image serves as the graffiti annotation mask. Each optical remote sensing image and a graffiti annotation mask constitute an optical remote sensing image training sample.

[0041] Step 2: Build a network architecture for the salient target detection model and use the optical remote sensing image training sample set to train the salient target detection model.

[0042] Traditional salient target detection methods rely on pixel-by-pixel annotated masks as supervision, directly modeling complex data distributions, with low computational efficiency and prone to mode collapse. Under weak supervision, computational efficiency is improved, but due to incomplete annotation information, it is often difficult to obtain fine boundaries and accurate target areas, which leads to low accuracy in the detection results of salient targets. Unlike traditional salient target detection methods, the diffusion model establishes a mapping relationship from noise to image data and uses iterative denoising to gradually restore salient targets. It can continuously optimize the prediction results in each denoising iterative process, and even if the initial annotation is inaccurate, it can gradually approach the accurate target during the iterative process. Therefore, in order to improve the accuracy of salient target detection results under weak supervision, this application constructs a salient target detection model for optical remote sensing images based on the diffusion model architecture.

[0043] Based on the diffusion model architecture, two branches are constructed: a conditional guidance module and a denoising module. One branch is used to generate prior knowledge, and the other branch uses this prior knowledge to guide the denoising process. To further improve computational efficiency, a knowledge distillation mechanism is introduced into the diffusion model architecture. This not only transfers knowledge between different branches but also enables self-distillation within the conditional guidance module, effectively combining the knowledge transfer mechanism of knowledge distillation with the gradual generation capability of the diffusion model. The salient object detection model includes a conditional guidance module, a denoising module, and a knowledge distillation module. The conditional guidance module is used to extract multi-scale features from optical remote sensing images and generate conditional guidance information and conditional prediction masks. The denoising module generates a noise mask and recovers salient objects from the noise mask based on the conditional guidance information to obtain a denoised prediction mask. The knowledge distillation module uses the conditional prediction mask generated by the conditional guidance module as a soft label to guide the denoising module in recovering salient objects and guide feature learning within the conditional guidance module.

[0044] Considering the huge size differences of objects in optical remote sensing images, such as buildings, roads, rivers, mountains, etc., multi-scale features are needed to capture details and semantic information at the same time. The Pyramid Vision Transformer (PVT) combines the global relationship modeling capability with the multi-scale computational efficiency of the pyramid, which can not only ensure the extraction of guiding features at multiple scales but also effectively improve the computational efficiency. This application verifies that the effect of 4 scales is better based on experiments. The network architecture of the salient target detection model built based on the 4-layer pyramid structure is as follows: Figure 1 shown.

[0045] As a prior knowledge generation module, the conditional guidance module plays a guiding role in the salient object detection results. Especially in the actual application of the trained salient object detection model, the model input is only the optical remote sensing image, and the model is required to directly recover the salient objects from the generated random noise mask. This requires the conditional guidance module to provide more accurate prior knowledge.

[0046] In one embodiment, the conditional guidance module includes a cascaded feature extraction encoder and a convolutional decoder; the feature extraction encoder is designed based on the pyramid vision Transformer, and the feature extraction encoder includes multiple stacked coding layers, and different coding layers extract guidance features of different scales of the optical remote sensing image and transmit them to the convolutional decoder, and the convolutional decoder includes decoding layers corresponding to the coding layers of the feature extraction encoder; the optical remote sensing image of the optical remote sensing image training sample is input into the feature extraction encoder, and the guidance features of multiple scales are extracted through multiple coding layers and then transmitted to the convolutional decoder, and the guidance features of multiple scales extracted by the feature extraction encoder are decoded through multiple decoding layers to obtain conditional guidance information and conditional prediction masks of multiple scales; as the number of coding layers increases, the scale of the extracted guidance features decreases, and as the number of decoding layers increases, the scale of the decoded conditional guidance information and conditional prediction masks increases.

[0047] Each decoding layer consists of two convolution blocks and a prediction output layer cascaded in sequence. Each convolution block includes a 1×1 convolution layer, a batch normalization layer, and a ReLU activation function layer connected in sequence from input to output. The convolution decoder adopts a top-down feature transfer mechanism. The first decoding layer inputs the guided features extracted by the last encoding layer into the two convolution blocks of the decoding layer to decode and obtain conditional guided information with the same scale as the guided features extracted by the last encoding layer. Starting from the second decoding layer, the conditional guided information output by the previous decoding layer is upsampled and fused with the guided features extracted by the corresponding encoding layer. The two convolution blocks of the input decoding layer are decoded to obtain conditional guided information with the same scale as the guided features extracted by the corresponding encoding layer. The prediction output layer consists of a cascaded 1×1 convolution layer and a Sigmoid activation function layer. The prediction output layer maps the conditional guided information and outputs it as a conditional prediction mask.

[0048] The conditional guidance module with 4 coding layers and 4 decoding layers has the following structure: Figure 2 shown. Figure 2 The feature extraction encoder in the image processing unit has 4 encoding layers. The scale of the guided features extracted by the feature extraction encoder decreases by 2 times as the number of layers increases. The convolution decoder has 4 decoding layers, which correspond to the encoding layers of the feature extraction encoder respectively. The scale of the conditional guided information and conditional prediction mask decoded by the decoding layer of the convolution decoder increases by 2 times as the number of layers increases. The 1st encoding layer corresponds to the 4th decoding layer, the 2nd encoding layer corresponds to the 3rd decoding layer, the 3rd encoding layer corresponds to the 2nd decoding layer, and the 4th encoding layer corresponds to the 1st decoding layer. The scale of the guided features extracted by the 1st encoding layer is the same as the scale of the input optical remote sensing image. The scale of the conditional guided information and conditional prediction mask decoded by the 4th decoding layer is the same as the scale of the input optical remote sensing image. The 4 scale guided features extracted by the feature extraction encoder are Convolutional decoder guides features from the deepest layer Start upsampling layer by layer and fuse it with the guidance features extracted from the corresponding coding layer. The four scale condition guidance information obtained by decoding are C4, C3, C2, and C1 respectively. The specific decoding calculation process is:

[0049]

[0050] Among them, the decoding layer i∈{1,2,3}, Up(*) represents the bilinear interpolation upsampling operation, and the spatial resolution expansion factor is 2; Concat(*) represents the feature connection operation along the channel dimension, Conv1(*) and Conv2(*) represent two cascaded convolution blocks respectively. The four decoding layers are processed by the conditional output layer to obtain the conditional prediction masks of four scales respectively.

[0051] The denoising module generates a noise mask and recovers the salient targets from the noise mask based on the conditional guidance information. During the model training process, the noise mask is generated according to the graffiti annotation mask through the forward process of the diffusion model; during the inference process of the model application, the noise mask is a randomly generated Gaussian noise mask. Recovering the salient target mask from the noise mask must ensure that the position of the detected salient target is accurate and that the details of the detected salient target are accurate. This requires the denoising model to have global relationship modeling capabilities and local feature extraction capabilities. The characteristic of the Transformer model is that it has excellent global modeling capabilities, and the characteristic of the CNN convolutional neural network is that it has high detail extraction capabilities. Therefore, this application uses a combination of Transformer and CNN for denoising.

[0052] In one embodiment, the denoising module includes a noise mask generation module, a denoising encoder and a denoising decoder; the noise mask generation module gradually adds Gaussian noise to the graffiti annotation mask x0 of the optical remote sensing image training sample according to the iterative time step based on the forward process of the diffusion model to obtain multiple noise masks, each noise mask corresponds to an iterative time step; the noise mask corresponding to any generated iterative time step t is Among them, ò~N(0,I) is standard Gaussian noise, is the noise coefficient. This process is to even out the data distribution so that the final noise distribution is easier to model.

[0053] The denoising encoder is designed based on the Transformer encoder. The denoising encoder includes multiple stacked denoising coding layers, and the denoising coding layers correspond one-to-one to the coding layers of the feature extraction encoder of the conditional guidance module; the denoising decoder includes a denoising decoding layer that corresponds one-to-one to the denoising coding layers of the denoising encoder; Figure 1Taking the four-layer structure shown as an example, the first denoising coding layer of the denoising coding layer corresponds to the first coding layer of the feature extraction encoder of the conditional guidance module, and the fourth denoising coding layer corresponds to the fourth coding layer of the feature extraction encoder of the conditional guidance module; the first denoising decoding layer of the denoising decoder corresponds to the fourth denoising coding layer, and the fourth denoising decoding layer of the denoising decoder corresponds to the first denoising coding layer.

[0054] To increase the amount of information input to the model, the optical remote sensing image and noise mask of the optical remote sensing image training sample are concatenated and input into the denoising encoder. After being extracted through multiple denoising encoding layers, noise coded features at multiple scales are transmitted to the denoising decoder. The noise coded features extracted at multiple scales by the denoising encoder are decoded through multiple denoising decoding layers to produce denoising prediction masks at multiple scales. As the number of denoising encoding layers increases, the scale of the extracted noise coded features decreases, while the scale of the decoded denoising prediction masks increases. Similarly, the scale changes by a factor of 2 at each layer.

[0055] The RGB three-channel optical remote sensing image and noise mask x of the optical remote sensing image training sample t Splicing in the channel dimension to construct a four-channel input I input ∈R H×W×4 Then the input I input The basic feature F is extracted through a convolution layer base Then, F base The noise coding features of 4 scales are obtained by inputting the 4-layer denoising coding layer for feature extraction.

[0056] Each denoising coding layer includes multiple denoising coding modules, and each denoising coding module includes a Transformer feature extraction module, a conditional enhancement module, a multi-head attention mechanism module, a layer normalization layer, and a feedforward network processing module connected in sequence from input to output. The conditional enhancement module integrates the iterative time step information into the noise features extracted by the Transformer feature extraction module, and fuses it with the guided features extracted by the conditional guidance module to obtain the conditional enhancement features. The conditional enhancement features are processed by the multi-head attention mechanism module, the layer normalization layer, and the feedforward network processing module to obtain denoising coding features with time and space perception characteristics; the iterative time step information indicates the iterative time step corresponding to the noise mask. The structure of the denoising coding module is as follows: Figure 3 As shown, the specific processing process is:

[0057] (1) In the diffusion model, the iteration time step t is an important parameter in the model generation process, which represents the progress of the diffusion process. For each iteration time step t, the model needs to know which step the current diffusion has reached so that the denoising operation can be adjusted in time. Therefore, the iteration time step information is particularly important. By encoding the iteration time step into the iteration time step embedding t emb To represent the iterative time step information. For any denoising coding layer i, the conditional enhancement module incorporates the iterative time step information into the noise features extracted by the Transformer feature extraction module and the guidance features extracted by the conditional guidance module Fusion obtains conditional enhancement features This preprocessing method ensures that temporal information and conditional semantics can effectively guide the subsequent denoising process.

[0058] (2) Enhanced conditional enhancement features After layer normalization, it is fed into the multi-head attention mechanism to capture the long-range dependencies between different spatial positions, and the output features are:

[0059]

[0060] Among them, MSA (*) represents the multi-head self-attention mechanism, and LayerNorm (*) represents the layer normalization operation. The multi-head self-attention mechanism and layer normalization operation adopt conventional model structure and processing methods, and the specific content is not repeated in this application.

[0061] (3) The output feature A is again processed by layer normalization and feedforward network to obtain the denoised coding feature for:

[0062]

[0063] Among them, FFN(*) represents a feedforward network, which consists of two linear layers plus a GELU activation function.

[0064] pass Figure 3 The design of the network structure shown in the figure enables the denoising coding module to fully perceive the temporal conditions and conditional semantic guidance, thus improving the model's ability to model complex scenes. The denoising coding features extracted by the denoising encoder with 4 denoising coding layers are

[0065] Based on the multi-scale denoising coding features extracted by the denoising encoder, a multi-layer superimposed denoising decoding layer is constructed to gradually upsample and restore the spatial resolution of the output results. Each denoising decoding layer includes two 3×3 convolutional layers and a denoising output layer connected sequentially from input to output. The first denoising decoding layer fuses the denoising coding features extracted by the last denoising coding layer with the conditional guidance information extracted by the last decoding layer of the convolution decoder of the conditional guidance module. The two 3×3 convolutional layers input to the denoising decoding layer decode the denoising decoding layer to obtain denoising decoding features of the same scale as the denoising coding features extracted by the last denoising coding layer. Starting from the second decoding layer, the denoising decoding features output by the previous denoising decoding layer are upsampled and fused with the denoising coding features extracted by the corresponding denoising coding layer and the conditional guidance information extracted by the corresponding decoding layer of the convolution decoder of the conditional guidance module. The two 3×3 convolutional layers input to the denoising decoding layer decode the denoising decoding layer to obtain denoising decoding features of the same scale as the denoising coding features extracted by the corresponding denoising coding layer. The denoising output layer includes a cascaded convolutional prediction head and a Sigmoid activation function layer. The denoising output layer outputs the denoising decoding feature map as a denoising prediction mask. The processing process of the denoising decoding layer is as follows:

[0066] (1) First, information fusion is performed based on the denoising coding features extracted by the denoising coding layer and the conditional guidance information extracted by the decoding layer of the convolution decoder of the conditional guidance module to obtain the fused features:

[0067]

[0068] in, It is the fusion feature of the first denoising decoding layer. Concat(*) indicates the channel dimension splicing processing. The fusion features obtained after information fusion of the four denoising decoding layers are processed by two 3×3 convolution modules respectively, and the denoising decoding features of the four scales are obtained, namely D4, D3, D2, and D1.

[0069] (2) Connect the convolution prediction head to the denoising decoding feature and generate the denoising prediction mask through the Sigmoid activation function:

[0070]

[0071] Among them, Conv pred (*) represents the convolution prediction head, and σ(*) represents the Sigmoid activation function. The denoised decoding features of the four scales are processed by the denoising output layer to obtain the denoising prediction masks of the four scales.

[0072] Because the diffusion model relies on multi-step iterative denoising, while it offers high accuracy, the repeated iterations also increase the computational workload. To further improve the model's computational efficiency, this application incorporates knowledge distillation into the diffusion model. This knowledge distillation provides high-quality soft labels for the denoising module, compensating for the lack of graffiti annotation information while further accelerating denoising convergence.

[0073] In one embodiment, the knowledge distillation module includes a conditional denoising distillation module and a conditional guided self-distillation module. The conditional prediction mask output by the conditional guided module is directly used as a soft label for knowledge distillation in the denoising module, accelerating denoising model convergence and improving prediction accuracy. A self-distillation process is also incorporated within the conditional guided module to enhance its own representation capabilities and improve the quality of soft label generation.

[0074] The conditional denoising distillation module uses the conditional prediction mask generated by the last decoding layer of the conditional guidance module as a soft label. Based on the difference between the soft label and the denoising prediction mask generated by each decoding layer except the first decoding layer in the denoising module, it guides the denoising module to recover the fine structure of the salient object.

[0075] The conditional guided self-distillation module uses the conditional prediction mask generated by the last decoding layer of the conditional guided module as a soft label. Based on the difference between the soft label and the conditional prediction mask generated by each decoding layer except the first and last decoding layers in the conditional guided module, it guides the features of each decoding layer within the conditional guided module to be consistent.

[0076] Based on the above-mentioned salient object detection model, the specific method for training the salient object detection model using the optical remote sensing image training sample set is as follows:

[0077] The optical remote sensing image training samples are input into the network architecture of the established salient target detection model. The conditional prediction mask and denoising prediction mask output by the salient target detection model are combined with the graffiti annotation mask of the optical remote sensing image training samples to calculate the joint loss function L. total =L c +L d +L mocd , and the model is trained using optical remote sensing image training samples based on the joint loss function.

[0078] Among them, the conditional guidance loss L c Used to measure the difference between the conditional prediction mask generated by the conditional guidance module and the graffiti annotation mask of the optical remote sensing image training sample; denoising loss L d Used to measure the difference between the denoising prediction mask generated by the denoising module and the graffiti annotation mask of the optical remote sensing image training sample; knowledge distillation loss L mocdIt is used to measure the difference between the conditional prediction mask generated by the conditional guidance module and the denoised prediction mask generated by the denoising module, as well as the difference between the internal features of the conditional guidance module.

[0079] Since the knowledge distillation module includes a conditional denoising distillation module and a conditional guided self-distillation module, the knowledge distillation loss is calculated from the conditional denoising loss and the self-distillation loss. In one embodiment, the feature extraction encoder of the conditional guided module includes M stacked encoding layers, and the denoising encoder of the denoising module includes M stacked denoising encoding layers;

[0080] Knowledge Distillation Loss in, is the conditional prediction mask generated by the last decoding layer of the conditional guidance module The denoising prediction mask generated by the M+1-i denoising decoding layer of the denoising module Conditional denoising loss, is the conditional prediction mask generated by the last decoding layer of the conditional guidance module The conditional prediction mask generated by the M+1-jth decoding layer of the conditional guidance module The self-distillation loss is , i is an integer parameter and 1≤i≤M-1, j is an integer parameter and 2≤j≤M-1.

[0081] This application minimizes the information difference between the true distribution and the predicted distribution by minimizing the KL divergence of the output distribution of the conditional guidance module and the denoising module, thereby causing the denoising prediction mask output by the denoising module to approach the soft label output by the conditional guidance module. Specifically, the denoising prediction mask and soft label output by each denoising decoding layer of the denoising module except the first denoising decoding layer are calculated separately. The KL divergence loss between them is calculated, and the corresponding loss weight is assigned according to the accuracy of the prediction results of each denoising decoding layer to obtain the conditional denoising loss.

[0082] In one embodiment, the conditional prediction mask generated by the last decoding layer of the conditional guidance module is The denoising prediction mask generated by the M+1-i denoising decoding layer of the denoising module Conditional denoising loss in, It is the loss weight of the M+1-ith denoising decoding layer of the denoising module. Since the deeper the denoising decoding layer, the closer the image resolution of the denoising prediction mask is to the original optical remote sensing image, the deeper the denoising decoding layer, the lower the corresponding loss weight. The specific loss weight value is customized according to actual needs. For the denoising module with a conditional guidance module of 4 decoding layers and 4 denoising decoding layers, this application sets the loss weight of the 4th denoising decoding layer Loss weight of the 3rd denoising decoding layer Loss weights for the second denoising decoding layer

[0083] Similarly, this application minimizes the information difference between the true distribution and the predicted distribution by minimizing the KL divergence of the output distribution of each decoding layer in the conditional guidance module, thereby prompting the conditional guidance module to output higher quality soft labels. Specifically, the conditional prediction mask and soft label output of each decoding layer except the first and last decoding layers of the conditional guidance module are calculated separately. The KL divergence loss between them is calculated, and the corresponding loss weight is assigned according to the accuracy of the prediction results of each decoding layer to obtain the conditional self-distillation loss.

[0084] In one embodiment, the conditional prediction mask generated by the last decoding layer of the conditional guidance module is The conditional prediction mask generated by the M+1-jth decoding layer of the conditional guidance module Self-distillation loss in, It is the loss weight of the M+1-jth decoding layer of the conditional guidance module. The specific loss weight value is customized according to actual needs. For the conditional guidance module with 4 decoding layers, this application sets the loss weight of the 3rd denoising decoding layer Loss weights for the second denoising decoding layer By setting the KL divergence constraint, the conditional guidance module forms consistency constraints between different levels, which can enhance the representation learning ability of the conditional guidance module, thereby generating more discriminative feature representations, ensuring the generation of accurate prior knowledge, and helping to improve the accuracy of the denoising module's prediction results.

[0085] For the training process of the conditional guidance module, cross entropy loss and structure perception loss are used for supervision to optimize the network parameters of the conditional guidance module. The specific calculation method of cross entropy loss and structure perception loss can refer to the existing technology and will not be described in detail in this application. c The calculation formula is:

[0086]

[0087] in, is the true value of the conditional prediction mask, Scr is the pixel set of the graffiti annotation mask, G is the structure-aware gating mechanism, and d is and The partial derivatives in the direction, x, y are the width and height directions of the image, a is a constant coefficient, and I is the unit matrix.

[0088] For the training process of the denoising module, cross entropy loss and structure perception loss are also used for supervision to optimize the network parameters of the denoising module. The denoising loss L d The calculation formula is:

[0089]

[0090] in, is the true value of the denoised prediction mask, It is the denoising prediction mask output by the last denoising decoding layer of the denoising module.

[0091] During the training process, the optical remote sensing image training sample set is divided into a training set and a validation set. This application uses a batch size of 4 training sample data for iteration and uses the Adam optimizer for parameter update. The initial learning rate is set to 2×10 -5 , the learning rate is dynamically adjusted using the cosine annealing strategy, and a total of 25 cycles of training are performed. Finally, the model parameters with the best performance on the validation set are selected as the final training result to obtain the salient object detection model.

[0092] Step 3: Use the trained salient object detection model to detect the optical remote sensing image to obtain the salient object detection result. The denoising prediction mask output by the denoising decoder in the last layer of the denoising module in the salient object detection model is the final salient object detection result.

[0093] The effectiveness of the method of this application is further verified by setting up a comparative experiment. The public datasets of ORSSD and EORSSD remote sensing images with graffiti annotation are used to perform significant target detection using the method of this application. Among them, the ORSSD dataset contains 800 optical remote sensing images, of which 600 are used for training and 200 are used for testing. The dataset covers a variety of complex scenes and target types. The EORSSD dataset is an extended version of ORSSD, containing 2000 optical remote sensing images, of which 1400 are used for training and 600 are used for testing, covering a more diverse range of land features. The specific experimental environment is: Python3.9, CPU: i9-12900K, GPU: GTX-3090Ti, and 64GB of memory.

[0094] The comparative experiment selected 23 advanced algorithms from the existing salient target detection methods, classified them into traditional optical remote sensing image salient target detection methods, deep learning-based fully supervised optical remote sensing image salient target detection methods, and deep learning-based weakly supervised optical remote sensing image salient target detection methods. Experiments were conducted on the EORSSD and ORSSD datasets, and the experimental data were compared with the proposed method. The final salient target detection results were evaluated by calculating four indicators: mean absolute error (MAE), structural metric (S α ), F-measure (F β ) and E-metric (E ξ The specific experimental results are shown in the following table:

[0095] Table 1 Comparison of salient object detection results between this application and other methods

[0096]

[0097] Traditional optical remote sensing image salient object detection algorithm

[0098]

[0099]

[0100] Weakly supervised salient object detection algorithm in optical remote sensing images based on deep learning

[0101]

[0102] The table shows the overall performance quantitative comparison results of the method of the present application and other methods on the benchmark dataset of salient target detection in optical remote sensing images, where Ours is the method of the present application. According to the evaluation index data of the comparative experiment in the table, in comparison with the traditional optical remote sensing image salient target detection algorithm, the method of the present application achieved the best results in all indicators on the two benchmark datasets. Compared with the weakly supervised optical remote sensing image salient target detection algorithm based on deep learning, the four indicators of the method of the present application on the ORSSD dataset are all better than the existing methods; on the EORSSD dataset, only in F β The performance index is slightly lower than GLSIN by 0.6%, and the other three indicators are better than the existing methods. Compared with the fully supervised optical remote sensing image salient target detection algorithm based on deep learning, under the weak supervision annotation training conditions of incomplete pixel-level annotation and imprecise boundary supervision, the proposed method has an E ξ 、F β and S α The indicators are still better than half of the fully supervised methods, and the MAE indicator is better than most of the fully supervised methods.

[0103] The comparison results of the PR curve (precision-recall curve) of this application method and other existing methods are as follows: Figure 4 As shown in the figure, the red curve is the method of the present application. It can be seen that the method of the present application can well handle salient targets in complex scenes under different threshold ranges. The visual comparison results of the salient target mask of the salient target detection results of the method of the present application and other existing methods are shown in the figure. Figure 5 As shown in the figure, compared with other methods, the method of this application can handle various challenging scenarios and generate prediction result maps with more accurate target positioning and finer boundaries.

[0104] The above description is only a preferred embodiment of the present application, and the present application is not limited to the above embodiments. It is understood that other improvements and variations directly derived or imagined by those skilled in the art without departing from the spirit and concept of the present application should be considered to be included in the scope of protection of the present application.

Claims

1. A weakly supervised salient object detection method for optical remote sensing images, characterized in that: The salient object detection method comprises: Acquire an optical remote sensing image training sample set, where each optical remote sensing image training sample includes an optical remote sensing image and a graffiti annotation mask thereof, where the graffiti annotation mask indicates a salient target in the optical remote sensing image; A network architecture for a salient object detection model is constructed, and the salient object detection model is trained using an optical remote sensing image training sample set. The salient object detection model includes a conditional guidance module, a denoising module, and a knowledge distillation module. The conditional guidance module is used to extract multi-scale features of the optical remote sensing image and generate conditional guidance information and a conditional prediction mask. The denoising module generates a noise mask and recovers salient objects from the noise mask based on the conditional guidance information to obtain a denoised prediction mask. The knowledge distillation module is used to use the conditional prediction mask generated by the conditional guidance module as a soft label to guide the denoising module to recover salient objects and guide feature learning within the conditional guidance module. The trained salient object detection model is used to detect the optical remote sensing image to obtain the salient object detection results.

2. The salient object detection method according to claim 1, wherein: The conditional guidance module includes a cascaded feature extraction encoder and a convolutional decoder; the feature extraction encoder is designed based on the pyramid vision Transformer, and the feature extraction encoder includes multiple stacked encoding layers, and different encoding layers extract guidance features of different scales of the optical remote sensing image and transmit them to the convolutional decoder. The convolutional decoder includes decoding layers that correspond one to one with the encoding layers of the feature extraction encoder; the optical remote sensing image of the optical remote sensing image training sample is input into the feature extraction encoder, and the guidance features of multiple scales are extracted by the multiple encoding layers and then transmitted to the convolutional decoder. The guidance features of multiple scales extracted by the feature extraction encoder are decoded by multiple decoding layers to obtain conditional guidance information and conditional prediction masks of multiple scales; As the number of coding layers increases, the scale of the extracted guided features decreases, while as the number of decoding layers increases, the scale of the decoded conditional guided information and conditional prediction mask increases. Each decoding layer consists of two convolution blocks and a prediction output layer cascaded in sequence. Each convolution block includes a 1×1 convolution layer, a batch normalization layer, and a ReLU activation function layer connected in sequence from input to output. The convolution decoder adopts a top-down feature transfer mechanism. The first decoding layer inputs the guided features extracted by the last encoding layer into the two convolution blocks of the decoding layer to decode and obtain conditional guided information with the same scale as the guided features extracted by the last encoding layer. Starting from the second decoding layer, the conditional guided information output by the previous decoding layer is upsampled and fused with the guided features extracted by the corresponding encoding layer. The two convolution blocks of the input decoding layer are decoded to obtain conditional guided information with the same scale as the guided features extracted by the corresponding encoding layer. The prediction output layer consists of a cascaded 1×1 convolution layer and a Sigmoid activation function layer. The prediction output layer maps the conditional guided information and outputs it as a conditional prediction mask.

3. The salient object detection method according to claim 2, wherein: The denoising module includes a noise mask generation module, a denoising encoder, and a denoising decoder. The noise mask generation module gradually adds Gaussian noise to the graffiti annotation mask of the optical remote sensing image training sample according to the forward process of the diffusion model according to the iterative time step to obtain multiple noise masks, each noise mask corresponding to one iterative time step. The denoising encoder is designed based on the Transformer encoder. The denoising encoder includes multiple stacked denoising coding layers, which correspond one-to-one to the coding layers of the feature extraction encoder of the conditional guidance module; the denoising decoder includes denoising decoding layers that correspond one-to-one to the denoising coding layers of the denoising encoder. The optical remote sensing image and noise mask of the optical remote sensing image training sample are spliced and input into the denoising encoder. After being extracted through multiple denoising encoding layers, the noise coding features of multiple scales are obtained and then transmitted to the denoising decoder. The noise coding features of multiple scales extracted by the denoising encoder are decoded through multiple denoising decoding layers to obtain denoising prediction masks of multiple scales. As the number of denoising coding layers increases, the scale of the extracted noise coding features decreases, and as the number of denoising decoding layers increases, the scale of the denoising prediction mask obtained by decoding increases.

4. The salient object detection method according to claim 3, wherein: Each denoising coding layer includes multiple denoising coding modules, and each denoising coding module includes a Transformer feature extraction module, a conditional enhancement module, a multi-head attention mechanism module, a layer normalization layer, and a feedforward network processing module connected in sequence from input to output; The conditional enhancement module integrates the iterative time step information into the noise features extracted by the Transformer feature extraction module, and fuses it with the guidance features extracted by the conditional guidance module to obtain the conditional enhancement features. The conditional enhancement features are processed by the multi-head attention mechanism module, the layer normalization layer and the feedforward network processing module to obtain the denoising coding features with temporal and spatial perception characteristics. The iteration time step information indicates the iteration time step corresponding to the noise mask.

5. The salient object detection method according to claim 4, characterized in that: Each denoising decoding layer consists of two 3×3 convolutional layers and a denoising output layer connected sequentially from input to output; The first denoising decoding layer fuses the denoising coding features extracted by the last denoising coding layer and the conditional guidance information extracted by the last decoding layer of the convolution decoder of the conditional guidance module, and then decodes the two 3×3 convolution layers of the input denoising decoding layer to obtain denoising decoding features with the same scale as the denoising coding features extracted by the last denoising coding layer. Starting from the second decoding layer, the denoising decoding features output by the previous denoising decoding layer are upsampled and fused with the denoising coding features extracted by the corresponding denoising coding layer and the conditional guidance information extracted by the corresponding decoding layer of the convolution decoder of the conditional guidance module. Then, the two 3×3 convolution layers of the input denoising decoding layer are decoded to obtain denoising decoding features with the same scale as the denoising coding features extracted by the corresponding denoising coding layer. The denoising output layer consists of a cascaded convolutional prediction head and a Sigmoid activation function layer. The denoising output layer outputs the denoised decoded feature map as a denoising prediction mask.

6. The salient object detection method according to claim 3, wherein: The knowledge distillation module includes a conditional denoising distillation module and a conditional guided self-distillation module; The conditional denoising distillation module distills the conditional prediction mask generated by the last decoding layer of the conditional guidance module As soft labels, guiding the denoising module to recover the fine structure of the salient object based on the difference between the soft labels and the denoising prediction masks generated by each decoding layer except the first decoding layer in the denoising module; The conditional guided self-distillation module generates the conditional prediction mask generated by the last decoding layer of the conditional guided module. As a soft label, according to the difference between the soft label and the conditional prediction mask generated by each decoding layer except the first decoding layer and the last decoding layer in the conditional guidance module, the features of each decoding layer in the conditional guidance module are guided to be consistent.

7. The salient object detection method according to claim 1, wherein: The training of the salient target detection model includes: The optical remote sensing image training sample is input into the network architecture of the established salient target detection model. The joint loss function L is calculated based on the conditional prediction mask and denoising prediction mask output by the salient target detection model and the graffiti annotation mask of the optical remote sensing image training sample. total =L c +L d +L mocd , and performing model training using the optical remote sensing image training samples according to a joint loss function; Among them, the conditional guidance loss L c Used to measure the difference between the conditional prediction mask generated by the conditional guidance module and the graffiti annotation mask of the optical remote sensing image training sample; denoising loss L d Used to measure the difference between the denoising prediction mask generated by the denoising module and the graffiti annotation mask of the optical remote sensing image training sample; knowledge distillation loss L mocd It is used to measure the difference between the conditional prediction mask generated by the conditional guidance module and the denoised prediction mask generated by the denoising module, as well as the difference between the internal features of the conditional guidance module.

8. The salient object detection method according to claim 7, wherein: The feature extraction encoder of the conditional guidance module includes M stacked coding layers, and the denoising encoder of the denoising module includes M stacked denoising coding layers; The knowledge distillation loss in, is the conditional prediction mask generated by the last decoding layer of the conditional guidance module The denoising prediction mask generated by the M+1-i denoising decoding layer of the denoising module Conditional denoising loss, is the conditional prediction mask generated by the last decoding layer of the conditional guidance module The conditional prediction mask generated by the M+1-jth decoding layer of the conditional guidance module The self-distillation loss is , i is an integer parameter and 1≤i≤M-1, j is an integer parameter and 2≤j≤M-1.

9. The method for detecting salient objects according to claim 8, wherein: The conditional prediction mask generated by the last decoding layer of the conditional guidance module The denoising prediction mask generated by the M+1-i denoising decoding layer of the denoising module Conditional denoising loss in, is the loss weight of the M+1-i denoising decoding layer of the denoising module.

10. The method for detecting salient objects according to claim 8, wherein: The conditional prediction mask generated by the last decoding layer of the conditional guidance module The conditional prediction mask generated by the M+1-jth decoding layer of the conditional guidance module Self-distillation loss in, is the loss weight of the M+1-jth decoding layer of the conditional guidance module.

Citation Information

Patent Citations

  • Remote sensing image saliency target detection method based on conditional diffusion model

    CN118015297A

  • Method for generating scene graph and electronic equipment

    CN118115627A

  • Weak supervision saliency target detection method based on enhanced graffiti annotation

    CN118230106A

  • Image segmentation method and device, computer equipment, storage medium and program product

    CN118840376A