A method and system for SAR image denoising

By combining a twin encoder and a dual-branch recovery network, the blob representation of SAR images is extracted and embedded spatial decoupling recovery is performed, which solves the problems of edge blurring and detail loss in the prior art and achieves efficient noise suppression and detail preservation.

CN122048715BActive Publication Date: 2026-07-21HUNAN NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN NORMAL UNIVERSITY
Filing Date
2026-04-20
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing SAR image denoising methods tend to cause edge blurring and detail loss when processing edges and areas with rich details, resulting in poor denoising performance.

Method used

A twin encoder is used to extract the blob representation of SAR noisy images, and a bi-branch recovery network is used to project it into the embedding space. The network is trained by a bi-branch attention network to generate low-frequency and high-frequency information embeddings. The total loss is calculated by combining a penalty term and a contrast loss, and the network is optimized to achieve denoising.

Benefits of technology

It significantly improves the denoising effect of SAR images, effectively suppresses speckle noise, and preserves the edge and texture details of the image while having high computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122048715B_ABST
    Figure CN122048715B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image processing, in particular to a SAR image denoising method and system. The method first extracts a speckle representation of a SAR noise image by using a twin encoder, and projects the speckle representation to an embedding space by a double-branch restoration network to generate low-frequency information embedding and high-frequency information embedding. Then, a double-branch attention network is constructed, and a loss function of the network is optimized by using the loss of the low-frequency and high-frequency information embedding, and then the network is trained by using the optimized loss function. Finally, the trained double-branch attention network is used for denoising the SAR noise image. The method effectively solves the technical problem that the denoising effect is poor due to the loss of edge details in the SAR image denoising process, can effectively suppress the speckle noise while reliably retaining the edge, texture and other key details of the image, and significantly improves the denoising quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a SAR image denoising method and system. Background Technology

[0002] Synthetic Aperture Radar (SAR) is a high-resolution imaging radar that can obtain high-resolution radar images similar to optical photography under extremely low visibility weather conditions. However, due to the limitations of the coherent imaging mechanism of SAR, complex ground environments often lead to noise speckle interference from electromagnetic wave reflections, thus affecting the imaging of the target area and subsequent image processing tasks. Therefore, effectively suppressing speckle noise in SAR images is an important task in the field of SAR image processing.

[0003] Existing patent application CN101833753A discloses a SAR image speckle removal method based on an improved Bayesian nonlocal mean filter. This method constructs a mean matrix and a point set, and applies a Bayesian nonlocal mean filter to each pixel in the input SAR image to achieve speckle removal. However, this method operates directly in the two-dimensional spatial domain of the image, and its denoising capability is mainly concentrated in homogeneous regions. When processing edge-rich and detail-rich regions, it easily causes edge blurring and detail loss, thus affecting the denoising effect. Summary of the Invention

[0004] Therefore, it is necessary to provide a SAR image denoising method and system to address the technical problem of poor denoising effect caused by the loss of edge details in the SAR image denoising process.

[0005] Firstly, this application provides a SAR image denoising method. The method includes:

[0006] Step S1: Acquire SAR noise image;

[0007] Step S2: Extract the blob representation of the SAR noise image using a twin encoder;

[0008] Step S3: Project the blob representation onto the embedding space using a dual-branch recovery network to generate a low-frequency information embedding space and a high-frequency information embedding space;

[0009] Step S4: Calculate the loss of low-frequency information embedding and high-frequency information embedding based on the penalty term, and calculate the total embedding loss based on the loss of low-frequency information embedding and high-frequency information embedding.

[0010] Step S5: Optimize the loss function of the dual-branch attention network based on the total embedded loss to obtain a joint optimized loss function, and train the dual-branch attention network based on the joint optimized loss function to obtain the trained dual-branch attention network; the dual-branch attention network includes a cascaded dual-branch attention group, a fully connected layer, a convolutional layer, and an upsampling layer, the dual-branch attention group includes at least one cascaded dual-branch attention block, and the dual-branch attention block includes a channel attention branch and a structural information branch;

[0011] Step S6: Denoise the SAR noise image using the trained dual-branch attention network.

[0012] Furthermore, the mathematical expression for the SAR noise image is:

[0013]

[0014] In the formula, This represents the SAR noise image. Represents the truth image. This represents non-stationary additive noise. This indicates multiplicative noise.

[0015] Furthermore, the twin encoder includes multiple cascaded residual blocks and a pooling layer.

[0016] Furthermore, the main path of each residual block includes two types of convolutional blocks. The first type of convolutional block includes a cascaded 3×3 convolutional layer, a normalization layer, a LeakyReLu activation layer, a 3×3 convolutional layer, and a normalization layer. The second type of convolutional block includes a cascaded 1×1 convolutional layer, a normalization layer, and a LeakyReLu activation layer.

[0017] Furthermore, the dual-branch recovery network includes a low-frequency recovery network branch and a high-frequency recovery network branch. The low-frequency recovery network branch includes cascaded 3×3 convolutional layers, nonlinear activation layers, pooling layers, flattening layers, and fully connected layers. The high-frequency recovery network branch includes cascaded 1×1 convolutional layers, nonlinear activation layers, pooling layers, flattening layers, and fully connected layers.

[0018] Furthermore, the formula for calculating the total embedding loss is as follows:

[0019]

[0020] In the formula, Represents the total loss of embedding. This represents the loss in embedding the low-frequency information. This represents the loss of the high-frequency information embedding. This indicates the weight of the low-frequency recovery branch. This indicates the weight of the high-frequency recovery branch.

[0021] Furthermore, the expression for the joint optimization loss function is:

[0022]

[0023]

[0024] In the formula, This represents the joint optimization loss of the dual-branch attention network. Represents the total loss of embedding. This represents the loss of the dual-branch attention network. Describes the L1 loss function. This represents the gradient mean squared error loss function. This represents the edge-preserving loss function. The weights represent the L1 loss function. The weights represent the gradient mean squared error loss function. This represents the weights of the edge-preserving loss function.

[0025] Furthermore, the formulas for calculating the losses of the low-frequency information embedding and the high-frequency information embedding are as follows:

[0026]

[0027]

[0028] In the formula, This represents the loss in the embedding of the low-frequency information or the embedding of the high-frequency information. This represents the neighbor loss of the k-th positive sample pair. This represents the inversion loss of the k-th positive sample pair. Indicates the first Embedded vectors With the Embedded vectors Contrast loss For temperature hyperparameters, Indicates when When, the value is 1, when The value is 0. Represents the i-th embedding vector With the j-th embedding vector Penalties between [these items].

[0029] Furthermore, the expression for the penalty term is:

[0030]

[0031] In the formula, This represents the i-th embedding vector. This represents the j-th embedding vector.

[0032] Secondly, this application provides a SAR image denoising system. The system includes:

[0033] The image acquisition module is used to acquire SAR noise images;

[0034] An image blob extraction module is used to extract blob representations from the SAR noise image using a twin encoder;

[0035] An embedding space generation module is used to project the blob representation onto the embedding space using a dual-branch recovery network to generate a low-frequency information embedding space and a high-frequency information embedding space.

[0036] An embedding loss calculation module is used to calculate the loss of the low-frequency information embedding and the high-frequency information embedding based on the penalty term, and to calculate the total embedding loss based on the loss of the low-frequency information embedding and the high-frequency information embedding.

[0037] The network training module is used to optimize the loss function of the dual-branch attention network based on the total embedded loss to obtain a joint optimized loss function, and to train the dual-branch attention network based on the joint optimized loss function to obtain the trained dual-branch attention network. The dual-branch attention network includes a cascaded dual-branch attention group, a fully connected layer, a convolutional layer, and an upsampling layer. The dual-branch attention group includes at least one cascaded dual-branch attention block, and the dual-branch attention block includes a channel attention branch and a structural information branch.

[0038] The image denoising module is used to denoise the SAR noisy image using the trained dual-branch attention network.

[0039] The aforementioned SAR image denoising method and system first extracts the speckle representation of the SAR image through a Siamese encoder and a contrastive learning mechanism, and then projects it into the embedding space using a dual-branch recovery network to achieve decoupling recovery of low-frequency and high-frequency components. This allows the feature extraction network to learn more discriminative feature representations, significantly improving the generalization ability of this method in complex real-world scenarios. Second, by constructing a dual-branch attention network that includes channel attention branches and structural information branches, and by weighting and fusing the contrastive loss of the Siamese encoder with the composite loss of this network to form a joint optimized loss function, the network is trained using this joint optimized loss function. This achieves synergistic optimization of feature learning and image reconstruction, effectively suppressing speckle noise while reliably preserving key details such as image edges and textures, and exhibiting high computational efficiency. Attached Figure Description

[0040] Figure 1 This is a flowchart illustrating a SAR image denoising method in one embodiment;

[0041] Figure 2 This is a schematic diagram of the structure of a twin encoder in one embodiment;

[0042] Figure 3 This is a schematic diagram of the twin encoder in another embodiment;

[0043] Figure 4 This is a schematic diagram of the low-frequency recovery network branch in one embodiment;

[0044] Figure 5 This is a schematic diagram of the structure of a high-frequency recovery network branch in one embodiment;

[0045] Figure 6 Here is a schematic diagram of the structure of a dual-branch attention network in one embodiment, wherein (a) is a schematic diagram of the overall structure of the dual-branch attention network, (b) is a schematic diagram of the structure of the dual-branch attention group, and (c) is a schematic diagram of the structure of the dual-branch attention block;

[0046] Figure 7 Here is a comparison of the effects of SAR noise image denoising before and after in one embodiment, where (a) is the SAR noise image before denoising and (b) is the SAR image after denoising;

[0047] Figure 8 Here is a comparison of the profiles of a SAR noise image before and after denoising in one embodiment, where (a) is a comparison of the range profiles of the SAR noise image before and after denoising, and (b) is a comparison of the azimuth profiles of the SAR noise image before and after denoising.

[0048] Figure 9 This is a comparison of the effects of denoising a SAR image containing multiple targets and multiple locations before and after denoising in one embodiment. (a) is the SAR image containing multiple targets and multiple locations before denoising, and (b) is the SAR image containing multiple targets and multiple locations after denoising.

[0049] Figure 10 This is a schematic diagram of the structure of a SAR image denoising system in one embodiment. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0051] Example 1

[0052] like Figure 1 As shown, this embodiment provides a SAR image denoising method, including:

[0053] Step S1: Obtain the SAR noise image.

[0054] Based on Goodman's theory, the mathematical expression for a SAR noisy image is shown in equation (1):

[0055] (1)

[0056] In the formula, Represents a SAR noise image. Represents the truth image. For element-wise product, Multiplicative noise is typically assumed to follow a gamma distribution with a mean of 1 and a variance of 1 / L in SAR noise images. Therefore, multiplicative noise... The probability distribution can be expressed as:

[0057] (2)

[0058] In the formula, Let L represent the probability distribution of multiplicative noise, and L represent the equivalent number of looks. For Gamma function. However, it is difficult to analyze and apply the multiplicative speckle model shown in Equation (1) directly. Therefore, after taking the logarithm of both sides of Equation (1), we obtain the additive noise model shown in Equation (3):

[0059] (3)

[0060] After logarithmic transformation, the additive noise model still needs exponential transformation to obtain the final denoised image. However, the exponential transformation process is prone to mean shift, which affects the final denoising accuracy. To avoid this problem, this embodiment constructs an additive noise model for SAR noise images as shown in equation (4):

[0061] (4)

[0062] In the formula, Represents a SAR noise image. Represents the truth image. Indicates multiplicative noise. This represents non-stationary additive noise, which is correlated with the signal and has a mean of 0. (SAR noise image) With the truth image The difference between them can be caused by non-stationary additive noise. Therefore, the model shown in Equation (4) transforms the denoising task into the estimation of non-stationary additive noise, avoiding the logarithmic-exponential transformation process and fundamentally avoiding the generation of mean drift, thus enabling the preservation of more image information while suppressing noise.

[0063] Since SAR noisy images lack clean, noise-free reference images, a training dataset is constructed by artificially adding noise to simulated target images. Simulated target images are clean, noise-free images, which can be optical images or simulated SAR images. For example, non-stationary additive noise in equation (4) is added to the simulated target image. The first step involved adding noise of the same magnitude as the peak intensity of the image. Subsequent steps reduced the noise intensity by 90% of the previous level. Datasets were generated based on the added noise intensity, with 100 images generated for each noise level, resulting in 16 different noise intensity levels and a total of 1600 images. All images were uniformly scaled to 512. Each image is 512 pixels long and stored in NumPy binary format (.npy). These images are then used as the training, validation, and test datasets in a 7:2:1 ratio. Those skilled in the art can modify the noise intensity step size, the total number of generated images, and the image storage format according to the specific application scenario.

[0064] Step S2: Use a twin encoder to extract the blob representation of the SAR noise image.

[0065] Specifically, firstly, query patches, positive samples, and negative samples are extracted from SAR noisy images, and then a twin encoder is used to extract blob representations of the query patches, positive samples, and negative samples. For example, first from each 512 A 64-pixel image was randomly cropped from a SAR noisy image of 512. 64 SAR small images, these SAR small images Batch set constituting the input N is the batch size, n=1, 2, ..., N; then from each SAR small image Two image patches are randomly selected from the image, denoted as... , , As a query patch , As positive samples Two image patches extracted from the same small SAR image are called a positive sample pair. Due to the self-similarity of spots in local regions, positive sample pairs are likely to have similar spot distributions. Accordingly, from the current batch set... Except Image patches extracted from the remaining N-1 images are used as negative samples. Negative samples originate from image regions that differ from the query patch in both image content and noise statistical distribution. With query patch As negative sample pairs.

[0066] The twin encoder E is used to extract blob representations of the input image patches. It consists of multiple cascaded residual blocks and a pooling layer. The input of the first residual block... The output of the last residual block is connected to the input of the pooling layer, and the output of the pooling layer is a speckle representation. .like Figure 2 As shown, each residual block includes a main path and a short-circuit path. The main path includes at least two convolutional blocks for learning the residual. Each convolutional block includes a convolutional layer, a normalization layer, and an activation layer. The short-circuit path directly passes the input to the output and adds it to the output of the main path.

[0067] Specifically, in a preferred embodiment, the structure of the twin encoder is as follows: Figure 3 As shown, it includes three residual blocks and one average pooling layer. The main path of each residual block includes two types of convolutional blocks. The first type of convolutional block includes a cascaded 3×3 convolutional layer, a normalization layer, a LeakyReLu activation layer, a 3×3 convolutional layer, and a normalization layer. The second type of convolutional block includes a cascaded 1×1 convolutional layer, a normalization layer, and a LeakyReLu activation layer. In the feature extraction process, the residual blocks in the front row (i.e., the shallow convolutional network) are dedicated to extracting low-level features such as local texture and edges of the input image blocks, while the residual blocks in the back row (i.e., the deep convolutional network) are dedicated to extracting high-level features such as global structure and semantics based on the low-level features. At the same time, the residual connections within each residual block realize the sharing of cross-layer features, which enhances the generalization ability of the encoder. The Siamese encoder E ends with an average pooling, and finally obtains the blob representation as shown in Equation (5):

[0068] (5)

[0069] In the formula, The dots represent the points. This represents the input of the twin encoder.

[0070] Specifically, regarding query patches Positive samples and negative samples The twin encoder E extracts features from each feature and outputs the corresponding blob representation. , and .

[0071] Step S3: The blob representation is projected into the embedding space using a dual-branch recovery network to generate a low-frequency information embedding space and a high-frequency information embedding space.

[0072] The dual-branch recovery network decouples high- and low-frequency information from the input feature map, aiming to focus on different frequency components of the image separately, thereby providing high-quality feature input for subsequent contrastive loss calculation. Specifically, the dual-branch recovery network consists of a low-frequency recovery network branch and a high-frequency recovery network branch connected in parallel. The low-frequency recovery network branch is configured with 3×3 convolutional layers to capture broader contextual information using its larger receptive field, thereby performing effective local smoothing and statistical feature aggregation. This branch is crucial for reconstructing flat regions and the overall structure of the image. The high-frequency recovery network branch is configured with 1×1 convolutional layers, focusing on inter-channel feature transformations to effectively extract and preserve high-frequency details such as edges and textures of the image without introducing additional spatial blur. At the same time, this branch can also compensate for and suppress the over-smoothing tendency that may be generated by the low-frequency branch.

[0073] In a preferred embodiment, such as Figure 4 As shown, the low-frequency recovery network branch includes cascaded 3×3 convolutional layers, nonlinear activation layers, pooling layers, flattening layers, and fully connected layers; as... Figure 5 As shown, the high-frequency recovery network branch includes cascaded 1×1 convolutional layers, nonlinear activation layers, pooling layers, flattening layers, and fully connected layers. Furthermore, the nonlinear activation layers can employ the LeakyReLu activation function, and the pooling layers can employ adaptive average pooling.

[0074] The bi-branch recovery network represents the speckle. After projection onto the embedding space, an embedding space as shown in Equation (6) is generated. Specifically, the low-frequency recovery network branch projects the blob representation onto the embedding space to generate a low-frequency embedding space representing low-frequency information. The high-frequency recovery network branch projects the blob representation onto the embedding space, generating a high-frequency embedding space that represents high-frequency information. .

[0075] (6)

[0076] In the formula, Indicates the embedding space. This indicates the projection of the project.

[0077] The blob representations input to both the low-frequency recovery network branch and the high-frequency recovery network branch include query patches. The spots indicate Positive samples The spots indicate and negative samples The spots indicate The low-frequency recovery network branch projects the blob representation of the input above, and the output is located in the low-frequency embedding space. Low-frequency information embedding , , The high-frequency recovery network branch projects the blob representation of the input above, and the output is located in the high-frequency embedding space. High-frequency information embedding , , .

[0078] Step S4: Calculate the loss of low-frequency information embedding and high-frequency information embedding based on the penalty term, and calculate the total embedding loss based on the loss of low-frequency information embedding and high-frequency information embedding.

[0079] This embodiment employs a contrastive learning strategy, aiming to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs in the embedding space. To achieve this goal, this embodiment defines a cosine similarity-based penalty term to measure the similarity between two embedding vectors. For any two embedding vectors... and The penalty term is defined as follows:

[0080] (7)

[0081] In the formula, This represents the penalty term between the i-th and j-th embedding vectors, with a value ranging from -1 to 1. A larger value indicates that the two embedding vectors are more similar. This represents the i-th embedding vector. This represents the j-th embedding vector.

[0082] Furthermore, normalized temperature-scaled cross-entropy (NT-Xent) is used as the contrastive loss function to constrain the Siamese encoder, ensuring that the similarity between each positive sample pair is as high as possible among the 2N embedding vectors, while the similarity with the remaining 2(N-1) negative sample embeddings is as low as possible. The expression for the contrastive loss function is:

[0083] (8)

[0084] In the formula, Indicates the first Embedded vectors With the Embedded vectors The comparative loss, For temperature hyperparameters, Indicates when When, the value is 1, when At that time, the value is 0. For each positive embedding pair in the same mini-batch The remaining 2(N-1) embeddings are considered as negative samples.

[0085] Based on the penalty term shown in equation (7) and the contrast loss function shown in equation (8), the losses for low-frequency information embedding and high-frequency information embedding are calculated respectively. Specifically, the loss calculation formulas for low-frequency information embedding and high-frequency information embedding are as follows:

[0086] (9)

[0087] For both low-frequency and high-frequency recovery network branches, the 2N embedding vectors they contain can generate N positive sample pairs. In the formula, This indicates the loss in low-frequency information embedding or high-frequency information embedding. This represents the neighbor loss of the k-th positive sample pair. Let represent the inversion loss of the k-th positive sample pair.

[0088] Furthermore, considering that the traditional NT-Xent loss function mainly constrains a single feature representation, this embodiment proposes a flexible and weighted dual-branch contrastive loss mechanism to adapt to the dual-branch architecture that decouples high- and low-frequency information. This mechanism can simultaneously receive and process two independent pairs of feature representations, low-frequency and high-frequency, ensuring that positive sample pairs within each branch are strengthened while achieving maximum separation from negative samples. The formula for calculating the total embedding loss is:

[0089] (10)

[0090] In the formula, Represents the total loss of embedding. This indicates the loss due to the embedding of low-frequency information. This represents the loss due to the embedding of high-frequency information. This indicates the weight of the low-frequency recovery branch. This indicates the weight of the high-frequency recovery branch.

[0091] Step S5: Optimize the loss function of the dual-branch attention network based on the total embedded loss to obtain the joint optimized loss function, and train the dual-branch attention network based on the joint optimized loss function to obtain the trained dual-branch attention network.

[0092] The input to the Dual-Branch Attention Network (DBAN) includes SAR-noise images. The blob representation of the output of the twin encoder The output is a denoised SAR image. DBAN incorporates an attention mechanism to fuse robust blob representations into the network, enabling it to generalize to different datasets. Specifically, the dual-branch attention network includes a cascaded dual-branch attention group (DBAG), fully connected layers, convolutional layers, and upsampling layers. The DBAG, as the core feature extraction module, includes at least one dual-branch attention block (DBAB). Each DBAB employs a dual-branch parallel design, specifically including a channel attention branch and a structural information branch. The channel attention branch pre-extracts features through stacked convolutional layers, such as two 3×3 convolutions, and the LeakyReLU activation function. It also introduces the core channel attention mechanism (CALayer), which dynamically learns the importance weights of each feature channel, adaptively enhancing feature channels beneficial to denoising while suppressing noise-dominated channels, significantly improving feature discrimination capabilities. The structural information branch focuses on preserving the image's geometric structure. Through a series of convolutional layers, such as five 3×3 convolutions, combined with residual connections, it enhances structural information such as edges, textures, and point objects. This branch provides direct prior guidance to the network, ensuring that key image structures are not overly smoothed during denoising. Through this collaborative mechanism, DBAN intelligently balances the two competing core objectives of noise suppression and detail preservation, ultimately achieving high-quality denoising results.

[0093] In a preferred embodiment, the overall structure of the DBAN is as follows: Figure 6 As shown in (a), "fully connected + 3×3 convolution" in the figure represents a cascaded 3×3 convolutional layer and a fully connected layer. Specifically, the SAR noise image X and the blob representation R are used as inputs to the DBAN, which are fed into multiple cascaded DBAGs. The output of the last DBAG is input to the cascaded 3×3 convolutional layer and the fully connected layer. The output of the fully connected layer is fused with the input of the DBAN and then input to the upsampling layer. Finally, the upsampling layer outputs a high-resolution denoised image. The structure of the DBAG is as follows: Figure 6 As shown in (b), the DBAB includes multiple cascaded connections, and the entire DBAB incorporates residual connections, directly adding the input to the output. The structure of the DBAB is as follows: Figure 6As shown in (c), "LeakyReLU activation + convolution" in the figure represents a cascaded LeakyReLU activation layer and a 3×3 convolutional layer. DBAB includes a channel attention branch and a structure information branch. The input to the channel attention branch passes sequentially through a 3×3 convolutional layer, a LeakyReLU activation layer, and another 3×3 convolutional layer before being fed into the channel attention layer. Residual connections are introduced within this branch. The input to the structure information branch passes sequentially through a 3×3 convolutional layer, a LeakyReLU activation layer, a 3×3 convolutional layer, a LeakyReLU activation layer, a 3×3 convolutional layer, and then the channel attention layer. Specifically, the output of the second 3×3 convolutional layer in the channel attention branch also serves as the input to the second LeakyReLU activation layer in the structure information branch.

[0094] Furthermore, the expression for the loss function of the dual-branch attention network is:

[0095] (11)

[0096] In the formula, This represents the loss of a two-branch attention network. Describes the L1 loss function. This represents the gradient mean squared error loss function. This represents the edge-preserving loss function. The weights represent the L1 loss function. The weights represent the gradient mean squared error loss function. This represents the weights of the edge-preserving loss function. In a preferred embodiment, , and Set them to 1.0, 0.1 and 0.1 respectively.

[0097] For SAR image denoising The loss function is robust to outliers and is defined as follows:

[0098] (12)

[0099] In the formula, This represents the input image for DBAN, i.e., the SAR noise image. This represents the output image of DBAN, i.e., the denoised SAR image. The width of the SAR noise image. The height of the SAR noise image. It is a loss function that combines the mean squared error between the gradient calculations of Y and X. Its core idea is to force the network to better preserve the edge and detail structure of the image by minimizing the mean squared error of the gradient between the output image and the input image. It is defined as follows:

[0100] (13)

[0101] In the formula, This represents the gradient of the output image Y. This represents the gradient of the input image X.

[0102] To further suppress noise and preserve the edge and detail information of targets in the image, this embodiment also introduces edge preservation loss. . The core idea is to predict images by penalizing them. With real images The difference in the gradient domain, used to smooth out image edges and discontinuities, is defined as follows:

[0103] (14)

[0104] In the formula, This represents the noiseless real image corresponding to the input image X, i.e., the clean label in the training dataset.

[0105] Based on equations (10)-(14), the joint optimization loss function is obtained as shown in equation (15):

[0106] (15)

[0107] In the formula, This represents the joint optimization loss of the dual-branch attention network.

[0108] Then, the dual-branch attention network is trained based on the joint optimization loss function shown in Equation (15) to obtain the trained dual-branch attention network. Specifically, firstly, the SAR noisy image X in the training dataset is input into the DBAN, and the predicted denoised image Y is obtained through forward propagation. Then, the joint optimization loss is calculated based on Equation (15). The backpropagation algorithm and AdamW optimizer were used to iteratively update the network parameters. The number of training iterations was set to 110,000, and the initial learning rate was set to 0.0001. The learning rate was decayed using a cosine annealing scheme, starting from 0.0001 and decreasing smoothly and monotonically to 1 according to a cosine function. 10 -7 End. After the loss function converges, save the network parameters to obtain the trained dual-branch attention network.

[0109] Step S6: Use the trained dual-branch attention network to denoise the SAR noisy image.

[0110] Specifically, the SAR noisy image to be processed is taken as input and fed into a pre-trained dual-branch attention network. During the forward propagation of the network, the Siamese encoder first extracts a robust blob representation from the input image; then, the dual-branch attention group, through the synergistic effect of the channel attention branch and the structural information branch, preserves the edge and texture details of the image while suppressing noise; finally, the image is gradually reconstructed through fully connected layers, convolutional layers, and upsampling layers to output a high-resolution, noise-free SAR image.

[0111] Figure 7 The image shows a comparison of the effects of the SAR noise image denoising method used in this embodiment before and after processing. Figure 7 (a) is the SAR noise image before denoising. Figure 7 (b) shows the denoised SAR image. It can be clearly seen from the image that the noise has been effectively suppressed while the target signal has been completely preserved. To further quantitatively evaluate the denoising effect of the SAR noise image denoising method of this embodiment, Figure 8 (a) shows a comparison of range profiles before and after denoising of a SAR noisy image. Figure 8 (b) shows a comparison of the azimuth profiles of the SAR image before and after denoising. The comparison shows that the edge features of the denoised SAR image are well preserved, and no obvious blurring or distortion is observed.

[0112] Figure 7 and Figure 8 The SAR image denoising method of this embodiment has been verified for its denoising performance on a single target at the center. However, real-world SAR images often contain multiple targets with complex distributions. To evaluate the model's generalization ability, this embodiment further constructs a test set containing multiple targets at multiple locations. Figure 9 This paper demonstrates a comparison of the denoising performance of the denoising method of this embodiment on SAR noisy images in the test set. Figure 9 (a) is a SAR image with multiple targets and locations before denoising. Figure 9 (b) is a SAR image containing multiple targets and locations after denoising. Figure 9 The results show that the denoising method of this embodiment can suppress speckle noise for targets with different spatial distributions, while significantly preserving texture details and edge structures among multiple targets without significant information loss. This result overcomes the limitations of fixed-location single-target verification, indicating that the denoising method of this embodiment has good potential in improving the practical value of SAR images.

[0113] To further verify the advancement of the SAR image denoising method in this embodiment, SAR-CAM (SAR Image Despeckling Using Continuous Attention Module) and MRDDANet (Multiscale Residual Dense Dual Attention Network) were selected as comparative algorithms and quantitatively evaluated on the same test dataset. Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Mean Squared Error (MSE) were introduced as evaluation metrics, and the quantitative results of different denoising methods are shown in Table 1.

[0114] Table 1

[0115] Among them, PSNR is used to estimate the image restoration capability; a higher value indicates a stronger ability of the denoising algorithm to suppress speckle noise. SSIM is used to evaluate the edge preservation degree of the denoised image; a higher value indicates a stronger edge preservation capability of the denoising algorithm. MSE calculates the squared mean of pixel-level errors; a smaller value indicates that the reconstructed image is closer to the original noise-free image. As shown in Table 1, the SAR noise image denoising method (Proposed) in this embodiment outperforms the comparison algorithm in all three metrics: PSNR, SSIM, and MSE. Quantitative results show that the denoising method in this embodiment has significant advantages in speckle suppression and edge preservation capabilities, verifying its advanced nature and effectiveness in denoising network design.

[0116] The SAR image denoising method in this embodiment introduces a twin encoder and a contrastive learning mechanism. First, it extracts speckle representations from the SAR noisy image, and then uses a bi-branch recovery network to project the speckle representations into the embedding space. In the embedding space, the low-frequency components and high-frequency components are decoupled and recovered, enabling the entire feature extraction network to learn more discriminative feature representations. This significantly improves the generalization ability of this method in real complex scenarios and effectively overcomes the problem of performance degradation of traditional methods on real SAR images.

[0117] Building upon this foundation, this method further constructs a dual-branch attention network comprising channel attention and structural information branches. This network performs depth optimization of features and image reconstruction through cascaded dual-branch attention groups, fully connected layers, convolutional layers, and upsampling layers. This effectively suppresses speckle noise while reliably preserving key details such as image edges, textures, and point objects. Furthermore, this method constructs a joint optimization loss function by weightedly fusing the contrastive loss of the Siamese encoder with the loss of the dual-branch attention network. This achieves collaborative optimization between the feature learning and image reconstruction stages. The contrastive loss ensures effective distinction between positive and negative samples in the embedding space, enhancing the discriminative power of features, while the composite loss constrains image reconstruction quality from multiple dimensions, including pixel-level, gradient-level, and structural-level constraints. This joint optimization mechanism enables the dual-branch attention network to achieve a better balance between the competing goals of noise suppression and detail preservation, significantly improving the final denoising performance.

[0118] Furthermore, this method employs an end-to-end training and inference architecture, which can ensure high denoising performance while maintaining high computational efficiency, thereby providing a clearer and more reliable data foundation for subsequent SAR image interpretation (such as target recognition, image segmentation, etc.).

[0119] In summary, the SAR image denoising method of this embodiment introduces a denoising mechanism based on contrastive learning, a bi-branch attention network, and a joint optimization loss function, which can adaptively remove speckle noise from the input image, improve the signal-to-noise ratio of the SAR image, and generate a high-resolution, clear image.

[0120] Example 2

[0121] like Figure 10 As shown, this embodiment provides a SAR image denoising system, including:

[0122] The image acquisition module is used to acquire SAR noise images;

[0123] The image blob extraction module is used to extract blob representations from SAR noisy images using a twin encoder;

[0124] An embedding space generation module is used to project the blob representation onto the embedding space using a dual-branch recovery network, generating a low-frequency information embedding space and a high-frequency information embedding space.

[0125] The embedding loss calculation module is used to calculate the loss of low-frequency information embedding and high-frequency information embedding based on the penalty term, and to calculate the total embedding loss based on the loss of low-frequency information embedding and high-frequency information embedding.

[0126] The network training module is used to optimize the loss function of the dual-branch attention network based on the total embedded loss, obtain the joint optimized loss function, and train the dual-branch attention network based on the joint optimized loss function to obtain the trained dual-branch attention network. The dual-branch attention network includes cascaded dual-branch attention groups, fully connected layers, convolutional layers, and upsampling layers. The dual-branch attention group includes at least one cascaded dual-branch attention block, and the dual-branch attention block includes a channel attention branch and a structural information branch.

[0127] The image denoising module is used to denoise SAR noisy images using a trained dual-branch attention network.

[0128] The system provided in this embodiment can implement the method described in embodiment 1, as detailed in embodiment 1, which will not be repeated here.

[0129] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0130] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A SAR image denoising method, characterized in that, The method includes: Step S1: Acquire SAR noise image; Step S2: Extract query patch, positive sample and negative sample from the SAR noise image, and use a twin encoder to extract the blob representation of the query patch, positive sample and negative sample; Step S3: Project the blob representation onto the embedding space using a dual-branch recovery network to generate a low-frequency information embedding space and a high-frequency information embedding space; Step S4: Calculate the loss of low-frequency information embedding and high-frequency information embedding based on the penalty term, and calculate the total embedding loss based on the loss of low-frequency information embedding and high-frequency information embedding. Step S5: Optimize the loss function of the dual-branch attention network based on the total embedded loss to obtain a joint optimized loss function, and train the dual-branch attention network based on the joint optimized loss function to obtain the trained dual-branch attention network; the dual-branch attention network includes a cascaded dual-branch attention group, a fully connected layer, a convolutional layer, and an upsampling layer, the dual-branch attention group includes at least one cascaded dual-branch attention block, and the dual-branch attention block includes a channel attention branch and a structural information branch; Step S6: Denoise the SAR noise image using the trained dual-branch attention network.

2. The SAR image denoising method according to claim 1, characterized in that, The mathematical expression for the SAR noise image is: In the formula, This represents the SAR noise image. Represents the truth image. This represents non-stationary additive noise. This indicates multiplicative noise.

3. The SAR image denoising method according to claim 1, characterized in that, The twin encoder includes multiple cascaded residual blocks and a pooling layer.

4. The SAR image denoising method according to claim 3, characterized in that, Each residual block's main path includes two types of convolutional blocks. The first type of convolutional block includes cascaded 3×3 convolutional layers, normalization layers, LeakyReLu activation layers, and normalization layers. The second type of convolutional block includes cascaded 1×1 convolutional layers, normalization layers, and LeakyReLu activation layers.

5. The SAR image denoising method according to claim 1, characterized in that, The dual-branch recovery network includes a low-frequency recovery network branch and a high-frequency recovery network branch. The low-frequency recovery network branch includes cascaded 3×3 convolutional layers, nonlinear activation layers, pooling layers, flattening layers, and fully connected layers. The high-frequency recovery network branch includes cascaded 1×1 convolutional layers, nonlinear activation layers, pooling layers, flattening layers, and fully connected layers.

6. The SAR image denoising method according to claim 1, characterized in that, The formula for calculating the total embedding loss is: In the formula, Represents the total loss of embedding. This represents the loss in embedding the low-frequency information. This represents the loss of the high-frequency information embedding. This indicates the weight of the low-frequency recovery branch. This indicates the weight of the high-frequency recovery branch.

7. The SAR image denoising method according to claim 1 or 6, characterized in that, The expression for the joint optimization loss function is: In the formula, This represents the joint optimization loss of the dual-branch attention network. Represents the total loss of embedding. This represents the loss of the dual-branch attention network. Describes the L1 loss function. This represents the gradient mean squared error loss function. This represents the edge-preserving loss function. The weights represent the L1 loss function. The weights represent the gradient mean squared error loss function. This represents the weights of the edge-preserving loss function.

8. The SAR image denoising method according to claim 1, characterized in that, The formulas for calculating the loss of the low-frequency information embedding and the high-frequency information embedding are as follows: In the formula, This represents the loss in the embedding of the low-frequency information or the embedding of the high-frequency information. This represents the neighbor loss of the k-th positive sample pair. This represents the inversion loss of the k-th positive sample pair. Indicates the first Embedded vectors With the Embedded vectors Contrast loss For temperature hyperparameters, Indicates when When, the value is 1, when The value is 0. Represents the i-th embedding vector With the j-th embedding vector Penalties between [these items].

9. The SAR image denoising method according to claim 1 or 8, characterized in that, The expression for the penalty term is: In the formula, This represents the i-th embedding vector. This represents the j-th embedding vector.

10. A SAR image denoising system, characterized in that, The system includes: The image acquisition module is used to acquire SAR noise images; An image blob extraction module is used to extract query patches, positive samples, and negative samples from the SAR noise image, and to extract blob representations of the query patches, positive samples, and negative samples using a twin encoder; An embedding space generation module is used to project the blob representation onto the embedding space using a dual-branch recovery network to generate a low-frequency information embedding space and a high-frequency information embedding space. The embedding loss calculation module is used to calculate the loss of the low-frequency information embedding and the high-frequency information embedding based on the penalty term, and to calculate the total embedding loss based on the loss of the low-frequency information embedding and the high-frequency information embedding. The network training module is used to optimize the loss function of the dual-branch attention network based on the total embedded loss to obtain a joint optimized loss function, and to train the dual-branch attention network based on the joint optimized loss function to obtain the trained dual-branch attention network. The dual-branch attention network includes a cascaded dual-branch attention group, a fully connected layer, a convolutional layer, and an upsampling layer. The dual-branch attention group includes at least one cascaded dual-branch attention block, and the dual-branch attention block includes a channel attention branch and a structural information branch. The image denoising module is used to denoise the SAR noisy image using the trained dual-branch attention network.