Self-supervised optical remote sensing image blind denoising method and system
The image pair set is constructed through non-local similarity sampling and neighborhood Bernoulli sampling, combined with parameterless attention calculation and partial convolution, the problem of poor real noise processing by self-supervised optical remote sensing image denoising method is solved, and efficient image denoising and computing resource optimization is achieved.
Patent Information
- Application Number
- CN202510620901.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing self-supervised optical remote sensing image denoising method is poor in processing real noise, especially the denoising effect of spatially related noise and non-independently distributed noise is limited, and the computing resource consumption is high, making it difficult to meet the processing needs of high-resolution images.
The image pair set is constructed using non-local similarity sampling and neighborhood Bernoulli sampling. Through self-supervised learning constraints, combined with parameterless attention calculation and partial convolution, the model parameter occupation is reduced, the noise spatial correlation is destroyed, and the image detail texture is restored.
Effectively process space-related noise, improve image denoising effect, reduce computing resource consumption, and adapt to the denoising needs of high-resolution optical remote sensing images.
Smart Images

Figure CN120495120A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image denoising, and in particular to a method and system for self-supervised blind denoising of optical remote sensing images. Background Art
[0002] In the field of computer vision, digital image denoising is a crucial step in the image preprocessing stage, and its processing quality plays a decisive role in the accuracy of downstream tasks. High-quality denoising provides a solid foundation for subsequent image analysis and processing. In recent years, the rapid development of deep learning has brought new solutions to the field of image denoising. Deep learning-based image denoising techniques have begun to be widely used in the field of optical remote sensing imagery.
[0003] Optical remote sensing images, as an important data source for obtaining surface information, cover a wealth of topographic features and information on various ground objects, including mountains, forests, buildings, roads, etc. Its downstream analysis and processing tasks, such as topographic mapping and land use classification, have extremely high requirements for the clarity of image detail textures. Compared with natural image denoising, optical remote sensing image denoising not only removes noise, but also needs to restore the image detail texture to the greatest extent possible to meet the needs of high-precision analysis. However, due to the complex and diverse spatial feature layout of optical remote sensing images and the extremely rich detail texture information they contain, accurately restoring the original detail texture from noise interference has become an extremely challenging task.
[0004] Throughout the development of optical remote sensing image denoising technology, deep learning-based denoising methods, with their superior denoising performance and flexible parameter settings, have gradually surpassed traditional filtering denoising methods and become the mainstream technology in this field. However, the development and optimization of deep learning denoising models rely heavily on large-scale training data. In particular, mainstream supervised learning methods require a large number of noise-free real images as training labels to ensure model effectiveness and generalization. However, in practical applications, obtaining large-scale paired "truth-noise" remote sensing images presents numerous challenges. Not only is the data collection process complex, but it is also costly. With limited training data, deep learning denoising models often fail to meet expected performance.
[0005] To address this issue, self-supervised learning methods have emerged. This method does not rely on real labeled data, effectively reducing its reliance on data. However, during training, to avoid overfitting, self-supervised learning methods typically discard some pixel information. This leads to certain limitations in the model's ability to restore image detail textures, making it difficult to achieve ideal denoising effects. Furthermore, image denoising, as a generative task, places more stringent demands on GPU memory than tasks such as image classification and object detection. When processing high-resolution optical remote sensing images, the model's memory usage and computational complexity increase exponentially, severely impacting the model's operational efficiency and posing significant challenges to both hardware and algorithm design.
[0006] Furthermore, most existing self-supervised denoising methods, such as Noise2Noise and Self2Self, have implicit theoretical assumptions: they assume that the noise in the noisy image is spatially independently distributed, or that the noise signal has zero mean. This makes these self-supervised denoising methods excellent for independently distributed noise, such as additive Gaussian noise, but leaves much room for improvement in denoising non-spatially independently distributed noise, such as Poisson distributed noise or photon noise introduced by optical sensors. Consequently, these methods struggle to effectively target the real noise signals present in optical remote sensing images, which impacts the overall image denoising performance. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this paper provides a self-supervised blind denoising method and system for optical remote sensing images. For a single optical remote sensing image with real mixed noise, the system destroys the spatial correlation of the noise by swapping similar non-neighborhood pixels to target spatially correlated components of the real noise. Random Bernoulli sampling is then used to mask a small number of pixels to construct a set of distinct sub-image pairs, targeting spatially independent noise components and implementing self-supervised learning constraints. Finally, the denoised image is obtained by decoding the denoised encoded feature maps and taking the mean of the resulting image sets.
[0008] On the one hand, a self-supervised blind denoising method for optical remote sensing images is provided, including: Acquire a remote sensing image to be denoised, perform non-local similarity sampling on the remote sensing image to be denoised, and obtain a first noise-free spatial similarity image; Performing neighborhood Bernoulli sampling on the first noise-space-independent similar images to obtain a first image pair set; The first image pair set is input into the trained denoising model to obtain a denoised remote sensing image.
[0009] On the other hand, a self-supervised optical remote sensing image blind denoising system is provided, including: An acquisition module is configured to: acquire a remote sensing image to be denoised, perform non-local similarity sampling on the remote sensing image to be denoised, and obtain a first noise-free spatial similarity image; A sampling module is configured to: perform neighborhood Bernoulli sampling on the first noise-space-independent similar image to obtain a first image pair set; The denoising module is configured to: input the first image pair set into the trained denoising model to obtain a denoised remote sensing image.
[0010] The above technical solution has the following advantages or beneficial effects: The present invention has a good processing effect on spatially correlated noise. This method is mainly used to process a single optical remote sensing image with real mixed noise. For the spatially correlated components in the real noise, it destroys the spatial correlation of the noise by exchanging non-neighborhood similar pixels. For the spatially independent noise components, random Bernoulli sampling is randomly used to mask a small number of pixels to construct different sets of sub-image pairs, thereby realizing self-supervised learning constraints. And through the parameter-free attention calculation method SimAM, the additional video memory occupancy caused by the model parameters is reduced, and the model is guided to pay more attention to the detailed texture features in the remote sensing image. Finally, the feature map after denoising encoding is decoded first, and then the mean of the image set results is calculated to obtain the final denoised image. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0012] Figure 1 This is a flow chart of the method of embodiment 1. DETAILED DESCRIPTION
[0013] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0014] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention. The terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0015] Example 1 This embodiment provides a self-supervised optical remote sensing image blind denoising method; like Figure 1 As shown in FIG, the self-supervised optical remote sensing image blind denoising method includes: S101: Acquire a remote sensing image to be denoised, perform non-local similarity sampling on the remote sensing image to be denoised, and obtain a first noise-free spatial similarity image; S102: performing neighborhood Bernoulli sampling on the first noise-independent similar images to obtain a first image pair set; S103: Input the first image pair set into the trained denoising model to obtain a denoised remote sensing image.
[0016] Furthermore, the step S101 of obtaining a remote sensing image to be denoised and performing non-local similarity sampling on the remote sensing image to be denoised to obtain a first noise-free spatial similarity image specifically includes: Calculate the cosine similarity between each pixel in the remote sensing image to be denoised and all other non-neighborhood pixels in the image; Sort all non-neighborhood pixels in descending order of similarity values and obtain the first m non-neighborhood pixels with the highest similarity; Each pixel in the remote sensing image to be denoised is regarded as a target pixel. From the first m non-neighborhood pixels with the highest similarity, a non-neighborhood pixel is randomly selected for each target pixel as a similar pixel of the target pixel. Then, the target pixel and the similar pixel are regarded as a matching pair. Finally, a matching pair set of all pixels corresponding to the remote sensing image to be denoised is obtained. Then, based on the set of matching pairs, similar pixels in each matching pair are replaced with the positions of corresponding target pixels to obtain a similar noise image; the size of the similar noise image is consistent with the size of the remote sensing image to be denoised; the similar noise image is the first noise space-independent similar image.
[0017] It should be understood that due to the complex distribution of real noise in remote sensing images, the noise signals are usually spatially correlated. In order to reduce the correlation between image content and noise signals, the present invention adopts a non-local pixel sampling and replacement method to calculate the cosine similarity between each pixel in the input image and all other non-neighboring pixels in the image. , the formula is: (1) in, is the current target pixel, After sorting the similarities from high to low, filter the first m non-neighborhood pixels with the highest similarity, and randomly record one of the non-neighborhood pixels similar to the current pixel (randomly select the first m non-neighborhood pixels). pixels, where and is an integer) is ,in is a randomly selected similar non-neighborhood pixel. After traversing all pixels in the input image as target pixels, a similar pixel matching set is obtained , is the number of pixels in the input image. Arrange each similar pixel to the position of the corresponding original pixel to obtain an image with the same size as the original input image. Same similar noise image .
[0018] Since similar noise images The spatial positions of the pixels in the original image are randomly arranged, which minimizes the spatial correlation of the noise signal in the new similar image as much as possible. By replacing similar pixels, the overall characteristics of the image are not significantly changed, ensuring the integrity of the image content and realizing the separation of the spatial correlation between the image content and the noise signal, laying the foundation for subsequent self-supervised denoising.
[0019] Furthermore, the step S102 of performing neighborhood Bernoulli sampling on the first noise-independent spatially similar image to obtain a first image pair set specifically includes: Similar images that are independent of the first noise space Repeated random Bernoulli sampling is performed to generate multiple pairs of images with different pixels masked. : ; (2-1) ; (2-2) in, and represents the complementary image generated after sampling, is the number of repeated samplings, Represents the corresponding multiplication operation of matrix elements, is the randomly generated masking matrix in Bernoulli sampling, Each element in The value rules are as follows: (3) The meaning of formula (3) is that for each element Randomly generate a random number between 0 and 1. If the random number is less than or equal to , then the element The value is 1; if the random number is greater than and less than or equal to , then the element The value of is 1.
[0020] Each element in the masking matrix has a probability The value of y may be 1, otherwise it is 0, thus forming a mask matrix B whose element value is only 1 or 0. After the mask matrix B is multiplied by the corresponding elements of the image y, in the new output image, the position where the element in B is 1 retains the pixel value of y, and the others (that is, the position where the element in B is 0) are set to 0.
[0021] It should be understood that in order to enable the denoising model to converge with only a single noisy image as input and avoid overfitting during training, the present invention adopts the neighborhood Bernoulli sampling method to randomly generate multiple pairs of "complementary" sub-image pairs from a single input image as training data for the model.
[0022] In the random Bernoulli sampling process of an image, each pixel is treated as an independent Bernoulli trial. For each pixel, a probability is set , probability Represents the probability that the current pixel remains unchanged. Correspondingly, is the probability that the pixel changes. It is used to decide for each pixel in the image whether it has changed and whether it should be kept or removed.
[0023] The current pixel has The probability of does not change, otherwise (i.e. The probability of the pixel being changed to 0. Specifically, a random number between 0 and 1 is generated for each pixel. If the random number is less than or equal to , it will not change; if the random number is less than , it changes to 0.
[0024] For each pixel in the input image, a set probability is used to decide whether to keep or modify the pixel. Random Bernoulli sampling introduces randomness while maintaining the basic structure of the image to avoid overfitting during self-supervised training. Since the masking matrix Generated randomly according to Bernoulli distribution, so in the set Each pair of images in The random Bernoulli sampling constrains the denoising model to learn the mapping of the input image itself, thus avoiding overfitting of the model training. In addition, during the sampling process, the probability of each pixel being covered is equal.
[0025] When the number of repetitions When large enough, the input image All pixels in the set are included in In the training, it can be considered that all the information in the original input image is involved in the training, and the image information loss caused by Bernoulli sampling is avoided as much as possible.
[0026] Furthermore, the step of obtaining the trained denoising model in S103 includes: S103-1: Acquire a remote sensing image dataset, add noise to all remote sensing images in the remote sensing image dataset, and obtain a noisy remote sensing image dataset; S103-2: performing non-local similarity sampling on each noisy remote sensing image in the noisy remote sensing image dataset to obtain a second noise spatially independent similar image; S103-3: performing neighborhood Bernoulli sampling on the second noise-space-independent similar images to obtain a second image pair set; S103-4: Input the second image pair set into the denoising model, train the model, and obtain a trained denoising model.
[0027] Furthermore, the step S103-1: obtaining a remote sensing image dataset, adding noise to all remote sensing images in the remote sensing image dataset to obtain a noisy remote sensing image dataset, includes: A known remote sensing image dataset is used, Gaussian-Poisson noise is added to the known remote sensing image dataset, and the image size is normalized to obtain a noisy remote sensing image dataset.
[0028] It should be understood that due to the lack of large-scale noisy remote sensing image datasets specifically designed for denoising tasks, in this paper, the UC-Merced and Optical-31 remote sensing image classification datasets were used to construct a noisy image dataset for training, with the addition of Gaussian-Poisson joint noise. This constructed noisy dataset was randomly divided into training and test sets at a ratio of 10:1. Due to the self-supervised learning nature of this invention, the ground-truth images in the original UC-Merced and Optical-31 datasets were not used in the training process and were only used as reference images for evaluation metrics.
[0029] Because some high-resolution remote sensing images are often large, computational resources often limit the original image size, requiring resizing and cropping before feeding it into the model. The model supports direct input of images up to 512×512 pixels. For high-resolution images significantly larger than this, they are cropped to create multiple subimages that meet this size. For input images smaller than this size or with an aspect ratio other than 1:1, they are upsampled and resized to fit within the 512×512 pixel input size.
[0030] Furthermore, the network structure of the trained denoising model includes: The encoders connected in sequence, the first 3 3 convolutional layers, the second 3 3 convolutional layers and decoder; The encoder comprises: a first encoding module, a second encoding module, a third encoding module, a fourth encoding module, a fifth encoding module and a sixth encoding module connected in sequence; The decoder includes: a first decoding module, a second decoding module, a third decoding module, a fourth decoding module, a fifth decoding module, a sixth decoding module, a first upsampling layer and a 1 convolutional layer; The first encoding module includes: a first convolutional layer, a second convolutional layer, a first activation function layer, a first attention mechanism layer, a first element-by-element multiplication unit, and a first maximum pooling layer connected in sequence; The second encoding module includes: a third convolutional layer, a fourth convolutional layer, a second activation function layer, a second attention mechanism layer, a second element-by-element multiplication unit, and a second maximum pooling layer connected in sequence; The third encoding module includes: a fifth convolutional layer, a sixth convolutional layer, a third activation function layer, a third attention mechanism layer, a third element-by-element multiplication unit, and a third maximum pooling layer connected in sequence; The fourth encoding module includes: a seventh convolutional layer, an eighth convolutional layer, a fourth activation function layer, a fourth attention mechanism layer, a fourth element-by-element multiplication unit, and a fourth maximum pooling layer connected in sequence; The fifth encoding module includes: a ninth convolutional layer, a tenth convolutional layer, a fifth activation function layer, a fifth attention mechanism layer, a fifth element-by-element multiplication unit, and a fifth maximum pooling layer connected in sequence; The sixth encoding module includes: an eleventh convolutional layer, a twelfth convolutional layer, a sixth activation function layer, a sixth attention mechanism layer, and a sixth element-by-element multiplication unit connected in sequence; The first decoding module includes: a seventh element-by-element multiplication unit, a thirteenth convolutional layer, a fourteenth convolutional layer, and a first adder connected in sequence; The second decoding module includes: a second upsampling layer, an eighth element-by-element multiplication unit, a fifteenth convolutional layer, a sixteenth convolutional layer, and a second adder connected in sequence; The third decoding module includes: a third upsampling layer, a ninth element-by-element multiplication unit, a seventeenth convolutional layer, an eighteenth convolutional layer, and a third adder connected in sequence; The fourth decoding module includes: a fourth upsampling layer, a tenth element-by-element multiplication unit, a nineteenth convolutional layer, a twentieth convolutional layer, and a fourth adder connected in sequence; The fifth decoding module includes: a fifth upsampling layer, an eleventh element-by-element multiplication unit, a twenty-first convolutional layer, a twenty-second convolutional layer, and a fifth adder connected in sequence; The sixth decoding module includes: a sixth upsampling layer, a twelfth element-by-element multiplication unit, a twenty-third convolutional layer, a twenty-fourth convolutional layer and a sixth adder connected in sequence.
[0031] The output of the first attention mechanism layer is connected to the input of the twelfth element-wise multiplication unit; The output of the second attention mechanism layer is connected to the input of the eleventh element-wise multiplication unit; The output of the third attention mechanism layer is connected to the input of the tenth element-wise multiplication unit; The output of the fourth attention mechanism layer is connected to the input of the ninth element-wise multiplication unit; The output of the fifth attention mechanism layer is connected to the input of the eighth element-wise multiplication unit; The output of the sixth attention mechanism layer is connected to the input of the seventh element-wise multiplication unit; The output end of the first convolutional layer is connected to the input end of the sixth adder; the output end of the second convolutional layer is connected to the input end of the fifth adder; the output end of the third convolutional layer is connected to the input end of the fourth adder; the output end of the fourth convolutional layer is connected to the input end of the third adder; the output end of the fifth convolutional layer is connected to the input end of the second adder; and the output end of the sixth convolutional layer is connected to the input end of the first adder.
[0032] For example, the model in the present invention is based on the U-Net structure. In the encoding stage of the present invention, the model consists of 6 encoder blocks (EBs). The first 5 EBs consist of two layers. The 6th EB consists of a partial convolutional layer + a LeakyReLu activation layer, a parameter-free attention mechanism layer Sim-AM, and a max pooling layer, while the 6th EB does not include a max pooling layer. In each EB, the attention map output by Sim-AM guides the model to focus more on the detailed features that are meaningful in the high-level semantics during the encoding process.
[0033] Furthermore, in the denoising model of this invention, the feature map input to the convolutional layer undergoes random Bernoulli sampling, causing some pixels to be masked and set to zero. Conventional convolution operations, when processing such feature maps, fail to distinguish between valid pixels and masked pixels with a value of zero, which can easily lead to undesirable results such as color discrepancies and artifacts. To overcome this problem, this method uses partial convolution to replace the original convolution operation in the encoder.
[0034] Partial Convolution only performs convolution operations on valid pixels and also has the function of automatically updating the mask to ensure that the convolution process will not be disturbed by pixels that are masked to 0, thereby effectively solving the problems caused by improper convolution operations in image completion tasks.
[0035] Compared with the usual convolutional neural network, the partial convolution requires the input of a mask matrix to avoid convolution on invalid pixels. Its formula is as follows: (4) in, is the convolution kernel, is the input image within the receptive field of the current convolution, is the mask matrix within the receptive field of the current convolution, Represents the size of the convolution kernel, Represents the receptive field of the convolution The area of the matrix where the value is 1.
[0036] In partial convolution, the mask matrix in each layer of convolution is updated as the number of layers increases. The update rule is as follows: if the convolution can calculate the output based on at least one valid input value, then in the next step, we will update the mask matrix of this pixel. Marked as 1. That is, the partial convolution The point in the layer The value is given by Within the receptive field of the layer For example: if Within the receptive field of the layer If the value is 1, then the point is in Layer partial convolution The value is 1, otherwise it is 0. It can be expressed as the following formula, For the updated The value of the element in: (5) Formula (5) is used to explain the update process of the mask matrix mask in each layer of partial convolution: It is an element in the mask matrix of the next layer of partial convolution. If at least one element in the mask matrix is 1 within the receptive field of the currently calculated convolution kernel in this layer, then the mask element value of the corresponding position of the convolution kernel at the same position in the next layer of convolution is still 1, otherwise it is 0.
[0037] Furthermore, in order to reduce the number of parameters when calculating the attention map to save memory space, the present invention uses the parameter-free attention mechanism Sim-AM to calculate the attention map. According to relevant research on neuroscience theory, if neurons have unique discharge patterns, they usually carry richer information and can efficiently encode and transmit complex information, thereby playing a key role in the visual processing process. In the process of neuronal activity, there is also a phenomenon called neuronal spatial inhibition, that is, active neurons have an inhibitory effect on the neurons around them. Specifically, an active neuron can reduce the activity level of nearby neurons. This feature provides important clues for studying the activity patterns of neurons. By estimating and calculating the energy function of neurons, it can become a key means to identify and locate activated neurons. In view of the existence of the neuronal spatial inhibition phenomenon, there are significant differences between activated neurons and surrounding neurons, so the activation degree of the target neuron can be reflected by calculating the linear separability of the target neuron and its neighboring neurons. Therefore, there is a neuron energy estimation function: (6) in, for Linear transformation of , for Linear transformation of .in is the current target neuron, are neurons other than the target neuron, is the number of all neurons . and are two different constants. For simplicity, and Defined as binary labels and The purpose of this formula is to find the linear separability between the current target neuron and other neurons. and When , in formula (6) Reaching the minimum value, at this time there is a target neuron The neuron that is most different from other neurons should carry richer effective feature information.
[0038] Furthermore, in order to simplify the amount of calculation in the process, the quick solution for formula (6) is: (7-1) (7-2) Among them, there are and .
[0039] Based on this, the minimum energy value of the target neuron is derived as follows: (8) In formula (8), The smaller the value, the greater the difference between the target neuron and other neurons, and more attention should be given to the target neuron in the denoising task.
[0040] Therefore, the inverse of the result of formula (8) is used As a measure of the importance of neurons. After traversing all neurons, all Arrange the neurons according to the positions of the original input feature map to obtain the output value attention map weight matrix of the attention mechanism layer .
[0041] In general, the encoding stage of the model in the present invention is a feature extraction network with 6 encoder blocks. Each encoder block contains a partial convolution with a convolution kernel of 3×3, and uses the Sim-AM attention module to guide the model to focus on detail features. At the same time, the structural characteristics of the U-Net network are used to cross-transfer the feature map and attention map in the encoder block of this level to the same level of the decoder. Specifically speaking of the differences in encoder structure, the first five encoders are composed of two convolution layers and one maximum pooling layer, respectively, and the sixth encoder has only two convolution layers without a pooling layer. All six encoder blocks contain attention modules. When the input image size is When , the feature map size of the final output of the model encoding stage is .
[0042] The decoding stage of the model in this invention consists of 6 decoder blocks and an output layer. The first decoder includes two layers of convolution kernels. The remaining decoders add an upsampling layer with a scaling factor of 2 to the input to gradually restore the size of the feature map. The final output stage consists of a convolution layer with a convolution kernel of 1×1, which compresses the 96 channels of the feature map to the number of channels of the input image (3 in this method) to ensure that the output image has the same size as the input image. In addition, the U-Net cascade operation combines the feature map passed from the encoder at the same level with the encoded attention. Figure 1 This design allows the attention map weights learned from the encoder to be effectively passed to the decoder, achieving effective transfer of attention features between the encoder and decoder, and ensuring the consistency and coherence of the feature extraction and reconstruction stages throughout the denoising task.
[0043] Furthermore, the loss function used in the training process of the trained denoising model is: Assume random Bernoulli sampling is , then the similar images of the first noise space are independent The two sub-images obtained after random sampling are , .structure As the input of the denoising model, the optimization objective of the denoising model is: (9) in, and is a complementary image, that is, it satisfies , is the masking matrix in Bernoulli sampling, is a sampling operation, It is the image output after y is sampled. is the denoising model.
[0044] At the same time, in order to avoid color deviation and distortion in the denoised image, a perceptual loss term is added to the target optimization problem to ensure the consistency of the denoised image in high-level features. The feature extraction model in the perceptual loss uses a pre-trained ResNet18 network, and its calculation formula is as follows: (10) in, For the pre-trained network Resnet18 The feature map output by the residual block, is the number of residual blocks of the currently used pre-trained network Resnet18. When the feature extraction model is ResNet18, .
[0045] The complete objective function of the denoising model is: (11) in, is a hyperparameter representing the weight coefficient of perceptual loss.
[0046] It should be understood that during the training phase, it is necessary to perform multiple random Bernoulli sampling on the similar images that are independent of the first noise space after similar replacement to obtain a set of multiple image pairs. As the model input, the denoising model needs to be optimized as a whole, and its objective function is: (12) in, is the masking matrix of the mth sampling, is the denoising model, are model parameters, is the number of repeated samplings, is the square of the 2-norm.
[0047] According to the above formula, only The covered pixels participate in the measurement of each pair of images Due to the loss Generated completely randomly according to Bernoulli distribution, when the number of repetitions When is large enough, the difference of all image pixels can be measured by the total loss of all pairs.
[0048] During the training process, the present invention uses the Adam optimizer , and use dynamic learning rate to optimize the network. The detailed settings of learning rate are as follows: the initial learning rate is , and decays in steps when the number of iterations reaches 20%, 40%, 60%, and 80% of the total number of iterations, and the decay coefficient is .
[0049] Hyperparameters in loss functions The calculation formula is as follows: (15) in, is the current iteration round, is the total number of iterations, is the attenuation coefficient, which is 2 and 0.5 respectively in the present invention. The random Bernoulli number of the model is , that is, a total of 100 pairs of randomly sampled subgraph pairs are used as training data, and the total number of iterations is When the model converges to the optimal state, that is, when the loss function reaches the minimum, the model of the current training round is used as the optimal model for the inference process.
[0050] In the inference phase, the input image Multiple random Bernoulli samplings are performed to obtain multiple sets of new images, and the newly generated image sets are input into the trained denoising model.
[0051] Finally, all the output denoised images The process of averaging to get the final result can be expressed as: (16) in, For the input image No. Example of the output after Bernoulli sampling.
[0052] The two multiple random Bernoulli samplings performed during training and inference are performed independently. Although this method is computationally expensive, it effectively prevents model overfitting during training. During actual training, this method further mitigates overfitting by performing data augmentation operations such as rotation and mirroring on the input images.
[0053] Example 2 This embodiment provides a self-supervised optical remote sensing image blind denoising system, including: An acquisition module is configured to: acquire a remote sensing image to be denoised, perform non-local similarity sampling on the remote sensing image to be denoised, and obtain a first noise-free spatial similarity image; a sampling module configured to: perform neighborhood Bernoulli sampling on the first noise-space-independent similar images to obtain a first set of image pairs; The denoising module is configured to: input the first image pair set into the trained denoising model to obtain a denoised remote sensing image.
[0054] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A self-supervised blind denoising method for optical remote sensing images, characterized by: include: Acquire a remote sensing image to be denoised, perform non-local similarity sampling on the remote sensing image to be denoised, and obtain a first noise-free spatial similarity image; Performing neighborhood Bernoulli sampling on the first noise-space-independent similar images to obtain a first image pair set; The first image pair set is input into the trained denoising model to obtain a denoised remote sensing image.
2. The self-supervised optical remote sensing image blind denoising method according to claim 1, wherein: Acquire a remote sensing image to be denoised, perform non-local similarity sampling on the remote sensing image to be denoised, and obtain a first noise-free spatial similarity image, specifically including: Calculate the cosine similarity between each pixel in the remote sensing image to be denoised and all other non-neighborhood pixels in the image; Sort all non-neighborhood pixels in descending order of similarity values and obtain the first m non-neighborhood pixels with the highest similarity; Each pixel in the remote sensing image to be denoised is regarded as a target pixel. From the first m non-neighborhood pixels with the highest similarity, a non-neighborhood pixel is randomly selected for each target pixel as a similar pixel of the target pixel. Then, the target pixel and the similar pixel are regarded as a matching pair. Finally, a matching pair set of all pixels corresponding to the remote sensing image to be denoised is obtained. Then, based on the set of matching pairs, similar pixels in each matching pair are replaced with the positions of corresponding target pixels to obtain a similar noise image; the size of the similar noise image is consistent with the size of the remote sensing image to be denoised; the similar noise image is the first noise space-independent similar image.
3. The self-supervised optical remote sensing image blind denoising method according to claim 1, wherein: The first noise-space-independent similar images are subjected to neighborhood Bernoulli sampling to obtain a first image pair set, specifically including: Similar images that are independent of the first noise space Repeated random Bernoulli sampling is performed to generate multiple pairs of images with different pixels masked. : ; ; in, and represents the complementary image generated after sampling, is the number of repeated samplings, Represents the matrix element corresponding multiplication operation, is the randomly generated masking matrix in Bernoulli sampling, Each element in The value rules are as follows: ; For each element Randomly generate a random number between 0 and 1. If the random number is less than or equal to , then the element The value is 1; if the random number is greater than and less than or equal to , then the element The value of is 1.
4. The self-supervised optical remote sensing image blind denoising method according to claim 1, wherein: The steps to obtain the trained denoising model include: A remote sensing image dataset is obtained, and noise is added to all remote sensing images in the remote sensing image dataset to obtain a noisy remote sensing image dataset; For each noisy remote sensing image in the noisy remote sensing image dataset, non-local similarity sampling is performed to obtain a second noise space-independent similar image; Performing neighborhood Bernoulli sampling on the second noise-independent similar images to obtain a second set of image pairs; The second image pair set is input into the denoising model, and the model is trained to obtain a trained denoising model.
5. The self-supervised optical remote sensing image blind denoising method according to claim 1, wherein: The trained denoising model has a network structure comprising: The encoders connected in sequence, the first 3 3 convolutional layers, the second 3 3 convolutional layers and decoder; The encoder comprises: a first encoding module, a second encoding module, a third encoding module, a fourth encoding module, a fifth encoding module and a sixth encoding module connected in sequence; The decoder includes: a first decoding module, a second decoding module, a third decoding module, a fourth decoding module, a fifth decoding module, a sixth decoding module, a first upsampling layer and a 1 convolutional layer; The first encoding module includes: a first convolutional layer, a second convolutional layer, a first activation function layer, a first attention mechanism layer, a first element-by-element multiplication unit, and a first maximum pooling layer connected in sequence; The second encoding module includes: a third convolutional layer, a fourth convolutional layer, a second activation function layer, a second attention mechanism layer, a second element-by-element multiplication unit, and a second maximum pooling layer connected in sequence; The third encoding module includes: a fifth convolutional layer, a sixth convolutional layer, a third activation function layer, a third attention mechanism layer, a third element-by-element multiplication unit, and a third maximum pooling layer connected in sequence; The fourth encoding module includes: a seventh convolutional layer, an eighth convolutional layer, a fourth activation function layer, a fourth attention mechanism layer, a fourth element-by-element multiplication unit, and a fourth maximum pooling layer connected in sequence; The fifth encoding module includes: a ninth convolutional layer, a tenth convolutional layer, a fifth activation function layer, a fifth attention mechanism layer, a fifth element-by-element multiplication unit, and a fifth maximum pooling layer connected in sequence; The sixth encoding module includes: an eleventh convolutional layer, a twelfth convolutional layer, a sixth activation function layer, a sixth attention mechanism layer and a sixth element-by-element multiplication unit connected in sequence.
6. The self-supervised optical remote sensing image blind denoising method according to claim 5, wherein: The first decoding module includes: a seventh element-by-element multiplication unit, a thirteenth convolutional layer, a fourteenth convolutional layer, and a first adder connected in sequence; The second decoding module includes: a second upsampling layer, an eighth element-by-element multiplication unit, a fifteenth convolutional layer, a sixteenth convolutional layer, and a second adder connected in sequence; The third decoding module includes: a third upsampling layer, a ninth element-by-element multiplication unit, a seventeenth convolutional layer, an eighteenth convolutional layer, and a third adder connected in sequence; The fourth decoding module includes: a fourth upsampling layer, a tenth element-by-element multiplication unit, a nineteenth convolutional layer, a twentieth convolutional layer, and a fourth adder connected in sequence; The fifth decoding module includes: a fifth upsampling layer, an eleventh element-by-element multiplication unit, a twenty-first convolutional layer, a twenty-second convolutional layer, and a fifth adder connected in sequence; The sixth decoding module includes: a sixth upsampling layer, a twelfth element-by-element multiplication unit, a twenty-third convolutional layer, a twenty-fourth convolutional layer, and a sixth adder connected in sequence; The output end of the first attention mechanism layer is connected to the input end of the twelfth element-by-element multiplication unit; the output end of the second attention mechanism layer is connected to the input end of the eleventh element-by-element multiplication unit; the output end of the third attention mechanism layer is connected to the input end of the tenth element-by-element multiplication unit; the output end of the fourth attention mechanism layer is connected to the input end of the ninth element-by-element multiplication unit; the output end of the fifth attention mechanism layer is connected to the input end of the eighth element-by-element multiplication unit; and the output end of the sixth attention mechanism layer is connected to the input end of the seventh element-by-element multiplication unit. The output end of the first convolutional layer is connected to the input end of the sixth adder; the output end of the second convolutional layer is connected to the input end of the fifth adder; the output end of the third convolutional layer is connected to the input end of the fourth adder; the output end of the fourth convolutional layer is connected to the input end of the third adder; the output end of the fifth convolutional layer is connected to the input end of the second adder; and the output end of the sixth convolutional layer is connected to the input end of the first adder.
7. The self-supervised optical remote sensing image blind denoising method according to claim 1, wherein: The loss function used in the training process of the trained denoising model is: Assume random Bernoulli sampling is , then the similar images of the first noise space are independent The two sub-images obtained after random sampling are , ;structure As the input of the denoising model, the optimization objective of the denoising model is: ; in, and is a complementary image, that is, it satisfies , is the masking matrix in Bernoulli sampling, yes The image output after the sampling operation is is the denoising model; ; in, For the pre-trained network Resnet18 The feature map output by the residual block, is the number of residual blocks of the currently used pre-trained network Resnet18; The complete objective function of the denoising model is: ; in, is a hyperparameter representing the weight coefficient of perceptual loss.
8. The self-supervised optical remote sensing image blind denoising method according to claim 7, wherein: Hyperparameters in loss functions The calculation formula is as follows: ; in, is the current iteration round, is the total number of iterations, is the attenuation coefficient.
9. The self-supervised optical remote sensing image blind denoising method according to claim 4, wherein: A remote sensing image dataset is obtained, and noise is added to all remote sensing images in the remote sensing image dataset to obtain a noisy remote sensing image dataset, including: using a known remote sensing image dataset, adding Gaussian-Poisson noise to the known remote sensing image dataset, and normalizing the image size to obtain the noisy remote sensing image dataset.
10. A self-supervised blind denoising system for optical remote sensing images, characterized by: include: An acquisition module is configured to: acquire a remote sensing image to be denoised, perform non-local similarity sampling on the remote sensing image to be denoised, and obtain a first noise-free spatial similarity image; A sampling module is configured to: perform neighborhood Bernoulli sampling on the first noise-space-independent similar image to obtain a first image pair set; The denoising module is configured to: input the first image pair set into the trained denoising model to obtain a denoised remote sensing image.
Citation Information
Patent Citations
Feature weighted connection guided single-image remote sensing image denoising method
CN117876692A
Real image denoising algorithm based on adaptive feature fusion
CN118314040A
Single-input optical remote sensing image self-supervision blind denoising method and system
CN119067877A