Self-supervised remote sensing image strip noise suppression method and system based on frequency domain structure prior
By employing a self-supervised learning method with frequency domain structure priors and cross-band consistency self-supervised constraints, this method addresses the shortcomings of existing techniques in remote sensing image stripe noise suppression, achieving effective suppression of stripe noise and protection of the true structure, thereby enhancing the generalization ability and interpretability of the method.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHEAST FORESTRY UNIV
- Filing Date
- 2026-02-28
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies for suppressing strip noise in remote sensing images suffer from several drawbacks: frequency domain detection lacks cross-band structure discrimination mechanisms, fixed filtering methods lack adaptability, deep learning methods rely on clean ground truth supervision, making it difficult to train stably on real remote sensing data, and there is a lack of a unified framework that integrates prior information on frequency domain structure with cross-band consistency information.
A self-supervised remote sensing image strip noise suppression method based on frequency domain structure prior is adopted. By modeling multi-band strip observations, frequency domain structure priors and cross-band consistency self-supervised constraints are constructed. A deep suppression network is established and trained using a self-supervised learning framework to achieve adaptive suppression of strip noise and protection of the real structure.
It significantly reduces strip frequency band energy, improves the fidelity of the real structure, enhances generalization ability, provides interpretable suppression results, and is applicable to large-format remote sensing image processing with high computational efficiency.
Smart Images

Figure CN122066604A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of digital image processing and remote sensing image analysis, and in particular to a method for suppressing strip noise in remote sensing images. Background Technology
[0002] Striping noise is a long-standing and typical problem in remote sensing image processing, especially prevalent in pushbroom or multi-detector array imaging systems. Striping noise typically originates from inconsistencies in detector response, radiometric calibration errors, scanning mechanism deviations, or electronic system interference, manifesting in the spatial domain as a striped structure extending along a fixed direction. This noise exhibits significant directional and periodic characteristics, affecting not only visual quality but also introducing systematic errors into downstream tasks such as surface parameter inversion, change detection, and classification.
[0003] Existing technologies for dealing with stripe noise mainly include the following types of methods:
[0004] The first category is destriping methods based on frequency domain filtering. These methods perform Fourier transform on the image to detect the directional frequency components corresponding to the stripes and then suppress or weaken those frequency bands. While simple to implement and computationally efficient, these methods have two significant drawbacks: first, the stripe frequencies may alias with the frequencies of the actual ground structures, easily leading to the false suppression of the real structures; second, fixed filters struggle to adapt to changes in stripe direction and intensity under different scenarios, resulting in insufficient robustness.
[0005] The second category consists of strip decomposition methods based on wavelet or multi-scale transforms. These methods achieve noise suppression by separating strip components from structural components in sub-bands of different scales and directions. However, these methods still rely on fixed rules or threshold selection, lack utilization of cross-band structural consistency, and are prone to oversmoothing or residual striping in complex scenarios.
[0006] The third category comprises strip removal methods based on variational models or low-rank decomposition. These methods model stripes as low-rank or directionally sparse perturbations and optimize the separation of the real image from the strip components. Although theoretically sound, these methods suffer from high model complexity, parameter sensitivity, and significant computational overhead on large-format remote sensing imagery.
[0007] In recent years, deep learning methods have also been applied to stripe noise removal, typically using supervised learning to train convolutional neural networks to recover clean images. However, such methods face two key problems: First, there is a lack of large-scale real-world stripe-free ground truth data, and training data often relies on synthetic stripe noise, which limits the model's generalization ability on real data; second, deep models are mostly black-box structures, lacking interpretability and physical constraints, making it difficult to guarantee the protection of the real structure during stripe suppression.
[0008] Furthermore, in multi-band remote sensing data, the spatial geometry of ground features is consistent across different bands, while stripe noise often exhibits band specificity or significant intensity differences. However, most existing stripe suppression methods are designed for single-band processing and do not fully utilize cross-band structural consistency information, thus resulting in shortcomings in stripe discrimination and suppression intensity control.
[0009] In summary, the existing technology has the following main drawbacks:
[0010] 1. Strip frequency domain detection lacks a cross-band structure discrimination mechanism, which can easily mistake the real structure for a strip;
[0011] 2. Fixed filtering or rule-based methods lack adaptability and are insufficient for generalization to different scenarios;
[0012] 3. Deep learning methods rely on clean ground truth supervision, making it difficult to train stably on real remote sensing data;
[0013] 4. There is a lack of a unified framework for integrating prior information on frequency domain structure and cross-band consistency information. Summary of the Invention
[0014] The purpose of this invention is to address the technical problems of insufficient generalization ability, structural fidelity and interpretability of existing technologies, and to provide a self-supervised remote sensing image stripe noise suppression method and system based on frequency domain structural priors.
[0015] The technical solution adopted by this invention to solve the above problems is: a self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior, the method comprising:
[0016] Step 1: Multi-band strip observation modeling, specifically including:
[0017] Let the image of the k-th band of the n-th scene be... ,in Representing spatial pixel coordinates, the stripe contamination model is the original stripe-free image. With strip perturbation term Superposition:
[0018]
[0019] in, It exhibits directional consistency and spatial correlation, assuming that the stripes have a concentrated energy distribution in the frequency domain;
[0020] In the frequency domain, for Performing a two-dimensional discrete Fourier transform yields ,in Frequency coordinates;
[0021] Before performing the Fourier transform, the images of each band are subjected to uniform grid processing and amplitude normalization.
[0022] Within the multi-band joint framework, the directional energy distribution is calculated for each band:
[0023]
[0024] in, Represents the direction in the frequency domain The corresponding set of angular frequencies;
[0025] Subsequently, robust aggregation of directional energy across multiple wavebands is performed if directional energy exists. Make If it is significantly higher than other directions, it is determined to be a candidate for the main direction of the strip;
[0026] Estimate the frequency bandwidth range corresponding to the stripe, in the direction The frequency amplitude distribution in the vicinity is analyzed to identify frequency ranges significantly higher than the radial average spectral energy; by comparing directional band energy with full-frequency domain statistics, a strip candidate frequency domain mask is constructed. The frequency points in the frequency domain that belong to stripe interference are identified.
[0027] The frequency domain structure is constructed by assuming that the real ground structure usually exhibits a certain degree of spatial consistency in multiple bands, and that its corresponding frequency components have similar energy distributions in multiple bands; and that strip noise has band specificity or intensity differences.
[0028] Step 2: Establish cross-band consistency self-supervised constraints to construct reproducible and interpretable self-supervised training signals for the stripe suppression network without requiring clean truth values;
[0029] Step 3: Establish a frequency domain structure prior-guided deep suppression network for stripe noise suppression. This network follows two constraints: the network output must be directly constrained by the self-supervised loss of Step 2 and stably backpropagate; and the network should explicitly utilize the frequency domain stripe candidate mask and directional narrowband prior obtained in Step 1.
[0030] Step 4: Train the deep inhibition network from Step 3 using a self-supervised learning framework;
[0031] Step 5: First, perform uniform meshing and normalization according to the same preprocessing protocol used during training, and estimate the main direction of the stripes and the candidate frequency band mask at the scene scale using the method from Step 1. ;
[0032] The image is then input into the network to obtain a destriped output;
[0033] The image is segmented into overlapping patches, striped separately, and then stitched together using a weighted fusion method to avoid block boundary artifacts.
[0034] The final output includes destriped images for each band, as well as optional frequency domain suppression intensity maps or network internal gain response maps.
[0035] Furthermore, the process of performing unified grid processing and amplitude normalization on each band image specifically includes:
[0036] All bands are first resampled to the same spatial resolution and pixel grid, and then cropped to a common coverage area. Subsequently, each band image undergoes mean-reduction processing to eliminate the dominance of the DC component on low-frequency energy, and a fixed window function is applied to reduce spectral leakage caused by boundary effects. If clouds, shadows, or invalid regions are present, a reliability mask is constructed.
[0037]
[0038] in, = Represents unreliable pixels, where H and W are the length and width of the image space;
[0039] A soft-weighted approach is used in the window function processing to avoid high-frequency artifacts generated by the mask boundary in the frequency domain.
[0040] Furthermore, in step 1, the method for robustly aggregating the directional energy of multiple bands includes: taking the median or a weighted average.
[0041] Further, step 2 specifically includes:
[0042] Specifically, it includes:
[0043] Let the image of the k-th band of the n-th scene be... The destriping result output by the network is ;
[0044] To avoid clouds, shadows, and invalid pixels misleading consistency constraints, a reliability mask is introduced:
[0045]
[0046] Used to exclude unreliable regions; and defines pixel weights. All losses in consistency and fidelity are within Weighted calculation;
[0047] Defining cross-band structural consistency specifically includes: assuming the structural operator is... Take the gradient magnitude, edge intensity, or structure energy map; calculate the structure representation for each destriped output. Robust normalization was performed on the structure diagram of each band to obtain... The cross-band structure consistency loss is defined as the dispersion of the structure map of each band relative to its cross-band central trend. ;
[0048] The structural consistency loss is:
[0049] ;
[0050] strip candidate frequency domain mask It identifies narrow band regions in the frequency domain that may be dominated by stripes, and outputs the destriping results for each band. Calculate its spectrum And define the stripe energy penalty as the spectral energy within the mask region:
[0051] ;
[0052] Introducing data consistency constraints, let the spectrum of the original observations be... ,definition:
[0053] ;
[0054] The training objective is:
[0055] in, , which is a weighting coefficient used to balance structural consistency, stripe suppression, and data fidelity.
[0056] Furthermore, step 3 specifically includes:
[0057] Let the image of the k-th band of the n-th sample be... , The basic output of the network is the destriped result. ;
[0058] To suppress the interference of clouds, shadows, or invalid pixels on network training, a reliability mask is introduced:
[0059]
[0060] in Indicates unreliable pixels. Represents reliable pixels, used for weighting in loss calculation;
[0061] backbone network Predict strip noise or strip residuals The final output is:
[0062] ;
[0063] Let the intermediate characteristics of the main trunk be... The frequency domain module first performs a differentiable Fourier transform on it to obtain the frequency domain features. Based on the strip candidate frequency domain mask obtained in step 1 Constructing learnable band modulation gain ,in For module parameters, ,in It is Sigmoid. It is a lightweight learnable function;
[0064] Multiply the frequency domain features by the gain and inversely transform back to the spatial domain:
[0065]
[0066] in, This is a two-dimensional discrete Fourier inverse transform; and... The data is then fed into subsequent convolutional layers to continue predicting the strip residuals.
[0067] Furthermore, step 3 also includes:
[0068] Introducing structural protection mechanisms at the network level, specifically including:
[0069] Multi-scale jump connections are used to ensure the direct transmission of low-frequency ground feature structures;
[0070] Gain in the frequency domain modulation module Apply smoothing constraints to avoid generating new frequency domain artifacts;
[0071] The strip suppression residual is restricted to a directional narrow-band structure;
[0072] During training, network parameters and Joint optimization of self-supervised loss constructed in step 2.
[0073] Further, step 4 specifically includes:
[0074] Select K band images for each scene After completing the unified meshing and amplitude normalization in step 1, a size of [size missing] is cropped from the same spatial location. The patch is used to obtain the observation image of the p-th spatial image patch in the n-th scene on the k-th spectral band. Where K is the total number of spectral bands;
[0075] To prevent clouds, shadows, or invalid pixels from violating self-supervised consistency constraints, a reliability mask is introduced during training:
[0076]
[0077] in, Indicates unreliable pixels. Represents a reliable pixel; during training, it is cropped into a mask corresponding to the patch, and weights are defined. This is used to mask unreliable regions in loss calculations;
[0078] Pre-compute stripe candidate frequency domain masks on each scene or each patch. The network maintains a consistent frequency domain coordinate convention throughout training and inference; during training, the network independently feeds forward to output destriping results for each band patch. Then, calculate the cross-band self-supervised loss within the multi-band group of the same scene;
[0079] The self-supervised loss consists of three parts and is weighted over a reliable region, specifically including:
[0080] Cross-band structural consistency loss is used to promote uniformity of destriped output across the structural domain;
[0081] Strip frequency band energy penalty, constraining the network in The indicated direction reduces energy in the narrow frequency domain, thus enabling the network to learn the prior knowledge that the stripes are mainly concentrated in a specific frequency band;
[0082] Data fidelity constraints on non-strip frequencies limit excessive modifications to non-strip frequencies by the network, preventing training from degenerating into full-frequency smoothing.
[0083] The weighted sum of the three losses constitutes the overall training objective:
[0084]
[0085] in, Main loss weight, For cross-band structural consistency loss, To suppress energy loss in the strip frequency band, For non-strip band fidelity loss, This is a stability regularization term used to suppress numerical oscillations in the early stages of training.
[0086] Secondly, the present invention provides a system for a self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior. The system has a program module corresponding to the steps of the self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described above, and executes the steps in the self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior during runtime.
[0087] Thirdly, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, it performs the steps of a self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described above.
[0088] Fourthly, the present invention provides a computer-readable storage medium for storing a computer program that executes a self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described above.
[0089] The beneficial effects of this invention are:
[0090] (1) The energy of the strip frequency band is significantly reduced.
[0091] In terms of frequency domain statistics, the energy residual of this invention within the candidate frequency band is significantly lower than that of traditional fixed frequency domain filtering methods. In typical Landsat 8 scenario experiments, compared with fixed bandpass filtering methods, this invention can reduce the proportion of energy residual in the frequency band by approximately 20%–35%, and remains stable even in complex terrain scenarios. This effect directly addresses the shortcomings of existing technologies where "fixed frequency domain filtering may inadvertently damage the real structure or fail to adequately suppress it."
[0092] (2) The fidelity of the real structure is significantly improved.
[0093] By constraining cross-band structural consistency, this invention significantly reduces the damage to real edges and texture structures during destripping compared to traditional frequency domain filtering or pure CNN denoising methods. Structural fidelity metrics (e.g., gradient energy change rate) within reliable regions are significantly superior to the comparative methods, demonstrating that this invention achieves a balance between "stripping suppression without destroying the real structure." This effect resolves the problems of "oversmoothing" and "edge damage" in existing methods.
[0094] (3) No need for clean truth value, improving generalization ability
[0095] Traditional deep learning stripe suppression methods rely on clean, manually or synthetic images as supervision signals, resulting in insufficient generalization ability on real remote sensing data. This invention constructs training signals through cross-band consistency self-supervised constraints, eliminating the need for clean ground truth images and allowing direct training based on real multi-band data, thus maintaining better stability across different scenarios. This overcomes the limitation of existing supervised learning methods that depend on high-quality labeled data.
[0096] (4) Provide interpretable suppression and consistent output.
[0097] This invention outputs a frequency domain suppression intensity map and a consistency index map, making the stripe suppression results interpretable. In complex scenarios, users can use the suppression map to locate potentially oversuppressed regions, thereby performing quality assessments or subsequent corrections. This effect overcomes the lack of interpretability in existing black-box depth methods.
[0098] (5) The computational efficiency is controllable and it is suitable for large-format remote sensing images.
[0099] This invention employs a residual learning and frequency domain local modulation mechanism. After training, the inference stage only requires one forward propagation and a finite frequency domain transformation. The computational complexity is lower than that of methods based on global optimization or low-rank decomposition, making it suitable for batch processing of large-format data such as Landsat.
[0100] In summary, this invention combines frequency domain structure prior with a cross-band consistency self-supervision mechanism, achieving synergistic optimization of strip noise suppression and real structure protection without requiring clean truth values. This overcomes the shortcomings of existing methods in terms of generalization ability, structure fidelity, and interpretability.
[0101] This invention is applicable to multispectral remote sensing data such as Landsat 8 / 9, and can also be extended to other image data processing scenarios with directional stripe interference. Attached Figure Description
[0102] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0103] Figure 1 This is a flowchart illustrating the self-supervised remote sensing image stripe noise suppression method based on frequency domain structure priors of the present invention. Detailed Implementation
[0104] Specific implementation method one, such as Figure 1 As shown in this embodiment, a self-supervised remote sensing image stripe noise suppression method based on frequency domain structure priors includes:
[0105] Step 1: Multi-band strip observation modeling, specifically including:
[0106] Let the image of the k-th band of the n-th scene be... ,in Representing spatial pixel coordinates, the stripe contamination model is the original stripe-free image. With strip perturbation term Superposition:
[0107]
[0108] in, It exhibits directional consistency and spatial correlation, assuming that the stripes have a concentrated energy distribution in the frequency domain;
[0109] In the frequency domain, for Performing a two-dimensional discrete Fourier transform yields ,in Frequency coordinates;
[0110] Before performing the Fourier transform, the images of each band are subjected to uniform grid processing and amplitude normalization.
[0111] Within the multi-band joint framework, the directional energy distribution is calculated for each band:
[0112]
[0113] in, Represents the direction in the frequency domain The corresponding set of angular frequencies;
[0114] Subsequently, robust aggregation of directional energy across multiple wavebands is performed if directional energy exists. Make If the energy is significantly higher than that of other directions, it is determined to be a candidate for the main direction of the strip; "significantly higher" means that the energy of the direction exceeds the sum of the average energy of all directions and its standard deviation.
[0115] Estimate the frequency bandwidth range corresponding to the stripe, in the direction Analyze the frequency amplitude distribution in the vicinity to find frequency ranges that are significantly higher than the radial average spectral energy; "significantly higher" means that the amplitude at a certain frequency point in the candidate strip direction is greater than the sum of the average spectral energy at the corresponding radius and its standard deviation.
[0116] By comparing directional band energy with full-frequency domain statistics, candidate frequency domain masks for stripes are constructed. The frequency points in the frequency domain that belong to stripe interference are identified.
[0117] The frequency domain structure is constructed by assuming that the real ground structure usually exhibits a certain degree of spatial consistency in multiple bands, and that its corresponding frequency components have similar energy distributions in multiple bands; and that strip noise has band specificity or intensity differences.
[0118] Step 2: Establish cross-band consistency self-supervised constraints to construct reproducible and interpretable self-supervised training signals for the stripe suppression network without requiring clean truth values;
[0119] Step 3: Establish a frequency domain structure prior-guided deep suppression network for stripe noise suppression. This network follows two constraints: the network output must be directly constrained by the self-supervised loss of Step 2 and stably backpropagate; and the network should explicitly utilize the frequency domain stripe candidate mask and directional narrowband prior obtained in Step 1.
[0120] Step 4: Train the stripe inhibition network using a self-supervised learning framework, which is the deep inhibition network from Step 3;
[0121] Step 5: First, perform uniform meshing and normalization according to the same preprocessing protocol used during training, and estimate the main direction of the stripes and the candidate frequency band mask at the scene scale using the method from Step 1. ;
[0122] The image is then input into the network to obtain a destriped output;
[0123] The image is segmented into overlapping patches, striped separately, and then stitched together using a weighted fusion method to avoid block boundary artifacts.
[0124] The final output includes destriped images for each band, as well as optional frequency domain suppression intensity maps or network internal gain response maps.
[0125] This implementation method performs frequency domain directional structure modeling on multi-band remote sensing images and constructs a self-supervised training signal by combining cross-band structure consistency. This achieves adaptive suppression of strip noise and structure protection, achieving a balance between strip suppression and real structure protection without requiring clean ground truth, and improving the interpretability and generalization ability of the method.
[0126] Implementation Method Two: This implementation method further defines step 2 in Implementation Method One. Step 2 specifically includes:
[0127] Specifically, it includes:
[0128] Let the image of the k-th band of the n-th scene be... The destriping result output by the network is ;
[0129] To avoid clouds, shadows, and invalid pixels misleading consistency constraints, a reliability mask is introduced:
[0130]
[0131] Used to exclude unreliable regions; and defines pixel weights. All losses in consistency and fidelity are within Weighted calculation;
[0132] Defining cross-band structural consistency specifically includes: assuming the structural operator is... Take the gradient magnitude, edge intensity, or structure energy map; calculate the structure representation for each destriped output. Robust normalization was performed on the structure diagram of each band to obtain... The cross-band structure consistency loss is defined as the dispersion of the structure map of each band relative to its cross-band central trend. ;
[0133] The structural consistency loss is:
[0134] ;
[0135] strip candidate frequency domain mask Its identifier is a narrow band region in the frequency domain that may be dominated by stripes, for each output Calculate its spectrum And define the stripe energy penalty as the spectral energy within the mask region:
[0136] ;
[0137] Introducing data consistency constraints, let the spectrum of the original observations be... ,definition:
[0138] ;
[0139] The training objective is:
[0140] in, , which is a weighting coefficient used to balance structural consistency, stripe suppression, and data fidelity.
[0141] This implementation constructs a reproducible and interpretable self-supervised training signal for the stripe suppression network without requiring a clean ground truth (i.e., no "stripless reference image"). It is important to emphasize that the design of the self-supervised signal in this implementation follows two verifiable principles: first, the real ground structure should remain consistent across bands, therefore the structural consistency loss should decrease after striping; second, the narrow-band frequency domain energy corresponding to the stripes should be significantly suppressed, while the energy and phase structure of non-strip frequency bands should be preserved as much as possible. The training results will be subsequently validated based on frequency domain residuals, structural fidelity, and improved cross-band consistency to ensure that the self-supervised constraints do not degenerate into over-smoothing or structural distortion.
[0142] Implementation Method 3: This implementation method further defines the self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described above.
[0143] Step 3 specifically includes:
[0144] Let the image of the k-th band of the n-th sample be... , The basic output of the network is the destriped result. ;
[0145] To suppress the interference of clouds, shadows, or invalid pixels on network training, a reliability mask is introduced:
[0146]
[0147] in Indicates unreliable pixels. Represents reliable pixels, used for weighting in loss calculation;
[0148] backbone network Predict strip noise or strip residuals The final output is:
[0149] ;
[0150] Let the intermediate characteristics of the main trunk be... The frequency domain module first performs a differentiable Fourier transform on it to obtain the frequency domain features. Based on the strip candidate frequency domain mask obtained in step 1 Constructing learnable band modulation gain ,in For module parameters, ,in It is Sigmoid. It is a lightweight learnable function;
[0151] Multiply the frequency domain features by the gain and inversely transform back to the spatial domain:
[0152]
[0153] And The data is then fed into subsequent convolutional layers to continue predicting the strip residuals.
[0154] This implementation provides a deep network structure for stripe noise suppression, designed according to two constraints: First, the network output must be directly constrained by the self-supervised loss of step 2 and stably backpropagate; second, the network should explicitly utilize the frequency domain stripe candidate mask and directional narrowband prior obtained in step 1, so that it "primarily modifies the stripe frequency band," rather than degrading destripping to generalized denoising or over-smoothing. To this end, this method adopts a structure of "spatial domain residual backbone + frequency domain modulation module," wherein the frequency domain module applies a learnable suppression gain on the candidate stripe frequency band and is constrained by a structure fidelity mechanism to avoid inadvertently damaging the real structure.
[0155] Implementation Method 4: This implementation method further defines the self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described above.
[0156] Step 4 specifically includes:
[0157] Select K band images for each scene After completing the unified meshing and amplitude normalization in step 1, a size of [size missing] is cropped from the same spatial location. The patch, obtained ;
[0158] To prevent clouds, shadows, or invalid pixels from violating self-supervised consistency constraints, a reliability mask is introduced during training:
[0159]
[0160] in, Indicates unreliable pixels. Represents a reliable pixel; during training, it is cropped into a mask corresponding to the patch, and weights are defined. This is used to mask unreliable regions in loss calculations;
[0161] Pre-compute stripe candidate frequency domain masks on each scene or each patch. The network maintains a consistent frequency domain coordinate convention throughout training and inference; during training, the network independently feeds forward to output destriping results for each band patch. Then, the cross-band self-supervised loss is calculated within the multi-band group of the same scene.
[0162] The self-supervised loss consists of three parts and is weighted over a reliable region, specifically including:
[0163] Cross-band structural consistency loss is used to promote uniformity of destriped output across the structural domain;
[0164] Strip frequency band energy penalty, constraining the network in The indicated direction reduces energy in the narrow frequency domain, thus enabling the network to learn the prior knowledge that the stripes are mainly concentrated in a specific frequency band;
[0165] Data fidelity constraints on non-strip frequencies limit excessive modifications to non-strip frequencies by the network, preventing training from degenerating into full-frequency smoothing.
[0166] The weighted sum of the three losses constitutes the overall training objective:
[0167]
[0168] in, Main loss weight, This is a stability regularization term used to suppress numerical oscillations in the early stages of training.
[0169] This implementation uses a self-supervised learning framework to train the stripe suppression network, eliminating the need for clean, stripe-free ground truth images. The trainable data comes from a real multi-band image set of Landsat 8 / 9, and the training samples are constructed using "multi-band patch groups of the same scene" as the basic unit, thus making the cross-band consistency loss calculable and meaningful.
[0170] Implementation Method 5: This implementation method proposes a remote sensing image stripe noise suppression method based on frequency domain structure prior and cross-band consistency self-supervised constraints. Targeting the common directional stripe noise problem in multi-band remote sensing images, it constructs a self-supervised learning framework that does not require clean truth values by using frequency domain stripe signature detection and cross-band structure consistency modeling, thereby achieving adaptive suppression of stripe noise and protection of the true structure.
[0171] The overall process consists of three main stages: strip observation modeling and frequency domain signature detection, cross-band consistency self-supervised training, and strip suppression and result output in the inference stage.
[0172] First, in the strip observation modeling stage, Landsat 8 / 9 multi-band imagery was uniformly gridded and amplitude normalized, and a reliability mask was constructed to eliminate interference from clouds, shadows, and invalid pixels on statistical analysis. Subsequently, the directional energy distribution of each band was analyzed in the frequency domain to identify the main direction of the strip and its corresponding frequency bandwidth, forming a candidate frequency domain mask for the strip. This mask is only used to identify possible strip frequency regions, rather than directly performing filtering.
[0173] Secondly, a deep suppression network is constructed during the self-supervised training phase. The network takes a single-band image as input, predicts stripe perturbations through residual learning, and combines this with a frequency domain modulation module to perform learnable suppression within the candidate stripe frequency bands. During training, it does not rely on any clean image ground truth, but instead optimizes the network parameters through three types of self-supervised constraints: first, cross-band structural consistency constraints, making the multi-band structural representation more consistent after striping; second, stripe frequency band energy suppression constraints, causing the network to reduce spectral energy within the candidate frequency bands; and third, non-strip frequency band fidelity constraints, limiting the network's excessive modification of the true structural frequencies. Through this combination of constraints, label-free self-supervised training is achieved.
[0174] Finally, during the inference phase, frequency domain stripe orientation detection and mask generation are performed on new remote sensing scenes, consistent with the training process. The images are then input into the trained network for stripe suppression. Since remote sensing images are typically large, a sliding window overlap strategy is used for inference, and weighted fusion is employed to avoid block boundary artifacts. The output includes a destripped image and an optional frequency domain suppression response map for subsequent quality verification.
[0175] It is important to emphasize that the method described in this implementation belongs to a self-supervised deep learning framework. During the training phase, cross-band consistency constraints are constructed using real multi-band remote sensing images, without requiring clean, stripe-free ground truth values. The testing phase quantitatively verifies the method using stripe frequency band energy residuals, structure fidelity indices, and the degree of cross-band consistency improvement, and compares it with traditional frequency domain filtering or wavelet destriping methods.
[0176] By combining frequency domain structure priors with a cross-band consistency self-supervised mechanism, this method effectively suppresses directional strip noise while maintaining the true ground object structure, and is suitable for large-scale automatic processing of multi-band remote sensing images such as Landsat 8 / 9.
[0177] Step 1: Strip Observation Modeling and Frequency-domain Structural Prior Construction:
[0178] In multi-band remote sensing images, stripe noise typically originates from inconsistencies in detector array response, scanning system residuals, or radiometric calibration errors. In the spatial domain, it manifests as periodic or quasi-periodic intensity perturbations along a fixed direction. For pushbroom imaging systems such as Landsat 8 / 9, the stripes often exhibit an approximately parallel column or row distribution. Unlike random noise, stripe noise possesses a clear directionality and frequency domain concentration, thus it can be characterized in the frequency domain as a narrow-band energy anomaly along a specific direction.
[0179] Let the image of the k-th band of the n-th scene be... ,in Represents spatial pixel coordinates. Strip contamination can be modeled as the original stripe-free image. With strip perturbation term Superposition:
[0180]
[0181] in It exhibits directional consistency and spatial correlation. This model does not assume that the strips are purely periodic signals, but it does assume that they have a concentrated energy distribution in the frequency domain.
[0182] In the frequency domain, for Performing a two-dimensional discrete Fourier transform yields ,in Here, we have the frequency coordinates. If strip noise exhibits periodic variation along the vertical direction in the spatial domain, it will manifest as energy concentration along frequency directions orthogonal to the strip direction in the frequency domain. In other words, if the strip varies periodically along the spatial direction... If the distribution is such that its frequency domain anomalous energy is concentrated in the direction... nearby.
[0183] To ensure the stability of the frequency domain analysis, a unified grid and amplitude normalization are performed on each band image before Fourier transform. All bands are first resampled to the same spatial resolution and pixel grid, and then cropped to a common coverage area. Subsequently, each band image undergoes mean removal to eliminate the dominance of the DC component on low-frequency energy, and a fixed window function (e.g., Hann window) is applied to reduce spectral leakage caused by boundary effects. If clouds, shadows, or invalid regions exist, a reliability mask is constructed.
[0184]
[0185] in This represents unreliable pixels. To avoid strong high-frequency artifacts in the frequency domain caused by the mask boundaries, a soft-weighted approach is used in the window function processing, rather than direct hard clipping.
[0186] Within a multi-band joint framework, strip orientation detection should not rely solely on single-band spectra. The directional energy distribution is calculated for each band:
[0187]
[0188] in Represents the direction in the frequency domain The corresponding set of angular frequencies. Then, robust aggregation of the directional energy across multiple bands is performed, for example, by taking the median or a weighted average:
[0189]
[0190] If a direction exists Make If the direction is significantly higher than other directions, it can be identified as a candidate for the main direction of the strip. This cross-band aggregation strategy can suppress the interference of occasional texture directions in single bands on strip detection.
[0191] After determining the stripe direction, it is necessary to estimate the corresponding frequency bandwidth range. Therefore, in terms of direction... By analyzing the frequency amplitude distribution in the vicinity, frequency intervals significantly higher than the radial average spectral energy are identified. By comparing the directional band energy with full-frequency domain statistics, candidate frequency domain masks for the strip can be constructed. The frequency points identified in the frequency domain may belong to stripe interference.
[0192] Unlike traditional fixed-band filtering, this application does not directly target... Complete suppression will be performed, and the frequency will be treated only as a "strip candidate frequency set". In the next stage, this set will be further screened and adaptively suppressed with intensity modulation, based on cross-band structural consistency criteria.
[0193] Furthermore, this implementation establishes a "frequency domain structure prior" assumption: real ground structures typically exhibit a certain degree of spatial consistency across multiple bands, with their corresponding frequency components having similar energy distributions across multiple bands; while strip noise often exhibits band specificity or intensity differences. Therefore, the performance of strip noise in the frequency domain can be considered as a combination of "directional narrowband perturbation + cross-band structural inconsistency." This frequency domain structure prior provides a theoretical basis for subsequent self-supervised constraints on cross-band consistency.
[0194] Step 2: Cross-band Consistency-based Self-supervised Constraint Design
[0195] The goal of this step is to construct reproducible and interpretable self-supervised training signals for the stripe suppression network without requiring a clean ground truth (i.e., without a "strip-free reference image"). The core idea is that real-world ground structures exhibit strong geometric consistency across multiple bands (despite differences in radiation amplitude), while stripe noise often manifests as highly directional perturbations that are unstable or exhibit significant intensity differences across bands. Therefore, if a destriping model can suppress stripe frequency domain components while making the multi-band structural representation more consistent and maintaining the necessary data consistency with the original observations, it can be self-supervised trained without a clean ground truth.
[0196] Let the image of the k-th band of the n-th scene be... The destriping result output by the network is To avoid clouds, shadows, and invalid pixels misleading the consistency constraints, a reliability mask is introduced:
[0197]
[0198] And define pixel weights and define pixel weights, all consistency and fidelity losses are in Weighted calculation. This mask is used only to "exclude unreliable regions" in this method and does not provide any destriping truth supervision.
[0199] First, we define cross-band structural consistency. To mitigate the impact of cross-spectral radiation differences, this method compares bands in the structural domain, not the intensity domain. Let the structural operator be... Gradient magnitude, edge intensity, or structural energy map can be used, for example. Where X is the input image, S() is the smoothing operator (e.g., Gaussian filtering), and ∇() is the spatial gradient operator. The structural representation is calculated for each destriped output. Since the structural response amplitudes may still differ across different bands, robust normalization (e.g., normalization based on quantile scales within the reliable region) is further performed on the structural diagrams of each band to obtain... This makes cross-band structure comparisons comparable. Then, the cross-band structure consistency loss is defined as the dispersion of each band structure map relative to its cross-band central trend, for example, let...
[0200]
[0201] Here, ref stands for reference, indicating a reference structure, and s is the sample or scene index.
[0202] The structural consistency loss can then be written as:
[0203]
[0204] The meaning of this loss is that the multi-band structure representation after striping should be spatially geometrically consistent, thereby suppressing structural interferences (typically stripes) that appear strongly in a single band but are unstable across bands.
[0205] Relying solely on structural consistency can lead to excessive smoothing of the network in an attempt to "force consistency." Therefore, it is necessary to introduce prior constraints on the frequency domain structure, ensuring that the network focuses on suppressing candidate frequency bands rather than arbitrarily weakening the structure. This is achieved through a frequency domain mask for candidate frequency bands. It identifies a narrow band region in the frequency domain that may be dominated by stripes. For each output Calculate its spectrum And define the stripe energy penalty as the spectral energy within the mask region:
[0206]
[0207] This approach encourages networks to reduce energy within stripe candidate frequency bands, explicitly promoting "de-striping" rather than generalized denoising from a frequency domain perspective.
[0208] Meanwhile, to avoid unnecessary modifications to non-strip frequency bands or the overall radiation trend by the network, a data consistency constraint is introduced to ensure that the output retains as much information as possible from the original observations in the "non-strip frequency bands". A reproducible implementation is to constrain the spectral preservation of the non-strip frequency bands: let the original observation spectrum be... Then define
[0209]
[0210] This constraint in the frequency domain states that "the network mainly changes the candidate frequency bands of the strips, while trying not to change other frequency bands," thereby reducing the risk of over-smoothing and structural damage, and keeping the training objective consistent with the frequency domain prior in step 1.
[0211] Based on the above three types of self-supervised constraints, the training objective of the method in this application is defined as follows:
[0212]
[0213] in The weighting coefficients are used to balance structural consistency, stripe suppression, and data fidelity. This loss function does not require any clean ground truth image; its training signal comes entirely from structural consistency and stripe frequency domain prior constraints among multi-band observations. This is due to the structure operator... Both Fourier transform and mask weighting can be implemented as differentiable operators (or implemented using differentiable approximations during training). This self-supervised objective can optimize network parameters through standard backpropagation, thereby achieving strip suppression learning without ground truth on the Landsat8 / 9 public dataset.
[0214] The design of the self-supervised signal follows two verifiable principles: first, the actual ground structure should remain consistent across bands, therefore the loss of structural consistency should decrease after destriping; second, the narrow-band frequency domain energy corresponding to the stripes should be significantly suppressed, while the energy and phase structure of non-striped frequency bands should be preserved as much as possible. The training results will be validated based on frequency domain residuals, structural fidelity, and improved cross-band consistency to ensure that the self-supervised constraints do not degenerate into over-smoothing or structural distortion.
[0215] Step 3: Frequency-aware StripeSuppression Network (SSNS) guided by frequency domain structure priors:
[0216] This step presents a deep network structure for stripe noise suppression, designed under two constraints: First, the network output must be directly constrained by the self-supervised loss from step 2 and stably backpropagate; second, the network should explicitly utilize the frequency-domain stripe candidate mask and directional narrowband prior obtained in step 1, enabling it to "primarily modify the stripe frequency band," rather than degrading destripping to generalized denoising or over-smoothing. To this end, this method employs a structure of "spatial-domain residual backbone + frequency-domain modulation module," where the frequency-domain module applies a learnable suppression gain to the candidate stripe frequency band and is constrained by a structure fidelity mechanism to avoid inadvertently damaging the real structure.
[0217] Let the image of the k-th band of the n-th sample be... , The basic output of the network is the destriped result. Regarding the input format, this method provides two reproducible implementations: one is "single-band input, cross-band parameter sharing," which means inputting the network separately for each band but sharing the network parameters. The first method achieves consistent destriping behavior across different bands. The second method involves "multi-band joint input," where multiple bands are spliced together in the channel dimension as input, and the network simultaneously predicts the destriping output of each band. Considering that stripes often have band-specific intensities, while ground structures have cross-band commonalities, this method recommends adopting the form of "single-band inference with shared parameters + cross-band self-supervised constraints": the network structure is the same for each band, but during training, different bands are coupled through the cross-band consistency loss in step 2, thereby achieving a combination of "single-band inference and multi-band consistency training."
[0218] To suppress the interference of clouds, shadows, or invalid pixels on network training, a reliability mask is introduced:
[0219]
[0220] in Indicates unreliable pixels. This represents a reliable cell. This mask is not used as a supervisory label; it is only used for weighting in loss calculations to prevent the network from being influenced by anomalous statistics from unreliable regions.
[0221] The network employs a residual learning paradigm to enhance its structural protection capabilities. Specifically, the backbone network... Predict strip noise or strip residuals The final output is
[0222]
[0223] Alternatively, residuals can be directly predicted and low-frequency structures can be preserved by skipping connections. The advantage of residual learning is that when a region does not contain obvious stripes, the network output can naturally approach zero, thus avoiding unnecessary modifications to the real structure.
[0224] To explicitly introduce a frequency-domain structure prior, this method inserts a frequency-domain modulation module within the backbone network. Let the intermediate features of the backbone be... (Note that F here represents the feature map, not to be confused with the Fourier spectrum symbol.) The frequency domain module first performs a differentiable Fourier transform on it to obtain the frequency domain features. Based on the strip candidate frequency domain mask obtained in step 1 Constructing learnable band modulation gain ,in These are module parameters. To ensure "learnable suppression only within the candidate frequency band of the stripe," this gain is designed as a mask-gated form:
[0225]
[0226] in It is Sigmoid. It is a lightweight learnable function (which can be implemented by local convolution of frequency domain features or a small MLP), thus ensuring that it works in non-strip frequency bands. hour (Without modification), within the candidate frequency band, adaptive suppression is performed based on the learned intensity. Then, the frequency domain features are multiplied by the gain and inversely transformed back to the spatial domain:
[0227]
[0228] And The data is then fed into subsequent convolutional layers to continue predicting the stripe residuals. This structure achieves "learnable suppression within the frequency domain candidate bands" while structurally avoiding significant modifications to non-strip frequencies, thus supporting the frequency domain fidelity loss in step 2 from the network structure level.
[0229] To further avoid real structural damage caused by oversuppression, this method introduces a structural protection mechanism at the network level. First, multi-scale hop connections (U-Net or residual pyramid) are used to ensure the direct transmission of low-frequency ground feature structures; second, gain is adjusted in the frequency domain modulation module. First, smoothing constraints are applied (e.g., limiting its sharp oscillations in the frequency domain) to avoid generating new frequency domain artifacts. Second, the stripe suppression residual is restricted to a directional narrowband structure (achieved through frequency domain mask gating and stripe energy penalty in step 2). These mechanisms make the network more inclined to learn the removal of "directional narrowband perturbations" rather than reduce loss through full-frequency domain smoothing.
[0230] During training, network parameters and Joint optimization is achieved through the self-supervised loss constructed in step 2. The cross-band structural consistency loss makes the destriping results more consistent across different bands in the structural domain, the striped frequency band energy penalty suppresses energy within the candidate frequency band, and the non-striped fidelity constraint restricts the network's modifications to other frequency bands, thus collectively forming a stable self-supervised signal. Since both the frequency domain modulation module and the structural operator can be implemented as differentiable operators, the above losses can be backpropagated end-to-end through the network, ensuring that the training process is reproducible and consistent with the frequency domain structural prior.
[0231] Step 4: Training and Inference Pipeline
[0232] This step employs a self-supervised learning framework to train the stripe suppression network, eliminating the need for clean, stripe-free ground truth images. Training data is derived from a real multi-band image set from Landsat 8 / 9. Training samples are constructed using "multi-band patch groups of the same scene" as the basic unit, ensuring that the cross-band consistency loss is computationally calculable and meaningful. Specifically, K band images for each scene are selected. After completing the unified meshing and amplitude normalization in step 1, a size of [size missing] is cropped from the same spatial location. The patch is used to obtain the observation image of the p-th spatial image patch in the n-th scene on the k-th spectral band. Let n represent the nth remote sensing scene, p represent the pth image patch cropped from that scene, k represent the kth spectral band, and K be the total number of spectral bands. To prevent clouds, shadows, or invalid pixels from violating the self-supervised consistency constraint, a reliability mask is introduced during training:
[0233]
[0234] in Indicates unreliable pixels. This represents a reliable pixel. During training, it is cropped into a mask corresponding to the patch, and weights are defined. It is used to mask unreliable regions in loss calculation. To avoid frequency domain artifacts introduced by hard mask boundaries, the mask can be slightly dilated and smoothed during training, so that it participates in loss weighting in a "soft weight" manner.
[0235] To provide convergent training signals without clean ground truth, stripe candidate frequency domain masks are pre-computed on each scene or patch. (From the directional energy detection and band estimation in step 1), and maintaining consistent frequency domain coordinate conventions (whether to perform FFT shift, window function, mean removal, etc.) during training and inference. During training, the network independently feeds forward to output the destriped result for each band patch. (Sharing network parameters), and then calculating cross-band self-supervised loss within a multi-band group in the same scene. This organization method ensures "cross-band consistency for training coupling" while maintaining the engineering property of "single-band inference usability".
[0236] The self-supervised loss consists of three parts and is weighted in the reliable region: The first is a cross-band structural consistency loss, used to promote uniformity of the destriped output across the structural domain. To reduce the contamination of the consistency reference by anomalous bands, robust aggregation (such as pixel-wise median or truncated mean) is recommended for the cross-band reference structure, and the deviation of each band structure map from the reference structure is further measured. The second is a striped band energy penalty, constraining the network in... The first is to reduce energy in the narrow band frequency domain, allowing the network to learn the prior knowledge that "the stripes are mainly concentrated in a specific frequency band." The second is to impose data fidelity constraints on non-strip frequency bands, limiting excessive modification of these bands and preventing training from degenerating into full-frequency smoothing. The weighted sum of these three losses constitutes the overall training objective:
[0237]
[0238] in Main loss weight, Stability regularization terms (e.g., smoothing constraints on frequency domain gain or weak constraints on output residual energy) are used to suppress numerical oscillations in the early stages of training. During training, standard backpropagation is used to update network parameters, and the implementation details of FFT (window function, normalization, frequency domain symmetry) must be fixed to ensure reproducibility. If there are concerns that insufficient strip strength in real data may lead to a weak learning signal, "controllable synthetic strip enhancement" can be introduced as data augmentation during training: directional narrowband perturbations (consistent with the frequency domain prior in step 1) are superimposed on some patches to give the network a stronger strip suppression gradient. However, this enhancement is only used as a training stabilization method and does not change the applicability of the method to real data.
[0239] The inference phase does not rely on any annotations or ground truth. For a given Landsat 8 / 9 scene image, it first undergoes uniform meshing and normalization according to the same preprocessing protocol used during training. Then, step 1 estimates the main direction of the stripes and the candidate frequency band mask at the scene scale. (Alternatively, a mask can be obtained directly using detection rules fixed during the training phase). The image is then input into the network to obtain a destriped output. Since Landsat scenes are typically large, a sliding window slicing strategy is used for inference: the image is divided into overlapping patches, each destriped, and then stitched together using a weighted fusion method to avoid patch boundary artifacts. The stitching weights can use the same window function weights as during training, making the patch center contribute more and the edges contribute less, thus achieving a smooth transition. The final output includes the destriped image for each band, as well as an optional frequency domain suppression intensity map or network internal gain response map, used for subsequent quality verification and quality control analysis.
[0240] Through the above training and inference process, this method constructs a self-supervised signal based on cross-band structural consistency and frequency domain strip priors without requiring clean truth values, thereby achieving learning-based suppression of noise in real remote sensing strips. At the same time, it ensures reproducible inference and stable output of large-format Landsat 8 / 9 data through a fixed frequency domain detection protocol and sliding window stitching strategy.
[0241] Example: This example uses Landsat 8 / 9 multi-band remote sensing imagery as the data source, selecting scenes containing various land surface types (vegetation, water bodies, towns, bare soil) and varying degrees of stripe noise as experimental data. Multiple spectral bands (e.g., visible light, near-infrared, shortwave infrared, etc.) within the same scene are preferentially selected as the input basis for cross-band consistency constraints. Each band image undergoes unified grid processing, including resampling to the same spatial resolution, cropping to a common coverage area, and amplitude normalization to avoid the impact of dynamic range differences between different bands on subsequent frequency domain analysis. If the data includes quality layers or cloud / shadow markers, these are used to construct reliable region weights to avoid interference from clouds, shadows, or invalid pixels on self-supervised consistency statistics; these weights are not used as supervision labels but only for weighted calculation of training loss.
[0242] In the stripe prior construction stage, each band image is first processed by mean removal and windowing to reduce boundary effects. Then, a frequency domain transformation is performed, and the directional distribution of frequency domain energy is statistically analyzed to identify the main direction and frequency concentration range corresponding to stripe noise. To avoid false detections caused by random texture directions in a single band, a cross-band robust aggregation strategy is adopted for stripe direction determination. That is, only when multiple bands show consistent directional anomalous energy in similar directions is that direction considered as a candidate for the main stripe direction. This constructs a stripe candidate frequency band mask, which identifies the narrowband region most likely to be dominated by stripe noise in the frequency domain and serves as the structural prior input for the subsequent network frequency domain modulation module.
[0243] In terms of network structure, this embodiment employs a residual learning framework. The network input is a single-band image, and the output is a stripe residual estimate. The final destriping result is obtained by subtracting the stripe residual from the input. To explicitly introduce a frequency domain structure prior, a frequency domain modulation module is inserted into the middle layer of the network. This module performs a frequency domain transformation on the intermediate features and learns a band suppression gain under the constraint of the stripe candidate band mask. This allows the network to focus on suppressing perturbations in the stripe bands rather than smoothing the entire frequency band. This gain satisfies symmetry constraints in the frequency domain to ensure that the output after the inverse transformation is a real value and does not introduce additional artifacts. The remaining parts of the network employ a multi-scale skip-connection structure to enhance structural fidelity and avoid edge breaks or excessive texture attenuation during the destriping process.
[0244] The training phase employs self-supervised learning, eliminating the need for clean, stripe-free ground truth images. Training samples are constructed using "multi-band image patch groups at the same spatial location" as the basic unit. This involves randomly sampling several spatial locations within the same scene and cropping fixed-size multi-band image patches to form a training sample group. The network feeds forward the destripping result for each band within this group, then calculates the cross-band structural consistency loss within the group: first, it calculates the structural representation (such as gradient or structural energy) for each band output; then, it uses robust aggregation to form a reference structure, constraining the structural outputs of each band to tend towards spatial geometric consistency, thereby suppressing structural inconsistencies caused by band-specific stripes. Simultaneously, stripe frequency band energy suppression constraints are added during training to encourage the network to reduce energy within candidate stripe frequency bands; and fidelity constraints for non-stripe frequency bands are added to limit excessive modifications to the non-stripe frequency band structure, preventing the model from degenerating into global smoothness. These three types of constraints are combined with fixed weights to form the total loss, which is used to update the network parameters through standard backpropagation. To enhance training stability, a mild data augmentation strategy can be adopted: superimpose synthetic strip perturbations that conform to the strip prior (narrow-band orientation) onto some training image patches to improve the strip suppression gradient strength, but the main training data still comes from the real Landsat scene to ensure that the model is adapted to the real strip pattern.
[0245] The inference phase requires no labels or additional input. The Landsat scene imagery to be processed is first normalized using the same preprocessing protocol as in the training phase, followed by strip orientation detection and candidate frequency band mask generation. Due to the large size of the Landsat scene, a sliding window slicing strategy is used for inference: the image is divided into overlapping blocks, each block is input into the network for destriping, and then a weighted fusion method is used to stitch the output to avoid block boundary artifacts. Smooth window weights are used for stitching, giving higher weights to the center region and lower weights to the edges, thus ensuring spatial continuity. The final output is the destriped imagery for each band, and can also output a frequency domain suppression gain map or a strip frequency band energy residual map for subsequent quality checks and parameter adjustments.
[0246] In simulation experiments, the proposed method was compared with two contrasting approaches: one is a traditional fixed frequency domain filtering or wavelet-frequency domain combined filtering method, and the other is a conventional deep denoising network without a frequency domain modulation module (learning only in the spatial domain). Evaluation employed three metrics: first, comparing the energy residuals within the candidate frequency bands to verify whether the band energy was significantly reduced; second, comparing structure fidelity metrics (such as gradient energy change or structural similarity) to assess whether the real structure was excessively damaged; and third, comparing the degree of improvement in cross-band structural consistency to verify whether the self-supervised consistency constraint effectively reduced inter-band structural differences. Experimental results typically showed that the proposed method significantly reduced band energy while minimizing damage to edge and detail structures, and significantly improved multi-band structural consistency; compared to traditional frequency domain filtering methods, the proposed method was less likely to mistakenly eliminate real structures in complex terrain texture scenes; and compared to conventional spatial domain denoising networks, the proposed method was more targeted towards directional bands, resulting in fewer band residues.
[0247] Furthermore, the method of this application allows for the substitution of key steps to adapt to different sensors or different stripe morphologies. For example, structural representation can employ structural operators such as gradient magnitude, edge intensity, or phase consistency; stripe candidate frequency band detection can employ different directional energy statistical methods; the frequency domain modulation module can be replaced with a frequency domain attention or frequency band gating structure; and cross-band consistency reference can employ robust aggregation methods such as median or truncated mean. The above substitutions do not change the core idea of this application of "frequency domain structural prior + cross-band consistency self-supervision" and should all be considered as optional implementation methods of this application.
[0248] This application proposes a complete technical solution including frequency domain stripe candidate detection, cross-band structural consistency self-supervised loss construction, frequency domain modulation deep network design, and self-supervised training process, achieving adaptive stripe noise suppression without requiring clean ground truth values. Specifically:
[0249] 1. The mechanism for constructing and utilizing prior knowledge of striped frequency domain structures
[0250] This application first detects the main direction and narrow frequency range of the stripe in the frequency domain, and then constructs a candidate frequency domain mask for the stripe. This mask is not directly used for fixed filtering, but rather serves as a priori constraint for the deep network's frequency domain modulation module, limiting the frequency range for stripe suppression. This combination mechanism of "candidate frequency band + learnable modulation" is a key innovation that distinguishes it from traditional fixed frequency domain filtering methods, including: the stripe direction and frequency band detection method; the construction process of the candidate frequency domain mask; and the implementation method of coupling the mask with the network's frequency domain module.
[0251] 2. Cross-band structural consistency self-supervised constraint mechanism
[0252] This application utilizes the structural consistency among multi-band remote sensing images to construct a self-supervised training signal, eliminating the need for clean ground truth images. By defining a cross-band structural consistency loss, the network maintains the geometric consistency of the multi-band structure during destriping, thereby avoiding over-smoothing or unintended suppression of the true structure. This includes: a definition of the structural domain consistency metric; a joint optimization framework for self-supervised consistency loss and frequency domain striping loss; and a training mechanism that does not require clean ground truth data.
[0253] 3. Frequency Domain Modulation Deep Network Structure
[0254] This application designs a deep network structure including a frequency domain modulation module, enabling adaptive suppression of frequency components within the candidate strip frequency band while preserving non-strip frequencies as much as possible. This structure achieves "targeted suppression of directional narrowband disturbances" while protecting low-frequency ground features through residual learning and skip-connection structures. This includes: the embedding method of the frequency domain modulation module; the gating mechanism of the learnable frequency domain gain function and strip mask; and the joint design of residual learning and structural protection.
[0255] 4. Output interpretability enhancement mechanism
[0256] This application not only outputs destriped images but also frequency domain suppression intensity maps or consistency confidence maps, enabling the stripe suppression process to be interpretable and capable of quality assessment. This output mechanism differs from traditional black-box network denoising methods. It includes: the generation method of the suppression intensity map; the calculation method of the cross-band consistency improvement index; and a quality control process based on the consistency map.
[0257] Based on the above method, this application solves the following technical problem:
[0258] First, how to accurately model the directionality and narrowband characteristics of strip noise in the frequency domain, and distinguish between strip frequencies and real ground structure frequencies, so as to avoid the false suppression of real structures by traditional frequency domain filtering.
[0259] Secondly, how to construct an effective training signal without clean ground truth images, so that the deep model can learn stripe suppression capabilities while preserving the true structural information.
[0260] Furthermore, how can we utilize the structural consistency information among multi-band remote sensing images to construct cross-band self-supervised constraints, so that the stripe suppression process does not rely on manual annotation, but drives model learning through structural consistency?
[0261] Finally, how to design a network structure that includes both frequency domain structure priors and spatial domain structure fidelity, so that stripe suppression has directional targeting and structural protection capabilities, while ensuring the feasibility and stability of the algorithm on large-format remote sensing images.
Claims
1. A self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior, characterized in that, The method includes: Step 1: Multi-band strip observation modeling, specifically including: Let the image of the k-th band of the n-th scene be... ,in Representing spatial pixel coordinates, the stripe contamination model is the original stripe-free image. With strip perturbation term Superposition: in, It exhibits directional consistency and spatial correlation, assuming that the stripes have a concentrated energy distribution in the frequency domain; In the frequency domain, for Performing a two-dimensional discrete Fourier transform yields ,in Frequency coordinates; Before performing the Fourier transform, the images of each band are subjected to uniform grid processing and amplitude normalization. Within the multi-band joint framework, the directional energy distribution is calculated for each band: in, Represents the direction in the frequency domain The corresponding set of angular frequencies; Subsequently, robust aggregation of directional energy across multiple wavebands is performed if directional energy exists. Make If it is significantly higher than other directions, it is determined to be a candidate for the main direction of the strip; Estimate the frequency bandwidth range corresponding to the stripe, in the direction The frequency amplitude distribution in the vicinity is analyzed to identify frequency ranges significantly higher than the radial average spectral energy; by comparing directional band energy with full-frequency domain statistics, a strip candidate frequency domain mask is constructed. The frequency points in the frequency domain that belong to stripe interference are identified. The frequency domain structure is constructed by assuming that the real ground structure usually exhibits a certain degree of spatial consistency in multiple bands, and that its corresponding frequency components have similar energy distributions in multiple bands; and that strip noise has band specificity or intensity differences. Step 2: Establish cross-band consistency self-supervised constraints to construct reproducible and interpretable self-supervised training signals for the stripe suppression network without requiring clean truth values; Step 3: Establish a frequency domain structure prior-guided deep suppression network for stripe noise suppression. This network follows two constraints: the network output must be directly constrained by the self-supervised loss of Step 2 and stably backpropagate; and the network should explicitly utilize the frequency domain stripe candidate mask and directional narrowband prior obtained in Step 1. Step 4: Train the deep inhibition network from Step 3 using a self-supervised learning framework; Step 5: First, perform uniform meshing and normalization according to the same preprocessing protocol used during training, and estimate the main direction of the stripes and the candidate frequency band mask at the scene scale using the method from Step 1. ; The image is then input into the network to obtain a destriped output; The image is segmented into overlapping patches, striped separately, and then stitched together using a weighted fusion method to avoid block boundary artifacts. The final output includes destriped images for each band, as well as optional frequency domain suppression intensity maps or network internal gain response maps.
2. The self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described in claim 1, characterized in that: The process of performing uniform grid processing and amplitude normalization on each band image specifically includes: All bands are first resampled to the same spatial resolution and pixel grid, and then cropped to a common coverage area. Subsequently, each band image undergoes mean-reduction processing to eliminate the dominance of the DC component on low-frequency energy, and a fixed window function is applied to reduce spectral leakage caused by boundary effects. If clouds, shadows, or invalid regions are present, a reliability mask is constructed. in, = Represents unreliable pixels, where H and W are the length and width of the image space; A soft-weighted approach is used in the window function processing to avoid high-frequency artifacts generated by the mask boundary in the frequency domain.
3. The self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described in claim 1, characterized in that: In step 1, the method for robustly aggregating the directional energy of multiple bands includes taking the median or a weighted average.
4. The self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described in claim 1, characterized in that: Step 2 specifically includes: Specifically, it includes: Let the image of the k-th band of the n-th scene be... The destriping result output by the network is ; To avoid clouds, shadows, and invalid pixels misleading consistency constraints, a reliability mask is introduced: Used to exclude unreliable regions; and defines pixel weights. All losses in consistency and fidelity are within Weighted calculation; Defining cross-band structural consistency specifically includes: assuming the structural operator is... Take the gradient magnitude, edge intensity, or structure energy map; calculate the structure representation for each destriped output. Robust normalization was performed on the structure diagram of each band to obtain... The cross-band structure consistency loss is defined as the dispersion of the structure map of each band relative to its cross-band central trend. ; The structural consistency loss is: ; strip candidate frequency domain mask It identifies narrow band regions in the frequency domain that may be dominated by stripes, and outputs the destriping results for each band. Calculate its spectrum And define the stripe energy penalty as the spectral energy within the mask region: ; Introducing data consistency constraints, let the spectrum of the original observations be... ,definition: ; The training objective is: in, , which is a weighting coefficient used to balance structural consistency, stripe suppression, and data fidelity.
5. The self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described in claim 1, characterized in that: Step 3 specifically includes: Let the image of the k-th band of the n-th sample be... , The basic output of the network is the destriped result. ; To suppress the interference of clouds, shadows, or invalid pixels on network training, a reliability mask is introduced: in Indicates unreliable pixels. Represents reliable pixels, used for weighting in loss calculation; backbone network Predict strip noise or strip residuals The final output is: ; Let the intermediate characteristics of the main trunk be... The frequency domain module first performs a differentiable Fourier transform on it to obtain the frequency domain features. Based on the strip candidate frequency domain mask obtained in step 1 Constructing learnable band modulation gain ,in For module parameters, ,in It is Sigmoid. It is a lightweight learnable function; Multiply the frequency domain features by the gain and inversely transform back to the spatial domain: in, This is a two-dimensional discrete Fourier inverse transform; and... The data is then fed into subsequent convolutional layers to continue predicting the strip residuals.
6. The self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described in claim 5, characterized in that: Step 3 also includes: Introducing structural protection mechanisms at the network level, specifically including: Multi-scale jump connections are used to ensure the direct transmission of low-frequency ground feature structures; Gain in the frequency domain modulation module Apply smoothing constraints to avoid generating new frequency domain artifacts; The strip suppression residual is restricted to a directional narrow-band structure; During training, network parameters and Joint optimization of self-supervised loss constructed in step 2.
7. The self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described in claim 1, characterized in that: Step 4 specifically includes: Select K band images for each scene After completing the unified meshing and amplitude normalization in step 1, a size of [size missing] is cropped from the same spatial location. The patch is used to obtain the observation image of the p-th spatial image patch in the n-th scene on the k-th spectral band. Where K is the total number of spectral bands; To prevent clouds, shadows, or invalid pixels from violating self-supervised consistency constraints, a reliability mask is introduced during training: in, Indicates unreliable pixels. Represents a reliable pixel; during training, it is cropped into a mask corresponding to the patch, and weights are defined. This is used to mask unreliable regions in loss calculations; Pre-compute stripe candidate frequency domain masks on each scene or each patch. The network maintains a consistent frequency domain coordinate convention throughout training and inference; during training, the network independently feeds forward to output destriping results for each band patch. Then, calculate the cross-band self-supervised loss within the multi-band group of the same scene; The self-supervised loss consists of three parts and is weighted over a reliable region, specifically including: Cross-band structural consistency loss is used to promote uniformity of destriped output across the structural domain; Strip frequency band energy penalty, constraining the network in The indicated direction reduces energy in the narrow frequency domain, thus enabling the network to learn the prior knowledge that the stripes are mainly concentrated in a specific frequency band; Data fidelity constraints on non-strip frequencies limit excessive modifications to non-strip frequencies by the network, preventing training from degenerating into full-frequency smoothing. The weighted sum of the three losses constitutes the overall training objective: in, Main loss weight, For cross-band structural consistency loss, To suppress energy loss in the strip frequency band, For non-strip band fidelity loss, This is a stability regularization term used to suppress numerical oscillations in the early stages of training.
8. A self-supervised remote sensing image stripe noise suppression method system based on frequency domain structure prior, characterized in that: The system has a program module corresponding to the steps of any one of claims 1-7, and executes the steps in the self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior during runtime.
9. A computer device, characterized in that: It includes a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, it performs the steps of the self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The storage medium is used to store a computer program that executes a self-supervised remote sensing image stripe noise suppression method based on frequency domain structure prior, as described in any one of claims 1-7.