Land cover classification method based on frequency-space dual domain interaction and uncertainty weighting

By using cross-modal frequency domain guided fusion of frequency domain feature maps and uncertainty-weighted classification, the problems of high-frequency detail loss and artifact generation in modal missing completion of optical remote sensing images are solved, achieving accurate identification and high-precision results for land cover classification.

CN121564443BActive Publication Date: 2026-04-10湖南工商大学
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
湖南工商大学
Filing Date
2026-01-21
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing methods for completing modal gaps in optical remote sensing images suffer from high-frequency detail loss, edge blurring, and artifact generation, and lack the ability to model cognitive uncertainties, resulting in low accuracy in land cover classification.

Method used

By acquiring spatial domain feature maps of SAR and optical remote sensing images and projecting them into the frequency domain, cross-modal frequency domain guided fusion is performed. The amplitude spectrum information of the SAR image is used to accurately repair the damaged frequency components of the optical remote sensing image. The uncertainty distribution map is used to characterize the uncertainty of each pixel and perform reliability-weighted classification.

Benefits of technology

It effectively restores the geometric structure and high-frequency edge information of land features obscured by clouds, avoids image blurring and spectral distortion, prevents the spread of misclassification, and improves the accuracy of land cover classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564443B_ABST
    Figure CN121564443B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of remote sensing information, and provides a land cover classification method based on frequency-space double-domain interaction and uncertainty weighting, comprising: acquiring a SAR image and an optical remote sensing image of a target region; acquiring a spatial domain feature map of the SAR image and the optical remote sensing image, and acquiring a frequency domain feature map of the SAR image and the optical remote sensing image based on the spatial domain feature map; performing cross-modal frequency domain guided fusion based on the frequency domain feature map of the SAR image and the optical remote sensing image to obtain a fusion feature map; inputting the fusion feature map into a decoder model for processing to obtain a pixel-level uncertainty distribution map and a completed optical remote sensing image; performing reliability weighted classification based on the uncertainty distribution map and the completed optical remote sensing image to obtain a land cover classification result of the target region. The application can improve the land cover classification precision.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of remote sensing information, and particularly relates to a land cover classification method based on frequency-space dual-domain interaction and uncertainty weighting. BACKGROUND

[0002] With the rapid development of earth observation technology, optical remote sensing images play an irreplaceable role in land cover classification, urban planning, disaster monitoring and other fields. However, optical sensors are passive imaging systems, which are easily disturbed by atmospheric conditions such as clouds, fog, haze and the like. According to statistics, about 60%-70% of the earth's surface is covered by clouds all year round, resulting in that the acquired optical remote sensing images often have partial or even large-area data missing. This modality missing phenomenon seriously destroys the spectral and texture information of ground objects, and greatly limits the all-weather application ability of remote sensing data.

[0003] In order to solve the above problems, the existing technical solutions are mainly divided into the following two categories:

[0004] 1) Traditional interpolation and filtering methods: using the pixels around the missing area for interpolation or based on tensor completion algorithm for recovery. This kind of method is only suitable for small range missing, for large area cloud cover, the recovered image is often blurred and cannot recover the real ground texture.

[0005] 2) Generative method based on deep learning: In recent years, deep learning methods represented by generative adversarial network (GAN) and convolutional neural network (CNN) are widely used in image restoration. This kind of method usually uses auxiliary modal data (such as synthetic aperture radar (SAR)) which is not affected by cloud cover as guidance, and reconstructs the missing optical remote sensing image by learning the mapping relationship between "SAR-optical".

[0006] Although the existing deep learning method improves the completion effect to a certain extent, there are still the following three technical defects for high-precision remote sensing application:

[0007] Defect one: only based on spatial domain learning, which is easy to cause "high frequency detail loss" and "artifacts". Most of the existing mainstream networks (such as deep learning-based image conversion model) directly calculate the loss function in the pixel space (Spatial Domain). Since the remote sensing image contains complex texture and edge information (high frequency component), pure pixel-level optimization tends to generate smooth and blurred results, and it is difficult to reconstruct clear road boundaries or building outlines. In addition, GAN model is easy to produce hallucination effect (Hallucination), that is, to generate seemingly realistic but actually non-existent ground objects, which seriously misleads the subsequent interpretation.

[0008] Defect two: lack of modeling ability for "cognitive uncertainty", and unreliable fault tolerance mechanism. Most of the current completion methods are deterministic, or can only model the accidental uncertainty of data through simple probability distribution. However, they cannot effectively capture the cognitive uncertainty of the model, that is, the "ignorance" caused by the model not having seen a certain complex cloud and fog scene. In the case of extreme occlusion, this defect of being unable to distinguish the source of uncertainty will cause the model to give wrong high-confidence predictions, seriously misleading the subsequent classification network.

[0009] Defect three: lack of physical adaptability and frequency selectivity in cross-modal feature fusion. Existing technologies mostly use simple splicing or addition when fusing SAR and optical data, ignoring the essential differences in imaging mechanism and spectral distribution between the two. This rough fusion method cannot adaptively convert the structural information of SAR to fit the spectral characteristics of optical, and lacks a dynamic evaluation mechanism for the damage degree of different frequency components, resulting in distortion of the spectral characteristics of the fused image and poor physical interpretability.

[0010] As mentioned above, the existing optical remote sensing image modal missing completion method mainly has two core defects: one is the excessive dependence on local convolution operation in pixel space, which leads to the loss of high-frequency texture details in the completed image, blurred edges, and easy to produce false artifacts that do not conform to physical facts; the second is the lack of evaluation mechanism for the reliability of the completion result, which directly uses low-confidence data containing false artifacts for subsequent classification, seriously misleading the feature extraction network and reducing the accuracy of land cover identification.

[0011] In summary, due to the poor completion effect of optical remote sensing images, there is a problem of low accuracy of land cover classification. SUMMARY

[0012] The embodiment of the present application provides a land cover classification method based on frequency-space dual-domain interaction and uncertainty weighting, which can solve the problem of low accuracy of land cover classification.

[0013] The embodiment of the present application provides a land cover classification method based on frequency-space dual-domain interaction and uncertainty weighting, which can solve the problem of low accuracy of land cover classification.

[0014] Obtain a pair of multi-modal remote sensing data of a target area; the pair of multi-modal remote sensing data includes a SAR image and an optical remote sensing image with data missing;

[0015] Respectively obtain the spatial domain feature map of the SAR image and the spatial domain feature map of the optical remote sensing image, and respectively project the spatial domain feature map of the SAR image and the spatial domain feature map of the optical remote sensing image into the frequency domain to obtain the frequency domain feature map of the SAR image and the frequency domain feature map of the optical remote sensing image;

[0016] Cross-modal frequency domain guided fusion is performed based on the frequency domain feature maps of SAR images and optical remote sensing images to obtain a fused feature map;

[0017] The fused feature map is input into the trained decoder model for processing to obtain a pixel-level uncertainty distribution map and a completed optical remote sensing image; the uncertainty distribution map is used to characterize the degree of uncertainty of each pixel in the completed optical remote sensing image;

[0018] Based on the uncertainty distribution map and the completed optical remote sensing image, a reliability-weighted classification is performed to obtain the land cover classification results for the target area.

[0019] Optionally, cross-modal frequency domain guided fusion is performed based on the frequency domain feature maps of SAR images and optical remote sensing images to obtain a fused feature map, including:

[0020] The frequency domain feature map of the SAR image is decoupled into the amplitude spectrum and the phase spectrum, and the frequency domain feature map of the optical remote sensing image is also decoupled into the amplitude spectrum and the phase spectrum.

[0021] Using the amplitude spectrum corresponding to the SAR image as a structural prior, the amplitude spectrum corresponding to the optical remote sensing image is guided and fused to obtain the fused amplitude spectrum.

[0022] The fused amplitude spectrum is recombined with the phase spectrum corresponding to the optical remote sensing image to obtain the recombined frequency domain feature map.

[0023] The recombined frequency domain feature map is subjected to inverse fast Fourier transform to obtain the fused feature map, which is a spatial feature map containing complete texture information.

[0024] Optionally, the amplitude spectrum corresponding to the SAR image is used as a structural prior to guide the fusion of the amplitude spectrum corresponding to the optical remote sensing image, resulting in a fused amplitude spectrum, including:

[0025] The fused amplitude spectrum is calculated using the following formula. :

[0026] ;

[0027] in, This represents the amplitude spectrum corresponding to the optical remote sensing image. Represents a dynamic spectrum gating function;

[0028] ;

[0029] This represents the Sigmoid activation function. This indicates the use of frequency domain channel attention modules for... To process, This indicates a splicing operation. This represents the amplitude spectrum corresponding to the SAR image. Represents element-wise multiplication. Represents a structural adaptive transformation function;

[0030] ;

[0031] express Convolutional layer This represents the ReLU activation function.

[0032] Optionally, the fused amplitude spectrum and the phase spectrum corresponding to the optical remote sensing image are recombined to obtain a reconstructed frequency domain feature map, including:

[0033] The recombined frequency domain feature map is calculated using the following formula. :

[0034] ;

[0035] in, Represents the imaginary unit. This represents the phase spectrum corresponding to the optical remote sensing image.

[0036] Optionally, the loss function used during decoder model training is the evidence regularization loss function, the expression of which is:

[0037] ;

[0038] in, This represents the value of the evidence regularization loss function. This represents the expected Bayesian risk loss. Indicates the annealing coefficient. , Indicates the current iteration number. This indicates the preset number of annealing steps. This represents the Dirichlet KL divergence regularization term;

[0039] ;

[0040] ;

[0041] This represents the total number of training samples. Indicates the first The true label corresponding to each training sample The decoder model represents the first... The output value of each training sample. Indicates the first an uncertainty strength parameter corresponding to each training sample, the uncertainty strength parameter being a trainable parameter of the decoder model during the training process, denotes a total number of classification categories, denotes a gamma function, denotes a double gamma function, denotes a total evidence amount, denotes an evidence parameter corresponding to the i-th category in a Dirichlet distribution.

[0042] Optionally, a reliability weighted classification is performed based on the uncertainty distribution map and the completed optical remote sensing image to obtain a land cover classification result of the target region, including:

[0043] inputting the completed optical remote sensing image into a feature extractor to perform feature extraction, to obtain a classification feature map;

[0044] performing normalization and negation on the uncertainty distribution map to obtain a reliability weight map;

[0045] performing element-level multiplication on the classification feature map using the reliability weight map to obtain a weighted feature map;

[0046] inputting the weighted feature map into a Softmax classification head for processing to obtain the land cover classification result of the target region.

[0047] Optionally, normalization and negation are performed on the uncertainty distribution map to obtain a reliability weight map, including:

[0048] the reliability weight map is calculated by the following formula :

[0049] ;

[0050] wherein, denotes a Sigmoid activation function, denotes the uncertainty distribution map.

[0051] The above scheme of the present application has the following beneficial effects:

[0052] ​In the embodiment of the present application, the spatial domain feature maps of the SAR image and the optical remote sensing image of the target region are extracted, and are projected into the frequency domain to obtain the frequency domain feature maps, then the cross-modal frequency domain guided fusion is performed based on the frequency domain feature maps of the SAR image and the optical remote sensing image, and the fusion feature map obtained by the fusion is input into the trained decoder model for processing to obtain the pixel-level uncertainty distribution map and the completed optical remote sensing image, and finally the reliability weighted classification is performed based on the uncertainty distribution map and the completed optical remote sensing image to obtain the land cover classification result of the target region. Among them, since the cross-modal frequency domain guided fusion based on the frequency domain feature map can adaptively utilize the amplitude spectrum information of the SAR image to accurately repair the damaged frequency components of the optical remote sensing image, effectively restore the geometric structure and high-frequency edge information of the ground object obscured by the cloud, avoid image blur and spectral distortion, and since the uncertainty distribution map can represent the uncertainty degree of each pixel point in the completed optical remote sensing image, when the reliability weighted classification is performed based on the uncertainty distribution map and the completed optical remote sensing image, the unreliable completion region can be accurately identified, and the feature interference of these regions can be automatically ignored, thereby effectively preventing the error classification propagation caused by the completion artifact and the model overconfidence, and achieving the effect of effectively improving the land cover classification precision.

[0053] Other beneficial effects of the present application will be described in detail in the subsequent specific embodiments section. BRIEF DESCRIPTION OF DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0055] Figure 1 The flow chart of the land cover classification method based on frequency-space dual domain interaction and uncertainty weighting provided by an embodiment of the present application;

[0056] Figure 2 The block diagram of the overall network architecture provided by an embodiment of the present application;

[0057] Figure 3 The schematic diagram of the frequency-space dual domain interaction principle provided by an embodiment of the present application. DETAILED DESCRIPTION

[0058] In the following description, for purposes of explanation and not limitation, specific details are set forth such as particular architectures, techniques, etc. in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, circuits, and

[0059] It is to be understood that the terminology "includes", "has", "holds", "contains" or "comprises", "comprising", or "including" when used in this specification and in the following claims, specifies the presence of stated features, integers, steps, operations, elements, or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof.

[0060] It is also to be understood that the terminology "and / or" when used in this specification and in the following claims, refers to one and / or all possible combinations of one or more relevantly listed items.

[0061] As used in this specification and in the claims, the terms "if" and "when" can be interpreted to mean "upon" or "in response to a determination" or "in response to a detection" depending on the context. Similarly, the phrase "if it is determined" or "if a detection is made" can be interpreted to mean "upon a determination" or "in response to a determination" or "upon a detection" or "in response to a detection" depending on the context.

[0062] In addition, the terms "first", "second", "third", etc. in the description of the present application and in the following claims are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0063] Reference in the specification to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places in the specification are not necessarily all referring to the same embodiment, although it can. The terms "including", "containing", "having" and variations thereof are meant to encompass the terms "including but not limited to" unless otherwise indicated.

[0064] To solve the problem of low accuracy of land cover classification, the embodiment of the present application provides a land cover classification method based on frequency-space dual domain interaction and uncertainty weighting. The method extracts the spatial domain feature map of the SAR image and the optical remote sensing image of the target area, projects it to the frequency domain to obtain the frequency domain feature map, then performs cross-modal frequency domain guided fusion based on the frequency domain feature map of the SAR image and the optical remote sensing image, inputs the fusion feature map obtained by fusion into the trained decoder model for processing to obtain the pixel-level uncertainty distribution map and the completed optical remote sensing image, and finally performs reliability weighted classification based on the uncertainty distribution map and the completed optical remote sensing image to obtain the land cover classification result of the target area. Since the cross-modal frequency domain guided fusion based on the frequency domain feature map can adaptively use the amplitude spectrum information of the SAR image to accurately repair the damaged frequency components of the optical remote sensing image, effectively restore the geometric structure and high-frequency edge information of the ground object obscured by clouds, avoid image blurring and spectral distortion, and since the uncertainty distribution map can represent the uncertainty degree of each pixel point in the completed optical remote sensing image, when performing reliability weighted classification based on the uncertainty distribution map and the completed optical remote sensing image, the unreliable completion area can be accurately identified, and the feature interference of these areas can be automatically ignored, thereby effectively preventing the propagation of false classification caused by completion artifacts and model overconfidence, and achieving the effect of effectively improving the accuracy of land cover classification.

[0065] The land cover classification method based on frequency-space dual domain interaction and uncertainty weighting provided by the present application will be exemplarily described below in combination with specific embodiments.

[0066] As shown in Figure 1 The land cover classification method based on frequency-space dual domain interaction and uncertainty weighting provided by the present application includes the following steps:

[0067] Step 11, obtaining a pair of multi-modal remote sensing data of the target area, the pair of multi-modal remote sensing data including a SAR image and an optical remote sensing image with data missing.

[0068] The above target area is an area that needs to be classified for land cover. In some embodiments of the present application, the above SAR image can be obtained by a platform equipped with a synthetic aperture radar sensor, for example, by a satellite; and the above optical remote sensing image can be obtained by a multi-spectral scanner, a hyperspectral imager or the like. The optical remote sensing image is an optical remote sensing image obscured by clouds and fog, and therefore has data missing.

[0069] It can be understood that after obtaining the SAR image and the optical remote sensing image with data missing of the target area, it is necessary to perform a conventional preprocessing operation on it in order to perform the subsequent steps.

[0070] Step 12: Obtain the spatial domain feature maps of the SAR image and the optical remote sensing image respectively, and project the spatial domain feature maps of the SAR image and the optical remote sensing image to the frequency domain respectively to obtain the frequency domain feature maps of the SAR image and the optical remote sensing image.

[0071] In some embodiments of this application, SAR images can be extracted using a deep neural network including a dual-stream encoder. and optical remote sensing images The spatial domain feature map is obtained, and then it is projected onto the frequency domain to obtain the frequency domain feature map of the SAR image and the frequency domain feature map of the optical remote sensing image.

[0072] Specifically, the implementation of step 12 above includes the following steps 12.1 to 12.2:

[0073] Step 12.1, transfer the SAR image (like Figure 2 SAR images and optical remote sensing images (like Figure 2 The multi-cloud map in the image is input into two convolutional neural network encoders with the same structure (e.g., the multi-cloud map in the image). Figure 2 The encoders A and B in the image are processed to obtain the spatial domain feature map of the SAR image. Spatial domain feature map of optical remote sensing images .in, , , , , These represent the number of channels, height, and width, respectively.

[0074] Step 12.2, analyze the spatial domain feature maps respectively. and spatial domain feature map Perform a two-dimensional fast Fourier transform on each channel (e.g. Figure 2 The Fast Fourier Transform (FFT) maps the spatial domain feature map to the frequency domain, resulting in the frequency domain feature maps of the SAR image and the optical remote sensing image.

[0075] Step 13: Perform cross-modal frequency domain guided fusion based on the frequency domain feature map of the SAR image and the frequency domain feature map of the optical remote sensing image to obtain the fused feature map.

[0076] In some embodiments of this application, the frequency domain feature map can be decoupled into amplitude spectrum and phase spectrum first, and then the amplitude spectrum information of the SAR image can be used to accurately repair the damaged frequency components of the optical remote sensing image.

[0077] Specifically, the implementation of step 13 above includes the following steps 13.1 to 13.4:

[0078] Step 13.1: Decouple the frequency domain feature map of the SAR image into amplitude spectrum and phase spectrum, and decouple the frequency domain feature map of the optical remote sensing image into amplitude spectrum and phase spectrum.

[0079] In some embodiments of this application, amplitude spectrum is used to characterize the texture and geometric structure information of ground features, and phase spectrum is used to characterize the location and contour information of ground features. The calculation formula is as follows:

[0080] ;

[0081] ;

[0082] In the above formula, Indicates amplitude spectrum, Represents the phase spectrum. and These are the real and imaginary parts of the complex feature (i.e., the frequency domain feature map), respectively.

[0083] It should be noted that, after the above formula is used to calculate the frequency domain feature map of the SAR image, the frequency domain feature map of the SAR image can be decoupled into the amplitude spectrum and the phase spectrum (e.g., Figure 2 The amplitude and phase spectra corresponding to the SAR image branches. After the above formula is used to calculate the frequency domain feature map of the optical remote sensing image, the frequency domain feature map of the optical remote sensing image can be decoupled into amplitude and phase spectra (e.g., Figure 2 (Amplitude and phase spectra corresponding to the branches in the multi-cloud diagram).

[0084] Step 13.2: Using the amplitude spectrum corresponding to the SAR image as a structural prior, guide the fusion of the amplitude spectrum corresponding to the optical remote sensing image to obtain the fused amplitude spectrum.

[0085] In some embodiments of this application, considering the fundamental differences in imaging mechanisms between SAR images and optical remote sensing images, their amplitude spectrum distributions cannot be directly equivalent. Therefore, this embodiment proposes a nonlinear fusion strategy based on a structure-adaptive transformation-dynamic gated injection mechanism. This strategy aims to learn a nonlinear mapping from SAR frequency domain features to optical frequency domain features, and inject structural information in residual form only in the frequency region where the amplitude spectrum of the optical image is damaged through a dynamic spectral gating mechanism. In some embodiments, a cross-modal frequency domain residual gated fusion module can be designed to perform guided fusion. This module specifically calculates the fused amplitude spectrum using the following formula. :

[0086] ;

[0087] in, This represents the amplitude spectrum corresponding to the optical remote sensing image. The dimension is , denotes element-wise multiplication, denotes dynamic spectral gating function, where is the dynamic spectral gating term.

[0088] ;

[0089] denotes Sigmoid activation function, denotes processing of by frequency channel attention block, denotes concatenation operation, denotes amplitude spectrum corresponding to SAR image, has dimension . Sigmoid activation function is also commonly known as logistic activation function, which is used as a nonlinear activation function in neural networks to map input values to the interval (0, 1).

[0090] is a complex nonlinear attention mechanism, which outputs a gating tensor with value range between to determine how much SAR information needs to be injected at each frequency point and each channel. It first concatenates the amplitude spectra of the two modalities in the channel dimension (Concat), and then calculates the gating weight through a frequency channel attention block (FCAB, Frequency Channel Attention Block). The frequency channel attention block first performs global average pooling (GAP) on the input frequency domain features to obtain global spectral descriptors, and then captures the dependency between frequency channels through two fully connected layers (FC), which mathematically simulates the adaptive evaluation of the importance of different frequency components.

[0091] denotes structure adaptive conversion function, is the structure adaptive conversion term.

[0092] ;

[0093] In the above formula, denotes convolutional layer, denotes ReLU activation function. Among them, Two layers of convolutional layers in are used for channel transformation in the frequency domain, and Figure 2 is a learnable micro network, which is used to map the amplitude spectrum features of the SAR image to the feature space of the optical image (i.e. asThe amplitude spectrum corresponding to the multi-cloud SAR image is guided by the amplitude spectrum corresponding to the multi-cloud SAR image, the style difference between the modes is eliminated, and the spectral distribution of the optical image is more suitable.

[0094] Step 13.3, recombine the fused amplitude spectrum and the phase spectrum corresponding to the optical remote sensing image to obtain a recombined frequency domain feature map.

[0095] It is worth mentioning that recombining the fused amplitude spectrum and the phase spectrum corresponding to the optical remote sensing image can ensure that the generated image has clear structure provided by SAR and does not have spatial position offset. The recombined frequency domain feature map can be calculated by the following formula :

[0096] ;

[0097] Wherein, is an imaginary unit, is the phase spectrum corresponding to the optical remote sensing image.

[0098] Step 13.4, inverse fast Fourier transform of the recombined frequency domain feature map to obtain a fused feature map, the fused feature map is a spatial feature map containing complete texture information.

[0099] That is, in some embodiments of the present application, inverse fast Fourier transform can be performed on the recombined frequency domain feature map (such as inverse transform in Figure 2 ), the recombined frequency domain feature map is mapped back to the spatial domain to obtain a spatial feature map containing complete texture information.

[0100] Step 14, input the fused feature map into the trained decoder model for processing to obtain a pixel-level uncertainty distribution map and a completed optical remote sensing image, the uncertainty distribution map is used to represent the uncertainty degree of each pixel point in the completed optical remote sensing image.

[0101] In some embodiments of the present application, the above decoder model can be Unet decoder (UNet is a convolutional neural network architecture for image segmentation tasks). In some embodiments, the decoder model can be constructed based on evidence theory (Evidence Theory). The uncertainty branch of the decoder model no longer outputs a single variance value, but outputs a non-negative evidence vector (for regression completion tasks, the pixel value can be discretized into intervals, or a continuous version is used, here a more general discretization form is described to show the complexity, and in actual regression, it is usually simplified to a single evidence value ).

[0102] The network output is ensured to be positive by the activation function Softplus, denoted as evidence . According to the evidence theory, this corresponds to the parameters of the Dirichlet distribution , . The total evidence mass is . At this point, the uncertainty of the prediction is modeled as Epistemic Uncertainty, defined as the inverse of the total evidence mass : , denotes the total number of classification classes.

[0103] The loss function in the training process of the above decoder model is an evidence regularization loss function, which aims to minimize the Bayesian risk of prediction error while penalizing the model for giving wrong high-confidence evidence through KL divergence, a measure method for quantifying the difference between two probability distributions. The expression of the evidence regularization loss function is:

[0104] ;

[0105] wherein denotes the value of the evidence regularization loss function, denotes the expected Bayesian risk loss, is the expected prediction error under the Dirichlet distribution, denotes the annealing coefficient, , denotes the current iteration number, denotes the preset annealing step number, which is small in the early stage of training to stabilize the main task learning, and increases to 1 in the later stage of training to enhance the KL regularization constraint, denotes the Dirichlet KL divergence regularization term, which is used to constrain the output distribution of the model not to deviate too far from the uniform distribution, preventing the model from being overconfident when there is not enough evidence.

[0106] ;

[0107] ;

[0108] denotes the total number of training samples, denotes the true label corresponding to the th training sample, denotes the output value of the decoder model for the th training sample, denotes the The uncertainty intensity parameter corresponds to each training sample, and the uncertainty intensity parameter is a trainable parameter of the decoder model during the training process. This indicates the total number of categories. Represents the gamma function. Represents the double gamma function. Indicates the total amount of evidence. This represents the th element in the Dirichlet distribution. Evidence parameters corresponding to each category.

[0109] For the prediction error term, the total amount of evidence The larger, The smaller the value, the more confident the model is in its predictions.

[0110] like Figure 2 As shown, the fused feature map output from the inverter is input into the decoder model (e.g., Figure 2 The decoder in the model processes the data, and the decoder model can then output a pixel-level uncertainty distribution map (such as...). Figure 2 Uncertainty map in the image) and the completed optical remote sensing image (e.g. Figure 2 (The complete image in the image).

[0111] Step 15: Perform reliability-weighted classification based on the uncertainty distribution map and the completed optical remote sensing image to obtain the land cover classification results for the target area.

[0112] In some embodiments of this application, step 15 is specifically implemented as follows: steps 15.1 to 15.4:

[0113] Step 15.1: Input the completed optical remote sensing image into the feature extractor for feature extraction to obtain a classification feature map.

[0114] In some embodiments of this application, such as Figure 2 As shown, conventional feature extractors can be used to process the completed optical remote sensing image (such as...). Figure 2 Feature extraction is performed on the complete image in the image to obtain a classification feature map (such as...). Figure 2 (Feature maps in the graph). As a preferred example, the feature extractor can be a graph neural network.

[0115] Step 15.2: Normalize the uncertainty distribution map and then invert it to obtain the reliability weight map.

[0116] Specifically, the reliability weight map can be calculated using the following formula. :

[0117] ;

[0118] in, This represents the Sigmoid activation function. Uncertainty distribution diagram, reliability weight diagram The closer the value is to 1, the more reliable the region is; the closer the value is to 0, the more reliable the region is as an incomplete artifact.

[0119] Step 15.3: Multiply the classification feature map element-wise using the reliability weight map to obtain the weighted feature map.

[0120] Specifically, the weighted feature map can be calculated using the following formula. :

[0121] ;

[0122] In the above formula, This represents the classification feature map. By using a reliability weight map to perform element-wise multiplication of the classification feature map, we can achieve the following: for imputation regions that the network is "unsure about," their feature responses are automatically set to zero to prevent them from misleading the classifier.

[0123] Step 15.4: Input the weighted feature map into the Softmax classification head for processing to obtain the land cover classification result of the target area.

[0124] In some embodiments of this application, such as Figure 2 As shown, classification feature maps (such as...) can be used to classify feature maps. Figure 2 (Feature map in the image) and uncertainty map are compared. The result of the calculation is input into the Softmax classification header (e.g.) Figure 2 The classification header in the data is processed to obtain the land cover classification results for the target area (e.g., ...). Figure 3 (The classification results in the model). Softmax is a mathematical function commonly used for multi-class classification problems.

[0125] The land cover classification results are used to indicate the type of the target area, such as cultivated land, forest land, grassland, wetland, and water body. It is understood that the target area may be of one type or may include multiple types; for example, part of the area may be forest land and another part may be water body.

[0126] It is understood that the learnable networks, learnable models, and other structures involving learnable parameters in the embodiments of this application are all trained before actual application (specifically, conventional deep learning training methods, such as stochastic gradient descent), so that the corresponding learnable parameters are at their optimal values, ensuring the accuracy of land cover classification.

[0127] Based on the above explanation, as Figure 3As shown, the land cover classification method based on frequency-space dual-domain interaction and uncertainty weighting provided in this application classifies land cover by analyzing the spatial domain feature maps respectively. (like Figure 3 (Optical feature map of clouds) and spatial domain feature map (like Figure 3 The SAR feature map in the image is subjected to Fast Fourier Transform (FFT) to obtain the frequency domain feature map of the SAR image and the frequency domain feature map of the optical remote sensing image. Then, these two are decoupled to obtain the amplitude spectrum corresponding to the cloud optical feature map (e.g., ...). Figure 3 (damaged amplitude map) and phase spectrum (such as) Figure 3 The original phase spectrum), and the amplitude spectrum corresponding to the SAR feature map (e.g. Figure 3 The amplitude map and phase spectrum are then fused through a cross-modal frequency domain residual gated fusion module (i.e., Figure 3 The CFRG module in the image is used to splice the data, apply dynamic spectral gating, and perform structural adaptive transformation to obtain the fused amplitude spectrum (e.g., ...). Figure 3 The amplitude spectrum after fusion is then reconstructed from the original phase spectrum and subjected to inverse Fourier transform to obtain the fused feature map (e.g., the amplitude spectrum after fusion is reconstructed from the original phase spectrum). ​ The fused feature map is processed using a decoder model to obtain a pixel-level uncertainty distribution map and a completed optical remote sensing image. Finally, a reliability-weighted classification is performed based on the uncertainty distribution map and the completed optical remote sensing image to obtain the land cover classification result of the target area.

[0128] As mentioned earlier, cross-modal frequency domain guided fusion based on frequency domain feature maps can adaptively utilize the amplitude spectrum information of SAR images to accurately repair the damaged frequency components of optical remote sensing images, effectively restoring the geometric structure and high-frequency edge information of land features obscured by clouds, avoiding image blurring and spectral distortion. At the same time, since the uncertainty distribution map can characterize the degree of uncertainty of each pixel in the completed optical remote sensing image, when performing reliability-weighted classification based on the uncertainty distribution map and the completed optical remote sensing image, it can accurately identify unreliable completion areas and automatically ignore the feature interference of these areas, thereby effectively preventing the propagation of misclassification caused by completion artifacts and model overconfidence, and achieving the effect of effectively improving the accuracy of land cover classification.

[0129] In summary, the land cover classification method based on frequency-space dual-domain interaction and uncertainty weighting provided in this application has the following advantages:

[0130] 1) Break through the limitation of single spatial domain repair, significantly improve the texture fidelity and physical consistency of the image: This application introduces the frequency domain processing mechanism, uses the decoupling characteristics of amplitude spectrum and phase spectrum, solves the deficiency of traditional convolutional network in long distance dependence modeling. By designing cross-modal frequency domain residual gating mechanism, it can adaptively use the amplitude spectrum information of SAR data to accurately repair the damaged frequency components of optical image, more effectively recover the geometric structure and high frequency edge information of ground objects obscured by clouds, and avoid image blur and spectral distortion.

[0131] 2) Establish a "completion-classification" closed-loop fault-tolerant mechanism based on evidence reasoning, improve the robustness of the system: This application introduces the evidence reasoning theory based on Dirichlet prior, which can effectively quantify the cognitive uncertainty of the model. The risk aversion strategy constructed by this makes the subsequent classification network accurately identify the unreliable completion area (high cognitive uncertainty) caused by the lack of evidence, and automatically ignore the feature interference of these areas. This effectively prevents the spread of false classification caused by completion artifacts and model overconfidence, especially in extreme cloud and fog obscured scenes.

[0132] 3) Realize the physical level deep fusion of cross-modal information: Unlike simple channel splicing, this application is based on the remote sensing imaging mechanism (SAR reflects structure, optical reflects texture), and the features are interacted at the frequency domain level, so that the modal fusion is more physically interpretable. The generated image is not only visually realistic, but also has stronger spectral consistency.

[0133] The above is the preferred embodiment of the present application. It should be pointed out that for ordinary skilled persons in the technical field, without departing from the principles described in the present application, a number of improvements and refinements can be made, which should also be considered as the protection scope of the present application.

Claims

1. A land cover classification method based on frequency-space dual domain interaction and uncertainty weighting, characterized in that, The method comprises the following steps: obtaining a pair of multi-modal remote sensing data of a target area; the pair of multi-modal remote sensing data comprises a SAR image and an optical remote sensing image with data missing; respectively obtaining a spatial domain feature map of the SAR image and a spatial domain feature map of the optical remote sensing image, and respectively projecting the spatial domain feature map of the SAR image and the spatial domain feature map of the optical remote sensing image into a frequency domain to obtain a frequency domain feature map of the SAR image and a frequency domain feature map of the optical remote sensing image; performing cross-modal frequency domain guided fusion based on the frequency domain feature map of the SAR image and the frequency domain feature map of the optical remote sensing image to obtain a fusion feature map; inputting the fusion feature map into a trained decoder model for processing to obtain a pixel-level uncertainty distribution map and a completed optical remote sensing image; the uncertainty distribution map is used to represent the uncertainty degree of each pixel point in the completed optical remote sensing image; performing reliability weighted classification based on the uncertainty distribution map and the completed optical remote sensing image to obtain a land cover classification result of the target area; the cross-modal frequency domain guided fusion based on the frequency domain feature map of the SAR image and the frequency domain feature map of the optical remote sensing image to obtain a fusion feature map comprises: decoupling the frequency domain feature map of the SAR image into an amplitude spectrum and a phase spectrum, and decoupling the frequency domain feature map of the optical remote sensing image into an amplitude spectrum and a phase spectrum; taking the amplitude spectrum corresponding to the SAR image as a structure prior, and performing guided fusion on the amplitude spectrum corresponding to the optical remote sensing image to obtain a fused amplitude spectrum; recombining the fused amplitude spectrum and the phase spectrum corresponding to the optical remote sensing image to obtain a recombined frequency domain feature map; performing inverse fast Fourier transform on the recombined frequency domain feature map to obtain a fusion feature map, which is a spatial feature map containing complete texture information; the taking the amplitude spectrum corresponding to the SAR image as a structure prior, and performing guided fusion on the amplitude spectrum corresponding to the optical remote sensing image to obtain a fused amplitude spectrum comprises: The amplitude spectrum after fusion is calculated by the following equation : ; wherein, represents an amplitude spectrum corresponding to the optical remote sensing image, represents a dynamic a spectrum gating function; ; denotes a Sigmoid activation function, denotes processing with a frequency domain channel attention module denotes processing, denotes a concatenation operation, denotes an amplitude spectrum corresponding to the SAR image, denotes an element-wise multiplication, denotes a structure adaptive conversion function; ; denotes convolutional layer, denotes a ReLU activation function; the recombining the fused amplitude spectrum and the phase spectrum corresponding to the optical remote sensing image to obtain a recombined frequency domain feature map comprises: The recombined frequency domain feature map is calculated by the following formula : ; wherein, denotes the imaginary unit, denotes a phase spectrum corresponding to the optical remote sensing image.

2. The land cover classification method of claim 1, wherein, the loss function in the training process of the decoder model is an evidence regularization loss function, and the expression of the evidence regularization loss function is: ; wherein, denotes a value of the evidence regularized loss function, denotes the expected Bayesian risk loss, denotes an annealing coefficient, , denotes a current iteration number, denotes a preset annealing step number, denotes a Dirichlet KL divergence regularizer term; ; ; denotes the total number of training samples, denotes the true label corresponding to the th training sample, denotes the output value of the decoder model for the th training sample, denotes the uncertainty intensity parameter corresponding to the th training sample, the uncertainty intensity parameter being a trainable parameter of the decoder model during the training process, denotes the total number of classification categories, denotes the gamma function, denotes the digamma function, denotes the total evidence quantity, denotes the evidence parameter corresponding to the th category in the Dirichlet distribution.

3. The land cover classification method of claim 1, wherein, the performing reliability weighted classification based on the uncertainty distribution map and the completed optical remote sensing image to obtain a land cover classification result of the target area comprises: inputting the completed optical remote sensing image into a feature extractor for feature extraction to obtain a classification feature map; performing normalization on the uncertainty distribution map and then taking the inverse to obtain a reliability weight map; performing element-level multiplication on the classification feature map using the reliability weight map to obtain a weighted feature map; inputting the weighted feature map into a Softmax classification head for processing to obtain the land cover classification result of the target area.

4. The land cover classification method of claim 3, wherein, the performing normalization on the uncertainty distribution map and then taking the inverse to obtain a reliability weight map comprises: The reliability weight map is calculated by the following equation : ; wherein, denotes a Sigmoid activation function, denotes an uncertainty distribution map.

Citation Information

Patent Citations

  • Farmland parcel identification method based on optical-Ka frequency band SAR (Synthetic Aperture Radar) feature fusion

    CN120544048A

  • System for physical-virtual environment fusion

    US20200394409A1