Multi-class feature recognition method for multi-modal remote sensing images based on feature decoupling

By employing multimodal feature decoupling and dynamic prototype classification techniques, the problems of insufficient information and labeling dependence in multi-category land cover identification in remote sensing images are solved, achieving efficient and accurate multi-category land cover identification.

CN122116167APending Publication Date: 2026-05-29CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIV
Filing Date
2026-02-06
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing remote sensing image multi-category ground feature identification technologies rely on single-modal data, which suffers from insufficient information and environmental interference limitations. The dependence on labeled data leads to high costs and poor adaptability, making it difficult to achieve accurate identification under extreme conditions.

Method used

A feature-decoupling-based multimodal remote sensing image recognition method is adopted. By extracting ground feature features from different remote sensing modal data through a multimodal feature extraction network, and using feature decoupling loss and dynamic prototype classification technology for unsupervised learning, rapid ground feature type recognition of multimodal remote sensing images is achieved.

Benefits of technology

It significantly improves the accuracy and adaptability of multi-category land cover identification, reduces identification time consumption, and enhances the classification accuracy and category differentiation ability of complex land cover categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116167A_ABST
    Figure CN122116167A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of multi-modal remote sensing image multi-class feature recognition methods based on feature decoupling, belong to remote sensing image recognition and understanding field.The method includes the following steps: S1: obtaining multi-modal remote sensing image data;S2: multi-modal feature extraction, including convolution module, Fourier orthogonal attention module, center feature fusion module;S3: calculate multi-modal contrast decoupling loss;S4: feature fusion and feature recognition, obtain feature recognition result chart.The performance of the method and system described in the present application is better than other multi-class feature recognition methods, and the present method can better obtain and identify feature type, and has advantage in identifying multi-class feature than other methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image recognition and understanding, and in particular to a method for recognizing multiple types of ground features in multimodal remote sensing images based on feature decoupling. Background Technology

[0002] With the rapid development of remote sensing technology and artificial intelligence, the automated processing of remote sensing images is playing an increasingly crucial role in fields such as military monitoring, resource management, agricultural surveying, and urban planning. Among these, multi-category land cover identification is a core task of remote sensing image analysis, aiming to automatically extract information on different land covers from remote sensing images and achieve accurate classification. Currently, multi-category land cover identification in remote sensing images mainly relies on single-modality remote sensing images, most commonly hyperspectral remote sensing images, synthetic aperture radar, and lidar. Although single-modality hyperspectral remote sensing images have achieved a certain level of automation in land cover classification, the following major problems still exist: First, single-modal remote sensing images cannot comprehensively describe the characteristics of systematic ground features. For example, while hyperspectral remote sensing images can provide rich spectral information about ground features, they are often limited by spectral distortion caused by cloud cover. Similarly, although synthetic aperture radar (SAR) remote sensing images can acquire all-day climate structure features, their geometric ambiguity makes it difficult to accurately distinguish complex urban structures, thus limiting the accuracy of multi-feature identification under extreme conditions. These limitations highlight the necessity of utilizing multimodal data to improve the accuracy of multi-feature identification. This helps alleviate the problem of insufficient information from single-modal input, reduces environmental interference, and achieves more accurate ground feature identification.

[0003] Secondly, compared to traditional computer vision tasks, multi-object recognition requires labeling all pixels in an image, which is not only time-consuming and labor-intensive but also requires more specialized knowledge. This directly leads to the extreme difficulty and high cost of labeling remote sensing images. Furthermore, the high dependence of existing automated multi-object recognition technologies for remote sensing images on labeled data limits the adaptability of the models, thus hindering the further development and application of automated remote sensing image analysis.

[0004] Therefore, there is an urgent need for an unsupervised learning approach that can break through the excessive reliance of traditional supervised information on labeled data, and at the same time combine an effective multimodal remote sensing image feature extraction and fusion mechanism to develop efficient land cover recognition technology and expand the accuracy and adaptability of multimodal remote sensing image multi-class land cover recognition. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a method for multi-class land cover recognition based on feature decoupling in multimodal remote sensing images. This method extracts land cover features from different remote sensing modalities using a multimodal feature extraction network; it utilizes feature decoupling loss based on multimodal contrast loss to simultaneously constrain common and unique features among different modalities; and it employs dynamic prototype classification technology to achieve rapid land cover type recognition from large-scale remote sensing image data, significantly reducing the time consumption for land cover recognition. This method can better extract identifiable features from land covers that are conducive to classification, significantly improving the classification accuracy and class distinction ability of complex land cover categories. To achieve the above objectives, this invention provides the following technical solution: A method for multi-class land cover recognition based on feature decoupling in multimodal remote sensing images includes the following steps: S1: Acquire multimodal remote sensing image data, perform data preprocessing such as radiometric and geometric correction, cloud removal and noise reduction on the remote sensing images, and construct sample data from the remote sensing images centered on the target pixels according to the patch size; S2: Extract features from the sample data based on a multimodal feature extraction network to extract rich semantic information of land covers from different modalities; S3: Calculate the multimodal contrast decoupling loss and train the network in an unsupervised manner; S4: Fuse the modal features and further input them into a dynamic prototype classification module for fine land cover type recognition, output the classification results, and generate the final multi-class land cover recognition result of the remote sensing image.

[0006] Furthermore, in step S1, the land cover identification data of the multimodal remote sensing image data is prepared, specifically including: firstly, cutting all samples according to patch size and dividing them according to a certain ratio to obtain a multimodal dataset. N is the number of samples, and each pair of data contains a sample with m+1 modalities. and their corresponding tags , in and These represent the input space and the label space, respectively. This refers to data patches extracted from multimodal remote sensing images using a sliding window. Specifically, it involves acquiring spatial patch data corresponding to a target pixel within the entire image, centered on that pixel and with fixed length and width. The multimodal images specifically include hyperspectral images, multispectral images, SAR images, and LiDAR images. Image patches from different modalities are extracted using a feature extraction module with the same structure but independent parameters. Furthermore... This represents one-dimensional spectral information from hyperspectral remote sensing images that does not contain spatial information.

[0007] Furthermore, in step S2, the multimodal feature extraction network structure includes a convolutional feature extraction module, a Fourier forward duct attention module, and a center feature fusion module; firstly, for the v-th modality of the i-th sample in the dataset... Shallow ground features are captured through three cascaded 3×3 convolutions. The formula is as follows:

[0008] in , , These are the Batch Normalization function, the GELU function, and the convolution function, respectively.

[0009] Furthermore, in the Fourier positive traffic lane attention module, firstly... Perform convolution operations Orthogonal filters are used to compress the spatial information of each feature, as shown in the following formula:

[0010] in This represents a Gramm-Schmidt quadrature filter; for a specific The implementation process first targets the shallow features of the ground. Where C is the number of channels and the spatial size Random initialization yields a matrix with the same shape. Then calculate The Gram-Schmidt orthogonalization process is performed, and the specific formula is as follows:

[0011] in Represent the vector dot product, and then use right The specific formula for quadrature modulation is as follows:

[0012] therefore for Feature representation after Gramm-Schmidt orthogonalization; then... Optimal frequency information is obtained by analyzing the frequency domain after orthogonal spatial transformation, and by generating a mask. To preserve low-frequency components, thereby eliminating heterogeneous noise and enhancing the consistency of cross-modal characteristics, where H and W are the features of the current sample. Space size, To sort by value from smallest to largest and select values ​​of a specified length, This means retaining the first 20% of frequency components with the smallest amplitude. Subsequently, a Fourier quadrature filter is obtained. The formula is as follows:

[0013] Furthermore, the Fourier quadrature filter The vector is compressed into a channel vector, which is used to generate the attention weights A, as shown in the following formula:

[0014] in For the Sigmoid function, The ReLU activation function is used. , The weights are learnable matrices. Then, the attention vector A is multiplied by the input features, and residual connections are added to the resulting features to mitigate the potential vanishing gradient problem, as shown in the following formula:

[0015] Furthermore, regarding Convolution The output features of Fourier orthogonal attention are obtained. In the overall network framework, a three-stage feature extraction module is composed of three cascaded Fourier orthogonal attention modules. This is the feature output for the i-th stage. Furthermore, since remote sensing image classification is a pixel-level recognition task, the patch data for each pixel is cropped from the entire image at a certain size centered on that pixel. Therefore, features closer to the center of the patch have more discriminative features for pixel category. Thus, a center feature fusion module is designed to retain the most representative features by cropping the central region, as shown in the following formula:

[0016] in, This represents the floor function. Output features for the i-th stage feature extraction module The central features are then dynamically adjusted according to the number of layers in each stage, and they are concatenated along the channel dimension, as shown in the following formula:

[0017] in This represents a vector concatenation operation. To prevent overfitting of deep feature semantic information, appropriate weights are set to reset the credit for feature fusion. Specifically, the weights and credits are... n represents the total number of feature extraction stages. Let be the number of output feature dimensions in the k-th feature extraction stage. Let be the number of output feature dimensions in the i-th feature extraction stage. The fusion feature represents the central features of the three stages, where B is the batch size. Finally, the latent features of each modality are obtained through FFN, with a feature dimension of d. For spectral modes Where m is the dimension of the original spectral sequence corresponding to the target pixel in the hyperspectral image, we directly capture its global relevant information through channel attention. First, obtain the spectral information of the shallow layer. The specific formula is as follows:

[0018] in , A trainable parameter matrix is ​​then generated. Attention is calculated using the following formula:

[0019] in For the corresponding attention mechanism, the query, key, and value are... , , To obtain the corresponding trainable parameter matrix, the attention values ​​are then calculated as follows:

[0020] in For the Softmax function, the final The feature dimension is d-dimensional.

[0021] Furthermore, in step S3, the features of each modality embedding layer learned by the network are decomposed into two parts. dc represents the number of dimensions of the common features in the feature dimension, where This represents the modal common features (i.e., the first part of the overall embedding feature dimension d). dimension), Represents modality-unique features (i.e., the first-order features divided by the first-order features in the overall embedding feature dimension d). (All feature dimensions after the dimension). First, construct the joint feature set. ,in Let be the feature representation of the training batch samples of the i-th modality. For the feature representation of the training batch samples of the j-th modality, SimCLR is used as the contrastive loss to maximize the consistency of the same sample among different modalities in the latent space, as shown in the following formula:

[0022] in For the cosine similarity between samples, exp( ) is an exponential function.

[0023] Furthermore, the feature decoupling loss is calculated. This loss leverages the high consistency of common features across different modalities in the latent space, while preserving the specificity of unique features for each modality and eliminating redundancy. The specific formula is as follows:

[0024] in , It is the balance coefficient. It is the common feature alignment loss. It uses unique feature orthogonal loss. Specifically, it's based on the similarity matrix. The common feature alignment loss is calculated using the following formula:

[0025] Where I represents the element corresponding to the diagonal element; based on the Euclidean distance between unique features. The formula for calculating the unique feature orthogonal loss is as follows:

[0026] Finally, the multimodal representation framework in this paper is optimized using the total loss function, as shown in the following formula:

[0027] The overall network uses the Adam optimizer, is trained with a learning rate of 0.0001 and a batch size of 128, and undergoes 200 iterations, with dynamic prototype updates performed during each iteration.

[0028] Furthermore, in step S4, during each iteration, land cover identification is performed through dynamic prototype classification. First, the decoupled multimodal features after representation learning are concatenated to preserve the unique features of each modality as much as possible, as shown in the following formula:

[0029] Then, the K-means algorithm is used to perform cluster analysis on the overall sample, as shown in the following formula:

[0030] in, For the K-means clustering process, Let P be the set of feature representations for all samples after the learning process is complete, and let P be the set of cluster labels for the entire sample. =Based on the k-means clustering results P in the current iteration, 10% of the sample features are randomly selected. The formula for calculating the categorical prototype features is as follows:

[0031] in for In the example, the number of samples with cluster label c. For the corresponding sample features, The cluster center is the corresponding category c.

[0032] The overall feature dimension is obtained by concatenating the average of the common features of all modalities with the private features of all modalities. Then, the class prototype is momentum-modified in the following way:

[0033] in, For the category prototype updated during the t-th iteration, Let be the weighted sum of the category prototypes from iteration t-1 and the current category prototype. The weight coefficients range from 0 to 1, and the category prototype updated in the last iteration is used as the final category prototype. C, Then, by calculating the similarity between the features of the land cover samples and the category prototypes, the land cover category discrimination boundary is dynamically adjusted, and the classification result of multi-category land cover identification is output, as shown in the following formula:

[0034] in Let K be the prototype vector of the j-th category, K be the total number of categories, and finally output the land cover category recognition result map.

[0035] The beneficial effects of this invention are as follows: The Fourier orthogonal attention module and central feature fusion module in the feature extraction network employed in this invention can fully capture the rich and discriminative semantic information in multimodal remote sensing image data that is beneficial for land cover identification. Secondly, the feature decoupling contrastive loss focuses specifically on the common and unique features of each modality within the entire feature set. Through contrastive learning, the consistency of common features can be enhanced while retaining the unique features of different modalities, resulting in more discriminative feature representations. Then, a dynamic prototype classification technique is used to achieve rapid identification of multiple land cover classes in remote sensing images. This invention proposes a representation learning method for decoupling multimodal remote sensing land cover information, generating discriminative multimodal image representations. This invention proposes a dynamic prototype classification technique for multi-class land cover identification in remote sensing images. This network focuses on the feature representation of each class to obtain class features with strong representational capabilities. Experimental results on publicly available multimodal remote sensing image multi-class land cover identification datasets show that the proposed scheme outperforms state-of-the-art image land cover identification methods.

[0036] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0037] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 This is a flowchart of a multi-modal remote sensing image multi-category land cover identification method based on feature decoupling, as described in this invention. Figure 2 This is a framework diagram of a multi-modal remote sensing image multi-category land cover recognition method based on feature decoupling proposed in this invention. Figure 3 This is a feature extraction network diagram of the present invention; Figure 4 The t-SNE visualization on the Trento dataset is (a) iterated once, and (b) iterated 200 times. Figure 5 Visualizations of different methods on the Houston dataset, where (a) ground truth, (b) MFLVC, (c) GCFAGG, (d) EMVCC, (e) AMKSC, (f) DGLAP, (g) FPFC, and (h) FoFD-Clust. Detailed Implementation

[0038] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.

[0039] Figure 1 This is a flowchart of the method of the present invention. The present invention provides a multi-modal remote sensing image multi-category land cover recognition method based on feature decoupling. As shown in the figure, in the image acquisition stage, the remote sensing image undergoes data preprocessing such as radiometric and geometric correction, cloud removal, and noise reduction. Sample data is constructed from the remote sensing image according to the patch size and centered on the target pixel. The framework for multi-category land cover recognition is as follows: Figure 2 As shown, this method can learn detailed and discriminative features of ground features from multimodal remote sensing data, thereby achieving rapid and accurate identification of ground feature categories. First, a feature extraction network is used to extract ground feature characteristics. Then, the ground feature characteristics are decoupled to obtain common and unique features of ground features from different remote sensing data, and the feature contrast loss is calculated. Next, the learned features are fused, and a dynamic prototype classification technique is proposed to achieve rapid identification of ground feature types. The multimodal remote sensing image multi-category ground feature identification method proposed in this invention, based on feature decoupling, pays particular attention to common and unique features in the entire multimodal remote sensing data. Through contrastive learning and dynamic prototype classification techniques, it enhances the representation ability of ground features, achieving accurate and rapid identification.

[0040] Specifically, the technical solution of the present invention includes the following: 1. Acquire multimodal remote sensing image data: Perform data preprocessing streams such as radiometric and geometric correction, cloud removal, and noise reduction on the remote sensing images. Cut all samples according to patch size and divide them into multimodal datasets according to a certain ratio. N is the number of samples, and each pair of data contains a sample with m+1 modalities. and their corresponding tags , in and These represent the input space and the label space, respectively. This represents a data segment extracted from a multimodal remote sensing image using a sliding window. This represents spectral information from hyperspectral remote sensing images.

[0041] 2. Multimodal feature extraction: such as Figure 3 As shown, the v-th mode of the i-th sample in the dataset Inputting three concatenated 3×3 convolutions captures shallow ground features. The formula is as follows:

[0042] in , , These are the Batch Normalization function, the GELU function, and the convolution function, respectively. First, let's... Perform convolution operations Then, the spatial information of each feature is compressed using the orthogonal filter F in the Fourier orthogonal channel attention module, as shown in the following formula:

[0043] in Represent the Gramm-Schmidt orthogonalization process; subsequently, Optimal frequency information is obtained by analyzing the frequency domain after orthogonal spatial transformation, and by generating a mask. To preserve low-frequency components, among which , This means retaining the first 20% of frequency components with the smallest amplitude. Subsequently, a Fourier quadrature filter is obtained. The formula is as follows:

[0044] Next, the Fourier quadrature filter... The vector is compressed into a channel vector, which is used to generate the attention weights A, as shown in the following formula:

[0045] in For the Sigmoid function, The ReLU activation function is used. , The weights are learnable matrices. Then, the attention vector A is multiplied by the input features, and residual connections are added to the resulting features to mitigate the potential vanishing gradient problem, as shown in the following formula:

[0046] right Convolution The output features of Fourier orthogonal attention are obtained. In the overall network framework, a three-stage feature extraction module is composed of three cascaded Fourier orthogonal attention modules. For the feature output of the i-th stage, such as Figure 3 As shown, in the central feature fusion module, the most representative features are retained by cropping the central region, as shown in the following formula:

[0047] in, This represents the floor function, which dynamically adjusts its weights based on the number of layers in each stage, and then concatenates them along the channel dimension, as shown in the following formula:

[0048] in This represents a vector concatenation operation, where weights are reset to reflect the information. Finally, the latent features of each modality are obtained through FFN: For spectral modes Where m is the dimension of the original spectral sequence corresponding to the target pixel in the hyperspectral image, we directly capture its global relevant information through channel attention. First, obtain the spectral information of the shallow layer. The specific formula is as follows:

[0049] in , A trainable parameter matrix is ​​then generated. Attention is calculated using the following formula:

[0050] in For the corresponding attention mechanism, the query, key, and value are... , , To obtain the corresponding trainable parameter matrix, the attention values ​​are then calculated as follows:

[0051] in For the Softmax function, the final The feature dimension is d-dimensional.

[0052] 3. Calculate the multimodal contrast decoupling loss: Decompose the embedded layer features learned by the network into two parts. ,in Represents modal common features, To represent modality-unique features, a joint feature set is first constructed. ,in Let be the feature representation of the training batch samples of the i-th modality. For the feature representation of the training batch samples of the j-th modality, SimCLR is used as the contrastive loss to maximize the consistency of the same sample among different modalities in the latent space, as shown in the following formula:

[0053] in Let be the cosine similarity between samples, and 'Neg' represent negative sample pairs. Then, the feature decoupling loss is calculated, with the specific formula as follows:

[0054] in , It is the balance coefficient. It is the common feature alignment loss. It is a unique feature orthogonal loss. First, it is based on the similarity matrix. The common feature alignment loss is calculated using the following formula:

[0055] Where I represents the element corresponding to the diagonal element; secondly, based on the Euclidean distance between unique features... The formula for calculating the unique feature orthogonal loss is as follows:

[0056] Finally, the multimodal representation framework in this paper is optimized using the total loss function, as shown in the following formula:

[0057] The entire network uses the Adam optimizer, with a learning rate of 0.0001 and a batch size of 128, and is trained for 200 iterations, with dynamic prototype updates during each iteration.

[0058] 4. Feature Fusion and Land Feature Recognition: The decoupled multimodal features after representation learning are concatenated to preserve the unique features of each modality as much as possible. The formula is as follows:

[0059] Secondly, the K-means algorithm is used to perform cluster analysis on the randomly selected samples, as shown in the following formula:

[0060] Then, the class prototype is momentum-modified in the following way:

[0061] in =Based on the k-means clustering results P in the current iteration, 10% of the sample features are randomly selected. The formula for calculating the categorical prototype features is as follows:

[0062] in for In the example, the number of samples with cluster label c. For the corresponding sample features, The cluster center is the corresponding category c.

[0063] Then, the similarity between the features of the land cover samples and the category prototypes is calculated to dynamically adjust the land cover category discrimination boundary, and the classification results of multi-category land cover identification are output, as shown in the following formula:

[0064] Finally, a map showing the results of land cover category identification is generated.

[0065] like Figure 4 , Figure 5 This paper presents the visualization results of a multimodal remote sensing image multi-class land cover recognition method based on feature decoupling, as described in this invention, on the Terento and Houston datasets. The results show that the land cover features exhibit significant intra-class compactness and inter-class clarity, and the recognition results are accurate. The land cover recognition performance of this invention can be further illustrated through comparative experiments. The method of this invention is compared with other existing methods such as MFLVC, GCFAGG, EMVCC, MDC, AMKSC, DGLAP, FPFC, and FoFD-Clust on the Houston13, Muufl, and Trento datasets. Overall Accuracy (ACC), Kappa coefficient, PUR, and Recall (NMI) are calculated respectively. Table 1 shows the values ​​of each index for the detection results of different methods. Table 1 Comparison of various methods on different datasets

[0066] As can be seen, the method of this invention achieves the best accuracy on the three datasets. The method proposed in this invention can better extract the identifiability of ground features and has advantages over other methods in multi-category ground feature identification.

[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications should be covered within the scope of the claims of the present invention.

Claims

1. A method for identifying multiple categories of ground features in multimodal remote sensing images based on feature decoupling, characterized in that: The method includes the following steps: S1: Acquire multimodal remote sensing image data; S2: Multimodal feature extraction; S3: Calculate the multimodal contrast decoupling loss; S4: Feature fusion and ground feature recognition; In step S1, all samples are cut according to patch size and divided into multimodal datasets according to a certain ratio. N is the number of samples, and each pair of data contains a sample with m+1 modal images. and their corresponding tags , in and These represent the input space and the label space, respectively. This represents a data segment extracted from a multimodal remote sensing image using a sliding window. This represents spectral information from hyperspectral remote sensing images; In step S2, the multimodal feature extraction includes a convolutional feature extraction module, a Fourier forward channel attention module, and a center feature fusion module for extracting ground feature features of each modality from shallow to deep. In the convolutional feature extraction module, shallow ground feature features are captured by three cascaded 3×3 convolutions. The formula is as follows: ; in , , These are the Batch Normalization function, the GELU activation function, and the convolution kernel, respectively. The convolution function, Regarding the image data of the vth modality of the i-th sample These are shallow features obtained from the corresponding modal image data; right Perform convolution operations Then, in the Fourier orthogonal channel attention module, an orthogonal filter is used to compress the spatial information of each feature, as shown in the following formula: ; in This represents the Gramm-Schmidt orthogonalization process. for Feature representation after Gramm-Schmidt orthogonalization; Optimal frequency information is obtained by transforming orthogonal space to the frequency domain and generating a mask. This is done to preserve low-frequency components, thereby eliminating heterogeneous noise and enhancing the consistency of cross-modal image features, where , For floor operations, H and W are the current image features. Space size, To sort by value from smallest to largest and select values ​​of a specified length, This means retaining the first 20% of frequency components with the smallest amplitude to obtain a Fourier quadrature filter. The formula is as follows: ; in, The inverse Fourier transform function is used to transform the Fourier quadrature filter. The vector is compressed into a channel vector, which is used to generate the attention weights A, as shown in the following formula: ; in For the Sigmoid function, The ReLU activation function is used. , To create learnable matrix weights, the attention vector A is multiplied by the input features, and residual connections are added to the resulting features to mitigate the potential gradient vanishing problem, as shown in the following formula: ; right Perform convolution The output features of Fourier orthogonal attention are obtained. In the overall network framework, a three-stage feature extraction module is composed of three cascaded Fourier orthogonal attention modules. The feature output is for the i-th stage. Then, in the central feature fusion module, semantic features from different stages are extracted. In the central feature fusion module, the most representative features are retained by cropping the central region. The formula is as follows: ; in, This indicates the floor function. Output features for the i-th stage feature extraction module The central features are then dynamically adjusted according to the number of layers in each stage, and they are concatenated along the channel dimension, as shown in the following formula: ; in This represents a vector concatenation operation, where weights are reset to reflect the information. n represents the number of overall feature extraction stages. Let be the number of output feature dimensions in the k-th feature extraction stage. Let be the number of output feature dimensions in the i-th feature extraction stage. The fusion features of the central features of the three stages, where B is the batch size, are obtained through a feedforward network (FFN) to obtain the latent feature representation of each modality image. For spectral modes It directly captures globally relevant feature representations through channel attention. The specific formula is as follows: ; in , A trainable parameter matrix is ​​then generated. Attention is calculated using the following formula: ; in For the corresponding attention mechanism, the query, key, and value are... , , To obtain the corresponding trainable parameter matrix, the attention values ​​are then calculated as follows: ; in For the Softmax function, Feature dimensions and The feature dimensions are the same, both being d-dimensional.

2. The method for identifying multiple categories of ground features in multimodal remote sensing images based on feature decoupling according to claim 1, characterized in that: In step S3, the features learned in step S2 are decomposed into two parts. ,in Represents modal common features, Represents the unique features of a mode, among which Feature dimensions for common features; construct a joint feature set ,in Let be the feature representation of the training batch samples of the i-th modality. For the feature representation of the training batch samples of the j-th modality, SimCLR is used as the contrastive loss to maximize the consistency of the same sample among different modalities in the latent space, as shown in the following formula: ; in To calculate the feature decoupling loss based on the cosine similarity between samples, the specific formula is as follows: ; in , It is the balance coefficient. It is the common feature alignment loss. It is a unique feature orthogonal loss; based on the similarity matrix The common feature alignment loss is calculated using the following formula: ; Where I represents the element corresponding to the diagonal element; based on the Euclidean distance between unique features. The formula for calculating the unique feature orthogonal loss is as follows: ; Finally, the feature representation learning process is optimized using the total loss function, as shown in the following formula: ; in For the overall loss, the entire network is trained using the Adam optimizer with a learning rate of 0.0001 and a batch size of 128 for 200 iterations, with dynamic prototype updates performed during each iteration.

3. The method for identifying multiple categories of land features in multimodal remote sensing images based on feature decoupling according to claim 1, characterized in that: In step S4, the decoupled multimodal features after representation learning are concatenated to preserve the unique features of each modality as much as possible, as shown in the following formula: ; Where d represents the overall number of feature dimensions, the K-means algorithm is used to perform cluster analysis on randomly selected samples, and the formula is as follows: ; in, This refers to the K-means clustering process, where Q is the set of feature representations for all samples after representation learning, and P is the set of cluster labels for the entire sample. To randomly select 10% of sample features based on the k-means clustering result P during the current iteration. The calculated categorical prototype features are shown in the following formula: ; in for The number of samples with cluster label c. For the corresponding set of sample features, The cluster centers for category c are determined, and then the momentum of the category prototypes is updated using the following formula: ; in, For the category prototype updated during the t-th iteration, For the category prototype of iteration t-1 Compared to the current category prototype The weighted sum, The weight coefficients range from 0 to 1, and the category prototype updated in the last iteration is used as the final category prototype. C, Then calculate the features and category prototypes of the ground cover samples. C The similarity is dynamically adjusted to change the boundary of land cover category discrimination, and the classification results of multi-category land cover recognition are output. The formula is as follows: ; in Let K be the prototype vector of the j-th category, K be the total number of categories, and finally generate the land cover category recognition result map.

4. A method for identifying multiple categories of ground features in multimodal remote sensing images based on feature decoupling, characterized in that: The system employs the method described in any one of claims 1 to 3.