Remote sensing image domain adaptive classification method based on high-order prototype guidance
By constructing a high-order prototype-guided remote sensing image domain adaptation classification method, the performance degradation caused by domain differences in cross-domain applications of remote sensing images is solved, and high-precision cross-domain remote sensing image classification is realized, which is suitable for dynamic land and resources surveys, emergency disaster assessments and cross-border ecological monitoring.
Patent Information
- Application Number
- CN202510826410.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing remote sensing image classification methods have performance degradation due to domain differences in cross-domain applications, making it difficult to effectively capture the high-order semantic structural information of remote sensing images, and lack of causal prompt utilization of pre-training knowledge, resulting in a decrease in classification accuracy.
A remote sensing image domain adaptation classification method based on advanced prototype guidance is constructed. By constructing a cross-scene remote sensing image classification data set, high-order feature fusion and causal prompt learning are used to generate category prototypes and perform gradient descent optimization to achieve effective migration of the source domain model in the target domain.
It significantly improves the robustness and classification accuracy of the model under cross-scene conditions, can realize Lupin classification in unknown scenarios in the target domain, and reduces the implementation cost of remote sensing image classification.
Smart Images

Figure CN120356014A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing image domain adaptation classification method based on high-order prototype guidance. Background Art
[0002] The intelligent classification technology of remote sensing images has important application values in the fields of land cover monitoring, disaster assessment, resource investigation, etc. Deep learning methods have greatly improved the classification accuracy by extracting the deep features of images. However, in practical engineering applications, affected by different geographical environments, imaging sensors, lighting conditions, and spatio-temporal changes, there are significant domain differences between the source domain training scenario and the target domain application scenario. This difference leads to serious performance degradation when the model trained in the source domain is directly transferred to the target domain.
[0003] The current mainstream domain adaptation methods usually adopt adversarial learning or feature distribution alignment strategies, aiming to eliminate the domain differences to obtain domain-invariant feature representations. However, these methods often focus on the global alignment of shallow features and are difficult to effectively capture the high-order semantic structure information of categories in remote sensing images (such as discriminative patterns within classes and structured boundaries between classes). At the same time, the transfer methods based on pre-trained vision-language large models (such as CLIP) often adopt direct fine-tuning strategies and fail to fully decouple and utilize the cross-domain causal semantic relationships implicit in the multi-modal pre-trained knowledge, resulting in limited generalization ability for complex remote sensing scenarios.
[0004] When the existing technologies deal with cross-domain classification tasks, they have insufficient modeling of the deep semantic structure and lack the use of causal cues of pre-trained knowledge, which easily causes feature confusion and class misclassification in the target domain. For this reason, the present invention proposes a remote sensing image domain adaptation classification method based on high-order prototype guidance, which significantly improves the robustness of the model under cross-scene conditions by constructing a domain adaptation learning framework that collaboratively strengthens features and semantics. Summary of the Invention
[0005] The purpose of the present invention is to overcome the limitations of the existing methods, improve the performance of cross-scene remote sensing image classification results and the adaptability considering the style differences of remote sensing images in different scenes, and provide a remote sensing image domain adaptation classification method based on high-order prototype guidance.
[0006] To achieve the above purpose, the technical solution of the present invention is as follows:
[0007] A remote sensing image domain adaptation classification method based on high-order prototype guidance, comprising the following steps:
[0008] 11) Construct a cross-scene remote sensing image classification dataset: Obtain source domain remote sensing images and source domain remote sensing image segmentation labels , where Represent the source domain; obtain remote sensing images of the target domain , where represents the target domain;
[0009] 12) Construct a high-order class prototype generation network based on feature interaction: Construct a feature extractor for high-order feature fusion with the input of remote sensing images; generate class prototypes through the feature extractor of high-order feature fusion;
[0010] 13) Construct a high-order causal prompt generation network based on prompt learning: Learn the high-order causal prompts of the source domain and the target domain regarding classes through gradient descent by prompt learning based on decoupling; use a pre-trained multi-modal vision-language large model for prompt learning based on decoupling;
[0011] 14) Train and test the remote sensing image domain adaptation classification method based on high-order prototype guidance: Optimize the parameters of the high-order class prototype generation network based on feature interaction and update the high-order causal prompts through gradient descent, and test on the remote sensing images of the target domain to obtain the final classification result.
[0012] The construction of the cross-scene remote sensing image classification dataset includes the following steps:
[0013] 21) Obtain remote sensing images of the source domain , classification labels of remote sensing images of the source domain and remote sensing images of the target domain , where s represents the source domain and t represents the target domain;
[0014] 22) Perform dynamic domain homogenization exchange on the remote sensing images of the source domain and the remote sensing images of the target domain , and perform the following operations: Adopt adaptive histogram matching, and correct the radiation characteristics of the source domain images channel by channel based on the radiation distribution of the target domain images;
[0015] 23) Extract the normalized difference vegetation index NDVI, normalized difference water index NDWI, and normalized difference built-up index NDBI of the remote sensing images of the source domain and the remote sensing images of the target domain to construct a multi-dimensional spectral feature vector ; Calculate the spectral mean vector of each ground object category based on the classification labels of the remote sensing images of the source domain , where C represents the number of categories, and force the remote sensing images of the target domain to optimize the spectral distribution through the consistency loss under unsupervised conditions, and its specific calculation process is shown by the following formula:
[0016] ;
[0017] where Represents the calculation of the second norm;
[0018] 24) Use the pre-trained ViT model to extract the global context embedding of the image, and minimize the maximum mean discrepancy between the source domain and target domain embeddings.
[0019] The construction of the high-order category prototype generation network based on feature interaction includes the following steps:
[0020] 31) The feature extractor for high-order feature fusion includes a high-order interaction feature extraction module and generates category prototypes through the feature extractor of high-order feature fusion;
[0021] 32) The high-order interaction feature extraction module is shown by the following steps;
[0022] 321) Set the order order and input feature channel number dim of the module;
[0023] 322) Calculate the channel allocation list according to the order: Define the channel allocation list dims, which contains order elements. The i-th element is the integer obtained by taking the integer part of the input feature channel number dim divided by 2 to the power of (order - 1 - i), where i is an integer between 0 and (order - 1);
[0024] 323) Construct the constituent layers of the module:
[0025] 3231) Input feature transformation layer: Use a 1×1 two-dimensional convolutional layer with an input channel number of dim and an output channel number of 2×dim;
[0026] 3232) Depthwise convolutional layer: Use a depthwise separable convolutional layer with the sum of all elements in the dims list as its input channel number;
[0027] 3233) Construct (order - 1) pointwise convolutional projection layers: Each pointwise convolutional projection layer is a 1×1 two-dimensional convolutional layer. Define j as an integer between 0 and (order - 2), then the input channel number of the j-th projection layer is dims[j], and the output channel number is dims[j + 1];
[0028] 3234) Output feature transformation layer: Use a 1×1 two-dimensional convolutional layer with an input channel number of dim and an output channel number of dim;
[0029] 324) The operation process of the module is shown by the following steps;
[0030] 3241) The input feature map passes through the input feature transformation layer to obtain the initial transformed feature;
[0031] 3242) Split the initial transformation features along the channel dimension into two parts: the first part is the base features with the number of channels being dims[0]; the second part is the interaction features with the number of channels being the sum of all elements in the dims list;
[0032] 3243) Obtain the depth feature map by passing the interaction features through a depth convolution layer;
[0033] 3244) Split the depth feature map along the channel dimension into order number of feature maps, denoted as the split feature map list, and the number of channels of each feature map corresponds to each element in the dims list respectively, where the number of channels of the k-th split feature map is dims[k];
[0034] 3245) Initialize the interaction process: Multiply the base features element-wise with the 0-th split feature map in the split feature map list to obtain the current interaction features;
[0035] 3246) Perform order number of order iterations: For the m-th iteration, first transform the channels of the current interaction features through the (m - 1)-th pointwise convolution projection layer, and then multiply the transformed features element-wise with the m-th split feature map in the split feature map list to obtain the updated interaction features;
[0036] 3247) Obtain the module output by passing the finally updated interaction features through the output feature transformation layer;
[0037] 33) Generate class prototypes through a feature extractor with high-order feature fusion as shown in the following steps;
[0038] 331) Source domain remote sensing image Setting of the generation and storage framework for class prototypes: Construct a coupled calculation system for the depth features output by the high-order feature interaction module and the sample labels, and set the prototype repository to contain the prototype vectors of each semantic class and their covariance statistics. The framework performs the following operations in sequence:
[0039] 3311) Input the source domain remote sensing image to the high-order feature interaction module to extract d-dimensional depth features f;
[0040] 3312) Align the depth features f with the classification labels of the source domain remote sensing image and calculate the prototype vectors of each semantic class;
[0041] 3313) Generate the prototype covariance matrix for each class and compress it into a positive semi-definite tensor storage structure;
[0042] 332) Calculation and verification mechanism for class prototypes:
[0043] 3321) Set the calculation rule for class prototypes as:
[0044] For all deep feature vectors belonging to the k-th class Execute the following formula;
[0045] ;
[0046] where represents the high-order feature interaction function, is the number of samples in the k-th class;
[0047] 333) Compressed storage structure of covariance statistics:
[0048] 3331) Set the covariance matrix calculation formula as:
[0049] ;
[0050] 3332) Use Cholesky decomposition to compress and store the covariance matrix as an upper triangular matrix :
[0051] .
[0052] The construction of the high-order causal prompt generation network based on prompt learning includes the following steps:
[0053] 41) Load the pre-trained multi-modal vision-language large model CLIP as the backbone network; freeze the text encoder parameters of CLIP and denote them as , freeze the image encoder parameters of CLIP and denote them as ;
[0054] 42) Define the high-order causal prompt , whose dimension is , where is the word vector dimension, is the prompt length; define the source domain non-causal prompt and the target domain non-causal prompt , whose dimensions are both , where is the word vector dimension, is the prompt length;
[0055] 43) Define the set of category text templates ;
[0056] 44) Target domain pseudo-label generation includes the following steps:
[0057] 441) Input the unlabeled target domain remote sensing image to the image encoder , and obtain the feature ;
[0058] 442) Generate category text embeddings: For the category , the text embedding is , where uses the description of the corresponding category of the text template;
[0059] 443) Calculate zero-shot similarity ;
[0060] 444) Filter the pseudo-labels of the target domain remote sensing image classification through the threshold : :
[0061] ;
[0062] 45) Construct the causal and non-causal joint text embeddings for the source domain remote sensing image , and its specific steps are shown by the following formula: :
[0063] ;
[0064] where represents the vector concatenation operation, represents encoding the text into tokens through the token embedding layer;
[0065] 46) Construct the causal and non-causal joint text embeddings for the target domain remote sensing image , and its specific steps are shown by the following formula: :
[0066] ;
[0067] where represents the vector concatenation operation, represents encoding the text into tokens through the token embedding layer;
[0068] 47) Perform cross-domain forced matching through the causal consistency loss constraint, and its specific steps are shown by the following formula:
[0069]
[0070] where represents the natural exponential function; and represent the source domain high-order prototype and the target domain high-order prototype generated by the high-order category prototype generation network based on feature interaction described in 12), that is, the source domain high-order prototype and the target domain high-order prototype ; The temperature coefficient is set to 0.2;
[0071] 48) Perform in-domain matching through the non-causal decoupling loss constraint, and its specific steps are shown by the following formula:
[0072]
[0073] where represents the natural exponential function; and represent the source-domain high-order prototype and the target-domain high-order prototype generated by the high-order category prototype generation network based on feature interaction described in 12), that is, the source-domain high-order prototype and the target-domain high-order prototype ; The temperature coefficient is set to 0.2.
[0074] The method for training and testing the remote sensing image domain adaptation classification based on high-order prototype guidance includes the following steps:
[0075] 51) Optimize the parameters of the high-order category prototype generation network based on feature interaction through gradient descent and update the high-order causal hint, and test on the target-domain remote sensing image to obtain the final classification result; it includes the following steps:
[0076] 52) Initialize the high-order causal hint vector , the source-domain non-causal hint and the target-domain non-causal hint , and randomly initialize them using a Gaussian distribution;
[0077] 53) Calculate the total loss function , and the calculation formula is as follows:
[0078]
[0079] where is the hyperparameter for constraint importance set to 0.5, is the causal consistency loss, is the non-causal decoupling loss;
[0080] 54) Perform gradient backpropagation, and the gradient of the high-order causal hint parameter is calculated by ; the gradient of the source-domain non-causal hint is calculated by ; the gradient of the target-domain non-causal hint is calculated by ;
[0081] 55) Update the parameters, which are calculated by the following formula:
[0082]
[0083] wherein is the learning rate;
[0084] 56) Input the remote sensing image of the target domain to obtain the final classification result.
[0085] Advantageous effects: The remote sensing image domain adaptation classification method based on high-order prototype guidance proposed by the present invention significantly breaks through the limitations of the prior art by coordinating high-order semantic modeling and cross-domain causal prompting mechanisms. First, an innovative high-order prototype alignment architecture is constructed to synchronously model the intra-class consistency and inter-class discriminative boundaries of the source domain and the target domain in the feature space. This method overcomes the defect that traditional domain adaptation methods only focus on shallow feature alignment, and accurately captures the high-order semantic structures of ground object categories in remote sensing images (such as the regularity of farmland textures and the geometric features of building outlines), enabling the model to maintain the discriminative basis for pixel-level classification in cross-domain scenarios.
[0086] Secondly, a causal-guided pre-training prompt decoupling mechanism is designed to deeply explore the cross-domain generalization knowledge hidden in the vision-language large model. By decoupling the causal semantic factors in the multi-modal pre-training features and generating domain-invariant causal prompt templates, the interference of pseudo-correlated features caused by sensor differences or seasonal changes is effectively suppressed. This technology significantly improves the zero-shot adaptation ability of the model in unknown target domain scenarios, and can still achieve robust classification even when facing unlabeled disaster-damaged areas or new land use types.
[0087] The technical solution of the present invention has strong engineering practicability, can achieve end-to-end cross-domain migration without target domain labeled data, and greatly reduces the implementation cost of remote sensing image classification. It can be widely applied to scenarios such as dynamic general surveys of land resources, emergency disaster assessment, and cross-border ecological monitoring, effectively solving the problem of model failure caused by the mismatch between training data and application scenario domains, and providing reliable technical support for global-scale remote sensing intelligent analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 is a remote sensing image domain adaptation classification method based on high-order prototype guidance;
[0089] Figure 2 is the framework diagram of the high-order interaction feature extraction module involved in the present invention;
[0090] Figure 3 is the framework diagram of the high-order category prototype generation network based on feature interaction involved in the present invention;
[0091] Figure 4 is the framework diagram of the high-order causal prompt generation network based on prompt learning involved in the present invention; Detailed implementation manner
[0092] To have a further understanding and recognition of the structural features and achieved effects of the present invention, the following is a detailed description with preferred embodiments and accompanying drawings:
[0093] As Figure 1 shown, a remote sensing image domain adaptation classification method based on high-order prototype guidance according to the present invention is characterized by comprising the following steps:
[0094] The first step is to construct a cross-scene remote sensing image classification dataset:
[0095] The specific steps are as follows:
[0096] (1) Obtain source domain remote sensing images , source domain remote sensing image classification labels and target domain remote sensing images , where s represents the source domain and t represents the target domain;
[0097] (2) Perform dynamic domain homogenization exchange on the source domain remote sensing images and target domain remote sensing images , and perform the following operations: adopt adaptive histogram matching, and use the radiation distribution of the target domain images as a benchmark to correct the radiation characteristics of the source domain images channel by channel;
[0098] (3) Extract the normalized difference vegetation index NDVI, normalized difference water index NDWI and normalized difference built-up index NDBI of the source domain remote sensing images and target domain remote sensing images to construct a multi-dimensional spectral feature vector ; calculate the spectral mean vector of each ground object category based on the source domain remote sensing image classification labels , where C represents the number of categories, and force the target domain remote sensing images to optimize the spectral distribution through the consistency loss under unsupervised conditions, and its specific calculation process is shown by the following formula:
[0099] ;
[0100] where represents calculating the two-norm;
[0101] (4) Use the pre-trained ViT model to extract the global context embedding of the images, and minimize the maximum mean discrepancy between the source domain and target domain embeddings.
[0102] The second step is to construct a high-order class prototype generation network based on feature interaction:
[0103] The specific steps are as follows:
[0104] (1) The feature extractor for high-order feature fusion includes a high-order interaction feature extraction module and generates class prototypes through the feature extractor of high-order feature fusion;
[0105] (2) The high-order interaction feature extraction module is shown in the following steps;
[0106] (2-1) Set the order order and the input feature channel number dim of the module;
[0107] (2-2) Calculate the channel allocation list according to the order: Define the channel allocation list dims, which contains order elements. The i-th element is the integer obtained by taking the integer part of the input feature channel number dim divided by 2 to the power of (order-1-i), where i is an integer between 0 and (order-1);
[0108] (2-3) Construct the constituent layers of the module:
[0109] (2-3-1) Input feature transformation layer: Use a 1×1 two-dimensional convolutional layer with an input channel number of dim and an output channel number of 2×dim;
[0110] (2-3-2) Depthwise convolutional layer: Use a depthwise separable convolutional layer with the sum of all elements in the dims list as its input channel number;
[0111] (2-3-3) Construct (order-1) pointwise convolutional projection layers: Each pointwise convolutional projection layer is a 1×1 two-dimensional convolutional layer. Define j as an integer between 0 and (order-2), then the input channel number of the j-th projection layer is dims[j], and the output channel number is dims[j+1];
[0112] (2-3-4) Output feature transformation layer: Use a 1×1 two-dimensional convolutional layer with an input channel number of dim and an output channel number of dim;
[0113] (2-4) The operation process of the module is shown in the following steps;
[0114] (2-4-1) The input feature map passes through the input feature transformation layer to obtain the initial transformed feature;
[0115] (2-4-2) Split the initial transformed feature along the channel dimension into two parts: The first part is the base feature with a channel number of dims[0]; The second part is the interaction feature with the sum of all elements in the dims list as its channel number;
[0116] (2-4-3) Pass the interaction feature through the depthwise convolutional layer to obtain the depth feature map;
[0117] (2-4-4) Split the depth feature map along the channel dimension into several feature maps, denoted as the list of split feature maps. The number of channels of each feature map corresponds to each element in the dims list respectively, where the number of channels of the k-th split feature map is dims[k];
[0118] (2-4-5)Initialize the interaction process: Multiply the base feature element-wise with the 0-th split feature map in the list of split feature maps to obtain the current interaction feature;
[0119] (2-4-6)Perform order iterations: For the m-th iteration, first transform the channels of the current interaction feature through the (m - 1)-th pointwise convolutional projection layer, and then multiply the transformed feature element-wise with the m-th split feature map in the list of split feature maps to obtain the updated interaction feature;
[0120] (2-4-7)Obtain the module output by passing the finally updated interaction feature through the output feature transformation layer;
[0121] (3)Generate class prototypes through the feature extractor with high-order feature fusion as shown in the following steps;
[0122] (3-1)Source domain remote sensing image Setting of the generation and storage framework for class prototypes: Construct a coupled calculation system for the depth features output by the high-order feature interaction module and the sample labels. Set the prototype repository to contain the prototype vectors of each semantic class and their covariance statistics. The framework performs the following operations in sequence:
[0123] (3-1-1)Input the source domain remote sensing image to the high-order feature interaction module to extract the d-dimensional depth feature f;
[0124] (3-1-2)Align the depth feature f with the classification label of the source domain remote sensing image and calculate the prototype vector of each semantic class;
[0125] (3-1-3)Generate the prototype covariance matrix of each class and compress it into a positive semi-definite tensor storage structure;
[0126] (3-2)Calculation and verification mechanism for class prototypes:
[0127] (3-2-1)Set the calculation rule for class prototypes as:
[0128] For all depth feature vectors belonging to the k-th class execute the following formula;
[0129] ;
[0130] where Represents a high-order feature interaction function, is the number of samples in the k-th class;
[0131] (3-3)Compressed storage structure of covariance statistics:
[0132] (3-3-1)Set the covariance matrix calculation formula as:
[0133] ;
[0134] (3-3-2)Use Cholesky decomposition to compress and store the covariance matrix as an upper triangular matrix :
[0135] .
[0136] Third, construct a high-order causal prompt generation network based on prompt learning:
[0137] The specific steps are as follows:
[0138] (1)Load the pre-trained multi-modal vision-language large model CLIP as the backbone network; freeze the text encoder parameters of CLIP and denote them as , freeze the image encoder parameters of CLIP and denote them as ;
[0139] (2)Define the high-order causal prompt , its dimension is , where is the word vector dimension, is the prompt length; define the source domain non-causal prompt and the target domain non-causal prompt , both of their dimensions are , where is the word vector dimension, is the prompt length;
[0140] (3)Define the set of category text templates ;
[0141] (4)Target domain pseudo-label generation includes the following steps:
[0142] (4-1)Input the unlabeled target domain remote sensing image to the image encoder , and obtain the feature ;
[0143] (4-2)Generate category text embeddings: For the category , the text embedding is , where uses the description of the corresponding category of the text template;
[0144] (4 - 3) Calculate zero - shot similarity ;
[0145] (4 - 4)Filter the pseudo - labels of the target - domain remote - sensing image classification through the threshold : :
[0146] ;
[0147] (5)Construct the causal and non - causal joint text embedding for the source - domain remote - sensing image , and its specific steps are shown by the following formula:
[0148] ;
[0149] where represents the vector concatenation operation, represents encoding the text into tokens through the token embedding layer;
[0150] (6)Construct the causal and non - causal joint text embedding for the target - domain remote - sensing image , and its specific steps are shown by the following formula:
[0151] ;
[0152] where represents the vector concatenation operation, represents encoding the text into tokens through the token embedding layer;
[0153] (7)Perform cross - domain forced matching through the causal consistency loss constraint, and its specific steps are shown by the following formula:
[0154]
[0155] where represents the natural exponential function; and represent the source - domain high - order prototype and the target - domain high - order prototype generated by the high - order category prototype generation network based on feature interaction described in (12), that is, the source - domain high - order prototype and the target - domain high - order prototype ; The temperature coefficient is set to 0.2;
[0156] (8)Perform in - domain matching through the non - causal decoupling loss constraint, and its specific steps are shown by the following formula:
[0157]
[0158] wherein represents the natural exponential function; and represent the source domain high-order prototype and the target domain high-order prototype generated by the feature interaction-based high-order category prototype generation network described in (12), i.e., the source domain high-order prototype and the target domain high-order prototype ; The temperature coefficient is set to 0.2.
[0159] Step 4: Train and test the remote sensing image domain adaptation classification method guided by high-order prototypes:
[0160] The specific steps are as follows:
[0161] (1) Optimize the parameters of the feature interaction-based high-order category prototype generation network through gradient descent and update the high-order causal cues, and test on the target domain remote sensing image to obtain the final classification result; including the following steps:
[0162] (2) Initialize the high-order causal cue vector , the source domain non-causal cue and the target domain non-causal cue , and randomly initialize them using a Gaussian distribution;
[0163] (3) Calculate the total loss function , and the calculation formula is as follows:
[0164]
[0165] wherein is a hyperparameter for constraining the importance, set to 0.5, is the causal consistency loss, is the non-causal decoupling loss;
[0166] (4) Perform gradient backpropagation, and the gradient of the high-order causal cue parameter is calculated by ; the gradient of the source domain non-causal cue is calculated by ; the gradient of the target domain non-causal cue is calculated by ;
[0167] (5) Update the parameters, calculated by the following formula:
[0168]
[0169] wherein is the learning rate;
[0170] (6) Input target domain remote sensing image Get the final classification result.
[0171] The remote sensing image domain adaptation classification method based on high-order prototype guidance proposed in the present invention significantly enhances the generalization performance of the model in cross-domain scenarios and solves the problem of model migration failure caused by geographical environment, sensor differences and temporal and spatial changes. Compared with the prior art, the present invention breaks through the dual limitations of traditional methods in deep semantic modeling and pre-training knowledge transfer through the synergistic integration of high-order prototype alignment and causal prompt decoupling mechanism, and achieves a simultaneous leap in cross-domain classification accuracy and robustness. The method makes full use of the discriminative semantic structure and cross-domain causal generalization knowledge of remote sensing images, and can maintain stable classification capabilities even in extreme scenarios where target domain annotations are missing or unknown object categories exist. Through the innovative dual-network closed-loop optimization design, the present invention resolves the problem of feature confusion and category misclassification caused by domain differences, and improves the cross-domain adaptability of hyperspectral and multi-temporal remote sensing images to a new level. This solution provides high-reliability technical support for global land surveys, disaster emergency response and long-term ecological monitoring, and shows significant advantages and broad prospects in the large-scale application of remote sensing intelligent analysis.
[0172] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.
Claims
1. A remote sensing image domain adaptation classification method based on high-order prototype guidance, characterized in that It includes the following steps: 11) Construct a cross-scenario remote sensing image classification dataset: Obtain source domain remote sensing images and source domain remote sensing image segmentation labels , where represents the source domain; Obtain target domain remote sensing images , where represents the target domain; 12) Construct a high-order category prototype generation network based on feature interaction: Construct a feature extraction module for high-order feature fusion with the input being a remote sensing image; Generate category prototypes through the feature extraction module for high-order feature fusion; 13) Construct a high-order causal prompt generation network based on prompt learning: Learn the high-order causal prompts regarding categories in the source domain and the target domain through gradient descent by means of decoupled prompt learning; Decoupled prompt learning uses the pre-trained multi-modal vision-language large model CLIP; 14) Train and test the remote sensing image domain adaptation classification method based on high-order prototype guidance: optimize the parameters of the high-order class prototype generation network based on feature interaction through gradient descent and update the high-order causal cues, and test on the remote sensing images of the target domain to obtain the final classification result.
2. The remote sensing image domain adaptation classification method based on high-order prototype guidance according to claim 1, wherein, The construction of the cross-scene remote sensing image classification dataset includes the following steps: 21) Obtain the source domain remote sensing image , the classification label of the source domain remote sensing image and the target domain remote sensing image , where s represents the source domain and t represents the target domain; 22) For the source domain remote sensing image and the target domain remote sensing image perform dynamic domain homogenization exchange, and perform the following operations: adopt adaptive histogram matching, and use the radiation distribution of the target domain image as a benchmark to correct the radiation characteristics of the source domain image channel by channel; 23) Extract the source domain remote sensing image and the target domain remote sensing image to construct a multi-dimensional spectral feature vector by calculating the Normalized Difference Vegetation Index (NDVI), Normalized Difference Water Index (NDWI), and Normalized Difference Built-up Index (NDBI). Based on the classification labels of the source domain remote sensing image calculate the spectral mean vector of each ground object category , where C represents the number of categories, and force the target domain remote sensing image to optimize the spectral distribution through the consistency loss under unsupervised conditions . The specific calculation process is shown by the following formula: , wherein represents the square of the calculation of the two-norm; 24) Use the pre-trained vision transformer model to extract the global context embedding of the image, and minimize the maximum mean discrepancy between the source domain and target domain embeddings.
3. A remote sensing image domain adaptation classification method based on high-order prototype guidance according to claim 1, characterized in that The construction of the high-order category prototype generation network based on feature interaction includes the following steps: 31) The high-order category prototype generation network includes a high-order interaction feature extraction module and a high-order feature category prototype generation module; 32) Set the high-order interaction feature extraction module, and its specific steps are as follows; 321) Set the order order and the input feature channel number dim of the module; 322) Define the channel allocation list dims as follows according to the set order order and input feature channel number dim: , where i represents an integer index from 0 to (order - 1), represents the floor operation; 323) The constituent layers of the high-order interaction feature extraction module include an input feature transformation layer, a depth convolution layer, (order - 1) projection layers, and an output feature transformation layer, and their respective specific parameters are as follows: 3231) Input feature transformation layer: Adopt a 2D convolutional layer with a size of 1×1, the input channel number is dim, and the output channel number is 2×dim; 3232) Depth convolution layer: Adopt a depthwise separable convolution layer, and its input channel number is the sum of all elements in the dims list; Define a set of projection layers , where j represents an integer index from 0 to (order - 2); the j-th projection layer has an input channel defined in 322) and an output channel defined in 322) ; 3234) Output feature transformation layer: Adopt a 2D convolutional layer with a size of 1×1, the input channel number is dim, and the output channel number is dim; 324) Set the operation process of the module, and its specific steps are as follows; 3241) The input feature map passes through the input feature transformation layer to obtain the initial transformed feature; 3242) Split the initial transformed feature along the channel dimension into two parts: The first part is the base feature, and its channel number is dims[0]; The second part is the interaction feature, and its channel number is the sum of all elements in the dims list; 3243) Pass the interaction feature through the depth convolution layer to obtain the depth feature map; 3244) Split the depth feature map along the channel dimension into order number of feature maps, denoted as the split feature map list, and the channel numbers of each feature map respectively correspond to the elements in the dims list, where the channel number of the k-th split feature map is dims[k]; 3245) Initialize the interaction process: Multiply the base feature element-wise with the 0-th split feature map in the split feature map list to obtain the current interaction feature; Perform order "order" iterations: For the m-th iteration, first transform the channels of the current interaction feature through the (m - 1)-th pointwise convolution projection layer, and then element-wise multiply the transformed feature with the m-th segmentation feature map in the list of segmentation feature maps to obtain the updated interaction feature; 3247) Obtain the module output by passing the finally updated interaction feature through the output feature transformation layer; 33) Set up a high-order feature class prototype generation module, and its specific steps are as follows; 331) Set the source domain remote sensing image Category prototype generation and storage framework: Construct a coupling calculation system for the deep features output by the high-order feature interaction module and the sample labels. Set the prototype repository to contain the prototype vectors of each semantic category and their covariance statistics. The framework performs the following operations in sequence: 3311) Input source domain remote sensing image The high-order feature interaction module extracts d-dimensional depth features f; Align the depth feature f with the classification labels of the source domain remote sensing images and calculate the prototype vectors for each semantic category; 3313) Generate the prototype covariance matrix for each class and compress it into a positive semi-definite tensor storage structure; 332) Calculation and verification mechanism for class prototypes: 3321) Set the category prototype The calculation rule is as follows: For all deep feature vectors belonging to the k-th class perform an aggregation operation to calculate the class prototype , and the calculation formula is as follows: , Among them represents the high-order feature interaction function is the number of samples of the k-th class; 333) Set the covariance matrix calculation formula as: 。 4. A remote sensing image domain adaptation classification method based on high-order prototype guidance according to claim 1, characterized in that, The construction of the high-order causal cue generation network based on prompt learning includes the following steps: 41) Load the pre-trained multi-modal vision-language large model CLIP as the backbone network; freeze the text encoder parameters of the multi-modal vision-language large model CLIP and denote it as , freeze the image encoder parameters of the multi-modal vision-language large model CLIP and denote it as ; 42) Define high-order causal prompts , whose dimension is , where is the word vector dimension, is the prompt length; Define source domain non-causal prompts and target domain non-causal prompts , both of whose dimensions are , where is the word vector dimension, is the prompt length; 43) Define a set of category text templates ; 44) The generation of pseudo-labels for the target domain includes the following steps: 441) Input the target domain remote sensing image without labels To the image encoder to obtain the image embedding ; 442) Generate category text embeddings: For category , the text embedding is , where uses the description of the corresponding category of the text template; 443) For text embeddings and image embeddings calculate zero-shot similarity , where represents the vector norm operation; 444) Through a threshold Screening pseudo-labels for remote sensing image classification in the target domain : ; 45) For the source domain remote sensing image Construct causal and non-causal joint text embeddings , and its specific steps are shown by the following formula: , Among them represents a vector concatenation operation represents encoding the text into text tokens through a text encoding layer represents taking the k-th category in the text template 46) For the remote sensing image of the target domain Construct causal and non-causal joint text embeddings , and its specific steps are shown by the following formula: , Among them represents a vector concatenation operation, represents encoding the text into text tokens through a text encoding layer, represents taking the k-th category in the text template; 47) Through the causal consistency loss constraint for cross-domain forced matching, and its specific steps are shown by the following formula: , where represents the natural exponential function; and represent the source-domain high-order prototype and the target-domain high-order prototype generated by the feature-interaction-based high-order class prototype generation network described in (12), i.e., the source-domain high-order prototype and the target-domain high-order prototype ; The temperature coefficient is set to 0.2; 48) Through non-causal decoupling loss constraints are used to perform in-domain matching, and its specific steps are shown by the following formula: , where represents the natural exponential function; and represent the source domain high-order prototype and the target domain high-order prototype generated by the feature interaction-based high-order category prototype generation network described in (12), that is, the source domain high-order prototype and the target domain high-order prototype ; The temperature coefficient is set to 0.
2.
5. A remote sensing image domain adaptation classification method based on high-order prototype guidance according to claim 1, characterized in that, The training and testing of the remote sensing image domain adaptation classification method based on high-order prototype guidance includes the following steps: 51) Optimize the parameters of the high-order category prototype generation network based on feature interaction through gradient descent and update the high-order causal cues, and test on the remote sensing images in the target domain to obtain the final classification result; including the following steps: 52) Initialize the high-order causal hint vector , the non-causal hint of the source domain and the non-causal hint of the target domain , and randomly initialize them using a Gaussian distribution; 53) Calculate the total loss function , and the calculation formula is as follows: , Among them The hyperparameter for the constraint importance is set to 0.5, is the causal consistency loss, is the non-causal decoupling loss; 54) Gradient backpropagation, high-order causal hint parameter gradient Calculated by Source domain non-causal hint gradient Calculated by Target domain non-causal hint gradient Calculated by Calculated; 55) Parameter update, calculated by the following formula: , wherein is the learning rate; 56) Input the remote sensing image of the target domain Obtain the final classification result.
Citation Information
Patent Citations
Deep learning generalization method for remote sensing image land cover classification
CN113343775A
Cross-domain multi-modal remote sensing image classification method based on comparative learning
CN116912595A
Universal cross-domain image conversion method and system
CN118334458A
Image segmentation method of high-order interaction space U-shaped model based on mixed selectivity
CN119205798A
Remote sensing image unsupervised domain adaptation method based on comparative learning and multi-prototype alignment
CN119251646A
Cited By
Multi-modal remote sensing data cross-domain classification method based on channel-global perception
CN121353758A
Remote sensing image water body information extraction method combining weak supervision and sample migration
CN121837935A
An aircraft target domain adaptation detection method based on scene feature decoupling
CN122510880A