Cross-modal remote sensing image classification domain adaptation method based on bidirectional visual language prompt
Through the cross-modal domain adaptation method of bidirectional visual language prompts, the problems of scarcity and modal differences in labeled data in remote sensing image classification are solved, and efficient classification of remote sensing images and improvement of model generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510612291.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has the problem of labeled data scarcity in remote sensing image classification. Traditional UDA methods are difficult to effectively align the complex category structure and modal differences of remote sensing data, resulting in insufficient generalization capabilities of the model.
A cross-modal domain adaptation method based on bidirectional visual language prompt is adopted. Through the bidirectional prompt network BPGM and gating mechanism strategy, combined with the weight adjustment of visual and text features, an efficient cross-domain alignment loss function is designed to achieve accurate alignment and synergy between visual and text features.
It improves the accuracy of remote sensing image classification and the generalization ability of the model, can capture the semantic information of the target domain more accurately, reduces the distribution differences between different modalities and data sources, and enhances the cross-modal feature alignment effect.
Smart Images

Figure CN120339722A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image classification for cross-modal domain adaptation. Background Art
[0002] With the development of spaceborne technology, remote sensing satellites can acquire high-resolution, multi-modal data, such as optical images, SAR data, infrared, hyperspectral, and multi-spectral images, etc. These data are of great value in the fields of environmental monitoring, land use analysis, disaster assessment, etc. However, the scarcity of data annotation remains the main challenge restricting the application of deep learning models in the remote sensing field. Annotating remote sensing data is not only costly and time-consuming but also requires professional knowledge support. Due to the high resolution of remote sensing images and the complex types of ground objects, the annotation process often relies on experienced experts. Especially for SAR data, which is greatly affected by noise and strong reflections, it is difficult for non-professionals to accurately interpret. This leads to difficulties in obtaining annotated data, thereby affecting the training and generalization ability of data-driven deep learning models.
[0003] To address the problem of scarce annotation, unsupervised domain adaptation (UDA) technology has become one of the key methods. UDA establishes a connection between the source domain (annotated data) and the target domain (unannotated data) through feature transformation or alignment, enabling the model to transfer the knowledge of the source domain to improve the performance of the target domain. Traditional UDA methods mainly use means such as moment matching or adversarial learning to minimize the distribution difference between domains. However, this simple alignment method may lead to semantic structure distortion and weaken the category discrimination ability of feature representations. In addition, most methods rely on numerical labels for supervision and fail to fully utilize the semantic information of categories. In the case of complex category structures or significant distribution differences between domains, traditional methods are difficult to effectively align features, affecting the generalization ability and classification accuracy of the model.
[0004] In recent years, Visual-Language Models (VLMs) have provided new ideas for solving semantic alignment problems in UDA tasks with their multimodal learning capabilities. By jointly training visual and language modalities, VLMs learn shared multimodal representations to achieve efficient knowledge transfer. Compared with traditional methods, VLMs can guide the model to focus on the target semantic regions through language prompts, avoid semantic distortions caused by simple feature alignment, enhance the understanding of complex category structures, and improve the generalization ability of cross-domain tasks. Although VLMs have powerful knowledge transfer and cross-modal learning capabilities, they still face challenges in the field of remote sensing. Remote sensing data is highly heterogeneous, and there are modal differences in the data obtained by different sensors. Factors such as the noise of SAR data and the illumination changes of optical images increase the difficulty of feature alignment. In addition, remote sensing data annotation usually involves fine-grained land cover classification, which poses higher requirements for the text description ability of VLMs. Therefore, how to effectively utilize VLMs for unsupervised cross-modal domain adaptation and improve the performance of the model on remote sensing data remains an urgent problem to be solved. Summary of the Invention
[0005] The present invention aims to solve the above-mentioned deficiencies of the prior art and proposes a cross-modal remote sensing image classification domain adaptation method based on bidirectional visual language prompts, with the expectation of improving the accuracy of image classification in cross-modal scenarios.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A cross-modal domain adaptation remote sensing image classification method based on bidirectional visual language prompts of the present invention is characterized in that it is carried out according to the following steps:
[0008] Step 1: Obtain a source domain dataset with annotation information , where represents the th source domain sample, represents the corresponding label, represents the number of source domain samples;
[0009] Obtain a target domain dataset without annotation information , where represents the th target domain sample, represents the number of target domain samples, and both contain categories;
[0010] Step 2: Construct a cross-modal domain adaptation remote sensing image classification network, including: a text encoder with frozen parameters 、A vision encoder with parameter freezing and a bidirectional prompt network BPGM;
[0011] Step 2.1: Input into the vision encoder to obtain the global visual embedding of and the visual context information ;
[0012] Step 2.2: Input the label information of into the text encoder to obtain the global text embedding of and the text context information , where represents the global text embedding of the k-th class label;
[0013] Step 2.3: Use the gating mechanism strategy in the bidirectional prompt network BPGM to guide to obtain the final visual embedding of ;
[0014] Step 2.4: Use the gating mechanism strategy in the bidirectional prompt network BPGM to guide to obtain the final text embedding of ;
[0015] Step 3: Construct the overall loss function , and ; ;
[0016] Step 4: Use the gradient descent method to train the bidirectional prompt network BPGM and calculate the overall loss function to update the network parameters until the overall loss function converges, so as to obtain the optimal image classification model for predicting and classifying images in the target domain.
[0017] The feature of a cross-modal domain adaptation remote sensing image classification method based on bidirectional visual language prompts according to the present invention also lies in that Step 2.3 is carried out according to the following steps:
[0018] Step 2.3.1: Obtain the corresponding text gating weight value of according to Equation (3.1):
[0019] (3.1)
[0020] In formula (2.1), represents the concatenation operation, is the Sigmoid function, is the text gating weight matrix to be learned;
[0021] Step 2.3.2: Obtain the updated text context information according to formula (3.2) :
[0022] (3.2)
[0023] In formula (3.2), represents the set of text gating weights, and ;
[0024] Step 2.3.3: Obtain the fused visual cue according to formulas (3.3) - (3.5) ;
[0025] (3.3)
[0026] (3.4)
[0027] (3.5)
[0028] In formulas (3.3) - (3.5), , represents the total number of layers of the Transformer network, represents the th decoding layer of the Transformer network, and are two projection operations; represents the semantic feature representation, represents the initial visual feature representation, represents the th layer of visual feature representation, represents the th layer of visual feature representation;
[0029] Step 2.3.4: Obtain the final visual embedding of ;
[0030] (3.6)
[0031] In formula (3.6), is the weight factor of the visual cue to be learned.
[0032] Furthermore, step 2.4 is carried out as follows:
[0033] Step 2.4.1: Obtain the visual gating weight :
[0034] (4.1)
[0035] In formula (4.1), is the visual gating weight matrix to be learned;
[0036] Step 2.4.2: Obtain the updated visual context information , thereby obtaining the updated visual context information set ;
[0037] (4.2)
[0038] Step 2.4.3: Obtain the final text embedding after fusion of the k-th category according to formulas (4.3)-(4.5) ;
[0039] (4.3)
[0040] (4.4)
[0041] (4.5)
[0042] In formulas (4.3)-(4.5), represents the source domain sample set, represents the initial text feature representation of the k-th category; represents the visual feature representation, represents the text feature representation of the k-th category in the layer, represents the text feature of the k-th category in the layer, represents the text feature representation of the k-th category in the
[0043] Step 2.4.4: Obtain the final text embedding ;
[0044] (4.6)
[0045] In formula (3.6), is the weight factor of the text prompt to be learned.
[0046] Furthermore, step 3 is carried out as follows:
[0047] Step 3.1 constructs the contrastive loss according to Equation (5.1) :
[0048] (5.1)
[0049] In Equation (5.1), is the temperature parameter, represents the cosine similarity;
[0050] Step 3.2: Calculate the final cross-domain loss from Equation (5.2) :
[0051] (5.2)
[0052] In Equation (5.2), represents the cross-domain amplitude information and is obtained from Equation (5.3), is the cross-domain angle information and is obtained from Equation (5.2),
[0053] (5.3)
[0054] In Equation (5.3), is the function that maps the original variable to the reproducing kernel Hilbert space and is the global visual embedding;
[0055] (5.4)
[0056] In Equation (5.4), represents the cosine similarity function, and , represents the inner product operation, is the norm;
[0057] Step 3.3: Construct the overall loss function of the bidirectional prompt network BPGM using Equation (5.5) :
[0058] (5.5)
[0059] In Equation (5.5), is the balance factor.
[0060] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program that supports the processor to execute the cross-modal remote sensing image classification domain adaptation method, and the processor is configured to execute the program stored in the memory.
[0061] A computer-readable storage medium according to the present invention, characterized in that a computer program stored on the computer-readable storage medium, when run by a processor, executes the steps of the cross-modal remote sensing image classification domain adaptation method.
[0062] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0063] 1. The present invention proposes a cross-modal remote sensing image classification domain adaptation framework based on bidirectional visual language cues, and for the first time applies the image-text multimodal fusion paradigm to the cross-modal domain adaptation task of remote sensing images. This method enhances the synergistic effect between visual and text features through a bidirectional cue mechanism, enabling the model to capture the semantic information of the target domain more accurately. At the same time, a distribution alignment strategy is adopted to reduce the distribution differences between different modalities and data sources, thereby improving the alignment effect of cross-modal features.
[0064] 2. The present invention designs a bidirectional cue strategy based on a gating mechanism, which realizes the bidirectional guidance of visual and text cues by adjusting the weights of visual and text features. While strengthening the cross-modal synergistic effect, this mechanism accurately aligns semantic information, further improving the flexibility and generalization ability of the model.
[0065] 3. The present invention proposes a cross-modal domain distribution alignment method, which designs an efficient metric function by combining amplitude and spatial direction information, breaking through the limitation of traditional methods that only rely on distribution distance for alignment. This method not only pays attention to the overall deviation of feature distributions during the alignment process, but also makes full use of amplitude information to enhance the sensitivity to local features, and combines spatial direction information to ensure the consistency of feature mapping. Compared with traditional distribution matching methods, this strategy can capture the relationship between different modal features more accurately, greatly improving the accuracy of cross-modal visual feature alignment. Description of the Drawings
[0066] Figure 1 is a flowchart of the cross-modal domain adaptation remote sensing image classification method of bidirectional visual language cues and distribution alignment of the present invention;
[0067] Figure 2a is the confusion matrix of the CoOp method on the MRSSC dataset;
[0068] Figure 2b is the confusion matrix of the method of the present invention on the MRSSC dataset;
[0069] Figure 3 It is a visualization comparison chart of Grad-CAM of the present invention with different methods. Specific implementation manners
[0070] In this embodiment, as Figure 1 shown, the cross-modal remote sensing image classification domain adaptation method based on bidirectional visual language cues includes the following steps:
[0071] Step 1: Obtain a source domain dataset with annotation information , where represents the th source domain sample, represents the corresponding label, represents the number of source domain samples;
[0072] Obtain a target domain dataset without annotation information , where represents the th target domain sample, represents the number of target domain samples, and both contain categories.
[0073] Step 2: Construct a cross-modal domain adaptation remote sensing image classification network, including: a text encoder with frozen parameters , a visual encoder with frozen parameters and a bidirectional prompt network BPGM;
[0074] Input into the visual encoder to obtain the global visual embedding of and the visual context information respectively;
[0075] Input 's label information into the text encoder to obtain the global text embedding of and the text context information , where represents the global text embedding of the kth category label; in this embodiment, the size of all images after preprocessing is , and the batch size is 64.
[0076] Step 3: Use the gating mechanism strategy in the bidirectional prompt network BPGM for Perform guidance to obtain the final visual embedding ;
[0077] Step 3.1: Obtain the corresponding text gating weight value :
[0078] (3.1)
[0079] In formula (2.1), denotes the concatenation operation, is the Sigmoid function, is the text gating weight matrix to be learned.
[0080] Step 3.2: Obtain the updated text context information according to formula (3.2) :
[0081] (3.2)
[0082] In formula (3.2), , denotes the set of text gating weights; the obtained values within change between [0, 1].
[0083] Step 3.3: Obtain the fused visual cue according to formulas (3.3) - (3.5) ;
[0084] (3.3)
[0085] (3.4)
[0086] (3.5)
[0087] In formulas (3.3) - (3.5), , denotes the total number of layers of the Transformer network, denotes the th decoding layer of the Transformer network, which is responsible for adjusting the association between visual and text information through the cross-attention mechanism, and are two projection operations; denotes the semantic feature representation after being transformed by the projection operation, represents the projection operation and The corresponding initial visual feature representation, indicating the visual feature representation after passing through the th layer, indicating the visual feature representation after passing through the th layer.
[0088] Step 3.4: Obtain the final visual embedding of according to Equation (3.6); ;
[0089] (3.6)
[0090] In Equation (3.6), is the weight factor of the visual cue to be learned, and its initial value is set to 0.01.
[0091] Step 4: Use the gating mechanism strategy in the bidirectional prompt network BPGM to guide and obtain the final text embedding of; ;
[0092] Step 4.1: Obtain the visual gating weight of according to Equation (4.1): :
[0093] (4.1)
[0094] In Equation (4.1), is the visual gating weight matrix to be learned.
[0095] Step 4.2: Obtain the updated visual context information according to Equation (4.2): :
[0096] (4.2)
[0097] Step 4.3: Obtain the final text embedding after fusion of the k-th category according to Equations (4.3) - (4.5); ;
[0098] (4.3)
[0099] (4.4)
[0100] (4.5)
[0101] In Equations (4.3) - (4.5), represents the source domain sample set; Represents the visual context information after the overall update of the sample, and , represents the initial text feature representation after the projection operation and the corresponding initial text feature representation; represents the visual feature representation after being transformed by the projection operation, represents the text feature representation of the k-th category in the l-th layer, represents the text feature of the k-th category after passing through the l-th layer, represents the text feature representation of the k-th category in the l-th layer.
[0102] Step 4.4: Obtain the final text embedding of according to Equation (4.6); ;
[0103] (4.6)
[0104] In Equation (4.6), is the weight factor of the text prompt to be learned, and its initial value is set to 0.01.
[0105] Steps 3 - 4 are carried out under the bidirectional prompt network BPGM designed in the present invention. BPGM combines layer normalization, linear layer, and Transformer decoder layer. Among them, the normalization layer normalizes the input features to stabilize and accelerate the training process. The linear layer mainly performs linear transformation to map the input features to another dimension. The Transformer decoder layer is used to interact between context features and update features through self-attention and cross-attention mechanisms.
[0106] Step 5: Construct the overall loss function :
[0107] Step 5.1: Construct the contrastive loss according to Equation (5.1) :
[0108] (5.1)
[0109] In Equation (5.1), is the temperature parameter, and
[0110] Step 5.2: Calculate the final cross-domain loss from Equation (5.2) :
[0111] (5.2)
[0112] In Equation (5.2), represents the cross-domain amplitude information and is obtained from Equation (5.3), is the cross-domain angle information and is obtained from Equation (5.2),
[0113] (5.3)
[0114] In Equation (5.3), is a function that maps the original variables to the reproducing kernel Hilbert space and is the global visual embedding of
[0115] (5.4)
[0116] In Equation (5.4), represents the cosine similarity function, and , represents the inner product operation, is the norm length.
[0117] Step 5.3: Construct the overall loss function of the bidirectional prompt network BPGM using Equation (5.5) :
[0118] (5.5)
[0119] In Equation (5.5), is the balance factor.
[0120] Step 6: Train the bidirectional prompt network BPGM using the gradient descent method and calculate the overall loss function to update the network parameters until the overall loss function converges, thereby obtaining the optimal image classification model for predicting and classifying images in the target domain.
[0121] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0122] In this embodiment, a computer-readable storage medium stores a computer program. When the computer program is run by a processor, it executes the steps of the above method.
[0123] Related comparative experiments:
[0124] 1. Experimental settings:
[0125] The effectiveness of the present invention was verified on the public datasets MRSSC and Remote Sensing. The specific results are shown in Tables 1 to 2, and the best results are highlighted in bold. The MRSSC dataset contains 26,710 images from the Chinese manned spacecraft "Tiangong-2", covering 7 typical scene categories (including rivers, lakes, cities, farmlands, mountains, coasts, and deserts). Remote Sensing consists of 4 different remote sensing datasets in 4 different domains, with each domain containing five categories, including Agriculture, Forest, River, Residential, and Parking. All experiments were implemented on an Intel (R) Core(TM) i9-13900KF CPU, 64G of memory, and a GeForce RTX 3090 GPU. The model uses the Adadelta optimizer, with the initial value of lr set to 0.1, the number of iterations of the model set to 30, and the batch-size set to 64. The structures of the visual encoder and the text encoder are consistent with CLIP, and their parameters are fixed. To accelerate training and reduce video memory occupancy, we introduced mixed-precision training and trained the model with 16-bit floating-point numbers (FP16). In addition, to enhance the generalization ability of the model, we adopted a weight decay strategy during training, with the weight decay coefficient set to 0.01. For the sake of fair comparison, the parameter settings of the compared methods are all set to the optimal performance parameters.
[0126] Result analysis:
[0127] The comparison results of the present invention in the two datasets are shown in Tables 1 to 2 respectively:
[0128] Table 1: Comparison of the MRSSC dataset with the SOTA method using ViT as the backbone. The best results are highlighted in bold black.
[0129]
[0130] Table 2: Comparison of the Remote Sensing dataset with the SOTA method using ViT as the backbone. The best results are highlighted in bold black.
[0131]
[0132] In Table 1-2, the CLIP concept is derived from the literature [Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C] / / International conference on machine learning. PMLR, 2021: 8748-8763.], the MaPLe concept is derived from the literature [Khattak M U, Rasheed H, Maaz M, et al. Maple: Multi-modal prompt learning[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023: 19113-19122.], the CoOp concept is derived from the literature [Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. IJCV, 130(9):2337–2348, 2022.], the CoCoOp concept is derived from the literature [Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In CVPR, pages 16816–16825, 2022.], the TVT concept is derived from the literature [Yang J, Liu J, Xu N, et al. Tvt: Transferable vision transformer for unsupervised domain adaptation[C] / / Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision. 2023: 520-530.], the SSRT concept is derived from the literature [Sun T, Lu C, Zhang T, et al.Safe self-refinement for transformer-based domain adaptation[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022: 7191-7200.] The CMKD concept is derived from the literature [Zhou W, Zhou Z. Unsupervised Domain Adaption Harnessing Vision-Language Pre-training[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024.], the PDA concept is derived from the literature [Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C] / / International conference on machine learning. PMLR, 2021: 8748-8763.], the DAMP concept is derived from the literature [Du Z, Li X, Li F, et al. Domain-agnostic mutual prompting for unsupervised domain adaptation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024: 23375-23384.], and Ours is the method of the present invention. From the experimental results in Table 1 - Table 2, it can be seen that the non-cross-domain methods CLIP, MaPLe, CoOp, and CoCoOp generally perform weakly, and their accuracies on the MRSSC and RemoteSensing datasets are lower than those of most cross-domain methods. For example, on the MRSSC dataset, the average accuracy of CLIP is only 64.2%, which is significantly lower than that of SSRT (76.3%) and DAMP (79. based on cross-domain learning2%) methods. This is mainly because these methods do not consider the deviation of feature distributions between domains during the feature learning process, making it difficult for the model to effectively transfer in complex cross-domain scenarios, and ultimately resulting in limited generalization ability of the model on the target domain. In contrast, cross-domain methods such as TVT, SSRT, CMKD, PDA, and DAMP perform better in terms of generalization ability, indicating that explicitly aligning the feature distributions of the source domain and the target domain is the key to improving cross-domain classification performance. Especially on the Remote Sensing dataset, for example, SSRT and DAMP perform significantly better than non-cross-domain methods on the Remote Sensing dataset. Among them, the accuracy of DAMP (ResNet50) is as high as 95.4%, which is 14.6% higher than that of CoCoOp, further demonstrating the effectiveness of cross-domain methods in cross-modal tasks. However, the above cross-domain methods still have problems such as insufficient alignment of inter-domain information, insufficient utilization of class structure information, and limited cross-modal fusion ability. Among all cross-domain methods, the method of the present invention achieves optimal performance in most tasks on the MRSSC and Remote Sensing datasets through a more efficient cross-modal information fusion mechanism. For example, on the Remote Sensing dataset, the classification accuracy of BVPDA exceeds 95% in multiple tasks and outperforms methods such as DAMP, PDA, and CMKD in multiple categories.
[0133] In Figures 2a to 2b it, for the tasks in the MRSSC dataset, the confusion matrices of two methods (CoOp and Ours) are compared. There are many misclassification phenomena in the CoOp method in the classification task, while the classification accuracy of the Ours method is significantly higher than that of CoOp in all categories, indicating its obvious advantages in feature extraction and classification boundary optimization.
[0134] In Figure 3 it, the visualization results of three methods (CoOp, DAMP, and Ours) are compared, and their attention areas are reflected through heatmaps. The attention points of the CoOp method are relatively scattered, covering a large number of irrelevant areas, affecting the accuracy of target recognition; the DAMP method has improved in terms of the concentration and relevance of attention points, but there is still a problem that the target area is not fully covered. In contrast, the Ours method, which is the method of the present invention, effectively integrates visual and language modality information through a multi-modal two-way hint and distribution alignment strategy, enhancing the understanding and localization ability of key features.
Claims
1. A cross-modal domain adaptation remote sensing image classification method based on bidirectional visual language cues, characterized in that It is carried out according to the following steps: Step 1: Obtain a source domain dataset with annotation information , where represents the th source domain sample, represents the corresponding label, represents the number of source domain samples; Obtain a target domain dataset without annotation information , where represents the th target domain sample, represents the number of target domain samples, and both contain categories; Step 2: Construct a cross-modal domain adaptation remote sensing image classification network, including: a text encoder with frozen parameters , a visual encoder with frozen parameters and a bidirectional prompting network BPGM; Step 2.1: Input into the visual encoder to obtain the global visual embedding and the visual context information respectively; ; Step 2.2: Input the label information of into the text encoder to obtain respectively the global text embedding of and the text context information , where represents the global text embedding of the k-th category label; Step 2.3: Use the gating mechanism strategy in the bidirectional prompting network BPGM to guide to obtain the final visual embedding of ; Step 2.4: Use the gating mechanism strategy in the bidirectional prompting network BPGM to guide and obtain the final text embedding of ; Step 3: Based on , and construct the overall loss function ; Step 4: Train the Bidirectional Prompting Network BPGM using the gradient descent method and calculate the overall loss function to update the network parameters until the overall loss function converges, thereby obtaining the optimal image classification model for predicting and classifying the images in the target domain.
2. The cross-modal domain adaptation remote sensing image classification method based on bidirectional visual language cues according to claim 1, wherein, Step 2.3 is carried out according to the following steps: Step 2.3.1: Obtain according to Equation (3.1) The corresponding text gating weight value : (3.1) In formula (2.1), represents a splicing operation, is the Sigmoid function, is the text gating weight matrix to be learned; Step 2.3.2: Obtain the updated text context information according to Equation (3.2) : (3.2) In formula (3.2), represents the set of text gating weights, and ; Step 2.3.3: Obtain the fused visual cue according to Equations (3.3)-(3.5) ; (3.3) (3.4) (3.5) In formulas (3.3)-(3.5), , represents the total number of layers of the Transformer network, represents the th decoding layer of the Transformer network, and are two projection operations; represents the semantic feature representation, represents the initial visual feature representation, represents the visual feature representation of the th layer, represents the visual feature representation of the th layer; Step 2.3.4: Obtain the final visual embedding of according to Equation (3.6); (3.6) In Equation (3.6), is the weight factor of the visual cue to be learned.
3. A cross-modal domain adaptation remote sensing image classification method based on bidirectional visual language cues according to claim 2, characterized in that Step 2.4 is carried out according to the following steps: Step 2.4.1: Obtain the visual gating weight of according to Equation (4.1): (4.1) In Equation (4.1), is the visual gating weight matrix to be learned; Step 2.4.2: Obtain according to Equation (4.2) the updated visual context information , thereby obtaining the updated visual context information set ; (4.2) Step 2.4.3: Obtain the final text embedding after fusion of the k-th category according to Equations (4.3)-(4.5) ; (4.3) (4.4) (4.5) In Equations (4.3)-(4.5), represents the source domain sample set, represents the initial text feature representation of the k-th category; represents the visual feature representation, represents the text feature representation of the k-th category in the -th layer, represents the text feature of the k-th category in the -th layer, represents the text feature representation of the k-th category in the Step 2.4.4: Obtain the final text embedding according to Equation (3.6) of ; (4.6) In Equation (3.6), is the weight factor of the text prompt to be learned.
4. A cross-modal domain adaptation remote sensing image classification method based on bidirectional visual language cues according to claim 3, characterized in that Step 3 is carried out according to the following steps: Step 3.1 Construct the contrastive loss according to Equation (5.1) :[[-END]] (5.1) In Equation (5.1), is the temperature parameter, represents the cosine similarity; Step 3.2: Calculate the final cross-domain loss according to Equation (5.2) :[[]]END]] (5.2) In Equation (5.2), represents the cross-domain amplitude information and is obtained from Equation (5.3), is the cross-domain angle information and is obtained from Equation (5.2). (5.3) In Equation (5.3), is a function that maps the original variable to the reproducing kernel Hilbert space , and is the global visual embedding; (5.4) In formula (5.4), represents the cosine similarity function, and , represents the inner product operation, is the norm; Step 3.3: Construct the overall loss function of the Bidirectional Prompting Graph Model (BPGM) using Equation (5.5) :[[]]END]] (5.5) In formula (5.5), is the balance factor.
5. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program for supporting the processor to execute the cross-modal remote sensing image classification domain adaptation method described in any one of claims 1-4, and the processor is configured to execute the program stored in the memory.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the cross-modal remote sensing image classification domain adaptation method described in any one of claims 1-4.
Citation Information
Cited By
Probability alignment unsupervised domain adaptive remote sensing image cross-scene classification method and system
CN120747758A
Probabilistic alignment unsupervised domain adaptation method and system for cross-scene classification of remote sensing images
CN120747758B
Remote sensing target segmentation method, electronic equipment, storage medium and program product
CN120876863A