Visual language remote sensing image classification method based on entropy fusion and Gaussian discriminant analysis
By combining self-supervised learning and Gaussian discriminant analysis with entropy fusion, the problem of uncaptured texture features of objects in remote sensing image classification is solved, and the classification accuracy and adaptability are improved, especially under few-sample conditions.
Patent Information
- Application Number
- CN202510879567.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-21
AI Technical Summary
Existing large multimodal models of visual language fail to fully capture the texture features of objects in remote sensing image classification, resulting in insufficient classification accuracy and limited adaptability and generalization capabilities.
Self-supervised learning is used to extract the texture features of objects in remote sensing images, and Gaussian discriminant analysis is combined to optimize the category feature distribution and decision boundaries. Entropy fusion is used to dynamically adjust the model's focus on complex object areas, and a multi-branch training strategy is used to optimize model performance.
It improves the accuracy and adaptability of remote sensing image classification, especially in the case of few samples, and enhances the model's generalization ability and recognition ability of complex terrain areas.
Smart Images

Figure CN120823433A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing, and in particular relates to a visual language remote sensing image classification method based on entropy fusion and Gaussian discriminant analysis. Background Art
[0002] In recent years, large multimodal models of visual language have attracted widespread attention due to their excellent performance in zero-shot and few-shot classification tasks. However, when applied to remote sensing image classification, such models face the challenges of significant data distribution differences and limited training samples. Natural images are often captured from everyday scenes and contain intuitive semantic information, while remote sensing images are typically acquired by sensors onboard satellites or aircraft and contain surface features (such as farmland, urban buildings, roads, and water bodies). These features differ significantly from natural images in low-level textures. Therefore, if large multimodal models of visual language used in natural image classification tasks are directly transferred to remote sensing images, the classification results will be unsatisfactory. Due to the large size of multimodal models, directly fine-tuning them on new datasets is expensive, relying on manual annotation and incurring significant computational resource overhead.
[0003] To enhance the transfer learning capabilities of large multimodal models in low-sample scenarios, recent studies have proposed a variety of methods to improve performance. These methods can be roughly divided into the following three categories: (1) Prompt optimization: For example, the CoOp method dynamically adjusts the input template of the text encoder through learnable text prompts to improve the model's semantic understanding of remote sensing categories. (2) Cache optimization: To reduce computational requirements, the cache structure is combined with the zero-shot multimodal model to reduce gradient calculations while maintaining classification accuracy. On this basis, a dual cache structure is introduced, combined with DINO to achieve knowledge distillation, and GPT- and DALL-E are used to enhance text and visual features. (3) Prediction value correction: By directly adjusting the output layer bias of the visual language model, the classification decision boundary is optimized in low-sample scenarios.
[0004] However, most of these methods focus on simple adjustments to model text or visual single-modal features to improve performance. They are still insufficiently adaptable to remote sensing images, a special data type with unique ground texture features and spatial structure information. Among them, prompt optimization methods focus solely on learning text prompts, ignoring the spatial texture features of remote sensing images and failing to fully exploit their deep spatial characteristics. While cache optimization methods can reduce computational requirements to a certain extent, they offer limited improvement in classification accuracy for the complex and diverse ground object categories in remote sensing images. Prediction correction methods only correct bias terms and fail to model the distribution of class features, resulting in limited adaptability to imbalanced data. Summary of the Invention
[0005] The purpose of the present invention is to provide a visual language remote sensing image classification method based on entropy fusion and Gaussian discriminant analysis. In the case of zero sample and few samples, on the one hand, it is used to solve the technical problem that the existing technology fails to fully capture the unique texture features of ground objects in remote sensing images; on the other hand, it is used to solve the technical problems that the existing technology has insufficient accuracy in remote sensing image classification and limited adaptability and generalization capabilities.
[0006] The visual language remote sensing image classification method based on entropy fusion and Gaussian discriminant analysis includes the following steps:
[0007] S1. Prepare a remote sensing image dataset: The dataset includes a remote sensing image training set with a small number of samples whose category names are known and a remote sensing image dataset with a large number of samples whose category names are unknown;
[0008] S2. Self-supervised learning: Utilize pre-trained self-supervised visual encoders to access remote sensing image datasets. Without label supervision, the self-supervised learning model learns the texture features of objects in remote sensing images.
[0009] S3, Gaussian discriminant analysis: Based on the Gaussian distribution assumption, the category feature distribution and decision boundary of the final prediction value are optimized to perform a priori correction on the imbalanced categories predicted by the visual language model;
[0010] S4. Multi-branch training based on entropy fusion: Dynamic weights are assigned to high-entropy samples by weighting the entropy values, and combined with the power scaling strategy, the model is guided to focus on complex terrain areas.
[0011] Preferably, in step S2, a self-supervised visual encoder E is introduced S , the self-supervised visual encoder is trained through self-supervised learning on the target dataset, and its basic contrast loss function is as follows:
[0012]
[0013] Where q θ (x1) and q θ (x2) is the output of the self-supervised visual encoder of the online network, corresponding to the two enhanced views x1 and x2 of the input respectively; z ξ (x1) and z ξ (x2) is the embedded feature output of the two enhanced views x1 and x2 on the target visual encoder; sg(·) represents the operation of stopping the gradient, and the parameters of the target network do not participate in the back-propagation update.
[0014] Preferably, in step S3, the self-supervised feature f is output by processing the remote sensing image training set and the remote sensing image data set through the self-supervised visual encoder. s ; Self-supervised feature f s is sent to the self-supervising network adapter AS In the mapping of remote sensing image features, the self-supervised prediction value Logit is obtained s ; The self-supervised prediction value Logits s Logits of the predicted values of the visual language model with zero samples m Perform integration to obtain the final predicted value logits F , the predicted value Logits output using Gaussian correction is obtained through Gaussian discriminant analysis G , correct the final predicted value logits F The corresponding formula is as follows:
[0015] Logits F =(1-α)·Logits m +α·Logits s +β·Logits G ,
[0016] Where, the hyperparameter α∈[0,1] is used to balance the self-supervised prediction value Logits s And the predicted value Logits of the zero-shot visual language model m , β is the predicted value Logits corresponding to the output of Gaussian correction G hyperparameters.
[0017] Preferably, self-supervised prediction value Logits s The calculation formula is:
[0018]
[0019] in, and is the self-supervised adapter A during few-shot training s The learned parameters, the superscript T represents the transpose of a vector or matrix, ReLU represents the ReLU function, and σ represents the Sigmoid function.
[0020] Preferably, the predicted value Logits of the zero-shot visual language model m The calculation formula is W m and They represent the text encoder E of the visual language model respectively T and visual encoder E V The output of the corresponding expressions are: W m =E T (T0) and Among them, T0 and I0 are the text input and visual input of the zero-shot visual language model, respectively.
[0021] Preferably, in the Gaussian discriminant analysis method, the category prior probability is estimated by the proportion of samples belonging to the category:
[0022]
[0023] Where m is the total number of samples, is an indicator function, which takes the value 1 when the i-th sample belongs to category k, otherwise it takes the value 0; the category mean μ of category k k It is calculated by the following formula:
[0024]
[0025] Among them, x (i) represents the i-th sample;
[0026] Setting the weight matrix and the bias vector K is the total number of categories, D is the dimension of the feature vector, then: w k =Σ -1 μ k , Among them, w k is the weight vector of category k, b k is the bias vector of category k; the mean μ of category k is estimated by using a few randomly selected training data from the visual language feature space k and the precision matrix Σ -1 ; Use Gaussian correction to output predicted values Logits G The formula is as follows: Logits G =W m W T +b,W m Text encoder E representing the visual language model T Output.
[0027] Preferably, in step S4, an entropy-based confidence formula is used to quantify the uncertainty of the model prediction. First, the prediction probability of the visual language model for the i-th category is calculated by the Softmax function as follows:
[0028]
[0029] Where C represents the number of categories, represents the predicted value of the visual language model for category i, p i is the predicted probability of category i; the prediction uncertainty of each sample is calculated by entropy, and the formula is as follows:
[0030]
[0031] In the formula, the entropy value H is used to measure the degree of dispersion of the predicted distribution.
[0032] Preferably, a power scaling operation is introduced to calculate the uncertainty factor ξ of the entire sample set based on the entropy value of the sample. The formula is as follows:
[0033]
[0034] Where N is the total number of samples, H n is the entropy value of the nth sample, and ρ is the power scaling factor; the final prediction value based on entropy fusion is as follows:
[0035] Logits F =ξ·(1-α)·Logits m +α·Logits s +β·Logits G .
[0036] Preferably, in step S4, a multi-branch training strategy is used, under the C-way-N-shot setting, the feature is the self-supervised feature of the j-th sample of the i-th class, and the self-supervised prediction value of the j-th sample for: The model is optimized through the cross entropy loss function so that the output probability distribution of the self-supervised branch and the output probability distribution of the main branch are as close as possible to the distribution of the true label y. The corresponding loss function is as follows:
[0037]
[0038] Where g(·) is a softmax function, represents the final predicted value of the jth sample;
[0039] The total loss l for multi-branch training Total The loss function is expressed as: Total =(1-λ)l s +λl F , where λ∈[0,1] is a hyperparameter used to balance the self-supervised branch loss l S and the main branch loss l F impact.
[0040] The present invention is a self-supervised few-sample remote sensing image classification method based on visual language, which has the following beneficial effects:
[0041] 1) Few-label domain adaptation pre-training: This method uses self-supervised contrastive learning to extract spatial-spectral features of ground objects from sparsely labeled remote sensing images, and then cross-modally fuses them with the semantic prediction results of the visual language model through nonlinear mapping to construct a robust joint feature representation.
[0042] 2) Dynamic entropy weight fusion mechanism: This method designs a multi-branch collaborative network, quantifies sample uncertainty based on predicted entropy, dynamically allocates attention weights to high-entropy regions, and enhances the model's ability to discriminate complex landform boundary features; during model training, it guides the model to pay more attention to complex landform areas, thereby improving the classification effect of minority class targets, that is, improving the adaptability and classification performance of difficult-to-classify samples.
[0043] 3) Gaussian discriminant prior optimization: This method uses Gaussian discriminant analysis to establish a class-conditional probability model, performs prior correction on cross-modal features, and effectively suppresses the distribution shift error under few-sample conditions, thereby enhancing the model's generalization ability for few-sample tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a basic flow chart of a visual language remote sensing image classification method based on entropy fusion and Gaussian discriminant analysis of the present invention.
[0045] Figure 2 Schematic diagram of the self-supervised pre-training stage in the present invention.
[0046] Figure 3 This is a flowchart of the multi-branch few-sample training phase of entropy fusion and Gaussian discriminant analysis in the present invention.
[0047] Figure 4 This is the distribution map of remote sensing image features extracted by the self-supervised model in this invention.
[0048] Figure 5 This is a diagram of the training process of the present invention on the NWPU-RESISC45 remote sensing dataset.
[0049] Figure 6 This is the confusion matrix of the present invention on the NWPU-RESISC45 remote sensing dataset. DETAILED DESCRIPTION
[0050] The specific implementation methods of the present invention will be further explained in detail below through the description of embodiments with reference to the accompanying drawings, so as to help those skilled in the art to have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.
[0051] like Figures 1-6 As shown, the present invention provides a visual language remote sensing image classification method based on entropy fusion and Gaussian discriminant analysis, which includes the following steps.
[0052] S1. Prepare a remote sensing image dataset: The dataset includes a remote sensing image training set with a small number of samples whose category names are known and a remote sensing image dataset with a large number of samples whose category names are unknown.
[0053] This step also creates prompt templates for the remote sensing images in the remote sensing image dataset. The prompt templates are used to guide the visual language model to understand and classify the remote sensing images. They correspond to each remote sensing image and provide additional information. Figure 4 The textual prompts in form the textual input of the zero-shot visual language model. For a given remote sensing image input x, the image classification goal of this method is to predict the category label y∈{1,...,K} depicted in the image x.
[0054] In the embodiment, the remote sensing image datasets used in this method include NWPU-RESISC45, EuroSAT, PatternNet and UC Merced Land Use datasets, which are widely used in remote sensing scene classification tasks and have the characteristics of rich categories, wide coverage, and high image quality. Among them, NWPU-RESISC45 contains 31,500 images in 45 categories, EuroSAT is based on Sentinel-2 satellite images and contains approximately 27,000 images in 10 categories, PatternNet provides 38 categories with a total of approximately 30,400 images, and the UC Merced dataset contains 21 categories with a total of 2,100 images. All images are high-resolution remote sensing images, covering a variety of typical land feature scenes such as nature, city, agriculture, and transportation, providing a diverse testing environment and sufficient experimental data support for the present invention.
[0055] S2. Self-supervised learning: Use the pre-trained self-supervised visual encoder to access the remote sensing image dataset, and let the self-supervised learning model learn the texture features of the remote sensing image without label supervision.
[0056] This method introduces a self-supervised visual encoder E S , the self-supervised visual encoder is trained through self-supervised learning on the target dataset, and its basic contrast loss function is as follows:
[0057]
[0058] Where q θ (x1) and q θ (x2) is the output of the self-supervised visual encoder of the online network, corresponding to the two enhanced views x1 and x2 of the input, respectively. The two enhanced views come from two enhancements of the unlabeled remote sensing image x, z ξ (x1) and z ξ (x2) is the embedded feature output of the two enhanced views x1 and x2 on the target visual encoder; sg(·) represents the operation of stopping the gradient, that is, the parameters of the target network do not participate in the back-propagation update.
[0059] Both the online network and the target network consist of a feature extractor and a projection head, enabling the joint extraction of spectral and spatial features. The parameters of the target network are synchronized with the online network using an exponential moving average (EMA). This step utilizes a pre-trained self-supervised visual encoder as the feature extractor. Available encoder types include ResNet and VIT.
[0060] S3, Gaussian discriminant analysis: Based on the Gaussian distribution assumption, it optimizes the category feature distribution and decision boundary of the final prediction value, and performs a priori correction on the imbalanced categories predicted by the visual language model.
[0061] The self-supervised prediction value of the self-supervised visual encoder is combined with the prediction value of the pre-trained zero-shot visual language model to obtain the final prediction value. After the self-supervised training is completed, the remote sensing images in the remote sensing image training set and the remote sensing image dataset are processed by the self-supervised visual encoder. The self-supervised visual encoder is used as a feature extractor to output the self-supervised feature f s . Self-supervised feature f s It is then sent to the self-supervising network adapter A S In the mapping of remote sensing image features, the self-supervised prediction value Logit is obtained s The calculation formula of the self-supervised prediction value is:
[0062]
[0063] in, and is the self-supervised adapter A during few-shot training s The learning parameters, superscript T represents the transpose of a vector or matrix, ReLU represents the ReLU function, and σ represents the Sigmoid function. The self-supervised features of each remote sensing image are adjusted and mapped to the Logits space through the above formula to obtain the self-supervised prediction value Logits s .
[0064] The self-supervised prediction value Logits s The final prediction value is integrated with the prediction value of the zero-sample visual language model to obtain the final prediction value, the final prediction value Logits F The overall expression is: Logits F =(1-α)·Logits m +α·Logits s , where Logits m is the predicted value of the zero-sample visual language model, calculated as The hyperparameter α∈[0,1] is used to balance the two prediction values; W m and They represent the text encoder E of the visual language model respectivelyT and visual encoder E V The output of the corresponding expressions are: W m =E T (T0) and Among them, T0 and I0 are the text input and visual input of the zero-shot visual language model, respectively.
[0065] The pre-trained zero-shot visual language model extracts features of each category of remote sensing images. These features are unevenly distributed. This uneven distribution may lead to complex decision boundaries between categories and difficulty in accurate separation. The present invention corrects the final prediction value logits by Gaussian Discriminant Analysis (GDA). F This further improves the classification performance.
[0066] In the GDA method, we first need to estimate the parameters of each category, and the category prior probability is expressed as: k =P(y=k), prior probability φ k represents the probability that a sample comes from category k, y represents the category, and P represents the probability function, which is estimated by the proportion of samples belonging to the category:
[0067]
[0068] Where m is the total number of samples, Is an indicator function, which takes the value 1 when the i-th sample belongs to category k, otherwise it takes the value 0. The category mean μ of category k k It is calculated by the following formula:
[0069]
[0070] Among them, x (i) represents the i-th sample.
[0071] Then use the covariance matrix Σ to estimate the distribution range of each category feature. Each category feature shares the same covariance matrix Σ, and the weights and bias terms of the linear classifier are obtained by sorting the discriminant function. Set the weight matrix and the bias vector K is the total number of categories, D is the dimension of the feature vector, then: w k =Σ -1 μ k , Among them, w k is the weight vector of category k, b k is the bias vector of category k; x T Σ -1 μ k Represents the distance between sample x and the center of category k; Indicates the relative position of category k; logφ k Introducing prior information of categories to reflect the impact of category imbalance on classification.
[0072] This step estimates the mean μ of category k by using a few randomly selected training data from the visual language feature space. k and the precision matrix Σ -1 (the inverse matrix of the covariance matrix), and then use these estimated values to calculate the linear classifier (i.e., weight matrix) W and bias vector b, so as to obtain the corrected predicted value logits F Corrected predicted values logits F The corresponding formula is as follows:
[0073]
[0074] Among them, W m Text encoder E representing the visual language model T Output, Logits G Indicates the predicted value output using Gaussian correction, β is the corresponding Logits G This step enhances the model's generalization ability for few-shot tasks by performing a priori correction on the imbalanced categories of the predicted values.
[0075] S4. Multi-branch training based on entropy fusion: Dynamic weights are assigned to high-entropy samples by weighting the entropy values, and combined with the power scaling strategy, the model is guided to focus on complex terrain areas.
[0076] In order to quantify the uncertainty of the model prediction, this step uses the entropy-based confidence formula for calculation. First, the prediction probability of the visual language model for the i-th category is calculated using the Softmax function as follows:
[0077]
[0078] Where C represents the number of categories, represents the predicted value of the visual language model for category i, p i is the predicted probability of category i. On this basis, the prediction uncertainty of each sample is further calculated by entropy, and the formula is as follows:
[0079]
[0080] In the formula, the entropy value H is used to measure the degree of dispersion of the predicted distribution. That is, by calculating the entropy value of the sample, the uncertainty of the model's prediction of these categories is measured, thereby reflecting the model's adaptability to diverse landforms. To further enhance the model's attention to high-entropy samples, this step also introduces a power scaling operation. Based on the entropy value of the sample, the uncertainty factor ξ of the entire sample set is calculated. The formula is as follows:
[0081]
[0082] Where N is the total number of samples, H n is the entropy value of the nth sample, and ρ is a power scaling factor used to control the weight of high-entropy samples. This is particularly applicable to samples of unevenly distributed categories in remote sensing images. By assigning higher weights to high-entropy samples, this step can guide the model to pay more attention to complex terrain areas during training, thereby improving the classification effect of minority class targets, that is, improving the adaptability and classification performance of difficult-to-classify samples. The final prediction value based on entropy fusion is as follows:
[0083] Logits F =ξ·(1-α)·Logits m +α·Logits s +β·Logits G .
[0084] In order to fully exploit the self-supervised features and the prediction value characteristics based on entropy fusion, this method uses a multi-branch training strategy. Under the C-way-N-shot setting (i.e., there are C categories and each category has N samples), the feature is the self-supervised feature of the j-th sample of the i-th class, and the self-supervised prediction value of the j-th sample for:
[0085] This method optimizes the model through the cross entropy loss function, so that the output probability distribution of the self-supervised branch and the output probability distribution of the main branch are as close as possible to the distribution of the true label y. The corresponding loss function is as follows:
[0086]
[0087] Where g(·) is a softmax function, represents the final predicted value of the jth sample. Therefore, the total loss l for multi-branch training is Total The loss function is expressed as: Total =(1-λ)l s +λl F , where λ∈[0,1] is a hyperparameter used to balance the self-supervised branch loss l S and the main branch loss l FThe impact of self-supervised adapter A S Parameters and Through the total loss l Total The loss function is optimized.
[0088] In this embodiment, the training process of each class in the training set has only 16 samples. Figure 4 As shown in Table 1, the classification performance of the present invention is compared with three existing few-sample image classification methods.
[0089] Table 1: Performance comparison between the method used in the present invention and three prior arts
[0090]
[0091]
[0092] To further explore the contributions of each model component, we conducted ablation experiments, examining the following factors: whether multimodality (C) was used, whether a multi-branch training strategy based on entropy fusion (M) was used, and whether Gaussian discriminant analysis (G) was used. Using the NWPU-RESISC45 and Eurosat datasets as examples, the results are shown in Table 2. The baseline model (SSL) with a linear layer connected for 4-shot training on NWPU-RESISC45 achieved an accuracy of 69.56%. After adding multimodal features (C) and the multi-branch strategy based on entropy fusion (M), the accuracy increased to 78.68% and 80.74%, respectively. After introducing Gaussian discriminant analysis (G) correction, the accuracy was further improved to 81.43%. On the Eurosat dataset, the 4-, 8-, and 16-shot accuracies of the baseline model (SSL) were 87.60%, respectively. After introducing multimodal features (C), the accuracy actually decreased. However, after the multi-branch strategy based on entropy fusion (M) was introduced, the accuracy was improved. This also verifies that the entropy fusion proposed in the present invention can guide the model to focus on complex categories and further improve the classification accuracy. As can be seen from the table, the combination of multimodal features (C) and the multi-branch strategy based on entropy fusion (M) has an important contribution to the model performance. Gaussian discriminant analysis (G) plays a key role in multiple scenarios and further improves the generalization ability of the model.
[0093] Table 2: Analysis results of ablation experiments
[0094]
[0095] Analysis of the above charts demonstrates that the proposed method leverages the complementary advantages of self-supervised learning and multimodal models, achieving excellent performance on remote sensing datasets, particularly in scenarios with limited sample sizes. This approach, without requiring additional complex processing or resource investment, provides an efficient and robust solution for few-sample classification of remote sensing data, demonstrating significant practical application. Future research can further expand on this by optimizing feature extraction, enhancing class differentiation, and integrating multimodal information.
[0096] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as various non-substantial improvements are made using the inventive concept and technical solution of the present invention, or the inventive concept and technical solution are directly applied to other occasions without improvement, they are all within the scope of protection of the present invention.
Claims
1. A visual language remote sensing image classification method based on entropy fusion and Gaussian discriminant analysis, characterized by: The following steps are involved: S1. Prepare a remote sensing image dataset: The dataset includes a remote sensing image training set with a small number of samples whose category names are known and a remote sensing image dataset with a large number of samples whose category names are unknown; S2. Self-supervised learning: Utilize pre-trained self-supervised visual encoders to access remote sensing image datasets. Without label supervision, the self-supervised learning model learns the texture features of objects in remote sensing images. S3, Gaussian discriminant analysis: Based on the Gaussian distribution assumption, the category feature distribution and decision boundary of the final prediction value are optimized to perform a priori correction on the imbalanced categories predicted by the visual language model; S4. Multi-branch training based on entropy fusion: Dynamic weights are assigned to high-entropy samples by weighting the entropy values, and combined with the power scaling strategy, the model is guided to focus on complex terrain areas.
2. The method for visual language remote sensing image classification based on entropy fusion and Gaussian discriminant analysis according to claim 1, characterized in that: In step S2, a self-supervised visual encoder E is introduced S , the self-supervised visual encoder is trained through self-supervised learning on the target dataset, and its basic contrast loss function is as follows: Where q θ (x1) and q θ (x2) is the output of the self-supervised visual encoder of the online network, corresponding to the two enhanced views x1 and x2 of the input respectively; z ξ (x1) and z ξ (x2) is the embedded feature output of the two enhanced views x1 and x2 on the target visual encoder; sg(·) represents the operation of stopping the gradient, and the parameters of the target network do not participate in the back-propagation update.
3. The method for visual language remote sensing image classification based on entropy fusion and Gaussian discriminant analysis according to claim 2, characterized in that: In step S3, the remote sensing image training set and the remote sensing image dataset are processed by the self-supervised visual encoder to output the self-supervised feature f s ; Self-supervised feature f s is sent to the self-supervising network adapter A S In the mapping of remote sensing image features, the self-supervised prediction value Logit is obtained s ; The self-supervised prediction value Logits s Logits of the predicted values of the visual language model with zero samples m Perform integration to obtain the final predicted value logits F , the predicted value Logits output using Gaussian correction is obtained through Gaussian discriminant analysis G , correct the final predicted value logits F The corresponding formula is as follows: Logits F =(1-α)·Logits m +α·Logits s +β·Logits G , Where, the hyperparameter α∈[0,1] is used to balance the self-supervised prediction value Logits s And the predicted value Logits of the zero-shot visual language model m , β is the predicted value Logits corresponding to the output of Gaussian correction G hyperparameters.
4. The method for visual language remote sensing image classification based on entropy fusion and Gaussian discriminant analysis according to claim 3 is characterized by: Self-supervised prediction values Logits s The calculation formula is: in, and is the self-supervised adapter A during few-shot training s The learned parameters, the superscript T represents the transpose of a vector or matrix, ReLU represents the ReLU function, and σ represents the Sigmoid function.
5. The method for visual language remote sensing image classification based on entropy fusion and Gaussian discriminant analysis according to claim 3, characterized in that: Logits of the predicted values of the zero-shot visual language model m The calculation formula is W m and They represent the text encoder E of the visual language model respectively T and visual encoder E V The output of the corresponding expressions are: W m =E T (T0) and Among them, T0 and I0 are the text input and visual input of the zero-shot visual language model, respectively.
6. The method for visual language remote sensing image classification based on entropy fusion and Gaussian discriminant analysis according to claim 3, characterized in that: In the Gaussian discriminant analysis method, the category prior probability is estimated by the proportion of samples belonging to the category: Where m is the total number of samples, is an indicator function, which takes the value 1 when the i-th sample belongs to category k, otherwise it takes the value 0; the category mean μ of category k k It is calculated by the following formula: Among them, x (i) represents the i-th sample; Setting the weight matrix and the bias vector K is the total number of categories, D is the dimension of the feature vector, then: w k =Σ -1 μ k , Among them, w k is the weight vector of category k, b k is the bias vector of category k; the mean μ of category k is estimated by using a few randomly selected training data from the visual language feature space k and the precision matrix Σ -1 ; Use Gaussian correction to output predicted values Logits G The formula is as follows: Logits G =W m W T +b,W m Text encoder E representing the visual language model T Output.
7. The method for visual language remote sensing image classification based on entropy fusion and Gaussian discriminant analysis according to claim 3, characterized in that: In step S4, the entropy-based confidence formula is used to quantify the uncertainty of the model prediction. First, the prediction probability of the visual language model for the i-th category is calculated by the Softmax function as follows: Where C represents the number of categories, represents the predicted value of the visual language model for category i, p i is the predicted probability of category i; the prediction uncertainty of each sample is calculated by entropy, and the formula is as follows: In the formula, the entropy value H is used to measure the degree of dispersion of the predicted distribution.
8. The method for visual language remote sensing image classification based on entropy fusion and Gaussian discriminant analysis according to claim 7, characterized in that: The power scaling operation is introduced to calculate the uncertainty factor ξ of the sample set as a whole based on the entropy value of the sample. The formula is as follows: Where N is the total number of samples, H n is the entropy value of the nth sample, and ρ is the power scaling factor; the final prediction value based on entropy fusion is as follows: Logits F =ξ·(1-α)·Logits m +α·Logits s +β·Logits G 。 9. The method for visual language remote sensing image classification based on entropy fusion and Gaussian discriminant analysis according to claim 7, characterized in that: In step S4, a multi-branch training strategy is used, under the C-way-N-shot setting, the feature is the self-supervised feature of the j-th sample of the i-th class, and the self-supervised prediction value of the j-th sample for: The model is optimized through the cross entropy loss function so that the output probability distribution of the self-supervised branch and the output probability distribution of the main branch are as close as possible to the distribution of the true label y. The corresponding loss function is as follows: Where g(·) is a softmax function, represents the final predicted value of the jth sample; The total loss l for multi-branch training Total The loss function is expressed as: Total =(1-λ)l s +λl F , where λ∈[0,1] is a hyperparameter used to balance the self-supervised branch loss l S and the main branch loss l F impact.
Citation Information
Cited By
Incomplete multi-view multi-label classification method based on cross-view distillation and adaptive mask
CN121743990A