Local descriptor-based three-prototype correction network few-sample remote sensing scene classification method

Through the three-prototype correction network method based on local descriptors, the problems of irrelevant background interference and in-class differences in remote sensing scene classification are solved, and higher classification accuracy and effective classification of new categories of images are achieved.

CN119963897AActive Publication Date: 2025-05-09CHONGQING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510023170.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-09
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

The existing remote sensing scene classification method of few samples ignores the serious unrelated background interference in remote sensing images, resulting in poor classification results, especially when facing new categories of images.

Method used

A three-prototype correction network with few-sample remote sensing scene classification method based on local descriptors is proposed. Through feature embedding learning, three-prototype learning, three-prototype correction and classifier modules of local descriptors, it reduces background interference and alleviates classification confusion caused by in-class differences.

Benefits of technology

Effectively reduce interference with irrelevant background information, improve the classification accuracy of remote sensing scene images, and especially when facing new categories of images, it can be classified more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963897A_ABST
    Figure CN119963897A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of artificial intelligence, and particularly relates to a three-prototype correction network few-sample remote sensing scene classification method based on local descriptors, which comprises the following steps: downloading a remote sensing data set and dividing the remote sensing scene image data set into a training set and a verification set according to the category of remote sensing scene images in the remote sensing scene image data set; establishing a three-prototype correction network few-sample remote sensing scene classification model based on local descriptors; performing meta-training on the three-prototype correction network few-sample remote sensing scene classification model based on the local descriptor by using a task-based meta-learning training strategy, performing meta-verification on the three-prototype correction network few-sample remote sensing scene classification model based on the local descriptor through a verification task while training, and storing a model with optimal performance; and classifying the remote sensing images by using the model with the optimal performance. The method can effectively alleviate the problem that a single prototype network overuses irrelevant background information and cannot effectively capture scene image features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence, natural language processing and sentiment computing, and particularly relates to a local descriptor-based three-prototype correction network few-sample remote sensing scene classification method. Background Art

[0002] The purpose of remote sensing scene classification is to distinguish different scene types by analyzing the content of remote sensing images. It is widely used in natural disaster monitoring, urban development planning, land cover type classification and other fields. In recent years, deep learning methods have made significant progress in remote sensing scene classification, but these methods rely on a large number of labeled samples for training and cannot effectively classify remote sensing images of new categories, so they have great limitations. In contrast, few-shot learning methods can be trained with only a small amount of labeled data, thereby effectively handling new categories of image classification problems, which just makes up for the shortcomings of deep learning methods. The concept of few-shot learning has been widely used in natural image classification, but because remote sensing images are essentially different from natural images, it is often difficult to obtain ideal classification results when the few-shot learning method for natural image classification is directly applied to remote sensing images. Compared with natural images, remote sensing images have the following significant characteristics: 1) Complex background: Remote sensing images are usually taken from high altitudes and contain a large number of objects that are not related to the category, resulting in complex image backgrounds. 2) Large intra-class differences: Due to the diversity of objects and differences in lighting, scale and angle, scene images of the same category are visually very different. The above characteristics greatly increase the difficulty of the few-sample remote sensing scene classification task.

[0003] Existing few-sample remote sensing scene classification methods often ignore the fact that remote sensing scene images contain serious irrelevant background interference. They mostly use classifiers based on single prototypes. However, the single prototype obtained by simply averaging the features of images of each category retains more background information in the image that is irrelevant to the category semantics, which is insufficient to represent the semantic commonality of the category. That is, the classifier based on a single prototype is greatly negatively affected by the irrelevant background information in the image, and therefore the classification effect is poor. Summary of the invention

[0004] In order to effectively reduce the interference of irrelevant background information, alleviate the classification confusion problem caused by large intra-class differences in scene images, and improve the classification accuracy of remote sensing scene images, the present invention proposes a three-prototype correction network few-sample remote sensing scene classification method based on local descriptors, which specifically includes the following steps:

[0005] Download the remote sensing dataset and divide the remote sensing scene image dataset into a training set and a validation set according to the categories of the remote sensing scene images in the remote sensing scene image dataset;

[0006] Establish a local descriptor-based three-prototype correction network few-shot remote sensing scene classification model;

[0007] The task-based meta-learning training strategy is used to meta-train the local descriptor-based three-prototype correction network few-sample remote sensing scene classification model. During the training, the local descriptor-based three-prototype correction network few-sample remote sensing scene classification model is meta-verified through the verification task, and the model with the best performance is saved.

[0008] Classify remote sensing images using the best performing models.

[0009] Furthermore, the construction of the remote sensing scene image dataset includes: obtaining remote sensing images of N categories, obtaining K remote sensing images of each category as the support set of the category, and M remote sensing images as the query set of the category.

[0010] Furthermore, the local descriptor-based three-prototype correction network few-sample remote sensing scene classification model includes a local descriptor feature embedding learning module, a three-prototype learning module, a three-prototype correction module and a classifier module, wherein:

[0011] A feature embedding learning module for local descriptors to obtain embedded representations of input remote sensing images;

[0012] A three-prototype learning module is used to obtain three different prototype representations from the embedded representation of remote sensing images through three learnable weight matrices;

[0013] The three-prototype correction module is used to correct the three different prototype representations obtained by the three-prototype learning module and correct the local description of each query sample;

[0014] The classifier module is used to classify the query sample based on the local description and the similarity of each prototype of the support set image.

[0015] Furthermore, the process of obtaining three different prototype representations from the embedded representation of the remote sensing image through three learnable weight matrices includes:

[0016] Reshape the tensor of K samples in the support set into a feature map Y;

[0017] Input the new feature map into the spatial channel dual excitation module to obtain the feature map Z;

[0018] A single prototype representation for each class is obtained by summing the local descriptors at all locations:

[0019]

[0020] Among them, P n represents a single prototype representation of the nth class, represents the projection of the (i, j) position of the k-th support image in the support set of the n-th class, y (i,j) Represents the local descriptor at position (i, j) in the feature map after axis transformation.

[0021] Furthermore, the process of obtaining a new feature map by the spatial channel dual excitation module includes:

[0022] Combine the K samples of each category in the support set into a tensor;

[0023] The embedded representation of remote sensing images is X=[x1,...,x h×w ] The projection tensor v is obtained by compression using convolution operation. Each element in the tensor v is a linear combination of the elements at that position of all channels. Different weight matrices are used in the compression process to obtain different prototypes of a class.

[0024] The tensor v is converted into a weight between 0 and 1 using the Sigmoid activation function to weight the embedded representation of the remote sensing image to obtain the spatial position excitation;

[0025] The feature map is regarded as a tensor of different channels. For each channel tensor, global average pooling is performed in space, that is, spatial dimension compression. The feature map of each channel is used as a local descriptor. All local descriptors form a vector z = [z1, ..., z C ]∈R 1×1×C , z c represents the local descriptor of the cth channel, C is the number of channels of the feature map;

[0026] Use a weight matrix to map the vector z, then use the ReLU activation function to process it, and finally use another weight matrix to map it to get the dependency relationship between channels.

[0027] The dependency is transformed by the Sigmoid activation function Convert to [0,1] and weight each channel tensor of the input to obtain channel position excitation;

[0028] The spatial position excitation and the channel position excitation are added and fused to obtain a fused feature map.

[0029] Furthermore, correcting the three different prototype representations obtained by the three-prototype learning module includes the following steps:

[0030] The support set prototype P is obtained s Take the average value as the representative prototype of the support set The average of all the local descriptors of the M scene images in the query set Q is taken as the representative vector of the query set

[0031] The difference information between the representative vector of the query set and the representative prototype of the support set is obtained, and half of the difference information is used to correct the three prototypes in the support set and the query feature set in the query set respectively.

[0032] Furthermore, the classification process of the classifier module includes:

[0033] By calculating the cosine similarity between all local descriptors of the query image and the nearest prototype of the category and summing these similarities, the overall similarity s between the mth query image and the nth support category is obtained. n,m ;

[0034] Calculate the mth scene image in the support set and each supporting class, and sort them according to the most similar supporting class. Perform classification prediction. The three-prototype strategy proposed in the present invention uses multiple prototypes as representatives of the support set category features, which can effectively alleviate the problem that the single prototype network overuses irrelevant background information and cannot effectively capture the scene image class features; at the same time, the three prototypes of the support set are corrected to enhance the representativeness of each prototype, making each prototype closer to the actual distribution of the scene category, and further alleviating the interference of background information. In addition, the present invention also proposes an intra-class symmetric KL divergence constraint to make scene images of the same class more compact, thereby improving the classification effect of remote sensing scene images. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A schematic diagram of a method for classifying a few-sample remote sensing scene using a three-prototype correction network based on local descriptors according to the present invention;

[0036] Figure 2 It is a diagram of the space and channel dual excitation module of the present invention;

[0037] Figure 3 This is the overall framework diagram of the classification model of the present invention. DETAILED DESCRIPTION

[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0039] The present invention proposes a three-prototype correction network few-sample remote sensing scene classification method based on local descriptors, which specifically includes the following steps:

[0040] Download the remote sensing dataset and divide the remote sensing scene image dataset into a training set and a validation set according to the categories of the remote sensing scene images in the remote sensing scene image dataset;

[0041] Establish a local descriptor-based three-prototype correction network few-shot remote sensing scene classification model;

[0042] The task-based meta-learning training strategy is used to meta-train the local descriptor-based three-prototype correction network few-sample remote sensing scene classification model. During the training, the local descriptor-based three-prototype correction network few-sample remote sensing scene classification model is meta-verified through the verification task, and the model with the best performance is saved.

[0043] The remote sensing images are classified using the model with the best performance. In this embodiment, a specific implementation method of a three-prototype correction network small-sample remote sensing scene classification method based on local descriptors is proposed. In this embodiment, the effectiveness of the present invention is also tested through a test task, which specifically includes the following steps:

[0044] 1) Download the remote sensing dataset and divide the remote sensing scene image dataset into a training set, a validation set, and a test set according to the categories of the remote sensing scene images in the remote sensing scene image dataset; specifically, in this embodiment, the total category set C total The dataset D is divided into training set D train , validation set D val and the test set D test , each set contains C categories of remote sensing scene images train , C val and C test , that is, C train ∪C val ∪C test =C total and

[0045] 2) According to the divided training set, validation set and test set, corresponding training tasks, validation tasks and test tasks are constructed respectively; specifically, the training set C train For example, the process of constructing a training task is as follows: From the category set C of the training set train Randomly select N categories from the dataset, randomly select K+M remote sensing scene images from each category, and use the K remote sensing scene images from each category as the support set and the M remote sensing scene images as the query set. At this time, the remote sensing scene images required for a training task are constructed. The verification task and the test task are constructed from the corresponding verification / test scene image category C. val and C test Sampling scene images in the same way as constructing training tasks;

[0046] 3) Establish a three-prototype correction network few-sample remote sensing scene classification model based on local descriptors;

[0047] 4) Meta-train the local descriptor-based three-prototype correction network few-sample remote sensing scene classification model using a task-based meta-learning training strategy. During the training, meta-verify the local descriptor-based three-prototype correction network few-sample remote sensing scene classification model through the verification task, and save the model with the best performance;

[0048] 5) After the training and validation process is completed, the best performing model is meta-tested on the test task.

[0049] In this embodiment, the local descriptor-based three-prototype correction network few-sample remote sensing scene classification model includes: a local descriptor feature embedding learning module, a three-prototype learning module, a three-prototype correction module and a classifier module.

[0050] like Figure 3 In the feature embedding learning module of the local descriptor, the ResNet12 network is used as the feature embedding. Given an input scene image I, the output feature map of the ResNet12 network can be represented as a 3D tensor X∈R with h×w×c elements. h×w×d From another perspective, X can be viewed as a set of dc-dimensional (d = h × w) local descriptors, that is, X = [x1, ..., x d ]∈R d×c , where x i is the i-th depth local descriptor, i∈{1,2,…,d}, such a descriptor corresponds to a spatial local feature (i.e., local descriptor) on the image; denote one of the local descriptors as x (i,j) , where (i,j) represents a specific position in the feature map, i∈{1,...,h}, j∈{1,...,w}, x (i,j) ∈R 1×1×c , use F θ (.) represents the ResNet12 network. The above process can be expressed by the formula:

[0051] X=F θ (I) = [x (1,1) ,...,x (i,j) ]∈R d×c

[0052] like Figure 3 In the three-prototype learning module, this embodiment designs a novel three-prototype learning module Three prototypes for each category are learned adaptively.

[0053] Specifically, there are three prototype learning modules The core part is the spatial channel dual excitation module, which includes: spatial excitation and channel excitation modules. Figure 2 As shown in Figure 2, the spatial excitation module compresses the feature map X along the channel dimension and expands it along the spatial dimension. The channel compression is achieved through convolution operation, that is, v = W s *X, where W s ∈R 1×1×c×1 Projection tensor v∈R h×w , each projection v (i,j) Represents the linear combination of the elements at position (i, j) in all channels. The projection is converted to a weight of 0 to 1 using the Sigmoid activation function σ(.), which is used for the excitation X of the spatial position. The above process formula is expressed as:

[0054] X s =f s (X) = [σ(v (1,1) )x (1,1) ,...,σ(v (i,j) )x (i,j) ,...,σ(v (h,w) )x (h,w) ]

[0055] Among them, each σ(v (i,j) ) corresponds to the relative importance of the information at position (i, j) on the feature map.

[0056] In another channel, namely the channel excitation module, the feature map X=[x1,...,x k ,...,x C ] is regarded as a feature map composed of different channel tensors, and the tensor of the cth channel is represented by x c ∈R h×w , k∈{1,...,C}, C is the number of channels in the feature map, and global average pooling is performed on each channel tensor, that is, spatial dimension compression. At this time, the feature map becomes a vector composed of local descriptors of each channel, represented by z=[z1,...,z C ]∈R 1×1×C , z c represents the local descriptor of the cth channel, which embeds the global spatial information into the vector z. The process is expressed as:

[0057]

[0058] Among them, x k (i,j) represents the (i,j) position of the k-th channel tensor.

[0059] The vector z that embeds the global spatial information is then encoded as the dependencies between channels through two linear layers and an activation function, which can be expressed as:

[0060]

[0061] in, Represent the weights of the two fully connected layers, Represents a ReLU operation.

[0062] Finally, the Sigmoid activation function σ(.) is used to The dynamic variable range is converted to [0,1] to obtain the feature map X after re-channel excitation c , expressed as:

[0063]

[0064] Among them, activation represents the importance of the kth channel, which are reweighted. As the network learns, these activations are adaptively adjusted to ignore less important channels and emphasize important channels.

[0065] This embodiment combines the features of the spatial excitation and channel excitation outputs by element-by-element spatial and channel excitation addition to form the output of the final spatial channel dual excitation module: When the position (i, j, c) of the input feature map X gains high importance from channel reweighting and spatial reweighting, it is given higher incentive. This recalibration encourages the network to learn more meaningful feature maps that are both spatially and channel-wise relevant, which helps to weaken the distraction of the scene image background.

[0066] In a general N-way K-shot few-shot classification problem, each class in the support set contains K samples, and the feature map of each sample is a c×h×w tensor. The feature maps of each class can be combined as X∈R K×c×h×w Tensor; then, the feature map tensor X can be reshaped into a new feature map Y∈R through the axis transformation operation c×(K×h)×w , the height of the reshaped feature map is K×h and the width is w; this newly constructed feature map is directly input into the spatial channel dual excitation module to obtain the new feature map Z∈R selected by the attention mechanism c×(K×h)×w , the new feature map Z can be regarded as the set of local descriptor positions of all positions of the support set category, that is, Z = [Z (1,1) ,...,Z (i,j) ,...,Z (K×h,w) ], where Z (i,j) ∈R 1×1×CRepresents the local descriptor of the (i,j)th position, i∈{1,...,K×h}, j∈{1,...,w};

[0067] A single prototype representation for each class is obtained by summing the local descriptors at all locations:

[0068]

[0069] Among them, P n represents a single prototype representation of the nth class, Represents the projection of the (i, j) position of the k-th support image in the support set of the n-th class, y (i,j) represents the local descriptor at position (i, j) in the feature map after axis transformation. It can be seen that the prototypes obtained by a scene image support class are the weighted average of all local descriptors in the class. Specifically, the weight of each local descriptor in the spatial position is given by W s This essentially acts as a weighted sum of all channels in that spatial position. In this way, the importance of each local descriptor (i.e., spatial position) can be calculated by W s Then, for a support class scene image, this embodiment uses three different W s To obtain three different prototypes P s =[p1,p2,p3], where p1,p2,p3∈R c .

[0070] In the three prototype correction modules h δ Specifically, the obtained support set prototype P is corrected. s Take the average value as the representative prototype of the support set Using Q m =[q (1,1) ,...,q (i,j) ] indicates that the mth query image feature of the ResNet network is used to take the average value of the local descriptors of all positions of the m scene images of the query set Q as the representative vector of the query set.

[0071]

[0072] in, represents the i-th prototype of the k-th class in the support set, The local descriptor representing the (i, j)th position of the feature of the mth scene image in the query set. After that, the difference information I between the representative vector of the query set and the representative prototype of the support set can be obtained. bias :

[0073]

[0074] In order to reduce the distance between the support set and the query set, the difference information of the two sets is divided into the original difference information I bias Half of the image is used to correct the three prototypes supporting the scene image class, and the other half is used to correct the local descriptor of the query scene image. The formula of this process is expressed as:

[0075]

[0076] in, represents the corrected support set three prototypes, Represents the corrected query set feature of the mth query set.

[0077] like Figure 3 , by using the proposed three-prototype learning module and three prototype correction modules h δ , each query sample can be obtained Corrected local description and three corrected prototypes for each support class For simplicity Simplified to j∈{1,2,…,h×w}.

[0078] The classification process involves measuring and The similarity between them, and for this query sample Assign the most similar class in the support set. Specifically, this embodiment uses cosine similarity to first A local descriptor for a location in Find three prototypes of the nth class support set The closest prototype

[0079]

[0080] In this way, we can calculate the mth scene image in the support set and each supporting class, and sort them according to the most similar supporting class. Make classification predictions:

[0081]

[0082] Among them, O n,m represents the probability that the mth query image belongs to the nth support category. The loss function used to supervise the entire model consists of two parts. The first part is the cross entropy loss for classification:

[0083]

[0084] Among them, L cls represents the classification loss function of the setting, Represents the query set m correctly predicted to have label l m The overall similarity of , that is, the correct classification.

[0085] In addition, this embodiment also proposes a symmetric KL divergence loss function L for alleviating the intra-class distribution difference. intra :

[0086]

[0087] Among them, D KL (.) represents the KL divergence operation. They represent the predicted probability distribution of the i-th image of the n-th category in the support set and the probability distribution of the j-th image in the query set respectively. Finally, the loss function L of the three-prototype correction network few-sample remote sensing scene classification model based on local descriptors of the present invention is expressed as:

[0088] L=L cls +αL intra

[0089] Among them, α is a hyperparameter that can be adjusted.

[0090] This method uses an N-way K-shot task-based approach to train the local descriptor three-prototype correction network. During the training process, a fixed number of iterations is set, and a training task is selected in each iteration to perform forward propagation and calculate the loss, and then the model parameters are continuously optimized through the SGD algorithm. After each predetermined number of iterations, a model validation is performed, and several validation tasks are selected. The query set of each task is classified and predicted using the current model, and the average classification accuracy of all validation tasks is calculated as the result of this validation. Based on the validation results, the optimal model in the current iteration is saved.

[0091] By testing the optimal model of the current iteration: 1000 test tasks are randomly selected, and the query set images in each test task are classified and predicted using the best model obtained through training. The average classification accuracy of these 1000 test tasks is calculated as the result of the model test.

[0092] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A three-prototype correction network few-sample remote sensing scene classification method based on local descriptors, characterized in that: The specific steps include: Download the remote sensing dataset and divide the remote sensing scene image dataset into a training set and a validation set according to the categories of the remote sensing scene images in the remote sensing scene image dataset; Establish a local descriptor-based three-prototype correction network few-shot remote sensing scene classification model; The task-based meta-learning training strategy is used to meta-train the local descriptor-based three-prototype correction network few-sample remote sensing scene classification model. During the training, the local descriptor-based three-prototype correction network few-sample remote sensing scene classification model is meta-verified through the verification task, and the model with the best performance is saved. Classify remote sensing images using the best performing models.

2. According to claim 1, a local descriptor-based three-prototype correction network few-sample remote sensing scene classification method is characterized by: The construction of the remote sensing scene image dataset includes: obtaining remote sensing images of N categories, obtaining K remote sensing images of each category as the support set of the category, and M remote sensing images as the query set of the category.

3. According to claim 1, a local descriptor-based three-prototype correction network few-sample remote sensing scene classification method is characterized in that: The three-prototype correction network based on local descriptors is a few-shot remote sensing scene classification model, which includes a feature embedding learning module of local descriptors, a three-prototype learning module, a three-prototype correction module and a classifier module, wherein: A feature embedding learning module for local descriptors to obtain embedded representations of input remote sensing images; A three-prototype learning module is used to obtain three different prototype representations from the embedded representation of remote sensing images through three learnable weight matrices; The three-prototype correction module is used to correct the three different prototype representations obtained by the three-prototype learning module and correct the local description of each query sample; The classifier module is used to classify the query sample based on the local description and the similarity of each prototype of the support set image.

4. According to claim 3, a local descriptor-based three-prototype correction network few-sample remote sensing scene classification method is characterized in that: The process of obtaining three different prototype representations from the embedded representation of the remote sensing image through three learnable weight matrices includes: Reshape the tensor of K samples in the support set into a feature map Y; Input the new feature map into the spatial channel dual excitation module to obtain the feature map Z; A single prototype representation for each class is obtained by summing the local descriptors at all locations: Among them, P n represents a single prototype representation of the nth class, represents the projection of the (i, j) position of the k-th support image in the support set of the n-th class, y (i,j) Represents the local descriptor at position (i, j) in the feature map after axis transformation.

5. According to claim 3, a local descriptor-based three-prototype correction network few-sample remote sensing scene classification method is characterized in that: The process of obtaining a new feature map by the spatial channel dual excitation module includes: Combine the K samples of each category in the support set into a tensor; The embedded representation of remote sensing images is X=[x1,...,x h×w ] The projection tensor v is obtained by compression using convolution operation. Each element in the tensor v is a linear combination of the elements at that position of all channels. Different weight matrices are used in the compression process to obtain different prototypes of a class. The tensor v is converted into a weight between 0 and 1 using the Sigmoid activation function to weight the embedded representation of the remote sensing image to obtain the spatial position excitation; The feature map is regarded as a tensor of different channels. For each channel tensor, global average pooling is performed in space, that is, spatial dimension compression. The feature map of each channel is used as a local descriptor. All local descriptors form a vector z = [z1, ..., z C ]∈R 1×1×C , z c represents the local descriptor of the cth channel, C is the number of channels of the feature map; Use a weight matrix to map the vector z, then use the ReLU activation function to process it, and finally use another weight matrix to map it to get the dependency relationship between channels. The dependency is transformed by the Sigmoid activation function Convert to [0,1] and weight each channel tensor of the input to obtain channel position excitation; The spatial position excitation and the channel position excitation are added and fused to obtain a fused feature map.

6. The method for remote sensing scene classification based on a three-prototype correction network with few samples based on local descriptors according to claim 1 is characterized in that: Correcting the three different prototype representations obtained by the three-prototype learning module includes the following steps: The support set prototype P is obtained s Take the average value as the representative prototype of the support set The average of all the local descriptors of the M scene images in the query set Q is taken as the representative vector of the query set The difference information between the representative vector of the query set and the representative prototype of the support set is obtained, and half of the difference information is used to correct the three prototypes in the support set and the query feature set in the query set respectively.

7. The method for remote sensing scene classification based on a three-prototype correction network with few samples based on local descriptors according to claim 1 is characterized in that: The difference between the query set representative vector and the support set representative prototype is expressed as: Where M is the number of remote sensing images in the query set, and K is the number of remote sensing images in the support set; The local descriptor representing the (i, j)th position of the feature of the mth scene image in the query set; represents the i-th prototype of the k-th class in the support set.

8. The method for remote sensing scene classification based on three prototype correction networks with few samples based on local descriptors according to claim 3 is characterized in that: The classification process of the classifier module includes: Using cosine similarity first A local descriptor for a location in Find three prototypes of the nth class support set The closest prototype By calculating the cosine similarity between all local descriptors of the query image and the nearest prototype of the category and summing these similarities, the overall similarity s between the mth query image and the nth support category is obtained. n,m ; Calculate the mth scene image in the support set and each supporting class, and sort them according to the most similar supporting class. Make classification predictions.

9. The method for remote sensing scene classification based on a three-prototype correction network with few samples based on local descriptors according to claim 1 is characterized in that: When training the three-prototype correction network few-shot remote sensing scene classification model based on local descriptors, the loss function used includes the classification loss and the loss of the intra-class distribution difference of the link, which is expressed as: L=L cls +αL intra Among them, L is the total loss function; L cls is the classification loss function; N represents the number of remote sensing images in the support set, and M represents the number of remote sensing images in the query set; Indicates that the query set m is correctly predicted to have label l m The overall similarity of D KL (.) represents the KL divergence operation. represents the predicted probability distribution of the nth class and the i-th image in the support set S, represents the probability distribution of the j-th image in the query set Q.

Citation Information

Patent Citations

  • CT image few-sample segmentation method and system based on spatial position and prototype network

    CN112686850A

  • High-resolution SAR image scene classification method based on multi-prototype learning

    CN116580221A

  • Less-sample image classification method and device introducing local feature alignment and prototype correction mechanism

    CN117576471A

  • Small-sample remote sensing scene classification method based on subspace network

    CN117689958A