Remote Sensing Sample Classification Method Based on Transfer Learning and Bag of Visual Words

Through transfer learning and visual word package model, the problem of poor generalization in the classification of land use scenes of remote sensing images is solved, and a high-performance visual dictionary is generated, which improves the accuracy of remote sensing image classification.

CN115661504BActive Publication Date: 2025-07-04BEIJING DATA INTELLIGENCE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211019600.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2025-07-04
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

The existing remote sensing image land use scene classification method has poor generalization due to the lack of intermediate semantic descriptions, making it difficult to process scene images outside the training set, and supervised visual dictionary generation and calculation are complex and costly.

Method used

Transfer learning and visual word package models are adopted to extract local features through deep learning networks, and tag mapping relationships are constructed using transfer learning algorithms to generate visual dictionaries, and scene classification models are generated through clustering and low-dimensional statistics.

Benefits of technology

The accuracy of remote sensing image classification is improved, and the visual dictionary is generated through a supervised method, which enhances the performance and classification effect of the visual dictionary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661504B_ABST
    Figure CN115661504B_ABST
Patent Text Reader

Abstract

The present invention provides a remote sensing sample classification method based on transfer learning and visual word bags. The method includes: obtaining a sample set and a source domain set; constructing a combined model, including a deep learning network and a transfer learning algorithm; inputting the sample set and the source domain set into the combined model to obtain the labels of the sample set, and numbering the labels of the sample set based on the scene categories; performing secondary feature extraction on the sample set to obtain a feature set, and using a clustering method to cluster the feature set to obtain a visual dictionary; performing visual word mapping on each sample image in the sample set to generate a visual word bag distribution map, and performing low-dimensional statistics on the visual word bag distribution map to obtain a low-dimensional statistical representation of the visual word bag features; using the low-dimensional statistical representation of the visual word bag features and the label numbers as training data to train and obtain a scene classification model. The method of the present invention makes full use of the prior knowledge of the previous images, and the obtained scene classification model has a better classification effect and higher accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for classifying remote sensing samples, and particularly to a method for classifying remote sensing samples based on transfer learning and bag-of-visual-words, belonging to the field of remote sensing image classification. Background Art

[0002] With the development of remote sensing technology and the improvement of spatial and temporal resolutions, the data volume of remote sensing images, especially high-spatial-resolution remote sensing images, has increased rapidly, making the land use scenarios in the images contain various types of land cover types. In this case, using the method of manual visual interpretation to classify the land use scenarios of remote sensing images requires a large amount of time and workload, and limited experts cannot process the massive data in a timely manner. In view of the deficiencies of visual interpretation, using computer technology for automated and intelligent land use scenario classification has become a research hotspot in the current remote sensing field.

[0003] For the classification of land use scenarios in remote sensing images, traditional methods usually establish a land use scenario model for remote sensing images using low-level features such as color, texture, and shape, and use a classifier to deduce the high-level information of the scenario. However, the land use scenario classification method based on low-level feature description has poor generalization because of the lack of intermediate semantic image representation, and it is difficult to be used to process scenario images outside the training set. In order to overcome the gap between the low-level visual features and high-level semantics of remote sensing images, the method of semantic modeling and description of land use scenarios based on middle-level features has gradually received extensive attention. Especially in recent years, the bag-of-visual-words (BOVW) model has achieved great success in the application of image analysis and image classification, becoming a new and effective research idea for image content expression, and has achieved certain results in the classification of land use scenarios in remote sensing images. The advantage of the bag-of-visual-words model is that it does not need to analyze the specific target composition in the scenario image, but applies the overall statistical information of the image scenario, regards the quantified low-level features of the image as visual words, and expresses the image scenario content through the distribution of visual words in the image, providing basic data for the scenario classification of the image.

[0004] The image classification system based on the bag-of-visual-words model mainly consists of four parts, namely: image feature extraction, visual dictionary generation, visual vocabulary feature construction, and classifier. Among them, the quality of the visual dictionary directly affects the performance of the system. How to construct a visual dictionary with good discrimination and strong expressive ability has become the focus of image classification research based on the bag-of-visual-words model in recent years. According to whether information such as known class labels in the training set is used in the generation process of the visual dictionary, the visual dictionary generation methods can be divided into two categories: unsupervised visual dictionary generation and supervised visual dictionary generation. And due to the introduction of additional information to supervise the generation process of the visual dictionary, the performance of the visual dictionary generated by the supervised method is usually better than that of the unsupervised method. However, the supervised method for generating visual dictionaries also has disadvantages: one is the high computational complexity. The supervised method usually corresponds to solving an optimization problem and requires careful design of fast and effective iterative algorithms; the other is that a large number of labeled samples are required to complete the training process, but the manual labeling cost is very high and the quality of labeling is uneven. Summary of the Invention

[0005] Based on the above technical problems, the present invention proposes a remote sensing sample classification method based on transfer learning and bag-of-visual-words, which uses the transfer learning algorithm to transfer the rich label information and knowledge of the previous images to the current sample set, and uses the bag-of-visual-words model to perform scene classification on the sample set containing label semantic information, and the obtained scene classification model has a higher accuracy.

[0006] The present invention provides a remote sensing sample classification method based on transfer learning and bag-of-visual-words, and the method includes:

[0007] S1 Obtain a sample set, and at the same time screen the source domain set in the previous image library;

[0008] S2 Construct a combined model, including a deep learning network and a transfer learning algorithm;

[0009] S3 Input the sample set and the source domain set into the combined model, extract local features through the deep learning network, use the transfer learning algorithm to perform transfer on the source domain set based on the local features, and construct a mapping relationship between the local features and the labels, obtain the labels of the sample set, and number the labels of the sample set based on the scene categories;

[0010] S4 Perform secondary feature extraction on the sample set to obtain a feature set, use the clustering method to cluster the feature set, and the value of each clustering center and the corresponding label number form a visual word, and all the visual words form a visual dictionary;

[0011] S5 maps each sample image in the sample set to the visual words in the visual dictionary to generate a visual word bag distribution map, and performs low-dimensional statistics on the visual word bag distribution map to obtain a low-dimensional statistical representation of the visual word bag features;

[0012] S6 uses the low-dimensional statistical representation of the visual word bag features and the label numbers as training data, and cross-trains to obtain multiple classifiers based on different scenarios. All the classifiers form a scene classification model.

[0013] In a specific embodiment of the present invention, the method further includes:

[0014] Obtain a new sample, and according to steps S3 - S5, obtain a low-dimensional statistical representation of the visual word bag features of the new sample;

[0015] Input the low-dimensional statistical representation of the visual word bag features of the new sample into the scene classification model to obtain the scene category of the new sample.

[0016] In a specific embodiment of the present invention, the deep learning network is at least one of AlexNet, ResNet, VGGNet or other deep learning networks, and the transfer learning algorithm is a feature-based transfer learning algorithm.

[0017] In a specific embodiment of the present invention, step S4 includes:

[0018] Perform dense grid sampling on each sample in the sample set, and use the SIFT method to perform secondary feature extraction on each sampling area to obtain SIFT features;

[0019] Adopt E 2 Use the LSH clustering method to cluster the SIFT features. The value of each clustering center and the corresponding label number form a visual word, and all the visual words constitute a visual dictionary.

[0020] In a specific embodiment of the present invention, step S5 includes:

[0021] Calculate the Euclidean distance between the SIFT features of each sampling area and the feature values corresponding to each visual word in the visual dictionary;

[0022] Find the number of the visual word with the smallest Euclidean distance, and use it as the visual word mapping result of the corresponding sampling area to obtain the visual word bag distribution map of each sample image;

[0023] Perform LBP transformation on the visual word bag distribution map as an image to obtain the LBP histogram representation of the visual word bag features.

[0024] In a specific embodiment of the present invention, inputting the sample set and the source domain set into the combined model and extracting local features through the deep learning network includes:

[0025] Inputting the sample images in the sample set and the sample images in the source domain set into the deep learning network respectively. The deep learning network includes a pooling layer in the fifth layer and fully connected layers in the sixth and seventh layers.

[0026] Extracting the outputs of the two fully connected layers in the sixth and seventh layers to obtain two different high-level features.

[0027] Extracting the output of the pooling layer in the fifth layer and performing dimensionality reduction using the principal component analysis method to obtain a third high-level feature.

[0028] Fusing the three high-level features in a concatenated form to obtain a fused deep feature vector as the extracted local feature.

[0029] In a specific embodiment of the present invention, obtaining the labels of the sample set includes:

[0030] Constructing an objective function for the label y of the sample x in the prediction sample set D t of t ;

[0031] Formulating the objective function as an expected loss related to y and x t where the expected loss includes θ Y and a loss function related to θ Y and are model parameters with respect to the label space Y and the sample set D t ;

[0032] Formulating the loss function as a distance function between two models θ Y and ;

[0033] Adjusting the expected loss according to the distance function to obtain an objective function including the distance function.

[0034] Estimating the objective function including the distance function using the maximum a posteriori probability and approximate integration to obtain an estimated objective function.

[0035] When the estimated objective function meets the preset conditions, adjusting the estimated objective function to a difference function between the parameter distribution of the label space Y and the parameter distribution of the sample set x t ; ;

[0036] Obtain the labels of the sample set according to the objective function and the difference function.

[0037] In a specific embodiment of the present invention, the use of the transfer learning algorithm to transfer the source domain set based on local features includes:

[0038] Obtain The prior probability of the distribution in the feature space X t And The prior probability of the distribution in the feature space X t The prior probability of the distribution;

[0039] Determine The prior probability of the distribution in the feature space X t And The prior probability of the distribution in the feature space X t The distance between the prior probabilities of the distribution;

[0040] Transfer the source domain set based on local features according to the distance.

[0041] In a specific embodiment of the present invention, obtain The prior probability of the distribution in the feature space X t Includes:

[0042] Estimate the probability of translating the source domain set feature space X s To the sample set feature space X t The probability of any label y' in the source domain set feature space X s The probability of being the label y' and the probability of the label y' in the sample set feature space X t The probability of the distribution;

[0043] According to the probability of translating the source domain set feature space X s To the sample set feature space X t The probability of any label y' in the source domain set feature space X s The probability of the distribution, The probability of being the label y' and the probability of the label y' in the sample set feature space X t The probability of the distribution, obtain The prior probability of the distribution in the feature space X t

[0044] In a specific embodiment of the present invention, obtain The prior probability of the distribution in the feature space X t Includes:

[0045] Estimate any sample set x' t Translated to Xt the probability and estimate x' t the probability;

[0046] According to any of the sample sets x' t translate to X t the probability and estimate x' t the probability, to obtain the prior probability of distribution in the feature space X t in.

[0047] The beneficial effects of the present invention are as follows: A remote sensing sample classification method based on transfer learning and visual word bags is proposed. First, a combined model is constructed, which includes a deep learning network and a transfer learning algorithm. The local features of the sample set and the source domain set are extracted by using the deep learning network, and then the transfer learning algorithm is used to perform transfer based on the local features on the source domain set, making full use of the label information and prior knowledge of the previous images to obtain the labels of the sample set and the internal relationship between the features and the labels. Then, the visual word bag model is used to perform scene classification on the sample set containing label semantic information. Among them, by using rich label information, a visual dictionary is generated in a supervised manner, improving the performance of the visual dictionary. The classification effect of the scene classification model composed of multiple classifiers obtained is better and the accuracy is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 is the flowchart of the method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present invention.

[0051] The present invention provides a remote sensing sample classification method based on transfer learning and visual word bags, and the method includes:

[0052] S1 Obtain a sample set, and at the same time screen a source domain set in the previous image library;

[0053] S2 constructs a combined model, including a deep learning network and a transfer learning algorithm;

[0054] S4 inputs the sample set and the source domain set into the combined model, extracts local features through the deep learning network, uses the transfer learning algorithm to perform transfer on the source domain set based on the local features, constructs the mapping relationship between the local features and the labels, obtains the labels of the sample set, and numbers the labels of the sample set based on the scene categories;

[0055] S7 performs secondary feature extraction on the sample set to obtain a feature set, uses a clustering method to cluster the feature set, and the value of each clustering center and the corresponding label number form a visual word, and all the visual words form a visual dictionary;

[0056] S10 maps each sample image in the sample set to the visual words in the visual dictionary, generates a visual word bag distribution map, and performs low-dimensional statistics on the visual word bag distribution map to obtain a low-dimensional statistical representation of the visual word bag features;

[0057] S13 uses the low-dimensional statistical representation of the visual word bag features and the label numbers as training data, and cross-trains to obtain multiple classifiers based on different scenes, and all the classifiers form a scene classification model.

[0058] The remote sensing image is also a remote sensing picture. First, obtain the sample set of the remote sensing picture. The sample set contains multiple samples of remote sensing pictures. The sample set does not contain label information or only a small number of samples among them contain label information. At the same time, screen the source domain set from the previous image library. The previous image library contains a large number of remote sensing pictures with label information. The screening strategy is: manually select the remote sensing pictures in the previous image library whose similarity of the land cover type to the sample set reaches more than 70%, and form the source domain set.

[0059] Construct a combined model, including a deep learning network and a transfer learning algorithm, where the deep learning network is AlexNet, ResNet, VGGNet or other deep learning networks, and the transfer learning algorithm is a feature-based transfer learning algorithm.

[0060] This embodiment is a seven-layer deep learning network and has been trained. The first five layers are layer1, layer2, layer3, layer4, and layer5 respectively. Layer1, layer2, and layer5 each include a convolutional layer and a pooling layer. Layer3 and layer4 each have only one convolutional layer. The sixth and seventh layers are fully connected layers. Among them, the convolutional layer is used to extract feature maps, the pooling layer is used to compress the feature maps obtained by the convolutional layer, and the fully connected layer is used to convert the two-dimensional feature maps into one-dimensional vectors.

[0061] Input the sample set and the source domain set into the combined model, and extract local features through the deep learning network. The steps of extracting local features from the sample set and the source domain set are the same. Taking the extraction of local features from the sample set as an example:

[0062] Input the sample images in the sample set into the deep learning network, extract the outputs of the sixth and seventh fully connected layers to obtain two different high-level features. At the same time, extract the output of the fifth pooling layer and use the principal component analysis method for dimensionality reduction to obtain the third high-level feature. Concatenate the three high-level features to perform feature fusion, and finally obtain the fused deep feature vector as the extracted local feature. The three different high-level features extracted in this step cover complete image information and have strong discriminability. Fusing the three high-level features, the fused deep feature vector can further enhance the expressiveness and robustness of the features.

[0063] The transfer learning algorithm in this embodiment is a feature-based transfer learning algorithm. Two important concepts in transfer learning are domain and task. The domain includes the source domain and the target domain, and the task is the goal of transfer learning. Represent the source domain and the target domain as D s and D t . Use X to represent the feature space and Y to represent the label space. The data set in the source domain is the source domain set. In the source domain, any sample x s ∈D s can be represented as a feature vector where Similarly, the data set in the target domain is the sample set. In the target domain, any sample x t ∈D t can be represented as a feature vector where The samples in the source domain set contain rich prior knowledge, that is, the source domain set contains label information. Let where is 's label. When the sample set contains a small number of labeled samples, the sample set can be divided into two sets, namely the labeled set and the unlabeled set U. Among them, and usually m is very small; U contains k samples Construct an objective function, also known as the expected risk function, defined as T(y, x t ), which is the risk of classifying a sample x t into y. To predict the label of x t , one needs to find a label y that minimizes the objective function, i.e.:

[0064]

[0065] Formalize the objective function as an expected loss related to y and x t , i.e.:

[0066]

[0067] where r = 1 represents an event of "relevant", i.e., "y and x t are relevant" or "the label of x t should be y"; θ Y and are model parameters regarding Y and D t ; Θ Y and are two corresponding model spaces, including various possible generative models; it should be clear that θ Y is only related to y, and is only related to x t .

[0068] In the formula,[[]] is a loss function related to θ Y and . It can be formulated as the sum of the distances between two models θ Y and , i.e.:

[0069]

[0070] where is a distance function, for example, it can be the Kullback-Leibler distance. Substituting Equation (3) into Equation (2) gives:

[0071]

[0072] Use the maximum a posteriori probability to approximate the integral to calculate Equation (4):

[0073]

[0074] Among them

[0075]

[0076]

[0077] In formula (5), is the prior probability with respect to , which is used to describe the imbalance between classes. When it is assumed that there is no prior difference between classes, the objective function is:

[0078]

[0079] Among them represents the difference between two models and , and represent the parameter distribution of the label space Y and the parameter distribution of the sample set x t respectively. The parameter distributions of the two models are unified in the feature space X t of the sample set. Using the KL distance as the distance function, we get:

[0080]

[0081] Among them represents the prior probability of the distribution in the feature space X t , represents the prior probability of the distribution in the feature space X t .

[0082] Model estimation:

[0083]

[0084] Among them, p(X t |X s ) represents the probability used to describe translating X s to X t , p(X s |y′) represents the probability of the label y′ distributed in the feature space X s of the source domain set, represents the probability of estimating the label y′, p(X t |y′) is the probability of the label y′ distributed in the feature space X t of the sample set.

[0085] Use the translator φ to estimate p(X t |X s), the source domain set is used to estimate p(X s |y′), and the labeled set in the sample set is used to estimate p(X t |y′). It can be estimated as when y′ = y, otherwise λ is a balancing factor used to adjust the influence between the two feature spaces of the source domain set and the sample set.

[0086]

[0087] Among them, p(X t |x′ t ) can be estimated by a feature extractor, which is used to describe the probability of translating any sample set x′ t to X t ; when It can be estimated as otherwise which is used to describe the probability of estimating x′ t .

[0088] The labels of the sample set are obtained according to the translation learning algorithm: estimating a model based on Equation (10) For any x t ∈ U, estimating a model based on Equation (11) According to Equations (1) and (8), for any x t , predicting its label h t (x t ), learning features through translation learning, using the source domain set to help the sample set construct a higher feature space representation, transferring the knowledge of the source domain to the target domain. Through the transfer at the feature level, the transferable data surface is greatly expanded. This method does not require the feature spaces of the source domain and the target domain to have a common set. Even if the feature spaces of the source domain dataset and the target domain dataset are completely different, feature transfer learning can still be carried out through this method, thereby improving the adaptability of data transfer.

[0089] While obtaining the label information of the sample set, construct and save the mapping relationship between the local features and the labels. That is, construct a mapping relationship between the local features and the labels of a sample, and its mapping path is expressed as:[[]] Among them, y(x t ) represents the final label of the sample x t , that is, there is also a mapping relationship between the local features of the sample x t and its label:[[]] And form a mapping table of the mapping relationships of the sample set.

[0090] Perform dense grid sampling on each sample in the sample set, and use SIFT to extract features again for each sampling area to obtain a SIFT feature set. SIFT features can effectively describe the local area information of an image, are invariant to image rotation, brightness changes, and scale changes, and are also highly robust to affine changes, perspective changes, and noise.

[0091] Local features can characterize the underlying visual characteristics of an image and are widely used in image content analysis. However, most image local features are located in a high-dimensional space, which is not convenient for storage and subsequent calculations. In addition, high-dimensional vectors usually also face problems such as sparsity and noise, known as the "curse of dimensionality", which cause the performance of algorithms that perform well in a low-dimensional space to deteriorate sharply in a high-dimensional space. Therefore, it is necessary to map the high-dimensional local features of an image to a low-dimensional space for easy storage, indexing, and calculation. Mapping a large number of local features to a low-dimensional space to obtain the corresponding codes for the local features, and these codes are called visual words. All visual words constitute a visual dictionary.

[0092] In the embodiment of the present invention, the steps for constructing a visual dictionary are as follows:

[0093] (1) Number the labels of the sample set, record the local features in the form of feature values, and update the mapping table, which shows a triple relationship group of feature value - label - label number in the mapping table.

[0094] (2) Use the 2 LSH algorithm and the clustering ensemble algorithm to cluster the SIFT feature set.

[0095] The 2 LSH is a special case of random mapping. After mapping high-dimensional vectors to a low-dimensional space, it also completes the task of dimensionality reduction. At the same time, points in the same bucket are more similar than points in different buckets. Therefore, if the data points are grouped according to the bucket flag, the purpose of clustering the data points can also be achieved. When used for clustering, when the scale of the data set changes, it is not necessary to re-cluster the entire data set, and only the changed data points need to be calculated. That is to say, it can conveniently play the role of dynamic clustering. Coupled with its time advantage, it can be better used for incremental large data set clustering. The 2 basic principle of LSH is to use a locality-sensitive hashing function to map high-dimensional vectors to a low-dimensional space while preserving the distance.

[0096] This embodiment uses a clustering method based on the 2 LSH algorithm and the clustering ensemble algorithm to cluster the SIFT feature set. The basis of the clustering method is the 2 hashing function and clustering ensemble of LSH.

[0097] The2 The hash function of LSH is based on the p-stable distribution function, where p ∈ (0, 2]. Its hash function is defined as follows:

[0098]

[0099] Among them, v is the d-dimensional original data, w is the width value, a is an n-dimensional random vector generated by the p-stable distribution function, the inner product (a·v) randomly maps the SIFT feature points, and b adds an offset to the result after mapping. w needs to be adjusted manually, and b is randomly generated using a uniform distribution, and the range of its uniform distribution is [0, w].

[0100] The SIFT feature set, after being mapped by the E 2 LSH algorithm, can maintain the relative distance, class boundary, and angle.

[0101] Cluster ensemble is a part of ensemble learning and plays an important role in improving the clustering performance. This is because no single clustering algorithm can discover classes with different shapes and distributions, and the internal structures of different datasets are different. Cluster ensemble combines multiple base clustering schemes of the same dataset to obtain a unified clustering scheme, and the clustering quality of this scheme is better than that of most base clustering schemes.

[0102] The steps of clustering the SIFT by the described clustering method include: determining the number of clusters, generating the E 2 LSH clustering scheme, and cluster scheme integration.

[0103] First, determine the number of clusters k for the SIFT feature set I through three validity indicators SI, DB, and Wint.

[0104] Then generate an n×k random matrix A = (A1, A2,..., A n ) T and b, w, where A i is a k-dimensional vector and n = |I|.

[0105] Randomly map all SIFT feature points, and the mapping result of each feature point x i is the k-dimensional bucket label B i . Then the bucket labels of all feature points are an n×k-dimensional matrix B. Among them, each column of B is a mapping of the entire feature set.

[0106] After that, use the PAM algorithm to cluster each column of B into k classes, and the obtained result is the k partitions of the feature set.

[0107] Finally, use the cluster ensemble method to integrate the k base partitions and obtain the final partition.

[0108] After the clustering is completed, obtain the clustering centers of k classes from the final clustering result.

[0109] (3) Compare the values of the obtained k clustering centers with the mapping table to obtain the corresponding label numbers, and form a visual word by combining the value of each clustering center and its label number. All the visual words constitute a visual dictionary.

[0110] The clustering method adopted in the embodiment of the present invention combines the characteristics of the E 2 LSH algorithm and integrated clustering, and can efficiently process high-dimensional data. When performing high-dimensional space partitioning, it can effectively weaken problems such as visual word synonymy and ambiguity existing when generating a visual dictionary. Moreover, this embodiment is a supervised generation of a visual dictionary. When constructing the visual dictionary, the label information is fully utilized, and the obtained visual dictionary has stronger discrimination ability and expression ability, and can better describe remote sensing images.

[0111] After the visual dictionary is constructed, calculate the Euclidean distance between the SIFT feature of each sampling area and the feature value corresponding to each visual word in the visual dictionary, find the number of the visual word with the smallest Euclidean distance, and use it as the visual word mapping result of the corresponding sampling area, that is, each sampling area of each sample is assigned a visual word number, and finally obtain the visual word bag distribution map of each sample image.

[0112] Perform low-dimensional statistics on the visual word bag distribution map as an image to obtain a low-dimensional statistical representation of the visual word bag features. In this embodiment, the low-dimensional statistics is LBP histogram statistics, that is, perform LBP histogram transformation on the visual word bag distribution map as an image to obtain the LBP histogram representation of the visual word bag features. By using the visual vocabulary features to express the image, the complexity of remote sensing image classification can be reduced, and performing LBP histogram transformation on the visual word bag distribution map further improves the integrity of the visual vocabulary features expressing the image content and enhances the accuracy of remote sensing image classification.

[0113] Use the low-dimensional statistical representation of the visual word bag features and the label numbers as training data, and cross-train to obtain multiple classifiers based on different scenarios. In this embodiment, the SVM algorithm is used for cross-training. Specifically: each label number represents a class of scenarios, each class of scenarios corresponds to a subset of samples, and an SVM classifier is learned and generated between every two subsets of samples to obtain multiple SVM classifiers based on different scenarios. Finally, all the SVM classifiers are used as the obtained scenario classification model.

[0114] Obtain a new sample, and according to the above steps, obtain the low-dimensional statistical representation of the visual word bag features of the new sample, and input it into the scene classification model. Adopt a voting mechanism to determine the category of the new sample: if an SVM classifier determines that the low-dimensional statistical representation of the visual word bag features of the new sample belongs to the x-th category, it means that the x-th category has obtained one vote. Finally, the scene category with the most votes is the category to which the new sample image belongs.

[0115] The beneficial effects of the present invention are as follows: A remote sensing sample classification method based on transfer learning and visual word bag is proposed. First, a combined model is constructed, which includes a deep learning network and a transfer learning algorithm. The local features of the sample set and the source domain set are extracted by using the deep learning network, and then the transfer learning algorithm is used to perform transfer based on the local features on the source domain set, making full use of the label information and prior knowledge of the previous images to obtain the labels of the sample set and the internal relationship between the features and the labels. After that, the visual word bag model is used to perform scene classification on the sample set containing label semantic information. Among them, by using rich label information, a visual dictionary is generated in a supervised manner, improving the performance of the visual dictionary. The classification effect of the scene classification model composed of multiple classifiers obtained is better and the accuracy is higher.

Claims

1. A remote sensing sample classification method based on transfer learning and visual word bags, characterized in that The method includes: S1 Obtain a sample set, and simultaneously screen a source domain set from a pre - existing image database; S2 Construct a combined model, including a deep learning network and a transfer learning algorithm; S3 Input the sample set and the source domain set into the combined model, extract local features through the deep learning network, use the transfer learning algorithm to transfer the source domain set based on the local features, and construct a mapping relationship between the local features and labels to obtain the labels of the sample set, and number the labels of the sample set based on scene categories; S4 Perform secondary feature extraction on the sample set to obtain a feature set, and use a clustering method to cluster the feature set. The value of each clustering center and the corresponding label number form a visual word, and all the visual words form a visual dictionary; S5 Map each sample image in the sample set to the visual words in the visual dictionary to generate a visual word bag distribution map, and perform low - dimensional statistics on the visual word bag distribution map to obtain a low - dimensional statistical representation of the visual word bag features; S6 Use the low - dimensional statistical representation of the visual word bag features and the label numbers as training data, and cross - train to obtain multiple classifiers based on different scenes. All the classifiers form a scene classification model.

2. The remote sensing sample classification method based on transfer learning and visual word bag according to claim 1, characterized in that, The method further includes: Obtain new samples, and according to steps S3 - S5, obtain the low - dimensional statistical representation of the visual word bag features of the new samples; Input the low - dimensional statistical representation of the visual word bag features of the new samples into the scene classification model to obtain the scene categories of the new samples.

3. The remote sensing sample classification method based on transfer learning and visual word bag according to claim 1, characterized in that The deep learning network is at least one of AlexNet, ResNet, VGGNet or other deep learning networks, and the transfer learning algorithm is a feature - based transfer learning algorithm.

4. The remote sensing sample classification method based on transfer learning and visual word bag according to claim 1, characterized in that, Step S4 includes: Perform dense grid sampling on each sample in the sample set, and use the SIFT method to perform secondary feature extraction on each sampling area to obtain SIFT features; Using E 2 The LSH clustering method is used to cluster the SIFT features. The value of each clustering center obtained and the corresponding label number form a visual word, and all the visual words constitute a visual dictionary.

5. The remote sensing sample classification method based on transfer learning and visual word bag according to claim 4, characterized in that, Step S5 includes: Calculate the Euclidean distance between the SIFT features of each sampling area and the feature values corresponding to each visual word in the visual dictionary; Find the number of the visual word with the smallest Euclidean distance, and use it as the visual word mapping result of the corresponding sampling area to obtain the visual word bag distribution map of each sample image; Perform LBP transformation on the visual word bag distribution map as an image to obtain the LBP histogram representation of the visual word bag features.

6. The remote sensing sample classification method based on transfer learning and visual word bag according to any one of claims 1 to 5, characterized in that The inputting the sample set and the source domain set into the combined model and extracting local features through the deep learning network includes: Input the sample images in the sample set and the sample images in the source domain set into the deep learning network respectively. The deep learning network includes a pooling layer in the fifth layer and fully - connected layers in the sixth and seventh layers Extract the outputs of the sixth and seventh fully - connected layers to obtain two different high - level features; Extract the output of the fifth - layer pooling layer, and perform dimensionality reduction using the principal component analysis method to obtain a third high - level feature; Fuse the three high - level features in a concatenated form to obtain a fused deep feature vector as the extracted local feature.

7. The remote sensing sample classification method based on transfer learning and visual word bag according to any one of claims 1 to 5, characterized in that, The obtaining the labels of the sample set includes: Construct the prediction sample set D t sample x t objective function of label y Formalize the objective function as an expected loss related to y and x t The expected loss includes θ Y and The related loss function, θ Y and are model parameters with respect to the label space Y and the sample set D t ; Formulate the loss function as a distance function between two models θ Y and ; Adjust the expected loss according to the distance function to obtain an objective function including the distance function; Estimate the objective function including the distance function by using maximum a posteriori probability re-approximate integration to obtain an estimated objective function; When the estimated objective function satisfies a preset condition, adjust the estimated objective function to the parameter distribution of the label space Y and the sample set x t of the parameter distribution between the difference function; Obtain the label of the sample set according to the objective function and the difference function.

8. The remote sensing sample classification method based on transfer learning and visual word bag according to claim 7, characterized in that The migration of the source domain set based on local features by using the transfer learning algorithm includes: Obtain the prior probability distributed in the feature space X t and the prior probability distributed in the feature space X t ; Determine the prior probability distributed in the feature space X t and the distance between the prior probabilities distributed in the feature space X t ; Migrate the source domain set based on local features according to the distance.

9. The remote sensing sample classification method based on transfer learning and visual word bag according to claim 8, characterized in that Obtain the prior probability distributed in the feature space X t includes: Estimate the feature space X of the source domain set s Translate to the feature space X of the sample set t The probability, the probability of any label y' in the feature space X of the source domain set s The probability of the distribution in, The probability for the label y' and the probability of the distribution of the label y' in the feature space X of the sample set t The probability of the distribution in; According to the source domain set feature space X s Translate to the sample set feature space X t The probability, the probability of any label y' distributed in the source domain set feature space X s The probability in, The probability of the label y' and the probability of the label y' distributed in the sample set feature space X t Get The prior probability of the distribution in the feature space X t ​ 10. The remote sensing sample classification method based on transfer learning and visual word bag according to claim 8, characterized in that, Obtain the prior probability distributed in the feature space X t including: Estimate the probability of any sample set x′ t translated to X t and estimate the probability of x′ t ; According to any of the sample sets x′ t Translate to X t The probability of and Estimate x′ t The probability of, obtain The prior probability of the distribution in the feature space X t Among them.

Citation Information

Patent Citations

  • Remote sensing image land utilization scene classification method based on two-dimension wavelet decomposition and visual sense bag-of-word model

    CN103413142A

  • Animal behavior identification method and apparatus based on transfer learning

    CN106056043A