Remote sensing image fine-grained classification method and device based on small sample continuous learning
By training a classification model using a few-shot continuous learning method, extracting remote sensing image features and filtering background features, the problems of insufficient remote sensing image samples and small differences between categories are solved, thus improving the accuracy of fine-grained classification.
Patent Information
- Application Number
- CN202310114241.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-29
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-01-29
AI Technical Summary
The limited number of remote sensing image samples leads to overfitting issues, and the small differences between fine-grained categories result in low accuracy in fine-grained classification of remote sensing images.
A classification model is trained using a few-shot continuous learning method to extract image features and determine the target response region. Background features are filtered out through feature selection conditions to perform fine-grained classification.
It improves the fine-grained classification accuracy of remote sensing images, avoids interference from background features, and expands the classification range of the classification model.
Smart Images

Figure CN115937691B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of computer vision, and particularly relates to a small sample continual learning based remote sensing image fine-grained classification method, device and storage medium. BACKGROUND
[0002] With the development of earth observation technology, new ground object categories are constantly emerging in massive remote sensing data. In the related technology, a trained deep learning model is usually used to perform fine-grained classification on remote sensing images.
[0003] In the process of implementing the concept of the present disclosure, the inventors have found that, due to the small number of remote sensing image samples, overfitting occurs in the process of small sample training, and due to the small difference between fine-grained categories of classification objects in remote sensing images, there is a problem of low fine-grained classification accuracy when performing fine-grained classification on remote sensing images. SUMMARY
[0004] In view of the above problems, the present disclosure provides a small sample continual learning based remote sensing image fine-grained classification method, device and storage medium.
[0005] According to a first aspect of the present disclosure, a small sample continual learning based remote sensing image fine-grained classification method is provided, comprising: extracting image features of a first remote sensing image by using a trained classification model to obtain first remote sensing image features, wherein the classification model is trained by using a small sample continual learning method; obtaining a target response region in the first remote sensing image features according to the first remote sensing image features, wherein the target response region represents a feature response region corresponding to a target object in the first remote sensing image features, and the feature response region contains features of the target object and background features; obtaining a feature screening condition according to the target response region; performing a feature filtering operation on the first remote sensing image features based on the feature screening condition to obtain a plurality of target image features corresponding to the target object; and performing classification processing on the plurality of target image features to obtain a classification result corresponding to the first remote sensing image.
[0006] According to an embodiment of the present disclosure, obtaining a target response region in the first remote sensing image features according to the first remote sensing image features comprises: obtaining a correlation result according to the association relationship between the plurality of target image features; and dividing the feature response region according to the correlation result to obtain the target response region.
[0007] According to an embodiment of the present disclosure, obtaining a feature screening condition according to the target response region comprises: determining a plurality of feature values according to the plurality of target image features; and obtaining the feature screening condition according to the average value of the plurality of feature values.
[0008] According to an embodiment of the present disclosure, a training method of a classification model comprises: extracting image features of a sample remote sensing image to obtain a first feature dataset for training a preset model; training the preset model using the first feature dataset to obtain an intermediate model and a first classification result; fixing parameters of at least one target residual block in the intermediate model; training the intermediate model using a second feature dataset to obtain a classification model and a second classification result, wherein the second feature dataset is obtained by sampling the first feature dataset.
[0009] According to an embodiment of the present disclosure, the training of the intermediate model using the second feature dataset to obtain the classification model comprises: dividing the second feature dataset to obtain a third feature dataset, wherein the third feature dataset comprises at least two feature datasets of different categories; and training the intermediate model using the third feature dataset to obtain the classification model.
[0010] According to an embodiment of the present disclosure, the training of the preset model using the first feature dataset to obtain the intermediate model comprises: training the preset model using the first feature dataset and a cross-entropy loss function through a gradient back propagation method to obtain first parameter information; and obtaining the intermediate model based on the first parameter information.
[0011] According to an embodiment of the present disclosure, the classification processing of the plurality of target image features to obtain the classification result corresponding to the first remote sensing image comprises: determining a similarity between the plurality of image features and the second classification result using a cosine similarity classification function; and classifying the plurality of target image features according to the similarity to obtain the classification result.
[0012] According to an embodiment of the present disclosure, the training of the intermediate model using the third feature dataset to obtain the classification model comprises: training the intermediate model using the third feature dataset to obtain second parameter information; and obtaining the classification model according to the second parameter information.
[0013] A second aspect of the present disclosure provides an electronic device, comprising: one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above method.
[0014] A third aspect of the present disclosure further provides a computer-readable storage medium having stored executable instructions, which, when executed by a processor, cause the processor to execute the above method.
[0015] The small sample based continuous learning remote sensing image fine-grained classification method, device and storage medium provided by the present disclosure can determine a target response region corresponding to a target object, and then screen features corresponding to the target response region according to a feature screening condition, so as to screen the features of the target object and background features through the feature screening condition, perform a feature filtering operation on the screened background features, avoid the interference of the background features in the process of fine-grained classification of the target object, obtain a plurality of target image features corresponding to the target object, and then perform classification processing on the plurality of target image features without background features, thereby improving the fine-grained classification accuracy of the remote sensing image. BRIEF DESCRIPTION OF DRAWINGS
[0016] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description of embodiments of the present disclosure taken in conjunction with the accompanying drawings, in which:
[0017] Figure 1 An application scenario diagram of the remote sensing image fine-grained classification method according to an embodiment of the present disclosure is schematically shown;
[0018] Figure 2 A flowchart of the remote sensing image fine-grained classification method according to an embodiment of the present disclosure is schematically shown;
[0019] Figure 3 A flowchart of the training method of the classification model according to an embodiment of the present disclosure is schematically shown;
[0020] Figure 4 A schematic diagram of the training intermediate model according to an embodiment of the present disclosure is schematically shown;
[0021] Figure 5 A structural block diagram of the remote sensing image fine-grained classification device according to an embodiment of the present disclosure is schematically shown; and
[0022] Figure 6 A block diagram of an electronic device suitable for implementing the remote sensing image fine-grained classification method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0023] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. It should be understood, however, that the description is merely exemplary and is not intended to limit the scope of the present disclosure. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it will be apparent to one skilled in the art that one or more embodiments can be practiced without these specific details. In addition, in the following description, descriptions of well-known structures and techniques have been omitted to avoid unnecessarily obscuring the concept of the present disclosure.
[0024] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the terms "comprises", "comprising", "includes", "including" and the like are specifically intended to be open-ended and to mean that other features, steps, operations, and / or components can be added.
[0025] All terms used herein, including technical and scientific terms, have the meanings commonly understood by one of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning that is consistent with the context of the specification, and should not be interpreted in an idealized or overly formal way.
[0026] In the case of using expressions similar to "at least one of A, B, and C, etc.", it should generally be interpreted that the meaning of the expression is at least one of A, B, and C, etc. (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C, etc.).
[0027] In the technical solutions of the disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information comply with relevant laws and regulations, necessary security measures are taken, and public order and good customs are not violated.
[0028] In the technical solutions of the disclosure, the acquisition, collection, storage, use, processing, transmission, provision, disclosure, and application of data comply with relevant laws and regulations, necessary security measures are taken, and public order and good customs are not violated.
[0029] At present, most deep learning models can only be trained once using all training data, and when new data arrives, a new model can only be retrained. Therefore, the model can continuously learn new categories or new tasks through continuous learning.
[0030] Since the number of newly acquired ground object category samples is not sufficient, the model needs to have small sample learning ability while continuously learning, that is, small sample continuous learning. However, since the new category only contains a small amount of training samples, training the model through small sample continuous learning may have the problem of overfitting.
[0031] In addition, since the differences between fine-grained categories of remote sensing images are small and the background interference is large, more computer resources are needed in the case of classifying target information, and common small sample continuous learning methods are not suitable for being directly used in fine-grained classification of remote sensing images. In addition, when the model continuously learns new categories, it may affect the previously learned categories, making it difficult to meet the demand for the number of learned categories.
[0032] Therefore, the embodiment of the present disclosure provides a remote sensing image fine-grained classification method based on small sample continual learning, comprising:
[0033] Using the trained classification model, the image features of the first remote sensing image are extracted to obtain the first remote sensing image, wherein the classification model is trained by using the small sample continual learning method;
[0034] According to the first remote sensing image features, the target response area in the first remote sensing image features is obtained, wherein the target response area represents the feature response area corresponding to the target object in the first remote sensing image features, and the feature response area contains the features and background features of the target object;
[0035] According to the target response area, the feature screening condition is obtained;
[0036] Based on the feature screening condition, the feature filtering operation is performed on the first remote sensing image features to obtain a plurality of target image features corresponding to the target object;
[0037] The plurality of target image features are classified to obtain the classification result corresponding to the first remote sensing image.
[0038] Figure 1 The application scenario diagram of the remote sensing image fine-grained classification method according to the embodiment of the present disclosure is schematically shown.
[0039] As shown in Figure 1 The application scenario 100 according to the embodiment can include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is used as a medium to provide a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0040] The user can use the first terminal device 101, the second terminal device 102, the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0041] The first terminal device 101, the second terminal device 102, the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers and desktop computers, etc.
[0042] The server 105 can be a server providing various services, for example, a background management server (only for example) providing support for a website browsed by a user using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server can perform analysis and the like on received user requests and the like, and feed back the processing results (for example, a web page, information, or data, or the like obtained or generated according to a user request) to the terminal device.
[0043] It should be noted that the remote sensing image fine-grained classification method provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the remote sensing image fine-grained classification apparatus provided in the embodiments of the present disclosure can generally be arranged in the server 105. The remote sensing image fine-grained classification method provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Accordingly, the remote sensing image fine-grained classification apparatus provided in the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0044] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above-mentioned scenario is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks, and servers.
[0045] The remote sensing image fine-grained classification method of the embodiments of the present disclosure will be described in detail below based on the scenario described above. Figure 1 Figures 2 to 4 The remote sensing image fine-grained classification method of the embodiments of the present disclosure will be described in detail below based on the scenario described above.
[0046] Figure 2 A flowchart of the remote sensing image fine-grained classification method according to the embodiments of the present disclosure is schematically shown.
[0047] As shown in Figure 2 The remote sensing image fine-grained classification method of this embodiment includes operation S210 to operation S250.
[0048] In operation S210, the image features of the first remote sensing image are extracted using the trained classification model, to obtain the first remote sensing image features, wherein the classification model is trained using a small sample continual learning method.
[0049] According to an embodiment of the present disclosure, the classification model can be a model for fine-grained classification of a target object in the first remote sensing image. For example, an airplane can be included in the first remote sensing image, and the fine-grained classification of the airplane in the first remote sensing image can be performed by using the trained classification model, and a fine-grained classification result corresponding to the airplane can be obtained, and the classification result can include class information of each part of the airplane in the remote sensing image, and the class information can include information corresponding to a fuselage part, information corresponding to a wing part, and the like.
[0050] According to an embodiment of the present disclosure, for example, the first remote sensing image feature can be an image feature of the first remote sensing image, and can include target object and background features. The target object can represent object features that need to be classified, and the background features can represent features that do not need to be classified.
[0051] According to an embodiment of the present disclosure, since the model is trained by using a conventional small sample method, there can be a problem of overfitting, and therefore, the model can be trained by using a small sample continual learning method. For example, the training process can include: first training a preset model by using a sample set to obtain an intermediate model, and after completing the training process, performing small sample task sampling on the sample set; fixing part of the parameters of the intermediate model to reduce the influence of the classification ability of the intermediate model obtained in the previous training process on the subsequent training process; and then performing task continual training on the intermediate model by using the sampled sample set, which can maximize the classification range of the trained classification model and solve the problem of overfitting.
[0052] In operation S220, a target response region in the first remote sensing image feature is obtained according to the first remote sensing image feature, wherein the target response region represents a feature response region corresponding to the target object in the first remote sensing image feature, and the feature response region includes the feature of the target object and the background feature.
[0053] According to an embodiment of the present disclosure, the feature response region can be a region where the target object is located in the first remote sensing image feature.
[0054] According to an embodiment of the present disclosure, for example, the feature response region where the target object is located in the first remote sensing image feature can be determined as the target response region. The background feature in the target response region can be filtered to improve the accuracy of fine-grained classification of the target object.
[0055] In operation S230, a feature screening condition is obtained according to the target response region.
[0056] According to an embodiment of the present disclosure, the feature screening condition can be obtained according to the target object in the target response region. For example, the average value of the feature value corresponding to the feature of the target object can be determined as the feature screening condition, and then the background features are filtered according to the feature screening condition, and the features of the target object are retained. On this basis, the target object is further classified in a fine-grained manner, which can improve the classification accuracy of the target object.
[0057] In operation S240, a feature filtering operation is performed on the first remote sensing image features based on the feature screening condition, and a plurality of target image features corresponding to the target object are obtained.
[0058] According to an embodiment of the present disclosure, the feature filtering operation can be used to filter the background features. By filtering the background features based on the feature screening condition, a plurality of target image features corresponding to the target object can be obtained.
[0059] According to an embodiment of the present disclosure, the target image features can represent the features of the target object.
[0060] In operation S250, the plurality of target image features are classified to obtain a classification result corresponding to the first remote sensing image.
[0061] According to an embodiment of the present disclosure, the target image features can be classified to classify the target object, and a classification result of the target object in the first remote sensing image is obtained.
[0062] According to an embodiment of the present disclosure, for example, the classification result can include category information corresponding to the target object.
[0063] According to an embodiment of the present disclosure, since the target response region corresponding to the target object is determined, and the feature screening condition corresponding to the target response region is obtained, the features of the target object and the background features can be screened by the feature screening condition, and the feature filtering operation is performed on the screened background features to avoid the interference of the background features in the process of classifying the target object in a fine-grained manner, a plurality of target image features corresponding to the target object are obtained, and the plurality of target image features do not contain the background features. Classification processing is performed on the plurality of target image features, which improves the fine-grained classification accuracy of the remote sensing image.
[0064] According to an embodiment of the present disclosure, the target response region in the first remote sensing image features is obtained according to the first remote sensing image features, including:
[0065] According to the correlation relationship between the plurality of target image features, a correlation result is obtained.
[0066] According to the correlation result, the target response region is obtained by dividing the feature response region.
[0067] According to an embodiment of the present disclosure, the association relationship can be an association relationship between different target components on a target object corresponding to a plurality of target image features in a target image. For example, the target object can be an airplane, and the target components can include a fuselage and wings, etc. The relevance score can be determined according to the association relationship, and the relevance score of a plurality of target image features belonging to the same target component is higher, and the relevance score of a plurality of target image features not belonging to the same target component is lower. For example: the score of the association relationship between a plurality of target image features representing the fuselage part can be higher, and the score of the association relationship between a plurality of target image features representing the wing part can be higher. However, the score of the association relationship between the target image features representing the fuselage part and the target image features representing the wing part can be lower. The relevance result of the association relationship between a plurality of target image features can be obtained according to the score of the association relationship.
[0068] According to an embodiment of the present disclosure, for example, the classification model can have a plurality of output channels, and each output channel can output a plurality of corresponding image features. The relevance of the image features output from the plurality of output channels to each other can be determined to obtain a relevance result.
[0069] According to an embodiment of the present disclosure, for example, the relevance result can include the relevance of the image features output from the plurality of output channels to each other. According to the relevance result, the image features whose relevance meets a preset condition can be integrated to divide the feature response region. After integration, the region corresponding to the image features whose relevance meets the preset condition, i.e., the target response region, can be obtained.
[0070] According to an embodiment of the present disclosure, since the relevance result is obtained according to the association relationship between a plurality of target image features, and the target response region is obtained by dividing the feature response region according to the relevance result, the determination of the target response region where the target object is located is realized, and thus the accuracy of the fine-grained classification of the target object can be improved.
[0071] According to an embodiment of the present disclosure, the feature screening condition is obtained according to the target response region, including:
[0072] A plurality of feature values are determined according to a plurality of target image features.
[0073] The feature screening condition is obtained according to the average of a plurality of feature values.
[0074] According to an embodiment of the present disclosure, for example, a feature value corresponding to each of the plurality of target image features can be determined, and an average of the plurality of determined feature values can be taken as the feature screening condition. The background features in the target response region can be filtered according to the average, so as to improve the accuracy of the target response region and obtain the plurality of target image features. The classification accuracy can be improved by classifying the plurality of target image features after filtering the background features.
[0075] According to an embodiment of the present disclosure, for example, the plurality of target image features can correspond to features of the wing, feature values corresponding to the features of the wing can be determined, and an average of the plurality of determined feature values can be determined. Through the average, features in the target image features that do not belong to the wing part can be filtered to extract the features of the wing.
[0076] According to an embodiment of the present disclosure, since the plurality of feature values are determined according to the plurality of target image features, and the feature screening condition is obtained according to the average of the plurality of feature values, the background features in the target response region can be filtered, and the target image features are retained, so that the background features do not affect the fine-grained classification of the remote sensing image, and the accuracy of the fine-grained classification of the remote sensing image is improved.
[0077] Figure 3 A flowchart of a training method of a classification model according to an embodiment of the present disclosure is schematically shown.
[0078] As shown in Figure 3 The training method of the classification model of this embodiment includes operation S310 to operation S340.
[0079] In operation S310, image features of a sample remote sensing image are extracted to obtain a first feature data set for training a preset model.
[0080] According to an embodiment of the present disclosure, the sample remote sensing image can be an image sample used to train the classification model.
[0081] According to an embodiment of the present disclosure, the first feature data set can include sample data of all categories of target objects required for training.
[0082] According to an embodiment of the present disclosure, the preset model can be a ResNet18 (a kind of convolutional neural network model) backbone network model to be trained.
[0083] In operation S320, the first feature data set is used to train the preset model to obtain an intermediate model and a first classification result.
[0084] According to an embodiment of the present disclosure, the intermediate model can be a model trained by the first feature data set. The classification model can be obtained by training the intermediate model.
[0085] According to an embodiment of the present disclosure, the first classification result can be a classification result of the preset model classifying the target object in the first feature data set.
[0086] In operation S330, parameters of at least one target residual block in the intermediate model are fixed.
[0087] According to an embodiment of the present disclosure, for example, four residual blocks can be included in the intermediate model, parameters of one to three residual blocks of the four residual blocks can be fixed, and the intermediate model can be trained again so that the intermediate model retains the parameters of the fixed residual blocks and can continue to train the categories through the residual blocks with unfixed parameters, thereby reducing the influence of subsequent training on the classification ability of previous training.
[0088] According to an embodiment of the present disclosure, for example, the target residual blocks can be the first three residual blocks of the intermediate model through which the feature data set passes in the data transmission process. By fixing the parameters of the first three residual blocks, the intermediate model can avoid losing the previously trained categories and storing the subsequently trained categories to the greatest extent.
[0089] In operation S340, the intermediate model is trained using a second feature data set to obtain a classification model and a second classification result, wherein the second feature data set is obtained by sampling the first feature data set.
[0090] According to an embodiment of the present disclosure, for example, in the case of fixing the parameters of the above-mentioned first three residual blocks, the intermediate model can add a background weakening mechanism between the third residual block and the fourth residual block. For example, the background weakening mechanism can include: passing the feature layer output by the third residual block in the intermediate model through a channel integration mechanism to integrate the features output by multiple output channels with each other to obtain a target response region corresponding to the second feature data set. Then, according to the target image features in the target response region, a feature screening condition is determined, and then, according to the feature screening condition, the background features of the target response region are filtered out to improve the classification accuracy of the target object in the second feature data set.
[0091] According to an embodiment of the present disclosure, for example, by adding the above-mentioned background weakening mechanism in the process of training the intermediate model, the intermediate model can be prevented from being disturbed by the background features, the classification accuracy of the classification model obtained by training can be improved, and the effectiveness of the training process can be improved and the classification range of the classification model can be expanded.
[0092] Figure 4 A schematic diagram of training an intermediate model according to an embodiment of the present disclosure is schematically shown.
[0093] As Figure 4As shown, the intermediate model can include a first residual block 420, a second residual block 430, a third residual block 440, a background weakening mechanism 450, a fourth residual block 460, and a classifier 470, wherein the classifier 470 can be used to classify a plurality of target image features output by the fourth residual block 460. During the training process, the parameters of the first residual block 420, the second residual block 430, and the third residual block 440 are fixed, and only the parameters of the fourth residual block 460 are updated. The remote sensing image is input into the intermediate model 410, and is processed by the first residual block 420, the second residual block 430, the third residual block 440, the background weakening mechanism 450, the fourth residual block 460, and the classifier 470 in sequence, and a classification result 480 is output.
[0094] According to an embodiment of the present disclosure, the second classification result can be a classification result of the intermediate model classifying the second feature data set.
[0095] According to an embodiment of the present disclosure, since the classification model is trained by using the first feature data set to obtain the intermediate model and the first classification result, and then the parameters of the target residual block are fixed, and then the intermediate model is trained by using the second feature data set to train the residual block whose parameters are not fixed, the influence of the training by using the second feature data set on the training by using the first feature data set is reduced, and the training by using the first feature data set is maximally retained. Moreover, since the intermediate model is trained by using the second feature data set obtained according to the first feature data set while the training by using the first feature data set is maximally retained, the problem of overfitting in the traditional training process is avoided, and the classification range of the classification model obtained by training can be maximally expanded.
[0096] According to an embodiment of the present disclosure, training the intermediate model by using the second feature data set to obtain the classification model and the second classification result includes:
[0097] dividing the second feature data set to obtain a third feature data set, wherein the third feature data set includes at least two feature data sets of different classes;
[0098] training the intermediate model by using the third feature data set to obtain the classification model.
[0099] According to embodiments of this disclosure, a second feature dataset is divided based on a first classification result to obtain a third feature dataset. For example, at least two categories that failed training can be determined based on the first classification result. The second feature dataset can be divided based on the at least two categories that failed training to obtain feature datasets with at least two different categories. The third feature dataset, composed of at least two feature datasets with different categories, can be used to train an intermediate model to maximize the classification range of the trained classification model.
[0100] According to embodiments of this disclosure, since the second feature dataset is divided according to the first classification result to obtain a third feature dataset including at least two different categories, and the intermediate model is trained using the third feature dataset, the classification range of the trained classification model can be expanded to the greatest extent.
[0101] According to embodiments of this disclosure, the first feature dataset can correspond to the base dataset in meta-learning, and the second feature dataset can correspond to the few-sample dataset in meta-learning. The training process of the intermediate model can be task-based training.
[0102] For example, the few-shot continuous learning method of this disclosure embodiment is as follows:
[0103] The dataset obtained from the remote sensing images can be divided according to the number of samples in each category to obtain the first feature dataset. This first feature dataset can be represented as...
[0104] The first feature dataset can be used The second feature dataset {T1 T2 …T} is obtained by sampling a few samples. n In the second feature dataset, each task is training data in an N-way-K-shot pattern, meaning each task selects N categories, and each category contains K feature samples. Within each task, the selected N categories can be trained. Through this task-based training, the intermediate model can be trained. N and K are positive integers.
[0105] For the second feature dataset, the N categories can be divided into a group to obtain the third feature dataset, and K feature samples are selected from each category to form the training set for one task. The training set from this single task can be used to train the intermediate model once, enabling it to learn N categories. All tasks corresponding to the third feature dataset can be represented as: A third feature dataset of 5-way-5-shot mode can be utilized for training, i.e., N=5 and K=5. Through the task-based training, n training processes of the intermediate model can be completed, where n and i are positive integers, and i≤n.
[0106] During the training phase, the classes learned by each task can be represented as {C (0) C (1) … C (n)}, where C (0) is a class learned by using the first feature dataset, C (i) is a class learned by using the third feature dataset in the i-th task, and C (j) is a class learned by using the third feature dataset in the j-th task. The classes learned by different tasks do not overlap, i.e., when i≠j, C (i) ∩C (j) =φ, and i,j∈{0,1,…,n}.
[0107] During the testing phase, each task needs to evaluate all the classes learned by the task, so as to classify the test sample according to all the learned classes. The test sample can be a first remote sensing image feature extracted from the first remote sensing image. The test set can be composed of the test sample, and the test set of all tasks can be represented as Taking the i-th task as an example, the test set may contain test data of all classes in the first i tasks, and can be represented as
[0108] According to an embodiment of the present disclosure, the first classification result can include a first classification vector, which can be represented as formula (1), where P c 0 is a classification vector of the c-th class in the process of training the preset model, P c 0 can be obtained by calculating the representation average of all training samples of the class:
[0109]
[0110] wherein, is the total number of samples of the c-th class in the process of training the preset model, is the i-th training sample of the class, and F0(·) is a feature extractor in the process of training the preset model, where i and c are positive integers.
[0111] According to an embodiment of the present disclosure, the second classification result can include a second classification vector, which is represented as formula (2):
[0112]
[0113] wherein, is a classification vector of the c-th class in the i-th task, K represents the number of samples of the class, represents the i-th sample of the class, F1(·) represents a feature extractor in the process of training the intermediate model, wherein i and c are positive integers.
[0114] According to an embodiment of the present disclosure, the preset model is trained by using the first feature data set to obtain an intermediate model, comprising:
[0115] The preset model is trained by using the first feature data set and a cross-entropy loss function through a gradient back propagation method to obtain first parameter information;
[0116] The intermediate model is obtained based on the first parameter information.
[0117] According to an embodiment of the present disclosure, for example, the first parameter information can be obtained by training the preset model, and can be used to obtain the parameter information of the intermediate model. The first parameter information can be used in the preset model to obtain the intermediate model.
[0118] According to an embodiment of the present disclosure, for example, the ResNet18 network model can be used as the preset model. The first feature data set The ResNet18 network model is trained, and a loss function of the training process is calculated to optimize the parameters of the entire preset model through a gradient back propagation method. The loss function can be a cross-entropy loss function, which can be as shown in formula (3):
[0119]
[0120] wherein, B can be the total number of samples processed by the preset model in a training process, |C (0) | is the number of classes that need to train the preset model, and respectively, the real probability and the prediction probability of the preset model that the j-th sample belongs to the c-th class in the process of training the preset model, wherein j and c are positive integers.
[0121] According to an embodiment of the present disclosure, the preset model is trained by using the first feature data set and a cross-entropy loss function through a gradient back propagation method to obtain first parameter information, and the first parameter information is used in the preset model to obtain the intermediate model that meets the requirements.
[0122] According to an embodiment of the present disclosure, in the process of training the intermediate model, the parameters of the residual block not fixed by the parameters can also be optimized to expand the classification range of the classification model by calculating the cross-entropy loss function and using the gradient back propagation method.
[0123] According to an embodiment of the present disclosure, the plurality of target image features are classified to obtain a classification result corresponding to the first remote sensing image, including:
[0124] The cosine similarity classification function is used to determine the similarity between the plurality of image features and the second classification result.
[0125] According to the similarity, the plurality of target image features are classified to obtain a classification result.
[0126] According to an embodiment of the present disclosure, for example, by determining the similarity between the image feature and the second classification result, the classification result most similar to the image feature can be determined from the second classification result, and the category of the image feature is determined according to the category corresponding to the classification result most similar to the image feature.
[0127] According to an embodiment of the present disclosure, for example, the cosine similarity function can be used to classify the plurality of target image features. The first remote sensing image feature x j After inputting the classification model, the obtained target image feature can be represented as: v j =F1(x j ). The second classification result can be represented as:
[0128] The cosine similarity function is used to classify the target image feature, and the prediction result definition can be represented as formula (4):
[0129]
[0130] Wherein, pre can represent the final category predicted according to the cosine similarity score.
[0131] According to an embodiment of the present disclosure, since the cosine similarity classification function is used to determine the similarity between the plurality of image features and the second classification result, and then the plurality of target image features are classified according to the similarity, the accuracy of classifying the target object is improved.
[0132] According to an embodiment of the present disclosure, the intermediate model is trained using the third feature data set to obtain a classification model, including:
[0133] The third feature data set is used to train the intermediate model to obtain second parameter information;
[0134] The second parameter information is used to obtain a classification model.
[0135] According to an embodiment of the present disclosure, the second parameter information can be obtained by training the intermediate model, and can be used to obtain the parameter information of the classification model. The second parameter information can be used to obtain the classification model from the intermediate model.
[0136] According to an embodiment of the present disclosure, the classification model satisfying the requirement is obtained by training the intermediate model using the third feature data set, obtaining the second parameter information, and using the second parameter information for the intermediate model.
[0137] Based on the above remote sensing image fine-grained classification method, the present disclosure further provides a remote sensing image fine-grained classification device. The following will be described in detail Figure 5 The device is described in detail.
[0138] Figure 5 The structure block diagram of the remote sensing image fine-grained classification device according to an embodiment of the present disclosure is schematically shown.
[0139] As Figure 5 shown, the remote sensing image fine-grained classification device 500 of this embodiment includes an extraction module 510, a first acquisition module 520, a second acquisition module 530, a filtering module 540, and a classification module 550.
[0140] The extraction module 510 is configured to extract image features of the first remote sensing image using the trained classification model to obtain first remote sensing image features. In an embodiment, the extraction module 510 can be configured to perform the operation S210 described above, and details are not repeated here.
[0141] The first acquisition module 520 is configured to obtain a target response region in the first remote sensing image features according to the first remote sensing image features, wherein the target response region represents a feature response region corresponding to a target object in the first remote sensing image features, and the feature response region contains features of the target object and background features. In an embodiment, the first acquisition module 520 can be configured to perform the operation S220 described above, and details are not repeated here.
[0142] The second acquisition module 530 is configured to obtain a feature screening condition according to the target response region. In an embodiment, the second acquisition module 530 can be configured to perform the operation S230 described above, and details are not repeated here.
[0143] The filtering module 540 is configured to perform a feature filtering operation on the first remote sensing image features based on the feature screening condition to obtain a plurality of target image features corresponding to the target object. In an embodiment, the filtering module 540 can be configured to perform the operation S240 described above, and details are not repeated here.
[0144] The classification module 550 is configured to classify the plurality of target image features to obtain a classification result corresponding to the first remote sensing image. In an embodiment, the classification module 550 can be configured to perform the operation S250 described above, and details are not repeated here.
[0145] According to an embodiment of the present disclosure, the first obtaining module 520 includes a first obtaining sub-module and a division sub-module. The first obtaining sub-module is configured to obtain a correlation degree result according to the correlation between the plurality of target image features. The division sub-module is configured to divide the feature response region according to the correlation degree result to obtain the target response region.
[0146] According to an embodiment of the present disclosure, the second obtaining module 530 includes a determination sub-module and a second obtaining sub-module. The determination sub-module is configured to determine a plurality of feature values according to the plurality of target image features. The second obtaining sub-module is configured to obtain a feature screening condition according to an average value of the plurality of feature values.
[0147] According to an embodiment of the present disclosure, the extraction module 510 includes an extraction sub-module, a first training sub-module, a fixing sub-module, and a second training sub-module. The extraction sub-module is configured to extract image features of a sample remote sensing image to obtain a first feature data set for training a preset model. The first training sub-module is configured to train the preset model using the first feature data set to obtain an intermediate model and a first classification result. The fixing sub-module is configured to fix parameters of at least one target residual block in the intermediate model. The second training sub-module is configured to train the intermediate model using a second feature data set to obtain a classification model and a second classification result, wherein the second feature data set is obtained by sampling the first feature data set.
[0148] According to an embodiment of the present disclosure, the second training sub-module includes a division unit and a first training unit. The division unit is configured to divide the second feature data set to obtain a third feature data set, wherein the third feature data set includes at least two feature data sets of different categories. The first training unit is configured to train the intermediate model using the third feature data set to obtain the classification model.
[0149] According to an embodiment of the present disclosure, the first training sub-module includes a second training unit and an obtaining unit. The second training unit is configured to train the preset model using the first feature data set and a cross-entropy loss function by a gradient back propagation method to obtain first parameter information. The obtaining unit is configured to obtain the intermediate model based on the first parameter information.
[0150] According to an embodiment of the present disclosure, the classification module 550 comprises a determination sub-module and a classification sub-module. The determination sub-module is configured to determine the similarity between the plurality of image features and the second classification result by using a cosine similarity classification function. The classification sub-module is configured to classify the plurality of target image features according to the similarity to obtain a classification result.
[0151] According to an embodiment of the present disclosure, the first training unit comprises a training sub-unit and an acquisition sub-unit. The training sub-unit is configured to train the intermediate model by using the third feature data set to obtain second parameter information. The acquisition sub-unit is configured to obtain the classification model according to the second parameter information.
[0152] According to an embodiment of the present disclosure, any of the extraction module 510, the first acquisition module 520, the second acquisition module 530, the filtering module 540 and the classification module 550 can be combined in one module, or any of them can be split into multiple modules. Alternatively, at least part of the function of one or more of these modules can be combined with at least part of the function of other modules, and implemented in one module. According to an embodiment of the present disclosure, at least one of the extraction module 510, the first acquisition module 520, the second acquisition module 530, the filtering module 540 and the classification module 550 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of hardware or firmware that can be integrated or packaged, or any one of software, hardware and firmware or any appropriate combination of any of them. Alternatively, at least one of the extraction module 510, the first acquisition module 520, the second acquisition module 530, the filtering module 540 and the classification module 550 can be at least partially implemented as a computer program module which can perform corresponding functions when it is run.
[0153] Figure 6 The block diagram of an electronic device suitable for implementing the remote sensing image fine-grained classification method according to an embodiment of the present disclosure is schematically shown.
[0154] As Figure 6As shown, the electronic device 600 according to the embodiments of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 602 or a program loaded into a random access memory (RAM) 603 from a storage section 608. The processor 601 can include, for example, a general purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), and so on. The processor 601 can also include an on-board memory for cache use. The processor 601 can include a single processing unit or multiple processing units for executing different actions of the method processes according to the embodiments of the present disclosure.
[0155] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The processor 601 performs various operations of the method processes according to the embodiments of the present disclosure by executing the programs in the ROM 602 and / or the RAM 603. Note that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method processes according to the embodiments of the present disclosure by executing the programs stored in the one or more memories.
[0156] According to the embodiments of the present disclosure, the electronic device 600 can further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 can further include one or more of the following components connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as necessary. A removable recording medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 610 as necessary, so that a computer program read out therefrom is installed in the storage section 608 as necessary.
[0157] The present disclosure also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, which when executed, implement the method according to the embodiments of the present disclosure.
[0158] According to an embodiment of the present disclosure, the computer readable storage medium can be a nonvolatile computer readable storage medium, for example, can include but not limited to: a portable computer diskette, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), a portable compact disc read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any appropriate combination thereof. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer readable storage medium can include one or more memories such as the ROM 602 and / or the RAM 603 described above and / or one or more memory other than the ROM 602 and the RAM 603.
[0159] Embodiments of the present disclosure also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the remote sensing image fine-grained classification method provided by the embodiments of the present disclosure.
[0160] The above functions defined in the system / device of the embodiments of the present disclosure are performed when the computer program is executed by the processor 601. According to an embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by computer program modules.
[0161] In one embodiment, the computer program can rely on tangible storage media such as optical storage media, magnetic storage media, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of signals on network media, and be downloaded and installed through the communication part 609, and / or installed from the detachable medium 611. The program codes contained in the computer program can be transmitted by any appropriate network media, including but not limited to: wireless, wired, etc., or any appropriate combination thereof.
[0162] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the detachable medium 611. When the computer program is executed by the processor 601, the above functions defined in the system of the embodiments of the present disclosure are performed. According to an embodiment of the present disclosure, the system, device, apparatus, module, unit, etc. described above can be implemented by computer program modules.
[0163] According to embodiments of the present disclosure, program code of the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages, and specifically, these computer programs can be implemented using a high-level procedural and / or object-oriented programming language, and / or an assembly / machine language. The programming language includes, but is not limited to, a programming language such as Java, C++, Python, "C" language, or a similar programming language. The program code can be executed entirely on a user computing device, partially on a user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case involving a remote computing device, the remote computing device can be connected to the user computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, connected through the Internet by using an Internet service provider).
[0164] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the reverse order, depending on the functionality involved. It should also be noted that each block in the flowcharts or block diagrams, and combinations of blocks in the flowcharts or block diagrams, can be implemented by special-purpose hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0165] It should be noted that the operations shown in the flowcharts of the embodiments of the present disclosure can be executed in any order unless otherwise specified, or unless the context clearly indicates otherwise, or unless the execution of the operations in a certain order is technically necessary. Multiple operations can be executed simultaneously, or in parallel, or in any order.
[0166] Those skilled in the art can understand that the features described in the various embodiments of the present disclosure and / or claims can be combined or / and integrated, even if such combinations or integrations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of the present disclosure and / or claims can be combined and / or integrated in various combinations, without departing from the spirit and teachings of the present disclosure. All such combinations and / or integrations fall within the scope of the present disclosure.
[0167] The above describes embodiments of the present disclosure. However, these embodiments are merely for illustrative purposes, and are not intended to limit the scope of the present disclosure. Although each embodiment is described above separately, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Those skilled in the art can make various substitutions and modifications without departing from the scope of the present disclosure, and these substitutions and modifications should all fall within the scope of the present disclosure.
Claims
1. A small sample based on continuous learning remote sensing image fine-grained classification method, comprising: using a trained classification model to extract image features of a first remote sensing image, obtaining first remote sensing image features, wherein the classification model is trained using a small sample continuous learning method; obtaining a target response region in the first remote sensing image features according to the first remote sensing image features, wherein the target response region represents a feature response region corresponding to a target object in the first remote sensing image features, and the feature response region contains features of the target object and background features; obtaining a feature screening condition according to the target response region; performing a feature filtering operation on the first remote sensing image features based on the feature screening condition to obtain a plurality of target image features corresponding to the target object; performing classification processing on the plurality of target image features to obtain a classification result corresponding to the first remote sensing image; wherein the training method of the classification model comprises: extracting image features of sample remote sensing images to obtain a first feature data set for training a preset model; training the preset model using the first feature data set to obtain an intermediate model and a first classification result; fixing the parameters of at least one target residual block in the intermediate model; training the intermediate model using a second feature data set to obtain the classification model and a second classification result, wherein the second feature data set is obtained by sampling the first feature data set; wherein the training of the intermediate model using the second feature data set to obtain the classification model comprises: dividing the second feature data set to obtain a third feature data set, wherein the third feature data set includes at least two feature data sets of different categories; training the intermediate model using the third feature data set to obtain the classification model; wherein the training of the preset model using the first feature data set to obtain an intermediate model comprises: training the preset model using the first feature data set and a cross-entropy loss function through a gradient back propagation method to obtain first parameter information; obtaining the intermediate model based on the first parameter information.
2. The method of claim 1, wherein, The first remote sensing image features according to the first remote sensing image features, comprising: obtaining a correlation result according to the association between the plurality of target image features; dividing the feature response region according to the correlation result to obtain the target response region.
3. The method of claim 1, wherein, The feature screening condition obtained according to the target response region, comprising: determining a plurality of feature values according to the plurality of target image features; obtaining the feature screening condition according to the average of the plurality of feature values.
4. The method of claim 1, wherein, The classification processing on the plurality of target image features to obtain a classification result corresponding to the first remote sensing image, comprising: determining the similarity between the plurality of image features and the second classification result using a cosine similarity classification function; classifying the plurality of target image features according to the similarity to obtain the classification result.
5. The method of claim 1, wherein, The training of the intermediate model by using the third feature dataset to obtain the classification model comprises: training the intermediate model by using the third feature dataset to obtain second parameter information; obtaining the classification model according to the second parameter information. 6.An electronic device, comprising: one or more processors; a storage device for storing one or more programs, wherein the one or more programs, when executed by the one or more processors, enable the one or more processors to perform the method according to any one of claims 1-5. 7.A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1-5.
Citation Information
Patent Citations
Method and system for semantic segmentation interactive annotation based on AI image
CN114782690A
Data set construction method, mobile terminal and readable storage medium
WO2019233297A1