Feature decoupling-based small sample image classification method and device, and storage medium
By designing a dual-channel feature extraction network based on feature decoupling in the small sample image classification method, the problem of limited ability to extract tag information in the prior art is solved, and higher image classification accuracy and network convergence speed are achieved.
Patent Information
- Application Number
- CN202510079229.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing small sample learning methods have limited ability to extract tag information, making it difficult to achieve high-accurate image classification under limited training data.
A small sample image classification method based on feature decoupling is designed. Through a feature decoupling network extracted by dual-channel feature, one way to detect label information and the other way to detect semantic information, and the two ways to pass each other, ultimately realizing the decoupling of label information and semantic information, strengthening the extraction of label information.
Through the design of feature decoupling network, the accuracy of image classification is effectively improved, the convergence speed of the network is improved, and good classification performance is achieved in the small sample image classification task.
Smart Images

Figure CN119992193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a small sample image classification method, device and storage medium based on feature decoupling. Background Art
[0002] In the field of computer vision research, deep learning technology has become the main method to solve problems. However, deep learning requires a large amount of labeled data for training. It is very labor-intensive and resource-intensive to obtain and annotate a large amount of data. In addition, in some special scenarios, it is impossible to obtain a large amount of data. Therefore, it is crucial to adapt deep learning to small sample data. At present, the research on small sample learning is still in its infancy, and the performance gap with large sample conditions is still large. Among many related tasks, the small sample image classification task refers to the technology of classifying images through machine learning and deep learning methods under limited training data. In some specific application scenarios, some scarce categories may make it very difficult to obtain and annotate data due to factors such as privacy, security, and high labeling costs, such as military remote sensing detection, disease diagnosis, and defective product detection in industrial production. Therefore, this technology is widely used in medical imaging, object detection, bioinformatics, and other fields, and is an important research direction in small sample learning. Summary of the invention
[0003] In order to solve the above-mentioned problems existing in the prior art, the present invention provides a small sample image classification method, device and storage medium based on feature decoupling to solve the technical problem that the existing small sample learning method has limited ability to extract label information.
[0004] The present invention designs a small sample image classification method, device and storage medium based on feature decoupling, which integrates the focus on label specific information into the mechanism of the classifier and decouples it from semantic information; the present invention designs a feature decoupling network that can realize dual-path feature extraction, one path detects label information and the other path detects semantic information, the two paths of information will be transmitted to each other, and finally the label information and semantic information will be decoupled to strengthen the extraction of label information; through this method, the accuracy of classification is effectively improved, which has high scientific significance and practical value.
[0005] The small sample image classification method based on feature decoupling includes the following steps:
[0006] S1. Preprocess the image data and divide it into a support set and a query set, and then input it into a feature decoupling network for processing. The feature decoupling network includes a semantic convolutional encoder, a label convolutional encoder, a graph neural network and a decoder;
[0007] S2, the semantic convolution encoder extracts features from the support set image to obtain potential semantic information;
[0008] S3, label convolution encoder extracts features from the support set image and the query set image and combines them with the potential semantic information obtained in S2 to obtain potential label information;
[0009] S4. Based on the potential label information, label propagation is performed through the graph neural network to predict the query set label, and the classification loss is calculated based on the predicted label and the true label of the query set;
[0010] S5. The decoder reconstructs the image based on the potential semantic information and potential label information, and calculates the reconstruction loss based on the reconstructed image and the original image; the classification loss and reconstruction loss are integrated to obtain the total loss to optimize the feature decoupling network and determine the network parameters.
[0011] Furthermore, the S1 includes:
[0012] S101, preprocessing the images in the original data set to make their sizes match the feature decoupling network;
[0013] S102, dividing the processed data set into a training set, a validation set and a test set, each data set includes multiple small sample tasks, and each small sample task includes a support set and a query set. Further, the S2 includes:
[0014] S201, inputting the support set image into the semantic convolution encoder to extract features of the image;
[0015] S202, mapping the features extracted in 201 into a Gaussian distribution, using two fully connected layers to obtain the mean and variance of the Gaussian distribution, and then obtaining a Gaussian distribution that can be back-propagated through re-parameterization to obtain potential semantic information.
[0016] Furthermore, the S3 includes:
[0017] S301, input the support set and query set images into the label convolution encoder for feature extraction;
[0018] S302: Combine the features extracted in S301 with the potential semantic information obtained in S2, and map them into a Gaussian distribution form of label information. The mean and variance of the Gaussian distribution are obtained through two fully connected layers, and then the Gaussian distribution that can be back-propagated is obtained through re-parameterization to obtain the potential label information.
[0019] Furthermore, the specific process of S4 is as follows:
[0020] S401, calculating the weights of any two feature vectors in the potential label information obtained in S3 on the undirected graph, and obtaining an undirected graph consisting of a support set and a query set;
[0021] S402, label propagation is performed on the undirected graph, the known annotated support set is used as the known nodes of the graph, label propagation is performed through the label propagation algorithm, and the label of the query set is predicted;
[0022] S403: Calculate the cross entropy loss function according to the predicted labels and true labels of the query set to obtain the classification loss.
[0023] Furthermore, the S5 includes:
[0024] S501, accumulating the latent semantic information obtained in S2 and the latent label information obtained in S3 and inputting them into the decoder for image reconstruction;
[0025] S502, using the reconstructed image and the original image to calculate the mean square error as the reconstruction loss;
[0026] S503, calculating the KL divergence between the potential semantic information of S2 and the standard normal distribution, calculating the KL divergence between the potential label information of S3 and the standard normal distribution, as the distribution alignment loss in the variational autoencoder, the variational autoencoder including the semantic convolutional encoder and decoder;
[0027] S504: Add the classification loss in S4, the reconstruction loss in S502, and the distribution alignment loss in S503 as the total loss, then perform gradient descent and back propagation on the feature decoupling network to optimize the network parameters and make the network converge.
[0028] Small sample image classification device based on feature decoupling, including:
[0029] Memory: used for storing the computer program of the small sample image classification method based on feature decoupling;
[0030] Processor: used to implement a small sample image classification method based on feature decoupling when executing the computer program.
[0031] A computer-readable storage medium stores a computer program, which can implement a small sample image classification method based on feature decoupling when executed by a processor.
[0032] The beneficial effects of the present invention include:
[0033] The present invention adopts a feature decoupling network that can perform dual-path detection, one path for detecting label information and the other for detecting semantic information. The two paths of information will be transmitted to each other, ultimately achieving the decoupling of label information from semantic information to enhance the extraction of image label information.
[0034] The label propagation method is used to predict the classification results, which effectively utilizes the decoupled semantic information, improves the accuracy of image classification, and also improves the convergence speed of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a schematic diagram of the network involved in the embodiment of the present application.
[0036] Figure 2 This is a flowchart of a small sample image classification method based on feature decoupling involved in an embodiment of the present application.
[0037] Figure 3 This is a detailed flowchart of S1 involved in the embodiment of the present application.
[0038] Figure 4 It is a detailed flowchart of S2 involved in the embodiment of the present application.
[0039] Figure 5 This is a detailed flowchart of S3 involved in the embodiment of the present application.
[0040] Figure 6 This is a detailed flowchart of S4 involved in the embodiment of the present application.
[0041] Figure 7 This is a detailed flowchart of S5 involved in the embodiment of the present application. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0043] For existing small sample image classification methods, since the ultimate goal of training the network is to obtain the label information of the inferred image and achieve the effect of target classification, most methods focus on how to effectively extract the features of the target category with a small number of training samples, and then use the features for classification. However, images are composed of different features, such as style, design pattern and contextual information, and these attributes are not necessarily relevant discriminant features for classification. In this method, this information is called semantic information. On the other hand, there are some feature attributes (such as bird wings, elephant trunks, and humps on camels) that are crucial for classification, and these features do not need to consider contextual information. In this method, such features are called label information. The semantic information of the image plays a dominant role in the image, and the label information is mixed in the semantic information and is relatively weak, but the degree of extraction of these weak label information determines the effectiveness of the classification algorithm.
[0044] Small sample image classification methods based on feature decoupling, such as Figure 1-2 As shown, the following steps are included:
[0045] S1. Preprocess the image data and divide it into a support set and a query set, and then input it into a feature decoupling network for processing. The feature decoupling network includes a semantic convolutional encoder, a label convolutional encoder, a graph neural network and a decoder;
[0046] S2, the semantic convolution encoder extracts features from the support set image to obtain potential semantic information;
[0047] S3, label convolution encoder extracts features from the support set image and the query set image and combines them with the potential semantic information obtained in S2 to obtain potential label information;
[0048] S4. Based on the potential label information, label propagation is performed through the graph neural network to predict the query set label, and the classification loss is calculated based on the predicted label and the true label of the query set;
[0049] S5. The decoder reconstructs the image based on the potential semantic information and potential label information, and calculates the reconstruction loss based on the reconstructed image and the original image; the classification loss and reconstruction loss are integrated to obtain the total loss to optimize the feature decoupling network and determine the network parameters.
[0050] In another embodiment, if Figure 3 As shown, the S1 includes:
[0051] S101, preprocessing the images in the original data set to make their sizes match the feature decoupling network;
[0052] S102, divide the processed data set into a training set, a validation set and a test set, each data set includes multiple small sample tasks, and each small sample task includes a support set and a query set. Specifically, obtain a small sample data set, preprocess the data, and divide it into a form suitable for small sample learning, divided into a support set and a query set. In small sample learning, the data set is usually divided into a training task for training. Each training task includes a support set and a query set. The support set includes N categories, each category has K samples, called N-way K-shot. The query set category is the same as the support set, and each category has Q samples. In the field of deep learning, the data set is usually divided into a training set, a validation set, and a test set. The three types of data sets have no intersection. This is also true in small sample learning. In the training network stage, multiple small sample tasks are randomly sampled from the training set. In each task, the original network parameters are first copied. The copied network is first trained in the support set, and the network parameters are updated using the loss obtained as an inner loop. Then, the query set is trained again, and the parameters of the original network are updated using the parameters of the query set as an outer loop.
[0053] In another embodiment, if Figure 4 As shown, the S2 includes:
[0054] S201, inputting the support set image into the semantic convolution encoder to extract features of the image;
[0055] S202, map the features extracted in 201 into a Gaussian distribution, use two fully connected layers to obtain the mean and variance of the Gaussian distribution, and then use the re-parameterization technique to obtain a Gaussian distribution that can be back-propagated to obtain potential semantic information.
[0056] Specifically, the extraction of latent semantic information is mainly achieved through a variational autoencoder, which includes a convolutional encoder and a decoder for extracting semantic information. The image is input into a convolutional encoder for extracting semantic information, namely a semantic convolutional encoder, and the obtained latent features are mapped to a Gaussian distribution to obtain low-dimensional semantic information; wherein the convolutional encoder is a feature extraction module, and a feature extraction head of a 4-layer convolutional layer plus a pooling layer is used in this embodiment, and other types of feature extraction heads can also be used; the extracted features are mapped to a Gaussian distribution using two fully connected layers, and the extracted features are converted into the form of a Gaussian distribution to obtain two feature vectors σ and μ; then a reparameterization technique is used to obtain a Gaussian distribution that can be back-propagated as the latent semantic information z s ; The reparameterization is calculated as follows: s =μ+σ×ε
[0057] where z sis the extracted latent semantic information, μ is the mean vector obtained previously, σ is the standard deviation vector obtained previously, and ε is the noise sampled from the standard normal distribution.
[0058] In another embodiment, if Figure 5 As shown, the S3 includes:
[0059] S301, input the support set and query set images together into the label convolution encoder for feature extraction;
[0060] S302: Combine the features extracted in S301 with the potential semantic information obtained in S2, and map them into a Gaussian distribution form of label information. The mean and variance of the Gaussian distribution are obtained through two fully connected layers, and then the Gaussian distribution that can be back-propagated is obtained through the re-parameterization technique to obtain the potential label information.
[0061] Specifically, the extraction of potential label information takes the support set and the query set as input. After feature extraction, the obtained potential features are combined with the potential semantic information obtained in S2 and mapped into a Gaussian distribution to obtain low-dimensional label information z l ;
[0062] In another embodiment, if Figure 6 As shown, the specific process of S4 is as follows:
[0063] S401, for any two feature vectors in the potential label information obtained in S3, use the Gaussian similarity function to calculate their weights on the undirected graph, and obtain an undirected graph consisting of a support set and a query set;
[0064] S402, label propagation is performed on the undirected graph, the known annotated support set is used as the known nodes of the graph, label propagation is performed through the label propagation algorithm, and the label of the query set is predicted;
[0065] S403: Calculate the cross entropy loss function according to the predicted labels and true labels of the query set to obtain the classification loss.
[0066] Specifically, the classification loss is calculated by propagating labels on the potential label information obtained in S3 through a graph neural network composed of feature embedding and undirected graph construction to achieve prediction of the query set and obtain the classification loss. The specific implementation process is: for any two feature vectors in the potential label information obtained in S3, their weights on the undirected graph are calculated using the Gaussian similarity function to obtain an undirected graph composed of the support set and the query set, and then the label propagation algorithm is used to propagate labels to obtain the classification results of the query set; the cross entropy loss function is calculated based on the predicted results and the true labels to obtain the classification loss L class , the classification loss is calculated as follows:
[0067] n is the number of samples, k is the number of categories, y ic is the unique hot encoding of the sample target value, h θ (x i ) c is the observed sample x i The probability of belonging to class c.
[0068] It should be noted that, in this embodiment, the Gaussian similarity function is used as the weight calculation method, and other distance algorithms can also be used to calculate the weight between two features.
[0069] In another embodiment, if Figure 7 As shown, the S5 comprises:
[0070] S501, accumulating the latent semantic information obtained in S2 and the latent label information obtained in S3 and inputting them into the decoder for image reconstruction;
[0071] S502, using the reconstructed image and the original image to calculate the mean square error as the reconstruction loss;
[0072] S503, calculating the relative entropy between the latent semantic information of S2 and the standard normal distribution, also known as Kullback-Leibler divergence, hereinafter referred to as KL divergence, calculating the KL divergence between the latent label information of S3 and the standard normal distribution, adding the two KL divergences as the distribution alignment loss in the variational autoencoder, the variational autoencoder including the semantic convolutional encoder and decoder;
[0073] S504: Add the classification loss in S4, the reconstruction loss in S502, and the distribution alignment loss in S503 as the total loss, and then perform gradient descent and back propagation on the entire network to optimize the parameters of the network so that the network converges.
[0074] Specifically, the semantic features extracted from S2 and the label features extracted from S3 are used to reconstruct the image to obtain the reconstructed image With the original image y i Calculate the mean square error as the reconstruction loss L recon , calculated as follows:
[0075] Where n represents the total number of samples, y i represents the original image, represents the reconstructed image, L recon represents the reconstruction loss.
[0076] Calculate the latent semantic information z of S2 sThe relative entropy with respect to the standard normal distribution is also called the Kullback-Leibler divergence, hereinafter referred to as the KL divergence, denoted by L s-kl ; Calculate the potential label information z in S3 l KL divergence with the standard normal distribution, denoted by L l-kl , put L s-kl With L l-kl The sum is used as the distribution alignment loss in the variational self-encoder. The KL divergence with the standard normal distribution is calculated as follows:
[0077] The role of the variational autoencoder is to effectively extract features, because the encoder of the variational autoencoder actually maps the extracted features into a Gaussian distribution. This Gaussian distribution needs to be as similar as possible to the standard normal distribution to ensure the effect of feature extraction, so a distribution alignment loss is needed to ensure that the distribution is as similar as possible.
[0078] Where μ is the mean vector of the Gaussian distribution and σ is the variance of the Gaussian distribution.
[0079] The classification loss L in S4 class , reconstruction loss L recon , distribution alignment loss L kl Add them up as the total loss L.
[0080] The calculation method is as follows: L = α c L class +α r L recon +β s L s-kl +β l L l-kl
[0081] Where L class is the classification loss, L recon is the reconstruction loss, L s-kl is the distribution alignment loss of latent semantic information, L l-kl is the distribution alignment loss of the latent label information, α c For L class The weight, α r For L recon The weight, β s For L s-kl The weight, β l For L l-kl The weight of .
[0082] Then for Figure 1 The feature decoupling classification network shown performs back propagation and gradient descent to optimize the network parameters and achieve network convergence.
[0083] Using multiple types of losses to jointly act on the optimization goal of the entire network, the entire network can achieve better performance in both classification and reconstruction tasks. Ensure that both label information extraction and semantic information extraction achieve good results
[0084] The present invention has achieved good classification performance in small sample image classification tasks. As shown in Table 1, on the miniImage dataset, a classification accuracy of 90.77% was achieved on the 5-way, 1-shot task; and a classification accuracy of 96.12% was achieved on the 5-way, 5-shot task, reaching the advanced level in the current small sample classification model.
[0085] Table 1 Comparison of classification performance
[0086]
[0087] In another embodiment, a small sample image classification device based on feature decoupling is provided, comprising:
[0088] Memory: used for storing a computer program of the cross-domain small sample image classification method based on local association reasoning;
[0089] Processor: used to implement a small sample image classification method based on feature decoupling when executing the computer program.
[0090] In another embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a small sample image classification method based on feature decoupling can be implemented.
[0091] The above-mentioned embodiments only express the specific implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the protection scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the technical solution concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A small sample image classification method based on feature decoupling, characterized in that: The following steps are involved: S1. Preprocess the image data and divide it into a support set and a query set, and then input it into a feature decoupling network for processing. The feature decoupling network includes a semantic convolutional encoder, a label convolutional encoder, a graph neural network and a decoder; S2, the semantic convolution encoder extracts features from the support set image to obtain potential semantic information; S3, label convolution encoder extracts features from the support set image and the query set image and combines them with the potential semantic information obtained in S2 to obtain potential label information; S4. Based on the potential label information, label propagation is performed through the graph neural network to predict the query set label, and the classification loss is calculated based on the predicted label and the true label of the query set; S5. The decoder reconstructs the image based on the potential semantic information and potential label information, and calculates the reconstruction loss based on the reconstructed image and the original image; the classification loss and reconstruction loss are integrated to obtain the total loss to optimize the feature decoupling network and determine the network parameters.
2. The small sample image classification method based on feature decoupling according to claim 1 is characterized in that: The S1 includes: S101, preprocessing the images in the original data set to make their sizes match the feature decoupling network; S102, dividing the processed data set into a training set, a validation set and a test set, each data set includes multiple small sample tasks, and each small sample task includes a support set and a query set.
3. The small sample image classification method based on feature decoupling according to claim 1 is characterized in that: The S2 includes: S201, inputting the support set image into the semantic convolution encoder to extract features of the image; S202, mapping the features extracted in 201 into a Gaussian distribution, using two fully connected layers to obtain the mean and variance of the Gaussian distribution, and then obtaining a Gaussian distribution that can be back-propagated through re-parameterization to obtain potential semantic information.
4. The small sample image classification method based on feature decoupling according to claim 1, characterized in that: The S3 includes: S301, input the support set and query set images into the label convolution encoder for feature extraction; S302: Combine the features extracted in S301 with the potential semantic information obtained in S2, and map them into a Gaussian distribution form of label information. The mean and variance of the Gaussian distribution are obtained through two fully connected layers, and then the Gaussian distribution that can be back-propagated is obtained through re-parameterization to obtain the potential label information.
5. The small sample image classification method based on feature decoupling according to claim 1 is characterized in that: The specific process of S4 is as follows: S401, calculating the weights of any two feature vectors in the potential label information obtained in S3 on the undirected graph, and obtaining an undirected graph consisting of a support set and a query set; S402, label propagation is performed on the undirected graph, the known annotated support set is used as the known nodes of the graph, label propagation is performed through the label propagation algorithm, and the label of the query set is predicted; S403: Calculate the cross entropy loss function according to the predicted labels and true labels of the query set to obtain the classification loss.
6. The small sample image classification method based on feature decoupling according to claim 1, characterized in that: The S5 includes: S501, accumulating the latent semantic information obtained in S2 and the latent label information obtained in S3 and inputting them into the decoder for image reconstruction; S502, using the reconstructed image and the original image to calculate the mean square error as the reconstruction loss; S503, calculating the KL divergence between the potential semantic information of S2 and the standard normal distribution, calculating the KL divergence between the potential label information of S3 and the standard normal distribution, as the distribution alignment loss in the variational autoencoder, the variational autoencoder including the semantic convolutional encoder and decoder; S504: Add the classification loss in S4, the reconstruction loss in S502, and the distribution alignment loss in S503 as the total loss, then perform gradient descent and back propagation on the feature decoupling network to optimize the network parameters and make the network converge.
7. A small sample image classification device based on feature decoupling, characterized in that: include: Memory: used for storing the computer program of the small sample image classification method based on feature decoupling; Processor: used to implement a small sample image classification method based on feature decoupling when executing the computer program.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a small sample image classification method based on feature decoupling can be implemented.
Citation Information
Patent Citations
System and method for solving small sample image classification based on graph neural network mechanism of auto-encoder, equipment and storage medium
CN113592008A
Long-tail multi-label medical image recognition method based on label decoupling and reconstruction
CN118711183A
Masked modeling-based method for constructing 3D medical image segmentation model and application thereof
WO2024244261A1