Small sample expansion method and system based on prototype completion in image classification

By adopting a small sample expansion method based on prototype completion in image classification, using feature extraction and pseudo-sample generation technology, the problems of overfitting and poor universality in small sample training are solved, and the image classification effect with high accuracy and time-saving is achieved.

CN115393666BActive Publication Date: 2025-06-06GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210923952.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-02
Publication Date
2025-06-06
Estimated Expiration
2042-08-02

AI Technical Summary

Technical Problem

The existing methods of using small sample training neural networks to classify images are prone to introduce excessive network parameters and are poor in popularity. In the small sample setting, it is difficult to ensure the accuracy of image classification.

Method used

Using a small sample expansion method based on prototype completion, the original sample image data set is divided into base class data set, support set and query set, and the feature extractor is used to extract features, calculate the distance between the support set and the base class data set based on the spatial metric distance function, determine the most similar samples, and generate a pseudo-sample set for image classification.

Benefits of technology

This method avoids the introduction of additional network parameters, improves the accuracy and universality of image classification, saves image classification time, and has strong interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393666B_ABST
    Figure CN115393666B_ABST
Patent Text Reader

Abstract

The present invention proposes a small sample expansion method and system based on prototype completion in image classification, which relates to the technical field of sample processing in image classification. First, the original sample image data set is collected, and the original sample image data set is divided into a base class data set, a support set and a query set. The features of the three types of data sets are respectively extracted based on a feature extractor, and the samples are mapped from space to feature space. The prototype features are completed in the feature space based on the similarity of the sample features to obtain the prototype features. Overall, only a feature library and a feature extractor of a base class data set are required to complete the prototypes of a small number of samples and generate a pseudo sample set, which greatly compensates for the shortcomings of small sample image classification, avoids the introduction of additional network parameters, and saves the time of subsequent image classification, and has good universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sample processing in image classification, and more specifically, to a small sample expansion method and system based on prototype completion in image classification. Background Art

[0002] Image classification refers to the problem of outputting a classification description of the image content after an image is input. Distinguishing different categories of images based on the semantic information of the image is an important basic problem in computer vision, and is also the basis for other high-level visual tasks such as image detection, image segmentation, object tracking, and behavior analysis. Image classification has applications in many fields, including face recognition and intelligent video analysis in the security field, traffic scene recognition in the transportation field, content-based image retrieval and automatic album classification in the Internet field, and image recognition in the medical field.

[0003] In the past, image classification was based on artificial features or simple machine learning methods. The disadvantages of this method are that it is not accurate enough, requires a large number of samples, and requires some manual design. With the development of deep learning, the means used for image classification have made phased progress, but there has been no improvement in the demand for the number of samples. The deep learning method still requires a huge number of samples to meet the training requirements, because when the number of samples is sufficient, the performance is better and the classification accuracy is high, but when the samples are scarce, it will lead to overfitting and the classification accuracy will drop sharply. However, in many scenarios, the acquisition and labeling of samples requires too many resources, and sufficient training samples usually contain certain noise interference.

[0004] Under the setting of small samples, many methods use small samples to train a deep network. For example, the prior art discloses a method for constructing an image classification model and an image classification method. First, a neural network is established, and then a limited image data set is used to train the neural network to form an image classification model. Finally, the image classification model is used to complete the image classification task. However, this method introduces excessive network parameters to a certain extent, and the network accepts an input and then feeds back an output. It is essentially a nonlinear mapping from x to y, but it cannot be explicitly or implicitly expressed as a certain logical rule. Therefore, there is also the problem that the network is difficult to explain. In addition, these trained network parameters have poor universality, and in practical applications, a small number of training samples may also have problems of incorrect labeling or poor sample quality, and the accuracy of image classification cannot be guaranteed. Summary of the invention

[0005] In order to solve the problem that the existing method of using small samples to train neural networks for image classification is prone to introduce excessive network parameters and has poor universality, the present invention provides a small sample expansion method and system based on prototype completion in image classification, which completes the prototype of a small number of samples, avoids the introduction of additional network parameters, makes up for the defects of small sample image classification, saves the time of subsequent image classification, and has good universality.

[0006] In order to achieve the above technical effects, the technical solution of the present invention is as follows:

[0007] A small sample expansion method based on prototype completion in image classification, the method comprising the following steps:

[0008] S1. Collect the original sample image dataset, and divide the original sample image dataset into a base class dataset, a support set, and a query set according to the image data category in the original sample image dataset;

[0009] S2. Select a feature extractor and train the feature extractor using the base class data set, and use the trained feature extractor to extract the data features of the base class data set, the support set data features, and the query set data features;

[0010] S3. Based on the spatial metric distance function, calculate the distance between the support set data features and the data features of all base class data sets, and determine the K samples that are most similar to the support set data features according to the distance;

[0011] S4. Perform weighted calculation on the similarity between the K samples and the support set data features to obtain the prototype features of the support set;

[0012] S5. Generate a pseudo sample set based on the prototype features of the support set;

[0013] S6. Use the generated pseudo sample set and support set data features together as training data, and use the query set data features as test data to perform image classification.

[0014] In the technical scheme, the original sample image data set is first collected, and the original sample image data set is divided into a base class data set, a support set and a query set. The features of the three types of data sets are extracted respectively based on the feature extractor, and the samples are mapped from space to feature space. The prototype features are completed in the feature space based on the similarity of the sample features to obtain the prototype features. Overall, only a feature library and feature extractor of a base class data set are required to complete the prototypes of a small number of samples and generate a pseudo sample set, which greatly compensates for the shortcomings of small sample image classification, and does not introduce additional parameters that need to be trained, saving time and equipment requirements for subsequent image classification. Compared with the "black box" neural network model, it has extremely strong interpretability.

[0015] Preferably, in step S1, it is assumed that there are Z categories of image data in the original sample image dataset, the base dataset belongs to category X, the support set and the query set belong to category Y, and the following conditions are satisfied:

[0016] X+Y=Z,

[0017] Preferably, in step S2, the feature extractor is selected as a Wide ResNet network. In the process of training the Wide ResNet network using a base class data set, the loss function of the Wide ResNet network is set, and the weight parameters of the Wide ResNet network are updated by back propagation until the loss function converges. After the training is completed, the network parameters are fixed, and the trained feature extractor extracts the data features of the set target task.

[0018] Preferably, the mean of the data features of the i-th category of the base class data set is used as the general term for the data features of this category, then:

[0019]

[0020] Among them, μ 1i represents the mean value of the data feature of the i-th category of the base class data set, x j Represents the jth data feature in the i-th data feature of the base class data set; n i Represents the number of data features of the i-th category of the base class data set;

[0021] Assume that the mean of the data features of the support dataset is μ 2 , in μ 2 As a general term for the data features of the supporting dataset, let the spatial metric distance function be collectively referred to as f d (), then based on the spatial metric distance function, when calculating the distance between the support set data features and the data features of all base class data sets, it satisfies:

[0022] d i =f d (μ i ,x s )

[0023] Among them, d i represents the distance between the support set data feature and the i-th data feature of the base class data set; finally, the distance set D is obtained, which is expressed as:

[0024] D={d 1 ,d 2 ,...,d i ,...,d q}

[0025] Among them, q represents the number of types of data features of the base class data set. The values ​​of all distance elements in the distance set D are arranged in ascending order. According to the arrangement order, the data features of the base class data set corresponding to the first K distance elements are selected as the K samples with the most similar data features of the support set.

[0026] Preferably, after obtaining the distance between the support set data features and the data features of all base class data sets, the distance is normalized, and the expression is:

[0027]

[0028] or

[0029]

[0030] Among them, d' i represents the normalized value of the distance. represents the maximum value in the distance, Indicates the minimum value in distance.

[0031] Preferably, in step S4, the similarity between the K samples and the support set data features is the mean μ of the K samples and the support set data features. 2 The distance between d () is obtained. Under the premise of distance normalization, the expression for weighted calculation is:

[0032]

[0033] Among them, e is a natural constant, w s is the weight of the support set data feature, which is pre-set according to the number of samples in the support set, w' g represents the weight of the gth sample among K samples, d' g Represents the distance between the gth sample among K samples and the data feature of the support set;

[0034] The expression of the prototype characteristics is:

[0035]

[0036] Among them, n K represents the sum of the number of similar samples and the number of support set data features, μ' represents the prototype feature;

[0037] The covariance matrix C' corresponding to the prototype feature is:

[0038]

[0039] Among them, C g It represents the covariance matrix of the data features of the g-th similar sample among K samples, and α is a hyperparameter.

[0040] Preferably, in step S5, the pseudo sample set generated is D y , the expression is:

[0041]

[0042] Among them, y represents the newly generated category, Represents the newly generated pseudo samples, which obey the Gaussian distribution.

[0043] Preferably, in step S6, a classifier is selected, the generated pseudo sample set and support set data features are used together as training data to train the classifier, and then the query set data features are used as test data to test the trained classifier to complete image classification.

[0044] Preferably, the classifier is a linear regression classifier or a support vector machine.

[0045] This application proposes a small sample expansion system based on prototype completion in image classification, the system comprising:

[0046] The original sample image acquisition and division unit is used to acquire the original sample image data set, and divide the original sample image data set into a base class data set, a support set and a query set according to the image data category in the original sample image data set;

[0047] A feature extraction unit is used to select a feature extractor, train the feature extractor using a base class data set, and respectively extract data features of the base class data set, data features of the support set, and data features of the query set using the trained feature extractor;

[0048] The similar sample determination unit calculates the distance between the support set data features and the data features of all base class data sets based on the spatial metric distance function, and determines the K samples that are most similar to the support set data features according to the distance;

[0049] The prototype feature determination unit is used to perform weighted calculation on the similarity between the K samples and the support set data features to obtain the prototype features of the support set;

[0050] A pseudo sample set generation unit generates a pseudo sample set based on the prototype features of the support set;

[0051] The image classification unit uses the generated pseudo sample set and support set data features as training data, and uses the query set data features as test data to perform image classification.

[0052] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0053] The present invention proposes a small sample expansion method and system based on prototype completion in image classification. First, the original sample image data set is collected, and the original sample image data set is divided into a base class data set, a support set and a query set. The features of the three types of data sets are extracted respectively based on a feature extractor, and the samples are mapped from space to feature space. The prototype features are completed in the feature space based on the similarity of the sample features to obtain the prototype features. Overall, only a feature library and a feature extractor of a base class data set are required to complete the prototypes of a small number of samples and generate a pseudo sample set, which greatly compensates for the shortcomings of small sample image classification, avoids the introduction of additional network parameters, and saves the time of subsequent image classification, and has good universality. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 A schematic diagram showing a flow chart of a small sample expansion method based on prototype completion in image classification proposed in Embodiment 1 of the present invention;

[0055] Figure 2 A flowchart showing the execution process of small sample expansion based on prototype completion in image classification proposed in Example 1 of the present invention.

[0056] Figure 3 It represents the TSNE visualization diagram proposed in Example 2 of the present invention;

[0057] Figure 4 A structural diagram showing a small sample expansion system based on prototype completion in image classification proposed in Example 3 of the present invention. DETAILED DESCRIPTION

[0058] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0059] In order to better illustrate the present embodiment, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the actual size;

[0060] It is understandable to those skilled in the art that descriptions of certain well-known contents in the drawings may be omitted.

[0061] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0062] The positional relationships described in the drawings are only for illustrative purposes and should not be construed as limiting the present patent;

[0063] Example 1

[0064] like Figure 1 As shown, this embodiment proposes a small sample expansion method based on prototype completion in image classification, which includes the following steps:

[0065] S1. Collect the original sample image dataset, and divide the original sample image dataset into a base class dataset, a support set, and a query set according to the image data category in the original sample image dataset;

[0066] In step S1, the original sample image dataset, i.e., the small sample image dataset, is to be expanded. The “according to the image data categories in the original sample image dataset” means that there are image data of a certain number of categories in the collected original sample image dataset. For example, the original sample image dataset is a mobile phone image, and the collection contains mobile phone images of different brands. Each brand is a category. Suppose there are a total of Z categories of image data in the original sample image dataset, the base class dataset belongs to category X, and the support set and the query set belong to category Y, satisfying:

[0067] X+Y=Z,

[0068] Here, among all the categories Z in the original sample image dataset, the category to which the base dataset belongs does not belong to the same category as the categories described in the latter two, and the support set and the query set belong to the same category.

[0069] S2. Select a feature extractor and train the feature extractor using the base class data set, and use the trained feature extractor to extract the data features of the base class data set, the support set data features, and the query set data features;

[0070] In step S2, the feature extractor is selected as the Wide ResNet network. In the process of training the Wide ResNet network using the base class data set, the loss function of the Wide ResNet network is set, and the weight parameters of the Wide ResNet network are updated by back propagation until the loss function converges. After the training is completed, the network parameters are fixed, and the trained feature extractor extracts the data features of the set target task.

[0071] Here, considering the universality of practical applications, a pre-trained feature extractor is used for feature extraction.

[0072] S3. Based on the spatial metric distance function, calculate the distance between the support set data features and the data features of all base class data sets, and determine the K samples that are most similar to the support set data features according to the distance;

[0073] Taking the mean of the data features of the i-th category of the base class data set as the general term for the data features of this category, it satisfies:

[0074]

[0075] Among them, μ 1i represents the mean value of the data feature of the i-th category of the base class data set, x jRepresents the jth data feature in the i-th data feature of the base class data set; n i Represents the number of data features of the i-th category of the base class data set;

[0076] In addition, the corresponding covariance matrix is ​​expressed as:

[0077]

[0078] Assume that the mean of the data features of the support dataset is μ 2 , in μ 2 As a general term for the data features of the supporting dataset, let the spatial metric distance function be collectively referred to as f d (), then based on the spatial metric distance function, when calculating the distance between the support set data features and the data features of all base class data sets, it satisfies:

[0079] d i =f d (μ i ,x s )

[0080] Among them, d i represents the distance between the support set data feature and the i-th data feature of the base class data set; finally, the distance set D is obtained, which is expressed as:

[0081] D={d 1 ,d 2 ,...,d i ,...,d q}

[0082] Among them, q represents the number of types of data features of the base class data set. The values ​​of all distance elements in the distance set D are arranged in ascending order. According to the arrangement order, the data features of the base class data set corresponding to the first K distance elements are selected as the K samples with the most similar data features of the support set.

[0083] Here, the spatial metric distance function can select different and specific metric functions. Considering that this implementation process is established on the premise that different spatial metric distance functions are applicable, after obtaining the distance between the support set data features and the data features of all base class data sets, the distance is normalized. This implementation method is applicable to different spatial metric functions. Taking the L1 similarity function as an example, the higher the similarity, the larger the function value, and this formula is used, and the expression is:

[0084]

[0085] Taking the JS divergence function as an example, the higher the similarity, the lower the function value, so you can use:

[0086]

[0087] Among them, d' i represents the normalized value of the distance. represents the maximum value in the distance, Indicates the minimum value in distance.

[0088] S4. Perform weighted calculation on the similarity between the K samples and the support set data features to obtain the prototype features of the support set;

[0089] The softmax function is used to make the sum of the distances equal to 1. In step S4, the similarity between the K samples and the support set data features is the mean μ of the K samples and the support set data features. 2 The distance between d () is obtained. Under the premise of distance normalization, the expression for weighted calculation is:

[0090]

[0091] Among them, e is a natural constant, w s is the weight of the support set data feature. Generally, the support set data sample is set to be only one or five. In the case of only one support set sample, w s Set to 1.5, 5 support set samples to 2.5, pre-set according to the number of samples in the support set, w' g represents the weight of the gth sample among K samples, d' g Represents the distance between the gth sample among K samples and the data feature of the support set;

[0092] The expression of the prototype characteristics is:

[0093]

[0094] Among them, n K represents the sum of the number of similar samples and the number of support set data features, μ' represents the prototype feature;

[0095] The covariance matrix C' corresponding to the prototype feature is:

[0096]

[0097] Among them, C g It represents the covariance matrix of the data features of the g-th similar sample among K samples, and α is a hyperparameter.

[0098] S5. Generate a pseudo sample set based on the prototype features of the support set;

[0099] In step S5, the generated pseudo sample set is D y , the expression is:

[0100]

[0101] Among them, y represents the newly generated category, Represents the newly generated pseudo samples, which obey the Gaussian distribution.

[0102] S6. Use the generated pseudo sample set and support set data features together as training data, and use the query set data features as test data to perform image classification.

[0103] In this step, a classifier is selected. The classifier can be a linear regression classification or a support vector machine. In actual implementation, you can choose one of them. The generated pseudo sample set and the support set data features are used as training data to train the classifier. Then the query set data features are used as test data to test the trained classifier to complete the image classification.

[0104] In general, in this embodiment, see Figure 2 First, the original sample image dataset is collected and divided into a base class dataset, a support set, and a query set. Figure 2 The base class images, support set images and query set images in the dataset are extracted, and the features of the three types of data sets are extracted based on the feature extractor. The samples are mapped from the space to the feature space, and the prototype features are completed in the feature space based on the similarity of the sample features to obtain the prototype features. Overall, only a feature library and feature extractor of the base class data set are needed to complete the prototypes of a small number of samples and generate a pseudo sample set, which greatly compensates for the shortcomings of small sample image classification, and does not introduce additional parameters that need to be trained, saving time and equipment requirements for subsequent image classification. Compared with the "black box" neural network model, it has extremely strong interpretability.

[0105] Example 2

[0106] This embodiment verifies the effectiveness of the method proposed in this application through specific experiments. The miniImagenet and CUB datasets used in the experiment are small sample datasets. The miniImagenet dataset contains 100 classes, including 600 images of size 84×84 pixels. It is divided into 64 base classes, 16 validation classes, and 20 new classes. The CUB dataset is a bird image dataset containing 200 species of birds, with a total of 11,788 images of size 84×84 pixels. It is divided into 100 base classes, 50 validation classes, and 50 new classes.

[0107] The types of algorithms involved in the comparison are: optimization-based methods, metric-based methods, and generation-based methods. The comparison results are shown in Table 1.

[0108] Table 1

[0109]

[0110]

[0111] From the results in Table 1, it can be found that the classification performance of the method proposed in this application is better than that of other comparison methods. The effectiveness of the present invention can be verified through the above simulation experiments. Figure 3 This is a dimensionality reduction visualization diagram of the pseudo sample set generated by the method proposed in this application. Figure 3 It can be seen that the generated pseudo sample clusters are all Gaussian distributed around the samples in the support set, and there are obvious discrimination boundaries between classes.

[0112] Example 3

[0113] like Figure 4 As shown, this embodiment proposes a small sample expansion system based on prototype completion in image classification, and the system includes:

[0114] The original sample image acquisition and division unit is used to acquire the original sample image data set, and divide the original sample image data set into a base class data set, a support set and a query set according to the image data category in the original sample image data set;

[0115] A feature extraction unit is used to select a feature extractor, train the feature extractor using a base class data set, and respectively extract data features of the base class data set, data features of the support set, and data features of the query set using the trained feature extractor;

[0116] The similar sample determination unit calculates the distance between the support set data features and the data features of all base class data sets based on the spatial metric distance function, and determines the K samples that are most similar to the support set data features according to the distance;

[0117] The prototype feature determination unit is used to perform weighted calculation on the similarity between the K samples and the support set data features to obtain the prototype features of the support set;

[0118] A pseudo sample set generation unit generates a pseudo sample set based on the prototype features of the support set;

[0119] The image classification unit uses the generated pseudo sample set and support set data features as training data, and uses the query set data features as test data to perform image classification.

[0120] Obviously, the above embodiments of the present invention are only examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A small sample expansion method based on prototype completion in image classification, It is characterized in that The method comprises the following steps: S1. Collect the original sample image dataset, and divide the original sample image dataset into a base class dataset, a support set, and a query set according to the image data category in the original sample image dataset; S2. Select a feature extractor and train the feature extractor using the base class data set, and use the trained feature extractor to extract the data features of the base class data set, the support set data features, and the query set data features; S3. Based on the spatial metric distance function, calculate the distance between the support set data features and the data features of all base class data sets, and determine the K samples that are most similar to the support set data features according to the distance; S4. Perform weighted calculation on the similarity between the K samples and the support set data features to obtain the prototype features of the support set; In step S4, the similarity between the K samples and the support set data features is the mean μ of the K samples and the support set data features. 2 The distance between d () is obtained. Under the premise of distance normalization, the expression for weighted calculation is: Among them, e is a natural constant, w s is the weight of the support set data feature, which is pre-set according to the number of samples in the support set, w' g represents the weight of the gth sample among K samples, d' g Represents the distance between the gth sample among K samples and the data feature of the support set; The expression of the prototype characteristics is: Among them, n K represents the sum of the number of similar samples and the number of support set data features, μ' represents the prototype feature; The covariance matrix C' corresponding to the prototype feature is: Among them, C g Represents the covariance matrix of the g-th similar sample data features among K samples, and α is a hyperparameter; S5. Generate a pseudo sample set based on the prototype features of the support set; S6. Use the generated pseudo sample set and support set data features together as training data, and use the query set data features as test data to perform image classification.

2. The small sample expansion method based on prototype completion in image classification according to claim 1, It is characterized in that In step S1, suppose that there are Z categories of image data in the original sample image dataset, the base dataset belongs to category X, the support set and the query set belong to category Y, and they satisfy:

3. The small sample expansion method based on prototype completion in image classification according to claim 1, It is characterized in that In step S2, the feature extractor is selected as the Wide ResNet network. In the process of training the Wide ResNet network using the base class data set, the loss function of the Wide ResNet network is set, and the weight parameters of the Wide ResNet network are updated by back propagation until the loss function converges. After the training is completed, the network parameters are fixed, and the trained feature extractor extracts the data features of the set target task.

4. The small sample expansion method based on prototype completion in image classification according to claim 3, It is characterized in that Taking the mean of the data features of the i-th category of the base class data set as the general term for the data features of this category, it satisfies: Among them, μ 1i represents the mean value of the data feature of the i-th category of the base class data set, x j Represents the jth data feature in the i-th data feature of the base class data set; n i Represents the number of data features of the i-th category of the base class data set; Assume that the mean of the data features of the support dataset is μ 2 , in μ 2 As a general term for the data features of the supporting dataset, let the spatial metric distance function be collectively referred to as f d (), then based on the spatial metric distance function, when calculating the distance between the support set data features and the data features of all base class data sets, it satisfies: d i =f d (μ i ,x s ) Among them, d i represents the distance between the support set data feature and the i-th data feature of the base class data set; finally, the distance set D is obtained, which is expressed as: D={d 1 ,d 2 ,...,d i ,...,d q } Among them, q represents the number of types of data features of the base class data set. The values ​​of all distance elements in the distance set D are arranged in ascending order. According to the arrangement order, the data features of the base class data set corresponding to the first K distance elements are selected as the K samples with the most similar data features of the support set.

5. The small sample expansion method based on prototype completion in image classification according to claim 4, It is characterized in that After obtaining the distance between the support set data features and the data features of all base class data sets, the distance is normalized and the expression is: or Among them, d i ' represents the normalized value of the distance, represents the maximum value in the distance, Indicates the minimum value in distance.

6. The small sample expansion method based on prototype completion in image classification according to any one of claims 1 to 5, It is characterized in that In step S5, the generated pseudo sample set is D y , the expression is: Among them, y represents the newly generated category, Represents the newly generated pseudo samples, which obey the Gaussian distribution.

7. The small sample expansion method based on prototype completion in image classification according to claim 6, It is characterized in that In step S6, a classifier is selected, and the generated pseudo sample set and support set data features are used as training data to train the classifier. Then, the query set data features are used as test data to test the trained classifier, thereby completing image classification.

8. The small sample expansion method based on prototype completion in image classification according to claim 7, It is characterized in that The classifier is either linear regression or support vector machine.

9. A small sample expansion system based on prototype completion in image classification, It is characterized in that The system comprises: The original sample image acquisition and division unit is used to acquire the original sample image data set, and divide the original sample image data set into a base class data set, a support set and a query set according to the image data category in the original sample image data set; A feature extraction unit is used to select a feature extractor, train the feature extractor using a base class data set, and respectively extract data features of the base class data set, data features of the support set, and data features of the query set using the trained feature extractor; The similar sample determination unit calculates the distance between the support set data features and the data features of all base class data sets based on the spatial metric distance function, and determines the K samples that are most similar to the support set data features according to the distance; The prototype feature determination unit is used to perform weighted calculation on the similarity between the K samples and the support set data features to obtain the prototype features of the support set; The similarity between the K samples and the support set data features is the mean μ of the K samples and the support set data features. 2 The distance between d () is obtained. Under the premise of distance normalization, the expression for weighted calculation is: Among them, e is a natural constant, w s is the weight of the support set data feature, which is pre-set according to the number of samples in the support set, w' g represents the weight of the gth sample among K samples, d' g Represents the distance between the gth sample among K samples and the data feature of the support set; The expression of the prototype characteristics is: Among them, n K represents the sum of the number of similar samples and the number of support set data features, μ' represents the prototype feature; The covariance matrix C' corresponding to the prototype feature is: Among them, C g Represents the covariance matrix of the g-th similar sample data features among K samples, and α is a hyperparameter; A pseudo sample set generation unit generates a pseudo sample set based on the prototype features of the support set; The image classification unit uses the generated pseudo sample set and support set data features as training data, and uses the query set data features as test data to perform image classification.

Citation Information

Patent Citations

  • Small sample image classification method based on base class sample feature synthesis

    CN114387473A

  • Small sample image increment classification method and device based on embedding enhancement and self-adaption

    CN114549894A