Fine-grained small-shot classification method based on task-specific channel reconstruction network
By using task-specific channels to reconstruct networks and multi-dimensional dynamic convolution technology in fine-grained small sample visual classification tasks, the problem of data scarcity is solved and efficient fine-grained prediction accuracy is achieved.
Patent Information
- Application Number
- CN202310807415.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-03
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-07-03
AI Technical Summary
The existing technology is difficult to effectively solve the data scarcity problem in the visual classification task of fine-grained small samples. The existing deep learning methods require a large number of samples with labeled information for training, and it is impossible to quickly learn fine-grained concepts from a few samples with supervised information.
A fine-grained small sample classification method based on task-specific channel reconstruction network is proposed. A multi-dimensional dynamic feature extraction network is constructed through multi-dimensional dynamic convolution, the features of support sets and query sets are calculated, the features are reconstructed using task-specific channel attention weights, and finally the similarity score is used for classification using distance metrics.
This method can efficiently learn the concept of fine-grained size in a small number of samples, achieve a high precision of fine-grained prediction, and effectively solve the problem of data scarcity in the classification task of fine-grained small sample size.
Smart Images

Figure CN116843970B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a fine-grained small sample classification method based on a task-specific channel reconstruction network. Background Art
[0002] Compared with image data containing only coarse-grained categories, the cost and difficulty of collecting and annotating fine-grained visual data are greater, and often require expert-level knowledge in a specific field to complete. For fine-grained classification tasks, the available data for some fine-grained subcategories is limited. For example, when a new fine-grained subcategory of a species is discovered, its available data samples are often very scarce. Humans can abstract new fine-grained concepts from a few samples with supervised information and quickly distinguish subsequent new samples belonging to the same category, while the current fine-grained classification methods under deep learning still require a large number of samples with labeled information for training, and there is a significant difference between the two. Inspired by this, small sample methods have become an important means to solve the data scarcity problem faced by fine-grained tasks. Therefore, some researchers have recently paid attention to fine-grained small sample visual classification methods, expecting the model to learn the conceptual knowledge of visually similar fine-grained categories from a limited number of fine-grained labeled samples (for example, only using 1 or 5 fine-grained labeled samples).
[0003] The rapid development of technologies in the field of small sample learning has provided a solution to the fine-grained problem in small sample settings. However, fine-grained small sample visual classification tasks have the dual challenges of fine-grained classification and small sample learning. Simply using small sample classification methods cannot effectively solve the fine-grained small sample problem. The method design must be fully combined with the characteristics of fine-grained tasks. Summary of the invention
[0004] In view of this, and in view of the defects and shortcomings of the prior art, the purpose of the present invention is to propose a fine-grained small sample classification method based on a task-specific channel reconstruction network, which can accurately and effectively perform fine-grained small sample classification.
[0005] First, a fine-grained small sample classification dataset is obtained for data preprocessing and label extraction. The dataset is divided into three parts, and N-way K-shot small sample meta-task sampling is performed according to the episode mode to obtain query set images and support set images of N categories. Then, a multi-dimensional dynamic feature extraction network is constructed through multi-dimensional dynamic convolution. All support set images and query set images are input into the multi-dimensional dynamic feature extraction network to obtain the features of the support set and query set of each category, and the prototype features of each category of the support set are calculated. Then, a weight generation block is constructed to calculate the initial weights of the support set and query set. The initial weights are adaptively aggregated to obtain task-specific channel attention weights, and the support set and query set features are reconstructed using the task-specific channel attention weights. Finally, the distance metric is used to perform similarity scoring to obtain the category to which the query set image belongs, and the fine-grained small sample classification task is completed efficiently.
[0006] The technical solution specifically adopted by the present invention to solve the technical problem is:
[0007] A fine-grained small sample classification method based on a task-specific channel reconstruction network comprises the following steps;
[0008] Step S1: Obtain a fine-grained small sample classification dataset for data preprocessing and complete label extraction; divide the dataset into D base , D val and D novel In the third part, N-way K-shot small sample meta-task sampling is performed according to the episode mode to obtain query set images and support set images of N categories;
[0009] Step S2: construct a multi-dimensional dynamic feature extraction network through multi-dimensional dynamic convolution, input all support set images and query set images into the multi-dimensional dynamic feature extraction network, obtain the features of the support set and query set of each category respectively, and calculate the prototype features of each category of the support set;
[0010] Step S3: construct a weight generation block, and input the support set features of the ch_i-th class to obtain the initial weight of the support set of the ch_i-th class, input the query set features into the weight generation block to obtain the initial weight of the query set, adaptively aggregate the initial weight of the support set of the ch_i-th class and the initial weight of the query set to obtain the task-specific channel attention weight of the ch_i-th class, and reconstruct the support set features and the query set features using the task-specific channel attention weight to obtain the support set reconstruction features of the ch_i-th class and the query set reconstruction features of the ch_i-th class;
[0011] Step S4: Use distance metrics to perform similarity scoring to obtain the category to which the query set image belongs: iterative training is performed on multiple sampling meta-tasks according to the specified training parameters, the model parameters are updated by optimizing the combined loss, the optimal model is continuously saved according to the verification accuracy, and the average classification accuracy of the optimal model in all randomly sampled episodes in the meta-test phase is calculated.
[0012] Furthermore, step S1 specifically includes the following steps:
[0013] Step S11: Use a public fine-grained small sample classification data set to perform data preprocessing to complete label extraction;
[0014] Step S12: For a given fine-grained image dataset D, divide it into a basic dataset D base , validation data set D val And the new class dataset D novel Three parts, represented by D base ={(x base_i ,y base_i ), y base_i ∈Y base}、D val ={(x val_i ,y val_i ), y val_i ∈Y val} and D novel ={(x novel_i ,y novel_i ), y novel_i ∈Y novel}, where Y base ∪Y val ∪Y novel =Y, where Y represents the label space of the original dataset, Y base represents the label space of the basic dataset, Y val represents the label space of the validation dataset, Y novel Represents the label space of the new class dataset; the categories of each part do not intersect with each other, that is, Basic dataset D base Each category in contains more labeled samples than the validation dataset D val And the new class dataset D novel Sample labels included in ;
[0015] Step S13: Construct an N-way K-shot fine-grained small sample classification setting, where all meta-tasks are randomly sampled based on episodes; each episode consists of a support set S and a query set Q, which contains N fine-grained categories in total, and each category contains K+M samples; for a certain category, the corresponding support set contains only K labeled image samples, and the query set contains M unlabeled image samples; the support set S and query set Q in each episode are defined as In this way, we obtain the query set images and the support set images of N categories.
[0016] Furthermore, step S2 specifically includes the following steps:
[0017] Step S21: construct a multi-dimensional dynamic convolution method: for the input feature map fea_x, the output feature map extracted by the multi-dimensional dynamic convolution ODConv is represented as fea_x′, and the specific calculation method of fea_x′ is as follows: fea_x′=(a f ⊙a c ⊙a s ⊙CONV_W)*fea_x
[0018] Among them, a f is the attention weight of the convolution kernel output channel, which assigns different attention weights to different convolution kernels with different number of channels of the output feature map. a c is the attention weight of the convolution kernel input channel, which assigns different weights to different channels of the convolution kernel to dynamically extract the features of different channels of the input feature map. c out Indicates the number of output channels, c cin Indicates the number of input channels; a s is the spatial attention weight of the convolution kernel, which assigns different attention weights to different spatial positions of the convolution kernel. s ∈R k×k , k represents the spatial size of the convolution kernel; CONV_W represents the convolution kernel parameters, CONV_W conv_n Represents the parameters of the conv_nth convolution kernel, conv_n∈[1,c out ];
[0019] Step S22: using the dynamic feature extraction method in step S21, using multi-dimensional dynamic convolution ODConv to improve the two feature extraction network structures of Conv-4 and ResNet-12 respectively, and obtaining two multi-dimensional dynamic feature extraction networks; using a single convolution kernel of 3×3 size ODConv to replace the second, third and fourth 3×3 size two-dimensional convolutions in the Conv4 network, and using a single convolution kernel of 3×3 size ODConv to replace the 3×3 size two-dimensional convolution in the BasicBlock of ResNet-12, so as to use the optimized ODBasicBlock to replace the last three network modules of ResNet-12; the size of the feature map of each stage before and after the improvement remains unchanged;
[0020] Step S23: According to the requirements of the fine-grained small sample classification task, a multi-dimensional dynamic feature extraction network is selected; all support set images and query set images are input into the multi-dimensional dynamic feature extraction network to obtain the support set features and query set features of each category respectively, and the query image x of the current original task is Q and the im_jth support image of the ch_ith category The input parameters are respectively DEFN_θ, multi-dimensional dynamic feature extraction network f DEFN_θ (·) to generate the feature maps F of the query set and support set images Q and The specific calculation method is as follows:
[0021] F Q =f DEFN_θ (x Q )
[0022]
[0023] Step S24: Calculate the support set prototype feature of each category. The support set category prototype feature of the ch_i-th category is expressed as The specific calculation method is as follows:
[0024]
[0025] Among them, N ch_i Indicates the number of samples in the ch_i class.
[0026] Furthermore, step S3 specifically includes the following steps:
[0027] Step S31: The spatial information of the query set features and the support set prototype features are aggregated through the global average pooling operation GAP, and the query set features F Q and the support set features of the ch_i class Get the initial query set weights w respectively Qand the support weight of the ch_i class The specific calculation method is as follows:
[0028] w Q = GAP(F Q )
[0029]
[0030] Step S32: Construct a weight generation block FCB(·) to convert the initial query set weight w Q and the support weight of the ch_i class Input the weight generation block FCB(·) to generate the support set attention weight of the ch_i-th class and the query set attention weight of the ch_i-th class, and get the query attention weight w′ of the ch_i-th class Q and the support attention weight of the ch_i class The specific calculation method is as follows:
[0031]
[0032]
[0033] in, is the weight of the fully connected layer, δ(·) represents the ReLU function, and σ(·) represents the 1+Tanh function;
[0034] Step S33: The channel attention weight w′ obtained in step S32 Q and We further use task-specific adaptive aggregation to highlight the task-specific key semantics to the greatest extent possible, and obtain the task-specific channel attention weight tw of the ch_i class: ch_i , the specific calculation process is as follows:
[0035]
[0036] Among them, τ represents the learnable parameter, τ∈[0,1];
[0037] Step S34: Using the task-specific channel attention weight tw obtained in step S33 ch_i To reconstruct the supporting prototype feature map of the ch_i class and query feature graph F Q The feature channel of ch_i is used to obtain the channel reconstruction feature map of the ch_i class and The specific calculation is as follows:
[0038]
[0039]
[0040] in, express The value in the dim_jth dimension, and Respectively represent F Q and The ch_jth channel feature in the channel dimension, C represents the total number of channels of the feature map.
[0041] Furthermore, step S4 specifically includes the following steps:
[0042] Step S41: Use distance metric to perform similarity scoring to obtain the category to which the query set image belongs; the probability p(y=ch_i|x) that the query image x belongs to the ch_i-th category is calculated as follows:
[0043]
[0044] Where γ represents the learnable temperature parameter, dis(·,·) represents the distance metric function, and CLS_N represents the number of categories;
[0045] Step S42: Calculate the cross entropy loss function using the probability p(y=ch_i|x) obtained in step S41, set the specified training parameters, and continuously update the gradient to perform iterative training on all episodes in the meta-training phase;
[0046] Step S43: Calculate the average classification accuracy of the model in all randomly sampled episodes in the meta-test phase The calculation formula is as follows:
[0047]
[0048] Among them, EPI_N represents the total number of episodes in the test phase, Acc epi_i Represents the classification accuracy of all queries in the epi_ith episode;
[0049] Step S44: During the training process, the model is verified at intervals of the number of iterations according to the verification interval flag, and the optimal model is continuously saved. When the number of iterations reaches a preset maximum number of iterations threshold, the training process ends.
[0050] Compared with the prior art, the present invention and its preferred embodiment have the following beneficial effects:
[0051] 1. In view of the scarcity of data in fine-grained classification tasks, a fine-grained small sample classification method based on task-specific channel reconstruction network is constructed, which can achieve high fine-grained prediction accuracy using a small number of samples.
[0052] 2. Aiming at the characteristics of fine-grained small sample classification tasks, the existing commonly used feature extraction network for small sample classification is improved by utilizing multi-dimensional attention at the convolution kernel level to better adapt to fine-grained small sample visual classification tasks.
[0053] 3. We analyze the problem of task-specific category-level semantic activation differences in fine-grained small-sample visual classification and propose a task-specific channel attention method that uses channel attention weights to reconstruct feature channels of the query set and support set to achieve better metric matching between fine-grained subcategories.
[0054] 4. A general fine-grained small sample classification method is proposed at the cost of extremely small parameters and computational complexity. The two component methods it contains can be effectively applied to different metric learning methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0056] Figure 1 The figure is a schematic diagram of the principle and implementation flow of an embodiment of the present invention. DETAILED DESCRIPTION
[0057] In order to make the features and advantages of this patent more obvious and easy to understand, the following embodiments are specifically described in detail as follows:
[0058] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0059] like Figure 1 As shown, this embodiment provides a fine-grained small sample classification method based on a task-specific channel reconstruction network, which specifically includes the following steps:
[0060] Step S1: Obtain a fine-grained small sample classification dataset for data preprocessing and complete label extraction. Divide the dataset into D base , D val and D novel In the third part, N-way K-shot small sample meta-task sampling is performed according to the episode mode to obtain query set images and support set images of N categories;
[0061] Step S2: construct a multi-dimensional dynamic feature extraction network through multi-dimensional dynamic convolution, input all support set images and query set images into the multi-dimensional dynamic feature extraction network, obtain the features of the support set and query set of each category respectively, and calculate the prototype features of each category of the support set;
[0062] Step S3: construct a weight generation block, and input the support set features of the ch_i-th class to obtain the initial weight of the support set of the ch_i-th class, input the query set features into the weight generation block to obtain the initial weight of the query set, adaptively aggregate the initial weight of the support set of the ch_i-th class and the initial weight of the query set to obtain the task-specific channel attention weight of the ch_i-th class, and reconstruct the support set features and the query set features using the task-specific channel attention weight to obtain the support set reconstruction features of the ch_i-th class and the query set reconstruction features of the ch_i-th class;
[0063] Step S4: Use the distance metric to perform similarity scoring to obtain the category to which the query set image belongs, perform iterative training on multiple sampling meta-tasks according to the specified training parameters, update the model parameters by optimizing the combined loss, continuously save the optimal model based on the verification accuracy, and calculate the average classification accuracy of the optimal model in all randomly sampled episodes in the meta-test phase.
[0064] In this embodiment, step S1 specifically includes the following steps:
[0065] Step S11: Use a public fine-grained small sample classification data set to perform data preprocessing to complete label extraction;
[0066] Step S12: For a given fine-grained image dataset D, divide it into a basic dataset D base , validation data set D val And the new class dataset D novel Three parts, represented by D base ={(x base_i ,y base_i ), y base_i ∈Y base}、D val ={(x val_i ,y val_i ), y val_i ∈Y val} and D novel ={(x novel_i ,y novel_i ), y novel_i ∈Y novel}, where Y base ∪Y val ∪Y novel =Y, where Y represents the label space of the original dataset, Y base represents the label space of the basic dataset, Y val represents the label space of the validation dataset, Y novel Represents the label space of the new class dataset. Under this partition setting, the categories of each part do not intersect with each other, that is, Basic dataset D baseEach category in contains enough labeled samples, and the validation dataset D val And the new class dataset D novel Contains a small number of labeled samples;
[0067] Step S13: Construct an N-way K-shot fine-grained small sample classification setting, where all meta-tasks are randomly sampled based on episodes. Each episode consists of a support set S and a query set Q, which contains N fine-grained categories in total, and each category contains K+M samples. For a certain category, the support set of the category only contains K labeled image samples, and the query set contains M unlabeled image samples. The support set S and query set Q in each episode are defined as In this way, we obtain the query set images and the support set images of N categories.
[0068] In this embodiment, step S2 specifically includes the following steps:
[0069] Step S21: Construct a multi-dimensional dynamic convolution method. For the input feature map fea_x, the output feature map extracted by multi-dimensional dynamic convolution ODConv is represented as fea_x′. The specific calculation method of fea_x′ is as follows
[0070] fea_x′=(a f ⊙a c ⊙a s ⊙CONV_W)*fea_x
[0071] where a f is the attention weight of the convolution kernel output channel, which assigns different attention weights to different convolution kernels with different number of channels of the output feature map. a c is the attention weight of the convolution kernel input channel, which assigns different weights to different channels of the convolution kernel to dynamically extract the features of different channels of the input feature map. c out Indicates the number of output channels, c cin Indicates the number of input channels. s is the spatial attention weight of the convolution kernel, which assigns different attention weights to different spatial positions of the convolution kernel. s ∈R k×k , k represents the spatial size of the convolution kernel. CONV_W represents the convolution kernel parameters, CONV_W conv_n Represents the parameters of the conv_nth convolution kernel, conv_n∈[1,c out ];
[0072] Step S22: Use the dynamic feature extraction method ODConv in step S21 to improve the two feature extraction network structures of Conv-4 and ResNet-12 commonly used for small sample classification, and obtain two multi-dimensional dynamic feature extraction networks. Use a single convolution kernel of 3×3 size ODConv to replace the second, third and fourth 3×3 size two-dimensional convolutions in the Conv4 network, and use a single convolution kernel of 3×3 size ODConv to replace the 3×3 size two-dimensional convolution in the BasicBlock of ResNet-12, so as to use the optimized ODBasicBlock to replace the last three network modules of ResNet-12. The size of the feature map of each stage remains unchanged before and after the improvement;
[0073] Step S23: According to the requirements of the fine-grained small sample classification task, a multi-dimensional dynamic feature extraction network is selected. All support set images and query set images are input into the multi-dimensional dynamic feature extraction network to obtain the support set features and query set features of each category respectively, and the query image x of the current original task is Q and the im_jth support image of the ch_ith category The input parameters are respectively DEFN_θ, multi-dimensional dynamic feature extraction network f DEFN_θ (·) to generate the feature maps F of the query set and support set images Q and The specific calculation method is as follows
[0074] F Q =f DEFN_θ (x Q )
[0075]
[0076] Step S24: Calculate the support set prototype feature of each category. The support set category prototype feature of the ch_i-th category is expressed as The specific calculation method is as follows
[0077]
[0078] Among them, N ch_i Indicates the number of samples in the ch_i class;
[0079] In this embodiment, step S3 specifically includes the following steps:
[0080] Step S31: Gather the spatial information of query set features and support set prototype features through global average pooling (GAP) operation, and input query set features F Q and the support set features of the ch_i class Get the initial query set weights w respectively Qand the support weight of the ch_i class The specific calculation method is as follows
[0081] w Q = GAP(F Q )
[0082]
[0083] Step S32: Construct a weight generation block FCB(·) to convert the initial query set weight w Q and the support weight of the ch_i class Input the weight generation block FCB(·) to generate the support set attention weight of the ch_i-th class and the query set attention weight of the ch_i-th class, and get the query attention weight w′ of the ch_i-th class Q and the support attention weight of the ch_i class The specific calculation method is as follows
[0084]
[0085]
[0086] in, is the weight of the fully connected layer, δ(·) represents the ReLU function, and σ(·) represents the 1+Tanh function;
[0087] Step S33: The channel attention weight w′ obtained in step S32 Q and We further use task-specific adaptive aggregation to highlight the task-specific key semantics to the greatest extent possible, and obtain the task-specific channel attention weight tw of the ch_i class: ch_i The specific calculation process is as follows
[0088]
[0089] Among them, τ represents the learnable parameter, τ∈[0,1].
[0090] Step S34: Using the task-specific channel attention weight tw obtained in step S33 ch_i To reconstruct the supporting prototype feature map of the ch_i class and query feature graph F Q The feature channel of ch_i is used to obtain the channel reconstruction feature map of the ch_i class and The specific calculation is as follows
[0091]
[0092]
[0093] in, express The value in the dim_jth dimension, and Respectively represent F Q and The ch_jth channel feature in the channel dimension, C represents the total number of channels of the feature map;
[0094] In this embodiment, step S4 specifically includes the following steps:
[0095] Step S41: Use distance metric to perform similarity scoring to obtain the category to which the query set image belongs. The probability p(y=ch_i|x) that the query image x belongs to the ch_ith category is calculated as follows:
[0096]
[0097] Where γ represents the learnable temperature parameter, dis(·,·) represents the distance metric function, and CLS_N represents the number of categories;
[0098] Step S42: Use the probability p(y=ch_i|x) obtained in step S41 to calculate the cross entropy loss function, set the specified training parameters, and continuously update the gradient to iteratively train all episodes in the meta-training phase.
[0099] Step S43: Calculate the average classification accuracy of the model in all randomly sampled episodes in the meta-test phase The calculation formula is as follows
[0100]
[0101] Among them, EPI_N represents the total number of episodes in the test phase, Acc epi_i Represents the classification accuracy of all queries in the epi_ith episode;
[0102] Step S44: During the training process, the model is verified at a certain iteration interval according to the verification interval flag, and the optimal model is continuously saved. When the number of iterations reaches a preset maximum iteration threshold, the training process ends;
[0103] In summary, the present invention proposes an effective solution to the task specificity of subtle discriminative semantics in a small sample setting. Subtle discriminative feature learning is crucial for fine-grained visual classification tasks, but most existing small sample learning methods ignore this point and cannot be effectively applied to fine-grained visual classification. To this end, the present invention first optimizes the dense feature representation construction process for small sample learning, and uses a multi-dimensional dynamic feature extraction network to capture fine-grained image subtle difference information. On this basis, the present invention further proposes a task-specific channel attention method, which uses task-specific channel attention weights to activate key semantic information specific to fine-grained small sample tasks. The method proposed in the present invention can efficiently learn fine-grained conceptual knowledge from a small number of samples, thereby achieving a higher fine-grained prediction accuracy to meet the dual challenges of fine-grained classification and small sample learning.
[0104] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0105] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0106] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0108] This patent is not limited to the above-mentioned optimal implementation mode. Anyone can derive various other forms of fine-grained small sample classification methods based on task-specific channel reconstruction networks under the inspiration of this patent. All equal changes and modifications made according to the scope of the patent application of the present invention should be covered by this patent.
Claims
1. A fine-grained small sample classification method based on a task-specific channel reconstruction network, characterized in that: The steps include: Step S1: Obtain a fine-grained small sample classification dataset for data preprocessing and complete label extraction; divide the dataset into D base , D val and D novel In the third part, N-way K-shot small sample meta-task sampling is performed according to the episode mode to obtain query set images and support set images of N categories; Step S2: construct a multi-dimensional dynamic feature extraction network through multi-dimensional dynamic convolution, input all support set images and query set images into the multi-dimensional dynamic feature extraction network, obtain the features of the support set and query set of each category respectively, and calculate the prototype features of each category of the support set; Step S3: construct a weight generation block, and input the support set features of the ch_i-th class to obtain the initial weight of the support set of the ch_i-th class, input the query set features into the weight generation block to obtain the initial weight of the query set, adaptively aggregate the initial weight of the support set of the ch_i-th class and the initial weight of the query set to obtain the task-specific channel attention weight of the ch_i-th class, and reconstruct the support set features and the query set features using the task-specific channel attention weight to obtain the support set reconstruction features of the ch_i-th class and the query set reconstruction features of the ch_i-th class; Step S4: Use distance metrics to perform similarity scoring to obtain the category to which the query set image belongs: perform iterative training on multiple sampling meta-tasks according to the specified training parameters, update the model parameters by optimizing the combined loss, continuously save the optimal model according to the verification accuracy, and calculate the average classification accuracy of all randomly sampled episodes of the optimal model in the meta-test phase; Step S3 specifically includes the following steps: Step S31: The spatial information of the query set features and the support set prototype features are aggregated through the global average pooling operation GAP, and the query set features F Q and the support set features of the ch_i class Get the initial query set weights w respectively Q and the support weight of the ch_i class The specific calculation method is as follows: w Q =GAP(F Q ) Step S32: Construct a weight generation block FCB(·) to convert the initial query set weight w Q and the support weight of the ch_i class Input the weight generation block FCB(·) to generate the support set attention weight of the ch_i-th class and the query set attention weight of the ch_i-th class, and get the query attention weight w' of the ch_i-th class Q and the support attention weight of the ch_i class The specific calculation method is as follows: in, is the weight of the fully connected layer, δ(·) represents the ReLU function, and σ(·) represents the 1+Tanh function; Step S33: The channel attention weight w' obtained in step S32 Q and We further use task-specific adaptive aggregation to highlight the task-specific key semantics to the greatest extent possible, and obtain the task-specific channel attention weight tw of the ch_i class: ch_i , the specific calculation process is as follows: Among them, τ represents the learnable parameter, τ∈[0,1]; Step S34: Using the task-specific channel attention weight tw obtained in step S33 ch_i To reconstruct the supporting prototype feature map of the ch_i class and query feature graph F Q The feature channel of ch_i is used to obtain the channel reconstruction feature map of the ch_i class and The specific calculation is as follows: in, express The value in the dim_jth dimension, and Respectively represent F Q and The ch_jth channel feature in the channel dimension, C represents the total number of channels of the feature map.
2. The fine-grained small sample classification method based on task-specific channel reconstruction network according to claim 1 is characterized in that: Step S1 specifically includes the following steps: Step S11: Use a public fine-grained small sample classification data set to perform data preprocessing to complete label extraction; Step S12: For a given fine-grained image dataset D, divide it into a basic dataset D base , validation data set D val And the new class dataset D novel Three parts, represented by D base ={(x base_i ,y base_i ),y base_i ∈Y base }、D val ={(x val_i ,y val_i ),y val_i ∈Y val } and D novel ={(x novel_i ,y novel_i ),y novel_i ∈Y novel }, where Y base ∪Y val ∪Y novel =Y, where Y represents the label space of the original dataset, Y base represents the label space of the basic dataset, Y val represents the label space of the validation dataset, Y novel Represents the label space of the new class dataset; the categories of each part do not intersect with each other, that is, Basic dataset D base Each category in contains more labeled samples than the validation dataset D val And the new class dataset D novel A sample of labels included in ; Step S13: Construct an N-way K-shot fine-grained small sample classification setting, where all meta-tasks are randomly sampled based on episodes; each episode consists of a support set S and a query set Q, which contains N fine-grained categories in total, and each category contains K+M samples; for a certain category, the corresponding support set contains only K labeled image samples, and the query set contains M unlabeled image samples; the support set S and query set Q in each episode are defined as In this way, we obtain the query set images and the support set images of N categories.
3. The fine-grained small sample classification method based on task-specific channel reconstruction network according to claim 2 is characterized in that: Step S2 specifically includes the following steps: Step S21: Construct a multi-dimensional dynamic convolution method: For the input feature map fea_x, the output feature map extracted by multi-dimensional dynamic convolution ODConv is represented as fea_x'. The specific calculation method of fea_x' is as follows fea_x′=(a f ⊙a c ⊙a s ⊙CONV_W)*fea_x Among them, a f is the attention weight of the convolution kernel output channel, which assigns different attention weights to different convolution kernels with different number of channels of the output feature map. a c is the attention weight of the convolution kernel input channel, which assigns different weights to different channels of the convolution kernel to dynamically extract the features of different channels of the input feature map. c out Indicates the number of output channels, c cin Indicates the number of input channels; a s is the spatial attention weight of the convolution kernel, which assigns different attention weights to different spatial positions of the convolution kernel. s ∈R k×k , k represents the spatial size of the convolution kernel; CONV_W represents the convolution kernel parameters, CONV_W conv_n Represents the parameters of the conv_nth convolution kernel, conv_n∈[1,c out ]; Step S22: using the dynamic feature extraction method in step S21, using multi-dimensional dynamic convolution ODConv to improve the two feature extraction network structures of Conv-4 and ResNet-12 respectively, and obtaining two multi-dimensional dynamic feature extraction networks; using a single convolution kernel of 3×3 size ODConv to replace the second, third and fourth 3×3 size two-dimensional convolutions in the Conv4 network, and using a single convolution kernel of 3×3 size ODConv to replace the 3×3 size two-dimensional convolution in the BasicBlock of ResNet-12, so as to use the optimized ODBasicBlock to replace the last three network modules of ResNet-12; the size of the feature map of each stage before and after the improvement remains unchanged; Step S23: According to the requirements of the fine-grained small sample classification task, a multi-dimensional dynamic feature extraction network is selected; all support set images and query set images are input into the multi-dimensional dynamic feature extraction network to obtain the support set features and query set features of each category respectively, and the query image x of the current original task is Q and the im_jth support image of the ch_ith category The input parameters are respectively DEFN_θ, multi-dimensional dynamic feature extraction network f DEFN_θ (·) to generate the feature maps F of the query set and support set images Q and The specific calculation method is as follows: F Q =f DEFN_θ (x Q ) Step S24: Calculate the support set prototype feature of each category. The support set category prototype feature of the ch_i-th category is expressed as The specific calculation method is as follows: Among them, N ch_i Indicates the number of samples in the ch_i class.
4. The fine-grained small sample classification method based on task-specific channel reconstruction network according to claim 3 is characterized by: Step S4 specifically includes the following steps: Step S41: Use distance metric to perform similarity scoring to obtain the category to which the query set image belongs; the probability p(y=ch_i|x) that the query image x belongs to the ch_i-th category is calculated as follows: Where γ represents the learnable temperature parameter, dis(·,·) represents the distance metric function, and CLS_N represents the number of categories; Step S42: Calculate the cross entropy loss function using the probability p(y=ch_i|x) obtained in step S41, set the specified training parameters, and continuously update the gradient to perform iterative training on all episodes in the meta-training phase; Step S43: Calculate the average classification accuracy of the model in all randomly sampled episodes in the meta-test phase The calculation formula is as follows: Among them, EPI_N represents the total number of episodes in the test phase, Acc epi_i Represents the classification accuracy of all queries in the epi_ith episode; Step S44: During the training process, the model is verified at intervals of the number of iterations according to the verification interval flag, and the optimal model is continuously saved. When the number of iterations reaches a preset maximum number of iterations threshold, the training process ends.
Citation Information
Patent Citations
Strawberry malformation state detection method based on small sample fine-grained image analysis
CN111582337A
Few-sample fine-grained image classification method for pyramid separation of double attention
CN114792385A