Semantically guided small sample image classification method

Through the semantic-guided small sample image classification method, combined with the technology of data augmentation, two-stage semantic-guided data augmentation, cross-domain alternating training and feature fusion module, the problem of small sample image classification in data sparse scenarios is solved, and efficient classification accuracy and generalization ability are achieved.

CN119992218APending Publication Date: 2025-05-13CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510228812.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the field of computer vision, especially in new category recognition under limited annotation data, the prior art is difficult to effectively solve the problem of small sample image classification, especially in scenarios with sparse data.

Method used

The semantic-guided small sample image classification method is adopted to generate high-quality enhanced data sets through the combination technology of data augmentation, two-stage semantic-guided data augmentation, cross-domain alternating training and feature fusion modules, and improve the generalization ability and adaptability of the model.

Benefits of technology

This method can improve the classification accuracy and generalization ability of the model in a data scarce environment, support the dynamic growth of image sets, and omit the model training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992218A_ABST
    Figure CN119992218A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic-guided small sample image classification method. The method comprises the following steps: performing random data enhancement on a sample by using a data enhancement method; performing data enhancement on the sample by using a two-stage semantic guidance data enhancement method; performing basic training on the training set of the base domain; performing enhancement training on the training set of the enhancement domain; the cross-domain alternate training is adopted to alternately carry out the steps S300 and S400 to train the model, so that the adaptability and generalization ability of the model to different fields are improved, and meanwhile, the influence of overfitting is reduced; the model trained in the step S500 is used for training a feature fusion module; and after model training is completed, feature fusion is carried out by utilizing the enhanced test set and the support set, a support set prototype is obtained and used for carrying out small sample classification in the query set, and finally a classification result is obtained. By combining strong semantic information, cross-domain alternate training and feature fusion, an effective solution is provided for small sample learning, and the classification accuracy and generalization ability of the model can be improved in a data scarce environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically to a semantically guided small sample image classification method. Background Art

[0002] In the field of computer vision, especially in the study of FSL, how to learn and recognize new categories from limited annotated data is a long-standing challenge. In the context of FSL, model-agnostic meta-learning (MAML) proposed by Finn et al. in 2017 provides an effective tool for quickly adapting to new tasks. However, meta-learning methods rely on a large amount of data, which becomes a problem in scenarios with sparse data. To address this problem, Gao et al. introduced Meta-Adam in 2023, a momentum-based meta-learning optimizer designed for FSL tasks, which enhances the generalization and adaptability of the model. In the field of metric learning, the Siamese network proposed by Koch et al. in 2015 provides a powerful framework for metric learning, optimizing the spatial arrangement of sample features through contrastive learning. Oreshkin et al. introduced TADAM in 2018 to improve FSL results through task-dependent adaptive metric learning, demonstrating the flexibility of metric learning in adapting to new tasks. In terms of data augmentation, Yun et al. demonstrated in 2019 that data augmentation methods based on weak semantic guidance can enhance the support set, but still face performance bottlenecks. In contrast, methods based on strong semantic guidance, such as the style-based generative adversarial network (Style-based GAN) proposed by Karras et al. in 2019, can utilize rich semantic information, but may introduce samples that are inconsistent with the semantics of the original data, leading to potential bias and degradation of model performance. In the field of language-image pre-training, the CLIP model proposed by Radford et al. in 2021 demonstrated zero-sample transfer capabilities in open vocabulary visual recognition tasks by pre-training on a large number of language-image pairs, and revealed the intrinsic connection between text embeddings and image category prototypes. These studies provide a theoretical basis and technical support for the proposal of this invention. Summary of the invention

[0003] The technical problem to be solved by the present invention is to provide a semantically guided small sample image classification method, which locates video semantic objects by establishing a match between an input video and an image set, omits the model training process, and supports the dynamic growth of the image set.

[0004] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows.

[0005] A semantically guided small sample image classification method includes the following steps: S100. Perform random data enhancement on samples using data enhancement methods to improve sample diversity; S200. Use a two-stage semantically guided data enhancement method to enhance the data of samples to generate high-quality enhanced datasets; S300. Perform basic training on the training set of the base domain to train the initial performance of the model; S400. Perform enhanced training on the training set of the enhanced domain, so that the model obtains stronger generalization performance from the high-quality enhanced data set; S500. Using cross-domain alternating training to alternately train the model in steps S300 and S400 to improve the adaptability and generalization ability of the model to different fields while reducing the impact of overfitting; S600. Using the model trained in S500 to train the feature fusion module; S700. After the model training is completed, the enhanced test set and the support set are used for feature fusion to obtain the support set prototype, which is used to perform small sample classification in the query set to finally obtain the classification result.

[0006] To further optimize the technical solution, in step S100, random enhancement is used to perform weak data enhancement for each training sample, including the following steps: S101.AutoContrast: Automatically adjust the contrast of the image; S102.Brightness: adjust the brightness of the image; S103.Color: adjust the color intensity of the image; S104.Contrast: adjust the contrast of the image; S105.Cutout: Randomly select an area on the image and replace it with gray; S106.CutoutAbs: randomly selects an area on the image, allows you to specify the size of the area to be cut, and replaces it with gray; S107.Equalize: Apply histogram equalization to enhance details in low-contrast areas of the image; S108.Identity: Do not change the image, as a baseline; S109.Invert: Invert the colors of the image; S110.Posterize: Reduces the tones of an image to a specified amount; S111.Rotate: randomly rotate the image; S112.Sharpness: adjust the sharpness of the image; S113.ShearX: Shears the image along the X axis; S114.ShearY: shears the image along the Y axis; S115.Solarize: Flips the pixel values ​​of an image to create an overexposed effect; S116.SolarizeAdd: flips the pixel values ​​of the image and adds a threshold parameter to produce an overexposure effect; S117.TranslateX: translate the image along the X axis; S118.TranslateY: Translates the image along the Y axis.

[0007] To further optimize the technical solution, in step S200, the method of performing data enhancement on the sample using the two-stage semantic guidance data enhancement method includes the following steps: Step S210: In the training phase, a template is constructed using text labels, and a stable diffusion model is combined to generate an enhanced training set; Step S220: During the testing phase, a text inversion algorithm is used to invert the test support set into text embedding vectors and alternative labels to generate an enhanced training set.

[0008] To further optimize the technical solution, in step S201, during the training phase, all categories in the training set have text labels. , using text labels Building a template , and combined with the stable diffusion model to generate enhanced training sets, the formula is as follows:

[0009] Where N is the number of categories, The number of augmented samples generated for each class, It is a stable diffusion model, in which the specific formula of the text label template is as follows: , Data augmentation using text templates from multiple views to guide the stable diffusion model.

[0010] To further optimize the technical solution, in step S202, during the testing phase, all categories in the test set have text labels. , using the text reversal algorithm, the test support set Reverse to text embedding vector and alternative tags , to generate an enhanced training set, the formula is as follows:

[0011] Text Embedding Vector As a test support set The inversion target, alternative label As a text-guided alternative label, the two together represent the test support set The knowledge contained in the text label template The formula is as follows:

[0012] Using Text Label Templates Steady Diffusion Model And generate the enhanced training set, the formula is as follows: , Where T is the number of samples generated for each category.

[0013] To further optimize the technical solution, in step S300, the method for performing basic training on the training set of the base domain includes the following steps: S310. First in the base domain Perform enough rounds of pre-training on the network so that the model has good performance in the first place; S320. In each CDAT training round, only fine-tune the model without training a large number of rounds. This step is only for a single round in the base domain. Conduct basic training on the

[0014] To further optimize the technical solution, in step S600, the formula of the feature fusion module is as follows:

[0015] in, , is the support set from the test set, is the augmented support set from the augmented training set, represents the feature fusion module, represents the parameter vector of the feature fusion module, , The first K parameters of are the weight parameters of the support set, and the last M parameters are the weight parameters of the enhanced support set. Represents the fusion support set The prototype representation obtained by feature extraction of the model is as follows:

[0016] in, express All samples of a class in Indicates the text features corresponding to the category, Representation model feature extraction, the model extracts text features and image features to obtain the prototype representation of the support set .

[0017] To further optimize the technical solution, in step S600, in order to enable the module trained in step S500 to adaptively fuse samples of the enhanced domain and the base domain, the training method of the feature fusion module includes the following steps: S610. Use the model after step S500 to extract features from the entire training set samples and calculate the prototype representation of the training set category , C is the number of categories in the training set; S620.Adopt The loss function trains the feature fusion module. The loss function formula is as follows:

[0018] in For the fusion support set The prototype features of the N categories are used to back-propagate the parameters of the update module by calculating the mean distance between the prototype representation of the N category fusion support set and the prototype representation of the training set.

[0019] Further optimizing the technical solution, in step S700, for the input support set , each category has samples, The support samples come from different angles of the same category; the model extracts features and obtains the support set prototype representation , the dimension is , where the support set prototype formula is as follows:

[0020] in, Indicates the text features corresponding to the category, Representation model feature extraction; the feature fusion module adaptively fuses the features of samples in different domains of each category. The formula is as follows:

[0021] in, Represents the parameter vector of the feature fusion module, and obtains the fusion prototype , the dimension is , query set After feature extraction, the model obtains the query set feature vector , the dimension is ; Transform the query set feature vector Fusion Prototype Perform cosine measurement and use cosine similarity measurement to obtain the classification result. The specific formula is as follows:

[0022] in Indicates A sample of queries, Indicates Support samples, Indicates the classification result.

[0023] Due to the adoption of the above technical scheme, the technical progress achieved by the present invention is as follows.

[0024] The present invention provides a semantically guided small sample image classification method, which generates high-quality enhanced samples that are semantically consistent with the original samples by integrating the knowledge of multiple pre-trained models and the semantic cues of category labels, thereby effectively improving the generalization ability of the model; adopts a cross-domain alternating training strategy to enable the model to adapt to samples from different domains, significantly enhancing its generalization ability; in the testing phase, the features of the enhanced samples and the original samples are integrated through the feature fusion module to construct a robust category prototype, further improving the classification performance of the model. In summary, the present invention provides an effective solution for small sample learning by combining strong semantic information, cross-domain alternating training and feature fusion, which can improve the classification accuracy and generalization ability of the model in an environment where data is scarce. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flow chart of Embodiment 1 of the present invention; Figure 2 This is a flow chart of Embodiment 2 of the present invention; Figure 3 This is a flowchart of Embodiment 3 of the present invention; Figure 4 This is a flowchart of Embodiment 4 of the present invention; Figure 5 This is a flow chart of Embodiment 5 of the present invention. DETAILED DESCRIPTION

[0026] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] A semantically guided small sample image classification method, combining Figure 1 As shown, it includes the following steps: S100. Perform random data enhancement on samples using data enhancement methods to improve sample diversity.

[0028] Weak data augmentation is performed for each training sample using random augmentation, including the following steps: S101.AutoContrast: Automatically adjust the contrast of the image; S102.Brightness: adjust the brightness of the image; S103.Color: adjust the color intensity of the image; S104.Contrast: adjust the contrast of the image; S105.Cutout: Randomly select an area on the image and replace it with gray; S106.CutoutAbs: randomly selects an area on the image, allows you to specify the size of the area to be cut, and replaces it with gray; S107.Equalize: Apply histogram equalization to enhance details in low-contrast areas of the image; S108.Identity: Do not change the image, as a baseline; S109.Invert: Invert the colors of the image; S110.Posterize: Reduces the tones of an image to a specified amount; S111.Rotate: randomly rotate the image; S112.Sharpness: adjust the sharpness of the image; S113.ShearX: Shears the image along the X axis; S114.ShearY: shears the image along the Y axis; S115.Solarize: Flips the pixel values ​​of an image to create an overexposed effect; S116.SolarizeAdd: flips the pixel values ​​of the image and adds a threshold parameter to produce an overexposure effect; S117.TranslateX: translate the image along the X axis; S118.TranslateY: Translates the image along the Y axis.

[0029] For each selected step, a value v (between 1 and self.m) is randomly selected, and then the operation is applied with a probability of 50%. Finally, regardless of whether the previous operation is applied, the step S106: CutoutAbs operation is applied. Each enhancement operation randomly selects multiple operations for combination. This random combination method can increase the diversity of data and help improve the generalization ability of the model.

[0030] S200. Use a two-stage semantically guided data augmentation method to perform data augmentation on samples to generate high-quality augmented datasets.

[0031] The method of using the two-stage semantic guidance data enhancement method to enhance the data of samples includes the following steps: Step S210: During the training phase, a template is constructed using text labels, and a stable diffusion model is used to generate an enhanced training set.

[0032] Step S220: During the testing phase, a text inversion algorithm is used to invert the test support set into text embedding vectors and alternative labels to generate an enhanced training set.

[0033] S300. Perform basic training on the training set of the base domain to train the initial performance of the model.

[0034] The method of basic training on the training set of the base domain includes the following steps: S310. First in the base domain Perform enough rounds of pre-training on the dataset to make the model have good performance in the first place.

[0035] S320. In each CDAT training round, only fine-tune the model without training a large number of rounds. This step is only for a single round in the base domain. Conduct basic training on the

[0036] S400. Perform enhanced training on the training set of the enhanced domain, so that the model can obtain stronger generalization performance from the high-quality enhanced data set. A single round of enhanced training is performed on the CNN to further improve the generalization ability of the model.

[0037] S500. The model is trained alternately in steps S300 and S400 using cross-domain alternating training to improve the adaptability and generalization ability of the model to different fields while reducing the impact of overfitting.

[0038] S600. Use the model trained in S500 to train the feature fusion module.

[0039] These training steps further improve the performance of the model, improve the generalization ability of the model, and the ability of the training module to adaptively fuse samples from different domains.

[0040] S700. After the model training is completed, the enhanced test set and the support set are used for feature fusion to obtain the support set prototype, which is used to perform small sample classification in the query set to finally obtain the classification result.

[0041] In the present invention, the flow chart of the second embodiment is as follows Figure 2As shown, the difference between this embodiment and the first embodiment is that in step S210, in this embodiment, during the training phase, all categories in the training set have text labels. , using text labels Building a template , and combined with the stable diffusion model to generate enhanced training sets, the formula is as follows:

[0042] Where N is the number of categories, The number of augmented samples generated for each class, It is a stable diffusion model, in which the specific formula of the text label template is as follows: , Data augmentation using text templates from multiple views to guide the stable diffusion model.

[0043] In the present invention, the flowchart of embodiment 3 is as follows Figure 3 As shown, the difference between this embodiment and the first embodiment is that step S220 is that in this embodiment, during the test phase, all categories of the test set have text labels. , using the text reversal algorithm, the test support set Reverse to text embedding vector and alternative tags , to generate an enhanced training set, the formula is as follows:

[0044] Text Embedding Vector As a test support set The inversion target, alternative label As a text-guided alternative label, the two together represent the test support set The knowledge contained in the text label template The formula is as follows:

[0045] Using Text Label Templates Steady Diffusion Model And generate the enhanced training set, the formula is as follows: , Where T is the number of samples generated for each category.

[0046] In the present invention, the flowchart of embodiment 4 is as follows Figure 4As shown, the difference between this embodiment and the first embodiment is that in step S500, starting from S320, the proposed cross-domain alternating training method (Cross-Domain Alternating Training, CDAT) and feature fusion module training include the following steps: S320. Model in base domain Conduct basic training on the

[0047] like Figure 4 (a) Basic training shows that the model is based on the semantic hinting method (SP), that is, by adding spatial channel attention to the Transformer model, the text features are adaptively aligned with the image features, so as to effectively use the text features to hint the model and find the corresponding image features.

[0048] The text encoder of the CLIP model extracts features from the text template. The text template formula is as follows:

[0049] The text template The text encoder input to the CLIP model is used for feature extraction. The formula is as follows:

[0050] in represents the text encoder of the CLIP model. The backbone model takes The text features of the training set are fused with the image features to train the model. The formula is as follows: .

[0051] S400. Figure 4 (b) Enhanced training shows that the model has an enhanced domain after being processed by the semantic guidance data enhancement method in the training phase of step S210. A single round of enhanced training is performed on the model, and the training method is the same as step S320 to further improve the generalization ability of the model.

[0052] S500. Alternate the training of step S400 and step S320, so that the model can ensure adaptation to the basic image domain while improving the generalization ability of the model and reducing the impact of overfitting. In each round, the model is alternately trained on the training set of the base domain and the enhanced training set of the enhanced domain. The model first performs the enhanced training of step S400 to break the overfitting of the model to the base domain training set and accept high-quality enhanced training set information; then after step S320, it is trained again on the training set of the base domain to ensure the adaptability of the model to the base domain, rather than changing to adapt to the enhanced domain.

[0053] S600. The feature fusion module learns how to adaptively and effectively fuse images from the base domain and the enhanced domain to obtain a robust prototype representation. The formula of the feature fusion module is as follows:

[0054] in , is the support set from the test set, is the augmented support set from the augmented training set. Represents the feature fusion module. Represents the parameter vector of the feature fusion module, as follows Figure 4 (c) Detail of the fusion part, where , The first K parameters of are the weight parameters of the support set, and the last M parameters are the weight parameters of the enhanced support set. Represents the fusion support set The prototype representation obtained by feature extraction of the model is as follows:

[0055] in, express All samples of a certain class in Indicates the text features corresponding to the category, Representation model feature extraction. The model extracts text features and image features to obtain the prototype representation of the support set. .

[0056] In step S600, in order to enable the module trained in step S500 to adaptively fuse samples of the enhanced domain and the base domain, the training method of the feature fusion module includes the following steps: S610. Use the model after step S500 to extract features from the entire training set samples and calculate the prototype representation of the training set category , C is the number of categories in the training set; S620.Adopt The loss function trains the feature fusion module. The loss function formula is as follows:

[0057] in For the fusion support set The prototype features of the N categories are used to back-propagate the parameters of the update module by calculating the mean distance between the prototype representation of the N category fusion support set and the prototype representation of the training set.

[0058] In the present invention, the flowchart of embodiment 5 is as follows Figure 5 As shown, the difference between this embodiment and the first embodiment lies in step S700. Specifically, for the input support set , each category has samples, The support samples come from different angles of the same category; the model extracts features and obtains the support set prototype representation , the dimension is , where the support set prototype formula is as follows:

[0059] in, Indicates the text features corresponding to the category, Representation model feature extraction; the feature fusion module adaptively fuses the features of samples in different domains of each category. The formula is as follows:

[0060] in, Represents the parameter vector of the feature fusion module, and obtains the fusion prototype , the dimension is , query set After feature extraction, the model obtains the query set feature vector , the dimension is ; After dimensionality reduction, the dimension becomes ; Transform the query set feature vector Fusion Prototype Perform cosine measurement and use cosine similarity measurement to obtain the classification result. The specific formula is as follows:

[0061] in Indicates A sample of queries, Indicates Support samples, Indicates the classification result.

Claims

1. A semantically guided small sample image classification method, characterized by: It includes the following steps: S100. Perform random data enhancement on samples using data enhancement methods to improve sample diversity; S200. Use a two-stage semantically guided data enhancement method to enhance the data of samples to generate high-quality enhanced datasets; S300. Perform basic training on the training set of the base domain to train the initial performance of the model; S400. Perform enhanced training on the training set of the enhanced domain, so that the model obtains stronger generalization performance from the high-quality enhanced data set; S500. Using cross-domain alternating training to alternately train the model in steps S300 and S400 to improve the adaptability and generalization ability of the model to different fields while reducing the impact of overfitting; S600. Using the model trained in S500 to train the feature fusion module; S700. After the model training is completed, the enhanced test set and the support set are used for feature fusion to obtain the support set prototype, which is used to perform small sample classification in the query set to finally obtain the classification result.

2. The semantically guided small sample image classification method according to claim 1, characterized in that: In step S100, weak data enhancement is performed for each training sample using random enhancement, including the following steps: S101.AutoContrast: Automatically adjust the contrast of the image; S102.Brightness: adjust the brightness of the image; S103.Color: adjust the color intensity of the image; S104.Contrast: adjust the contrast of the image; S105.Cutout: Randomly select an area on the image and replace it with gray; S106.CutoutAbs: randomly selects an area on the image, allows you to specify the size of the area to be cut, and replaces it with gray; S107.Equalize: Apply histogram equalization to enhance details in low-contrast areas of the image; S108.Identity: Do not change the image, as a baseline; S109.Invert: Invert the colors of the image; S110.Posterize: Reduces the tones of an image to a specified amount; S111.Rotate: randomly rotate the image; S112.Sharpness: adjust the sharpness of the image; S113.ShearX: Shears the image along the X axis; S114.ShearY: shears the image along the Y axis; S115.Solarize: Flips the pixel values ​​of an image to create an overexposed effect; S116.SolarizeAdd: flips the pixel values ​​of the image and adds a threshold parameter to produce an overexposure effect; S117.TranslateX: translate the image along the X axis; S118.TranslateY: Translates the image along the Y axis.

3. The semantically guided small sample image classification method according to claim 1, characterized in that: In step S200, the method for performing data enhancement on the sample using the two-stage semantic guidance data enhancement method includes the following steps: Step S210: In the training phase, a template is constructed using text labels, and a stable diffusion model is combined to generate an enhanced training set; Step S220: During the testing phase, a text inversion algorithm is used to invert the test support set into text embedding vectors and alternative labels to generate an enhanced training set.

4. The semantically guided small sample image classification method according to claim 3, characterized in that: In the training phase, step S210 has text labels for the categories in the training set. , using text labels Building a template , and combined with the stable diffusion model to generate enhanced training sets, the formula is as follows: Where N is the number of categories, The number of augmented samples generated for each class, It is a stable diffusion model, in which the specific formula of the text label template is as follows: , Data augmentation using text templates from multiple views to guide the stable diffusion model.

5. The semantically guided small sample image classification method according to claim 3, characterized in that: In the test phase, step S220 has text labels for the categories in the test set. , using the text reversal algorithm, the test support set Reverse to text embedding vector and alternative tags , to generate an enhanced training set, the formula is as follows: Text Embedding Vector As a test support set The inversion target, alternative label As a text-guided alternative label, the two together represent the test support set The knowledge contained in the text label template The formula is as follows: Using Text Label Templates Steady Diffusion Model And generate the enhanced training set, the formula is as follows: , Where T is the number of samples generated for each category.

6. The semantically guided small sample image classification method according to claim 1, characterized in that: In step S300, the method for performing basic training on the training set of the base domain includes the following steps: S310. First in the base domain Perform enough rounds of pre-training on the network so that the model has good performance in the first place; S320. In each CDAT training round, only fine-tune the model without training a large number of rounds. This step is only for a single round in the base domain. Conduct basic training on the 7. The semantically guided small sample image classification method according to claim 1, characterized in that: In step S600, the formula of the feature fusion module is as follows: in, , is the support set from the test set, is the augmented support set from the augmented training set, represents the feature fusion module, represents the parameter vector of the feature fusion module, , The first K parameters of are the weight parameters of the support set, and the last M parameters are the weight parameters of the enhanced support set. Represents the fusion support set The prototype representation obtained by feature extraction of the model is as follows: in, express All samples of a class in Indicates the text features corresponding to the category, Representation model feature extraction, the model extracts text features and image features to obtain the prototype representation of the support set .

8. The semantically guided small sample image classification method according to claim 1, characterized in that: In step S600, in order to enable the module trained in step S500 to adaptively fuse samples of the enhanced domain and the base domain, the training method of the feature fusion module includes the following steps: S610. Use the model after step S500 to extract features from the entire training set samples and calculate the prototype representation of the training set category , C is the number of categories in the training set; S620.Adopt The loss function trains the feature fusion module. The loss function formula is as follows: in For the fusion support set The prototype features of the N categories are used to back-propagate the parameters of the update module by calculating the mean distance between the prototype representation of the N category fusion support set and the prototype representation of the training set.

9. The semantically guided small sample image classification method according to claim 1, characterized in that: In step S700, for the input support set , each category has samples, The support samples come from different angles of the same category; the model extracts features and obtains the support set prototype representation , the dimension is , where the support set prototype formula is as follows: in, Indicates the text features corresponding to the category, Representation model feature extraction; the feature fusion module adaptively fuses the features of samples in different domains of each category. The formula is as follows: in, Represents the parameter vector of the feature fusion module, and obtains the fusion prototype , the dimension is , query set After feature extraction, the model obtains the query set feature vector , the dimension is , after dimensionality reduction, the dimension becomes ; Transform the query set feature vector Fusion Prototype Perform cosine measurement and use cosine similarity measurement to obtain the classification result. The specific formula is as follows: in Indicates A sample of queries, Indicates Support samples, Indicates the classification result.