Small sample capacity training method based on deep learning

By employing cross-domain feature alignment, meta-knowledge distillation, and dynamic prototype library construction, combined with category attention and uncertainty quantization, the overfitting and insufficient generalization capabilities of deep learning models in few-sample scenarios are addressed, achieving efficient few-sample training results.

CN120996113APending Publication Date: 2025-11-21SUZHOU JIELIXUN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202511155534.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Deep learning models are prone to overfitting and have poor generalization ability in small sample scenarios. They also fail to transfer across domains, underutilize sample information, and lack dynamic optimization mechanisms, resulting in decreased model performance and large differences in accuracy between the training and validation sets.

Method used

By constructing a cross-domain hybrid training library, using GAN to achieve feature alignment, training a multi-task teacher model and performing meta-knowledge distillation, constructing a dynamic prototype library, introducing a category attention mechanism and Monte Carlo dropout to quantify uncertainty, and performing iterative optimization.

Benefits of technology

It significantly improves the model's generalization ability in small sample scenarios, increases accuracy by 15% to 30%, reduces the risk of overfitting, improves sample utilization efficiency, and adapts to changes in the distribution of the target task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996113A_ABST
    Figure CN120996113A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning and small sample learning, in particular to a small sample capacity training method based on deep learning, which comprises the steps of 1, cross-domain data adaptation and feature alignment, 2, meta-knowledge distillation and prototype enhancement, 3, attention-guided small sample fine adjustment, and 4, model uncertainty quantification and iterative optimization. According to the small sample capacity training method based on deep learning, through cross-domain feature alignment, meta-knowledge distillation, prototype enhancement and dynamic iterative optimization, the problems of model overfitting and weak generalization ability in a small sample scene are solved, high-precision model training when the sample size is less than or equal to 50 is realized, and the training efficiency is improved. The method is suitable for data scarce scenes such as medical images and minority language processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of deep learning and few-shot learning technology, and in particular to a few-shot capacity training method based on deep learning. Background Technology

[0002] The performance of deep learning models is highly dependent on large-scale labeled data. However, in practical applications, many scenarios suffer from data scarcity (such as novel disease samples and industrial anomaly detection), leading to models being prone to overfitting and exhibiting poor generalization ability. Existing small-sample training methods have the following limitations: 1. Cross-domain transfer failure: When using auxiliary datasets (such as public datasets) for transfer learning, the large differences in domain distribution can easily lead to "negative transfer" (degradation of model performance). 2. Insufficient utilization of sample information: In small sample scenarios, it is difficult to capture the distribution of class features by relying on only a limited number of samples, and traditional data augmentation (such as rotation and cropping) is prone to introducing noise; 3. High risk of overfitting: When the model is fine-tuned on a small number of samples, it tends to remember training details rather than learn general rules, and the difference in accuracy between the training set and the validation set can be more than 20%. 4. Lack of dynamic optimization mechanism: It takes into account the uncertainty of model prediction, and it is difficult to iteratively optimize the training strategy according to the real-time distribution quality.

[0003] To address the above problems, this invention proposes a small-sample-capacity training method based on deep learning. Summary of the Invention

[0004] The main objective of this invention is to provide a small-sample-capacity training method based on deep learning, which can effectively solve the problems in the background technology.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A deep learning-based method for training small sample sizes includes the following steps: Step 1: Cross-domain data adaptation and feature alignment: Construct a hybrid training library containing a small sample set of the target task (sample size ≤ 50) and an auxiliary task dataset (sample size ≥ 1000), wherein the auxiliary tasks are semantically related to the target task (such as different categories or different scenarios of the same type of task in the same domain). By using the domain discriminator of a Generative Adversarial Network (GAN), the feature distribution distance between the target task and the auxiliary task is minimized, thereby generating domain-fitted feature vectors. The feature alignment loss function is:

[0006] ( For domain discriminator, , These are the original features of the target task and the auxiliary task, respectively. Step 2, Meta-knowledge Distillation and Prototype Enhancement: Train a multi-task teacher model to learn a general feature extractor on the auxiliary task dataset. and task adapter head ( To assist in task indexing, knowledge distillation is used to transfer the meta-knowledge (including attention weights and intermediate layer feature distributions) of the teacher model to the student model. Based on a small sample of the target task, a dynamic prototype library is built: for each category Calculate the initial prototype For category (number of samples), and generate through a generative adversarial network. An enhanced prototype ( Ensure that the total number of prototypes in each category is ≥50). Step 3: Attention-guided fine-tuning with small samples: Freeze the first 80% of the parameters of the general feature extractor E, and only fine-tune the last 20% of the parameters and the target task header. ; A category attention mechanism is introduced to calculate the cosine similarity between sample features and similar prototypes in the prototype library, thereby generating attention weights. And weighted loss function: ( For cross-entropy loss, (for sample labels) Step 4: Model uncertainty quantification and iterative optimization: Calculating forecast uncertainty using Monte Carlo dropout Screening high-confidence samples ( , (for preset thresholds) After each round of fine-tuning, the prediction results of high-confidence samples are used as pseudo-labels to supplement the small sample set of the target task, and steps 2)-3) are repeated, for a total of [number of iterations]. Dynamically adjusted according to sample size (The initial number of samples for the target task).

[0007] Preferably, in step 1, semantic relevance is quantitatively evaluated by word vector cosine similarity (≥0.6) or domain knowledge graph relevance (≥0.5), and the number of categories in the auxiliary task dataset is 3 to 5 times the number of categories in the target task.

[0008] Preferably, in step 2, the multi-task teacher model adopts a Transformer architecture, containing 6-12 layers of self-attention modules, and knowledge distillation uses intermediate layer feature matching loss. .

[0009] Preferably, in step 2, the prototype is enhanced. Generates through a conditional generation network (such as Conditional GAN), with the following constraints: ( The prototype drift threshold is set to 0.1 to 0.3 times the original prototype standard deviation.

[0010] Preferably, in step 3, the category attention mechanism combines spatial attention and channel attention: spatial attention focuses on the key regional features of the sample, while channel attention strengthens discriminative feature channels (such as edge feature channels of images and semantic keyword channels of text).

[0011] Preferably, the fine-tuning process in step 3 uses a cosine annealing learning rate (initial learning rate). Minimum learning rate And an early stop strategy (stop if the accuracy of the validation set does not improve after 5 consecutive rounds).

[0012] Preferably, in step 4, the pseudo-label supplementation adopts a "confidence filtering + diversity screening" strategy: in addition to the high confidence condition, k-means clustering is used to ensure that the supplementary samples cover the data distribution of the target task (cluster center distance ≥ 0.5).

[0013] Preferably, it also includes a prior knowledge embedding step: encoding the domain rules of the target task (such as lesion location constraints in medical images, keyword weights in text classification) into learnable prior vectors. And it is integrated into the feature extractor through a gating mechanism: ( For the sigmoid function, (For gating weights).

[0014] Preferably, in step 1, when the target task is an image, the GAN uses the StyleGAN2 architecture to generate domain-adaptive features; when the target task is a text, the BERT-GAN architecture is used, and feature alignment is achieved through sentence vector mapping.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. Significantly improved generalization ability: Through cross-domain alignment and meta-knowledge distillation, the model accuracy in small sample scenarios (5-shot, 10-shot) is improved compared with traditional transfer learning, and the overfit index (training / validation set difference) is ≤8%.

[0016] 2. Sample utilization efficiency optimization: Prototype augmentation and attention mechanisms improve the information utilization rate of each small sample by 2 to 3 times without relying on large-scale data augmentation.

[0017] 3. Strong dynamic adaptability: Uncertainty quantification and iterative optimization enable the model to adapt to changes in the distribution of the target task and maintain stable performance even in scenarios with incremental samples.

[0018] 4. Good domain scalability: It supports multimodal few-sample tasks such as image, text, and voice, and can be quickly adapted to specific domains (such as medical and industrial inspection) through prior knowledge embedding. Attached Figure Description

[0019] Figure 1 This is an overall flowchart of a deep learning-based small-sample-capacity training method according to the present invention; Figure 2 This is a diagram illustrating the cross-domain feature alignment architecture of a deep learning-based small-sample-capacity training method according to the present invention. Figure 3 This is a flowchart of meta-knowledge distillation and prototype enhancement for a small-sample-capacity training method based on deep learning according to the present invention. Figure 4 This is a flowchart of attention fine-tuning and iterative optimization of a small-sample-capacity training method based on deep learning according to the present invention. Detailed Implementation

[0020] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0021] A deep learning-based method for training small sample sizes includes the following steps: Step 1: Cross-domain data adaptation and feature alignment Hybrid training library construction: collecting small sample sets for the target task (N≤50, (for labels) and auxiliary task datasets (M≥1000), requiring the semantic correlation between the auxiliary task and the target task to be ≥0.6 (quantified by cosine similarity of word vectors). For example, "skin cancer classification" (target task) can be associated with "general classification of skin diseases" (auxiliary task).

[0022] Adversarial domain alignment: Feature distribution adaptation is achieved using a GAN architecture, including: Feature extractor (e.g., ResNet-18, BERT): Maps input data to a feature space ; Domain discriminator To determine whether a feature originates from the target domain or the auxiliary domain, the loss function is: Alternate optimization during training (minimize (Confusion domain discriminator) and (maximize (Distinguishing domain origins) and ultimately generating domain adaptation features. This reduces cross-domain distribution distances (such as Wasserstein distance) by ≥40%.

[0023] Step 2: Meta-knowledge distillation and prototype enhancement Multi-task teacher model training: In Training teacher model Its structure is a "general feature extractor" Task adapter head ( To assist in the number of tasks, learn cross-task general features and task-specific knowledge (such as attention weights and category prototypes).

[0024] Knowledge distillation to the student model: transferring meta-knowledge from the teacher model to the student model. (and shared (Structure), employing double distillation losses:

[0025] in, To output the distribution matching loss, The intermediate layer feature matching loss is λ=0.5.

[0026] Building a dynamic prototype library: For each category of the target task Calculate the initial prototype: For category (number of samples); Generate using Conditional GAN An enhanced prototype ,constraint ( (The initial prototype standard deviation) ensures that the total number of prototypes for each category is ≥50, and expands the category feature distribution information.

[0027] Step 3: Small-sample fine-tuning guided by attention Parameter freezing and fine-tuning strategies: Freezing The first 80% of parameters (preserving general feature extraction capabilities) are adjusted, with only the last 20% of parameters and the target task header fine-tuned. Cosine annealing learning rate (initial) is used ,lowest ).

[0028] Category attention mechanism: Calculate the cosine similarity between the sample features and the prototype of the same class: ; Generate attention weights: Strengthen the training weights of key samples (those with high similarity); Weighted cross-entropy loss: This improves the efficiency of discriminative feature learning.

[0029] Step 4: Model Uncertainty Quantification and Iterative Optimization Uncertainty assessment: via Monte Carlo dropout (in and Add a dropout layer (sample 10 times during testing) to calculate the prediction variance: Screening high-confidence samples ( ).

[0030] Pseudo-label supplementation: For high-confidence unlabeled samples (such as unlabeled data of the target task), the prediction results are used as pseudo-labels to supplement the data. Furthermore, k-means clustering was used to ensure the diversity of supplementary samples (cluster center distance ≥ 0.5).

[0031] Dynamic iteration: Repeat steps 2-3, iteration number. ( (The initial number of samples) is used to update the prototype library after each iteration, gradually improving the model's adaptability to the target task.

[0032] Step 5: Embedding Prior Knowledge Encode domain prior rules (such as "lesions in medical images are mostly located in the lung region") into learnable vectors. Feature extraction is integrated through a gating mechanism: ( For the sigmoid function, (as gating weights), further constraining the direction of feature learning.

[0033] Example Example 1: 5-shot medical image classification (skin cancer identification) Objective: Classify images of 5 types of skin cancer, with only 5 labeled samples for each type (25 samples in total).

[0034] Auxiliary dataset: 10,000 images of common skin diseases (10 classes), with a semantic correlation of 0.72 with the target task.

[0035] Cross-domain alignment: ResNet-18 was used as the feature extractor and the GAN domain discriminator was a 3-layer MLP. After training for 50 rounds, the cross-domain Wasserstein distance decreased from the initial 0.82 to 0.35.

[0036] Knowledge distillation: The teacher model is trained to 92% accuracy on the auxiliary dataset. The student model is initialized with distillation loss, achieving an initial accuracy of 75% (compared to only 58% with traditional transfer learning).

[0037] Prototype enhancement: 45 enhanced prototypes are generated for each class (total 50), and prototype drift is controlled within 0.15σ.

[0038] Attention fine-tuning: Freeze the first 8 layers of ResNet-18, fine-tune the last 4 layers and the classification head, and adjust the learning rate from... Annealing to After 30 rounds of training, the accuracy rate reached 83%.

[0039] Iterative optimization: 4 iterations ( By adding 20 high-confidence pseudo-label samples, the final accuracy was improved to 88%, with an overfit index of 6%.

[0040] Results: The accuracy is improved by 22% and the risk of overfitting is reduced by 50% compared to traditional small-sample methods (such as Prototypical Networks).

[0041] Example 2: 10-shot Classification of Minor Language Texts (Sentiment Analysis of Icelandic Text) Objective: Classify Icelandic sentiment into three categories (positive / negative / neutral), with 10 samples per category (30 samples in total).

[0042] Auxiliary dataset: 10,000 English sentiment texts, semantically associated through bilingual dictionary mapping (association degree 0.65).

[0043] Cross-domain alignment: Using the BERT-GAN architecture with a sentence vector dimension of 768, the cross-domain cosine similarity after alignment was improved from 0.42 to 0.78.

[0044] Knowledge distillation: The teacher model (multilingual BERT) was trained on English data, and the student model achieved an initial accuracy of 62% after distillation (compared to 45% for traditional methods).

[0045] Prototype enhancement and fine-tuning: 40 enhanced prototypes are generated for each category, combined with text keyword attention (weighting sentiment words), and the accuracy is 79% after fine-tuning.

[0046] Iterative optimization: After 3 iterations and the addition of 15 pseudo-labeled samples, the final accuracy reached 85%, which is 18 percentage points better than existing small-sample text methods (such as Matching Networks).

[0047] In summary, this invention presents a deep learning-based method for training small-sample-capacity models, belonging to the field of small-sample learning. This method constructs a cross-domain hybrid training library and utilizes GANs to achieve feature alignment to reduce domain differences; trains a multi-task teacher model and transfers meta-knowledge to the student model through knowledge distillation; constructs a dynamic prototype library based on small samples of the target task and guides fine-tuning using a category attention mechanism; and quantifies uncertainty through Monte Carlo dropout, selecting high-confidence samples to supplement the training set and iteratively optimizing. This invention solves the problems of overfitting and weak generalization ability in small-sample scenarios, improving accuracy by 15%~30% when the sample size is ≤50, and is suitable for data-scarce scenarios such as medical imaging and processing of less commonly spoken languages.

[0048] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A small-sample-capacity training method based on deep learning, characterized in that, Includes the following steps: Step 1: Cross-domain data adaptation and feature alignment: Construct a hybrid training library containing a small sample set of the target task (sample size ≤ 50) and an auxiliary task dataset (sample size ≥ 1000). The auxiliary tasks are semantically related to the target task (such as different categories or different scenarios of the same type of task in the same domain). By using the domain discriminator of a Generative Adversarial Network (GAN), the feature distribution distance between the target task and the auxiliary task is minimized, thereby generating domain-fitted feature vectors. , The feature alignment loss function is: ( For domain discriminator, , These are the original features of the target task and the auxiliary task, respectively. Step 2, Meta-knowledge Distillation and Prototype Enhancement: Train a multi-task teacher model to learn a general feature extractor on the auxiliary task dataset. and task adapter head ( To assist in task indexing, knowledge distillation is used to transfer the meta-knowledge (including attention weights and intermediate layer feature distributions) of the teacher model to the student model. Based on a small sample of the target task, a dynamic prototype library is built: for each category Calculate the initial prototype For category (number of samples), and generate through a generative adversarial network. An enhanced prototype ( Ensure that the total number of prototypes in each category is ≥50). Step 3: Attention-guided fine-tuning with small samples: Freeze the first 80% of the parameters of the general feature extractor E, and only fine-tune the last 20% of the parameters and the target task header. ; A category attention mechanism is introduced to calculate the cosine similarity between sample features and similar prototypes in the prototype library, thereby generating attention weights. And weighted loss function: ( For cross-entropy loss, (for sample labels) Step 4: Model uncertainty quantification and iterative optimization: Calculating forecast uncertainty using Monte Carlo dropout Screening high-confidence samples ( , (for preset thresholds) After each round of fine-tuning, the prediction results of high-confidence samples are used as pseudo-labels to supplement the small sample set of the target task, and steps 2)-3) are repeated, for a total of [number of iterations]. Dynamically adjusted according to sample size (The initial number of samples for the target task).

2. The method for training small-sample-capacity deep learning based on claim 1, characterized in that: In step 1, semantic relevance is quantitatively evaluated using word vector cosine similarity (≥0.6) or domain knowledge graph relevance (≥0.5). The number of categories in the auxiliary task dataset is 3 to 5 times the number of categories in the target task.

3. The method for training small-sample-capacity deep learning based on claim 1, characterized in that: In step 2, the multi-task teacher model adopts a Transformer architecture, containing 6-12 layers of self-attention modules, and knowledge distillation uses intermediate layer feature matching loss. .

4. The method for training small-sample-capacity training based on deep learning according to claim 1, characterized in that: Step 2, enhancing the prototype Generates through a conditional generation network (such as Conditional GAN), with the following constraints: ( The prototype drift threshold is set to 0.1 to 0.3 times the original prototype standard deviation.

5. The method for training small-sample-capacity deep learning based on claim 1, characterized in that: In step 3, the category attention mechanism combines spatial attention and channel attention: spatial attention focuses on the key regional features of the sample, while channel attention strengthens the discriminative feature channels (such as the edge feature channel of the image and the semantic keyword channel of the text).

6. The method for training small-sample-capacity deep learning based on claim 1, characterized in that: The fine-tuning process in step 3 uses a cosine annealing learning rate (initial learning rate). Minimum learning rate And an early stop strategy (stop if the accuracy of the validation set does not improve after 5 consecutive rounds).

7. The method for training small-sample-capacity deep learning based on claim 1, characterized in that: In step 4, the pseudo-label supplementation adopts a "confidence filtering + diversity screening" strategy: in addition to the high confidence condition, k-means clustering is used to ensure that the supplementary samples cover the data distribution of the target task (cluster center distance ≥ 0.5).

8. The method for training small-sample-capacity deep learning based on claim 1, characterized in that: It also includes a prior knowledge embedding step: encoding the domain rules of the target task (such as lesion location constraints in medical images and keyword weights in text classification) into learnable prior vectors. And it is integrated into the feature extractor through a gating mechanism: ( For the sigmoid function, (For gating weights).

9. The method for training small-sample-capacity deep learning based on claim 1, characterized in that: In step 1, when the target task is an image, the GAN uses the StyleGAN2 architecture to generate domain-adaptive features. When the target task is text-based, the BERT-GAN architecture is adopted, and feature alignment is achieved through sentence vector mapping.

Citation Information

Cited By

  • Construction site safety supervision model non-inductive adaptation method, system and equipment

    CN121305286A

  • Remote sensing large model increment pre-training optimization system and method for small sample scene

    CN121330426A

  • Deep learning hidden danger automatic identification method and system with few samples and strong generalization

    CN121582551A

  • Fracture treatment scheme and operation scheme simulation system based on big data analysis

    CN121839028A

  • A fracture treatment scheme operation scheme simulation system based on big data analysis

    CN121839028B