A multi-teacher distillation domain knowledge memory and migration method using image recognition
Patent Information
- Application Number
- CN202311781461.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-22
AI Technical Summary
然而现有方法大多采用了从单一教师蒸馏知识的策略,忽视了多教师能够记忆更加多样的源域知识并提高模型对目标域的兼容性的能力
[0030]本发明所提出的基于多教师蒸馏的域知识记忆与迁移方法,能够从多方位对源域知识进行高效地记忆,并通过多教师蒸馏损失有效地将知识迁移到目标域上。同时,以分支的结构表达多个教师能够显著地降低推理时间以及存储上的消耗。
Smart Images

Figure CN117892803B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image recognition, and in particular to a memory and transfer method for domain knowledge based on multi-teacher distillation in image classification tasks. Background Technology
[0002] With the development and popularization of deep learning technology, numerous artificial intelligence applications have emerged in people's daily lives. Daily life scenarios are constantly changing, requiring deep learning-based models to continuously memorize knowledge from one image domain while simultaneously transferring that knowledge to new image domains. However, existing deep learning models often face problems such as semantic gaps, feature differences, and weak model generalization ability when transferring knowledge to new image domains. To improve the effective memorization of source image domain knowledge by deep learning models and promote knowledge transfer to the target image domain, this invention, compared to existing methods, adopts a multi-teacher distillation approach for memorizing and transferring domain knowledge.
[0003] Distillation is essentially a function regularization method that constrains the input-output mapping between the teacher and student models to remain unchanged. It is currently an effective strategy for knowledge transfer from the source domain to the target domain, improving the model's generalization ability in the target domain. However, most existing methods employ a strategy of distilling knowledge from a single teacher, neglecting the ability of multiple teachers to memorize more diverse source domain knowledge and improve the model's compatibility with the target domain. Although existing methods have adopted the idea of multi-teacher distillation, such as obtaining an additional teacher through independent training on a new task and maximizing the mutual information between the student model and the other two teacher models for knowledge transfer, this approach incurs significant time and storage costs. Therefore, knowledge memorization and transfer based on multi-teacher distillation needs to address two issues: how to effectively find multiple teachers and how to efficiently represent multiple teachers. Summary of the Invention
[0004] To memorize knowledge from the source domain, this invention effectively finds multiple teachers through weighted arrangement, feature perturbation, and diversity regularization, thus representing the knowledge in the source domain from multiple perspectives. To transfer knowledge to the target domain, knowledge distillation is used to transfer knowledge from the source domain to the target domain. To reduce the inference time and storage consumption of multiple teachers, each teacher is represented as a small sub-branch of the model.
[0005] This invention employs three strategies to effectively find multiple teachers: weight reordering, feature perturbation, and diversity regularization. To reduce the inference time and storage costs associated with multiple teachers, each teacher is represented as a small branch of the original model. Ultimately, knowledge in the source domain is memorized through multiple models, and knowledge is transferred to the target domain using distillation loss by treating the multiple models in the source domain as teachers. Therefore, the technical solution of this invention is: a method for multi-teacher distillation domain knowledge memorization and transfer using image recognition, the method comprising:
[0006] Step 1: Train a basic image recognition model using sample images from the source domain. The basic image recognition model includes: L concatenated convolutional layers and a classifier.
[0007] Step 2: Duplicate the last M convolutional layers and classifier of the trained base image recognition model n-1 times to obtain n-1 teacher models; the input of these n-1 teacher models is the output of the LM-th convolutional layer in the base model.
[0008] Step 3: Randomly rearrange the parameters of each convolutional layer in the n-1 teacher models obtained in Step 2 to obtain n-1 new teacher models;
[0009] Step 4: Introduce a perturbation for each new teacher model to change the input of each new teacher model;
[0010] Step 5: Randomly sample the same number of K samples for each class from all source domain sample images to construct a class-balanced subset of the source domain;
[0011] Step 6: Use on the category-balanced subset The loss is optimized for each teacher model; Losses include: diversity regularization loss Classification loss Migration loss
[0012] Step 7: Apply distillation loss to the target domain By transferring multi-teacher knowledge distillation to the basic image recognition model, a new image recognition model is obtained.
[0013] Step 8: Fine-tune the training of the new image recognition model;
[0014] Step 9: Use the finely tuned new image recognition model to recognize the new image.
[0015] Furthermore, the method for rearranging the parameters of each convolutional layer in step 3 is as follows:
[0016]
[0017] Among them, W′ l W represents the parameters of the l-th convolutional layer after the arrangement. l P represents the parameters of the l-th convolutional layer. l This represents the rearrangement matrix of the l-th layer. This represents the transpose of the rearrangement matrix of the (l-1)th layer.
[0018] Furthermore, the specific method for step 4 is as follows:
[0019]
[0020] Where, x i Let x represent the input to the i-th new teacher model. L-M This represents the output of the LM-th convolutional layer in the base model. It follows a normal distribution, α is the scaling factor, and δ i It is the perturbation factor.
[0021] Furthermore, the diversity regularization loss in step 6 Classification loss Migration loss The calculation method is as follows:
[0022]
[0023]
[0024]
[0025] in, Let x represent the 2-combination number of n new teacher models. i,L Let x represent the embedding of the i-th teacher with respect to sample x. j,L Let CE(.) represent the embedding of the j-th teacher for sample x, and let CE(.) be the cross-entropy loss. This represents the output of the base model. This represents the output of the model before fine-tuning, where y represents the correct label corresponding to sample x. This represents the output of the i-th teacher;
[0026] This represents the average output of the teacher model.
[0027] Furthermore, the distillation loss in step 7 The calculation method is as follows:
[0028]
[0029] in, Indicates the current model With multi-teacher model The KL divergence over the target domain, where x originates from the target domain. This is an image recognition model for the target domain.
[0030] The domain knowledge memorization and transfer method based on multi-teacher distillation proposed in this invention can efficiently memorize source domain knowledge from multiple perspectives and effectively transfer knowledge to the target domain through multi-teacher distillation loss. Furthermore, representing multiple teachers with a branched structure can significantly reduce inference time and storage consumption. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of domain knowledge memorization and transfer based on multi-teacher distillation. Detailed Implementation
[0032] Existing research shows that model parameters in different low-loss regions exhibit different inference mechanisms; and weight arrangement can effectively transfer model parameters from one low-loss region to another. The original model is represented as... The initial branch corresponding to the teacher model is represented as follows: It is a copy of the last M convolutional layers and the classifier of the original model, that is... for The calculations for the l-th and (l+1)-th layers can be formalized as follows:
[0033]
[0034] Where x l-1 It is the output of the (l-1)th layer, σ is the element-wise nonlinear activation function, and W l and W l+1 These are the parameters for the l-th and (l+1)-th layers, respectively. We omit the bias terms in the convolutional layers for simplicity. The weights are rearranged using a permutation matrix to rearrange the positions of the parameters in the layers. For the parameters of the l-th layer... For example, its permutation matrix P l By rearranging the identity matrix The permutation matrix P is obtained by processing the columns. l It is an orthogonal matrix that satisfies:
[0035] P T P = PP T =I
[0036] In order to make P l Applied to parameter W l Keeping the output unchanged, we can reformulate the calculation of the l-th and l+1-th layers as follows:
[0037]
[0038] Considering that layers (l-1) and (l+1) also have weighted arrangements, the parameters of layers l and (l+1) change after applying the weighted arrangement as follows:
[0039]
[0040] Arranged branches This is obtained by sequentially arranging the M convolutional layers in the initial branch, where the P layer of the l-th layer... l It is obtained by randomly arranging the columns of the identity matrix. Taking the original model as a teacher, if we need to obtain n teachers, we need to randomly arrange the initial branch n-1 times to obtain the branches corresponding to the remaining n-1 teachers.
[0041] While weighted arrangements can transfer parameters from one low-loss region to another, the invariance between the transferred and original parameters conflicts with our desired diversity of teachers. To overcome this invariance, different branches should have different optimization trajectories. One of the simplest methods is to have different inputs between branches. We employ a perturbation operation on the branch inputs, namely:
[0042]
[0043] in It follows a normal distribution, and α is the scaling factor.
[0044] To further enable teachers to have different inference mechanisms, this invention employs diversity regularization loss to constrain multi-teacher optimization. Significantly different inference mechanisms among teachers imply that their intra-layer features should be as different as possible, which manifests in the feature space as mutual orthogonality between teacher features. Based on this, the diversity regularization loss is implemented by minimizing the absolute cosine similarity of the embedding features among teachers. Assume the embedding of the i-th teacher is x. i,L It is the branch after the arrangement. The output of the penultimate layer, i.e.:
[0045]
[0046] Diversity regularization loss The definition is as follows:
[0047]
[0048] in It is a 2-group number of n teachers.
[0049] For finding multiple teachers, this invention first trains a basic teacher model on the source domain dataset using standard methods. Then, it replicates the last M convolutional layers and the classifier n-1 times, and rearranges the weights to obtain n-1 rearranged branches corresponding to the remaining teachers. Next, it optimizes the original model's classifier and rearranged branches on a representative, class-balanced subset of the source domain. The average output logistic value of the original model and the rearranged branches is expressed as:
[0050]
[0051] The first term for optimization loss is classification loss. Right now:
[0052]
[0053] Where CE represents the cross-entropy loss. The second term of the optimization loss is... Used to transfer knowledge from the base model to permutation branches.
[0054]
[0055] Where KL is the KL divergence. This is the original model before balancing optimization. Finally, the total loss during the balancing fine-tuning phase is:
[0056]
[0057] Where β is the knowledge transfer coefficient from the base model to the multi-teacher model, and ρ is the diversity regularization coefficient. After balanced optimization, we transfer knowledge from the source domain to the target domain using the following loss:
[0058]
[0059] The x in this context originates from the target domain. The model is for the target domain.
[0060] This invention was tested on commonly used image classification datasets such as CIFAR-100, ImageNet-100, and ImageNet-1000, and the results are shown in Table 1. The first half of the categories in each dataset were used to construct the source domain dataset, and the second half were divided into S equal parts to construct S target domain datasets. The experimental results represent the average performance across the S target domains. It can be seen that the multi-teacher method proposed in this invention can improve the performance of existing methods on various datasets. For example, on ImageNet-1000, with 25 target domains, it improves the performance by 3.57 percentage points compared to the existing method PODNet.
[0061] Table 1: Classification accuracy of existing methods and the method of this invention on the target domain. "MTD" represents the multi-teacher method proposed in this invention, and "PODNet w / MTD" represents the classification accuracy achieved by the existing method PODNet after using the multi-teacher method MTD. S represents the number of target domains;
[0062]
Claims
1. A multi-teacher distillation domain knowledge memorization and transfer method using image recognition, the method comprising: Step 1: Train a basic image recognition model using sample images from the source domain. The basic image recognition model includes: L concatenated convolutional layers and a classifier. Step 2: Finally, apply the trained basic image recognition model... Convolutional layers and classifier replication Next, get A teacher model; this The inputs to each teacher model are the first in the basic model. The output of each convolutional layer; Step 3: For the results obtained in Step 2... The parameters of each convolutional layer in the teacher model are randomly rearranged to obtain... A new teacher model; Step 4: Introduce a perturbation for each new teacher model to change the input of each new teacher model; Step 5: Randomly sample the same number of samples for each category from all source domain sample images. 100 samples are used to construct a class-balanced subset of the source domain; Step 6: Use on the category-balanced subset The loss is optimized for each teacher model; Losses include: diversity regularization loss Classification loss Migration loss ; Step 7: Apply distillation loss to the target domain By transferring multi-teacher knowledge distillation to the basic image recognition model, a new image recognition model is obtained. Step 8: Fine-tune the training of the new image recognition model; Step 9: Use the finely tuned new image recognition model to recognize the new image.
2. The method for multi-teacher distillation domain knowledge memorization and transfer using image recognition as described in claim 1, characterized in that, The method for rearranging the parameters of each convolutional layer in step 3 is as follows: ; in, Indicates the number after the arrangement Convolutional layer parameters, Indicates the first Convolutional layer parameters, Indicates the first Layer rearrangement matrix, Indicates the first Transpose of the rearrangement matrix at level -1.
3. The method for multi-teacher distillation domain knowledge memorization and transfer using image recognition as described in claim 1, characterized in that, The specific method for step 4 is as follows: ; in, This represents the input to the i-th new teacher model. In the basic model, the first The output of each convolutional layer, It follows a normal distribution. This is the scaling factor. It is the perturbation factor.
4. The method for multi-teacher distillation domain knowledge memorization and transfer using image recognition as described in claim 1, characterized in that, Diversity regularization loss in step 6 Classification loss Migration loss The calculation method is as follows: ; ; ; in, express The 2-combination number of a new teacher model Indicates the first A sample of teachers Embedded, Indicates the first A sample of teachers The embedding, CE(.) is the cross-entropy loss, This represents the output of the basic image recognition model. This represents the output of the model before fine-tuning. Indicates sample The correct corresponding label, Indicates the first The output of each teacher; This represents the average output of the teacher model.
5. The multi-teacher distillation domain knowledge memorization and transfer method using image recognition as described in claim 4, characterized in that, Distillation loss in step 7 The calculation method is as follows: ; in, Indicates the current model With multi-teacher model KL divergence over the target domain Originating from the target domain, This is a new image recognition model for the target domain.
Citation Information
Patent Citations
Multi-teacher self-adaptive joint knowledge distillation
CN112418343A
Double-distribution matching multi-domain adaptive segmentation method and system for water area scene
CN116468744A