Contrast-enhanced characterization standardized distillation algorithm

Through the standardized distillation algorithm of contrast-enhanced characterization (CERND), the output characteristics of the teacher network and student network are standardized, and a new loss function is established, which solves the problem of limited performance of existing knowledge distillation methods when deploying models, and realizes the effect that the student network performance is close to or exceeds the teacher network performance, and meets the need to deploy complex models on resource-limited devices.

CN120106137APending Publication Date: 2025-06-06XIAN UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510182548.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing knowledge distillation method has the problem of performance limitations when deploying models, especially when the gap between the student model and the teacher model is large, the performance of the student model will decline, making it difficult to deploy complex models on resource-limited devices.

Method used

A contrast-enhanced characterization standardized distillation algorithm (CERND) is proposed. By performing Z-score standardization and extreme difference standardization of the output characteristics of the teacher network and student network, a new loss function CERND Loss is established, and it is integrated into the CRD Loss to form a CERND loss function to improve the distillation effect.

Benefits of technology

By characterizing standardization and new loss functions, the information redundancy and negative impact of negative samples are effectively weakened, learning efficiency and model distinction capabilities are improved, and the performance of student networks is closer to or even greater than that of teacher networks, thus meeting the need to deploy complex models on resource-limited devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106137A_ABST
    Figure CN120106137A_ABST
Patent Text Reader

Abstract

The invention discloses a contrast enhancement characterization standardization distillation algorithm, which is characterized in that Z-score standardization and range standardization processing are carried out on output characteristics of a teacher network and a student network, learning is promoted by using differences between sample pairs, and new loss is established and integrated into CRD Loss. According to the method, the characterization standardization technology and the CERND loss function are combined, experimental verification is carried out on a plurality of reference data sets, and the performance superior to that of a traditional contrast characterization distillation algorithm and other advanced knowledge distillation methods is shown; the problem of information redundancy caused by a large number of negative sample pairs in the prior art can be solved, the distillation effect can be improved, and effective model compression is achieved; and after the model is compressed, the performance of the student network can be closer to or even better than that of the teacher network, so that the requirement of deploying a complex model on resource limited equipment is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural network model deployment, and relates to a contrast enhancement characterization standardization distillation algorithm. Background Art

[0002] With the development of deep neural networks, their models are becoming more and more complex and the number of parameters is increasing. It is a huge challenge to deploy these bulky deep models on devices with limited resources. To this end, researchers have begun to study various model compression and acceleration technologies. In the face of the demand for model deployment, lightweight technologies are gaining more and more attention. Lightweight technologies in the field of deep learning mainly include: model pruning, parameter sharing, knowledge distillation, etc.

[0003] Knowledge distillation is a typical model compression and acceleration method that can bring the performance of a small model closer to that of a large model by transferring "knowledge". In the process of knowledge distillation, in order to achieve full distillation from a certain perspective, many researchers have begun to look at knowledge from the perspective of information theory, believing that the teacher network is to impart the ability to analyze the distribution of "knowledge" to the student network. In the past, the teacher network acquired knowledge by using the maximum log-likelihood MLL or the maximum conditional mutual information MCMI to approximate the Bayesian conditional probability distribution of "knowledge"; in CRD, the lower bound of maximizing mutual information was proved, and the knowledge of contrastive learning was cited to propose CRDLoss. Now, many studies have been improved on the basis of the CRD algorithm.

[0004] The traditional method transfers the soft labels of the teacher model to the student model through a softmax function with a shared temperature. However, this method assumes that the teacher and student models have the same temperature parameters, which may lead to limited performance. To address this problem, some studies have proposed setting the temperature according to the weighted standard deviation of logits and performing Z-score normalization preprocessing. This method enables the student model to focus more on the core logical value relationship in the teacher model, thereby improving the distillation performance; and points out that when the gap between the student model and the teacher model is too large, the performance of the student model will decline. To alleviate this limitation, some studies have proposed a multi-teacher knowledge distillation method, which introduces an intermediate-sized network (teaching assistant model) to narrow the gap between the student model and the teacher model.

[0005] The principle of knowledge distillation is to minimize the KL divergence between the probability outputs of the teacher network and the student network. Geoffrey Hinton, Oriol Vinyals and Jeff Deant pioneered knowledge distillation. The authors explored in depth how to effectively compress the knowledge of the integrated model into a single model, and proposed an innovative integrated model-knowledge distillation framework. This pioneering work laid an important foundation for the subsequent series of knowledge distillation methods. In the process of development, knowledge distillation is mainly divided into relationship-based knowledge distillation, feature-based knowledge distillation, and logits-based knowledge distillation.

[0006] Recently, a study proposed a new method called "Mutual Information Maximization Knowledge Distillation (MIMKD)". This method aims to implement knowledge distillation from the feature level by simultaneously estimating and maximizing the mutual information lower bound of local and global feature representations between the teacher network and the student network through comparative learning objectives. This strategy effectively avoids the adverse effects that a large number of negative sample pairs may bring in contrastive response distillation. From the perspective of mutual information, many studies have demonstrated methods to improve distillation performance. Although many works have pointed out its reliance on a large number of negative sample pairs based on CRD, none of them have fully demonstrated the necessity of this reliance.

[0007] A plug-and-play Z-score preprocessing method was proposed in the literature, which normalizes logits before applying the softmax function and Kullback-Leibler divergence, and showed good performance in logits-based knowledge distillation. However, this method is not suitable for feature-based knowledge distillation. Summary of the invention

[0008] The problem solved by the present invention is to provide a contrast enhanced representation normalization distillation algorithm (CERND), which can improve the distillation effect and achieve effective model compression. The performance of the student network can be closer to or even exceed the performance of the teacher network, thereby meeting the needs of deploying complex models on resource-limited devices.

[0009] The present invention is achieved through the following technical solutions:

[0010] A contrast enhancement characterization normalization distillation algorithm, including the following operations:

[0011] 1) In a neural network including a teacher network and a student network, a data set including positive and negative samples is preprocessed and then input into an input layer;

[0012] 2) The teacher network and the student network are independently subjected to feature extraction. After feature extraction, the obtained features are subjected to two-channel normalization to obtain standardized features.

[0013] The two-channel standardization processing is to perform Z-score standardization processing on one channel and range standardization processing on the other channel, and combine the two channels to obtain standardized features;

[0014] Establish a new loss and integrate it into CRD Loss;

[0015] 3) The teacher network and the student network pair positive and negative samples based on standardized features, and use the differences between sample pairs to promote learning. During the pairing process, the CERND loss function is formed based on the CRD Loss function and the EN Loss function.

[0016] The EN Loss function is obtained based on the margin value and the maximum negative difference value; the margin value is a set value of the acceptable range of the difference between positive and negative samples, and the maximum negative difference value is the difference between a sample and its nearest negative sample;

[0017] 4) The student network imitates the output distribution of the teacher network by minimizing the CERND loss function. The prediction results of the student network are as close as possible to the prediction results of the teacher network, so that the student network can inherit the knowledge learned by the teacher network.

[0018] Further, the standardized features are:

[0019] New features = Z score(features) +Min_Max (features) ;

[0020] Among them, Z score(features) It is the feature obtained by Z-score standardization; Min_Max (features) It is the feature obtained by range standardization.

[0021] Furthermore, the CERND loss function is:

[0022] CERNDLoss=CRDLoss+ENLoss;

[0023] Among them, CRDLoss is the CRD Loss loss function, ENLoss is the EN Loss loss function;

[0024] EN Loss = Relu (margin-x max ), where margin is the margin value between positive and negative samples, x is the difference between the sample and its nearest negative sample, and xmax Represents the maximum value in the x vector;

[0025] x=log_D1-log_D0, log_D1 is the logarithm of the positive sample, and log_D0 is the logarithm of the negative sample.

[0026] Furthermore, in addition to imitating the teacher network, the student network also interacts with the true labels through prediction results to ensure that its learned knowledge is accurate and reliable.

[0027] Compared with the prior art, the present invention has the following beneficial technical effects:

[0028] The contrast enhancement representation standardization distillation algorithm provided by the present invention performs Z-score standardization and range standardization on the output features of the teacher network and the student network, including two strategies of Z-score and range standardization, to process the output features of the teacher network and the student network respectively, thereby effectively weakening the information redundancy and negative impact brought by negative sample pairs.

[0029] At the same time, based on the sufficiency of a large number of negative sample pairs for CRD training, the present invention increases the triplet loss, uses the sample pair difference to optimize the loss function, makes the positive sample pairs more closely clustered, and the negative sample pairs more significantly separated, improves the learning efficiency and the model's distinguishing ability, establishes a new loss and integrates it into the CRD Loss to form a CERND loss function to solve the negative impact of a large number of negative sample pairs, uses the two-channel normalization of the output features of the teacher network and the student network to reduce information redundancy, and enhances the utilization of negative sample pairs.

[0030] The present invention combines representation normalization technology and CERND loss function, and is experimentally verified on multiple benchmark data sets, showing performance superior to traditional comparative representation distillation algorithms and other advanced knowledge distillation methods. The present invention can solve the information redundancy problem caused by a large number of negative sample pairs in the prior art, improve the distillation effect, and achieve effective model compression; after model compression, the performance of the student network can be closer to or even exceed that of the teacher network, thereby meeting the needs of deploying complex models on resource-limited devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a schematic diagram of the classic CRD theory;

[0032] Figure 2 It is a schematic diagram of the flow of the CERND algorithm of the present invention. DETAILED DESCRIPTION

[0033] The present invention is further described in detail below in conjunction with the embodiments, which are intended to explain the present invention rather than to limit it.

[0034] Existing model deployment research focuses on architecture and loss function optimization, but pays insufficient attention to the impact of input sample level. We observed that a large number of negative sample pairs may cause information redundancy and affect the efficiency of knowledge transfer. To address this problem, we propose a representation standardization method to process the output features of the teacher network and the student network through Z-score and range standardization to reduce information redundancy.

[0035] At the same time, a new loss is established and integrated into CRD Loss to strengthen the similarity of positive sample pairs and increase the distance of negative sample pairs. The present invention proposes the CERND algorithm based on representation standardization and a new loss function, and experimentally verifies it on the CIFAR100 dataset. The results show that compared with the CRD algorithm and the latest knowledge distillation method based on features or logits, the CERND algorithm of the present invention exhibits better performance and obvious distillation effect.

[0036] The following is a detailed description of each part.

[0037] Figure 1 As shown in Figure 1, a schematic diagram of the classic CRD theory. Points of different shapes represent samples of different categories, and the black lines between different shapes represent the decision boundaries between samples (i.e., there is a distance between sample distributions).

[0038] The present invention proposes a CERND algorithm based on CRD. The teacher network transfers contrastive learning to the student network, uses the output features of the teacher network and the student network and performs two-channel normalization to reduce information redundancy, and designs a new loss to enhance the utilization of negative sample pairs.

[0039] 1. A large number of negative samples are sufficient for the needs

[0040] Since contrastive learning requires a large number of negative sample pairs for network training, when contrastive learning is combined with knowledge distillation, we will deeply analyze the role of negative sample pairs in the CRD algorithm. The traditional knowledge distillation (KD) objective function usually adopts a complete decomposition form, which has limitations in transferring structural knowledge. In order to overcome this shortcoming, the lower limit of the mutual information between the teacher network and the student network is introduced to achieve effective transfer of structural knowledge.

[0041] Assume a latent variable C, and use T to represent the teacher network and S to represent the student network. Construct the mutual information I(T, S) between the teacher network and the student network. The lower limit of the mutual information is as follows:

[0042] I(T, S)≥log N+E q (T,S|C=1)log q(C=1|T,S) (1)

[0043] Where N represents the number of negative sample pairs input. First, assume that q(T, S|C=1)=P(T, S) represents the joint distribution, q(C=0|T, S)=P(T)P(S) represents the marginal distribution, then the expectation of the joint distribution of the teacher network and the student network is expressed as E q (T, S|C=1); E q (T, S|C=1) is represented by E X , E X Between 0-1.

[0044] When a pair of identical pairs is input, the prior value of C is q(C=1)=1 / N+1, q(C=0)=N / N+1, and N is usually regarded as a parameter representing the number of negative sample pairs. However, this intuitive understanding may overlook an important fact: in practical applications, N is not directly equivalent to the absolute number of negative sample pairs required by the entire network. Since the Bayesian probability distribution is unknown, the conditional distribution q(c=1|T, S) in the lower bound is unknown, and it is impossible to directly determine whether it is affected by N. Therefore, we need to find a concise method to explicitly represent the number of negative sample pairs implied in the prior probability, thereby expressing the relationship between negative sample pairs and contrast representation distillation.

[0045] From mathematical statistics we can know:

[0046]

[0047] Among them, X represents a given sample, and θ is the parameter of the sample distribution. The relevant formula reveals the relationship between the posterior probability P(θ|X) and the prior probability p(θ) and the likelihood function P(X|θ). Specifically, the formula shows that the posterior probability is proportional to the product of the prior probability and the likelihood function, that is:

[0048] P(θ|X)∝P(θ)·P(X|θ) (3)

[0049] Using formula (3), we can rewrite the conditional probability q(c=1|T,S). Specifically, we replace q(c=1|T,S) in the original formula with q(c=1|T,S)·q(c=1). We can get:

[0050] I(T, S)≥logN+E X (logX+logq(C=1)) (4)

[0051] I(T, S)≥logN+E X logX-E X log(N+1) (5)

[0052]

[0053] Inequality (6) expresses the expected value of the joint distribution of the teacher network and the student network, which contains the term related to negative samples The larger the N value is, the larger the lower limit of the mutual information is. This means that the mutual information between the teacher network and the student network is improved, allowing the knowledge of the teacher network to be more fully imparted to the student network, reflecting the sufficiency of the number of negative sample pairs in improving the knowledge sharing between the teacher network and the student network.

[0054] 2. Characterization Standardization Knowledge Distillation

[0055] Based on the need for a large number of negative sample pairs in contrastive learning, we introduced a representation normalization method, which specifically covers two strategies: Z-score normalization and range normalization. Under this framework, the output features of the teacher network and the student network are standardized separately. This innovative approach can significantly reduce the information redundancy and potential negative impact brought by negative sample pairs, and provide strong support for the efficient implementation of contrastive learning.

[0056] New features = Z score(features) +Min_Max (features) (7)

[0057] Formula (7) is the process of processing the original features to obtain new features. The original features are processed by two-channel Z-score processing and range standardization processing (both standardization processing methods are common tools) to obtain new features. This two-channel feature processing method is called a representation standardization processing method.

[0058] 3. Propose new losses

[0059] If there is noise in the data sample (causing a deviation in the selected positive and negative samples), contrastive learning will introduce a large amount of additional noise while enhancing the data. At this time, based on the contrast loss function of the CRD algorithm, the present invention incorporates the triplet loss to deeply explore and fully utilize the differences between sample pairs to promote the learning process, thereby suppressing the impact of the introduced noise.

[0060] CERNDLoss=CRDLoss+ENLoss (8)

[0061] Formula (8) is a new loss composition method proposed by the present invention, which is composed of CRD Loss and EN Loss. Among them, CRD Loss is an existing contrast loss, and the main implementation steps of the EN Loss we proposed are as follows:

[0062] First, set a margin value, which is used to measure the acceptable range of the difference between positive and negative samples.

[0063] Next, the logarithmic difference between positive samples (log_D1) and negative samples (log_D0) is calculated. To evaluate the performance of the model in distinguishing positive and negative samples, we focus on the difference between each sample and its nearest negative sample (i.e., the maximum negative difference).

[0064] Finally, we calculate the loss based on the margin value and the maximum negative difference value.

[0065] Specifically, EN Loss can be reflected as follows:

[0066] margin=1#Adjust the distance between positive and negative samples according to actual conditions

[0067] x = log_D1-log_D0 # Calculate the logarithmic difference between positive samples (log_D1) and negative samples (log_D0)

[0068] EN Loss = Relu (margin-x max )#When the maximum difference is smaller than the boundary value, the difference may not be distinguishable and a loss is incurred by default.

[0069] where x max Represents the maximum value in the x vector. The ReLU function is a common mathematical function, and its mathematical expression is f(x) = max(0, x)

[0070] By optimizing the loss function, on the basis of the original contrast loss function, the present invention utilizes the sample pair difference to optimize the loss function, so that the positive sample pairs are more closely clustered and the negative sample pairs are more significantly separated, thereby improving the learning efficiency and the model's distinguishing ability; thereby achieving a closer clustering of the positive sample pairs and a more significant separation of the negative sample pairs in the feature space; this improvement not only improves the learning efficiency, but also further enhances the model's distinguishing ability.

[0071] The present invention combines the representation normalization technology with a newly designed loss function, and innovatively proposes a contrast enhanced representation normalization distillation algorithm (CERND for short).

[0072] like Figure 2 As shown, the contrast enhancement characterization standardization distillation algorithm of the present invention includes the following operations:

[0073] 1) In a neural network including a teacher network and a student network, a data set including positive and negative samples is preprocessed and then input into an input layer;

[0074] Specifically, public data sets are used in the network training process. Data preprocessing includes data cleaning, data normalization, and data standardization, and data enhancement techniques include rotation, scaling, flipping, cropping, adding noise, etc. The data is then converted into a format acceptable to the neural network (such as tensors) and passed to the network's input layer;

[0075] 2) The teacher network and the student network are independently subjected to feature extraction. After feature extraction, the obtained features are subjected to two-channel normalization to obtain standardized features.

[0076] The two-channel standardization processing is to perform Z-score standardization processing on one channel and range standardization processing on the other channel, and combine the two channels to obtain standardized features;

[0077] Establish a new loss and integrate it into CRD Loss;

[0078] Specifically, the standardized features are:

[0079] New features = Z score(features) +Min_Max (features) ;

[0080] Among them, Z score(features) It is the feature obtained by Z-score standardization; Min_Max (features)

[0081] It is the characteristic obtained by range standardization;

[0082] 3) The teacher network and the student network pair positive and negative samples based on standardized features, and use the differences between sample pairs to promote learning. During the pairing process, the CERND loss function is formed based on the CRD Loss function and the EN Loss function.

[0083] The EN Loss function is obtained based on the margin value and the maximum negative difference value; the margin value is a set value that measures the acceptable range of differences between positive and negative samples, and the maximum negative difference value is the difference between a sample and its nearest negative sample.

[0084] Furthermore, the CERND loss function is:

[0085] CERNDLoss=CRDLoss+ENLoss;

[0086] Among them, CRDLoss is the CRD Loss loss function, ENLoss is the EN Loss loss function;

[0087] EN Loss = Relu (margin-xmax ), where margin is the margin value between positive and negative samples, x is the difference between the sample and its nearest negative sample, and x max Represents the maximum value in the x vector;

[0088] x=log_D1-log_D0, log_D1 is the logarithm of the positive sample, and log_D0 is the logarithm of the negative sample.

[0089] The student network mimics the output distribution of the teacher network by minimizing the loss function, and the prediction results of the student network should be as close as possible to the prediction results of the teacher network. In this way, the student network can inherit the knowledge learned by the teacher network.

[0090] Furthermore, in addition to imitating the teacher network, the student network also needs to interact with the real labels through prediction results to ensure that the knowledge it has learned is accurate and reliable. The design of this loss function aims to guide the student network to move closer to the learning direction of the teacher network, thereby achieving effective knowledge transfer and improving model performance.

[0091] To comprehensively evaluate the performance of the CERND algorithm, we conducted extensive experimental validation on multiple benchmark datasets and compared its performance with traditional contrastive representation distillation algorithms and other advanced knowledge distillation methods.

[0092] A large number of negative sample pairs may also bring some negative effects, such as increasing computational complexity and possibly causing overfitting. To solve this problem, we proposed the CERND algorithm. This algorithm uses the output features of the teacher network and the student network to perform representation standardization to reduce information redundancy. At the same time, it establishes its own loss function and introduces additional constraints to enhance the use of negative sample pairs. In this way, we can fully utilize the benefits of negative sample pairs while mitigating their possible negative effects.

[0093] The present invention is compared and evaluated with the existing advanced knowledge distillation algorithms. Resnet56, wrn_40_2, vgg13, WRN40_2 are selected as teacher networks, resnet20, wrn_16_2, vgg8, WRN40_1 are selected as student networks, and some prior experiments are conducted on the CIFAR100 dataset to comprehensively evaluate the performance of the student network. In the experiment, we selected the knowledge distillation algorithm based on basic features and the knowledge distillation algorithm based on logits to compare with CERND; the results are shown in Table 1.

[0094] Table 1 Experimental comparison of Top-1 accuracy (%) of different distillation algorithms on CIFAR dataset

[0095]

[0096]

[0097] ReviewKD, OFD and other baseline methods improve the accuracy of the student network to different degrees, but the overall improvement is limited. The above networks are different classification network architectures. From the experimental results, we can see that the CERND algorithm we proposed can greatly improve the accuracy of the classification network, and can better make the small model approach the effect of the large model, or even surpass it.

[0098] The present invention adds representation standardization on the basis of CRD, and the student network achieves a good accuracy improvement under different combinations of student networks and teacher networks. This shows that feature standardization helps to further improve the effect of knowledge distillation. The CERND algorithm achieves the highest accuracy based on CRD+representation standardization, which shows that the CERND method has significant advantages in knowledge distillation.

[0099] Other combination methods such as CRD+CutMixPick, simKD, CAT-KD and other methods also showed different degrees of accuracy improvement, but they were basically lower than CERND and CRD+characterization normalization. Although kd+CTKD and its variants (such as adding Normalization) also provided a certain degree of accuracy improvement, the overall effect was not as good as the method proposed in the present invention.

[0100] The above embodiments are preferred examples for implementing the present invention, and the present invention is not limited to the above embodiments. Any non-essential additions and substitutions made by those skilled in the art based on the technical features of the technical solution of the present invention shall fall within the protection scope of the present invention.

Claims

1. A contrast enhancement characterization standardization distillation algorithm, characterized in that: The following operations are included: 1) In a neural network including a teacher network and a student network, a data set including positive and negative samples is preprocessed and then input into an input layer; 2) The teacher network and the student network are independently subjected to feature extraction. After feature extraction, the obtained features are subjected to two-channel normalization to obtain standardized features. The two-channel standardization processing is to perform Z-score standardization processing on one channel and range standardization processing on the other channel, and combine the two channels to obtain standardized features; Establish a new loss and integrate it into CRD Loss; 3) The teacher network and the student network pair positive and negative samples based on standardized features, and use the differences between sample pairs to promote learning; During the pairing process, the CERND loss function is formed based on the CRD Loss loss function and the EN Loss loss function; The EN Loss function is obtained based on the margin value and the maximum negative difference value; the margin value is a set value of the acceptable range of the difference between positive and negative samples, and the maximum negative difference value is the difference between a sample and its nearest negative sample; 4) The student network imitates the output distribution of the teacher network by minimizing the CERND loss function. The prediction results of the student network are as close as possible to the prediction results of the teacher network, so that the student network can inherit the knowledge learned by the teacher network.

2. The contrast enhancement characterization normalization distillation algorithm according to claim 1, characterized in that: The standardized features are: Newfeatures=Z score ( features )+Min_Max (features) ; Among them, Z score ( features ) is the feature obtained by Z-score standardization; Min_Max (features) It is the feature obtained by range standardization.

3. The contrast enhancement characterization normalization distillation algorithm of claim 1, wherein: The CERND loss function is: CERNDLoss=CRDLoss+ENLoss; Among them, CRDLoss is the CRD Loss loss function, ENLoss is the EN Loss loss function; EN Loss = Relu (margin-x max ), where margin is the margin value between positive and negative samples, x is the difference between a sample and its nearest negative sample, x max Represents the maximum value in the x vector; x=log_D1-log_D0, log_D1 is the logarithm of the positive sample, and log_D0 is the logarithm of the negative sample.

4. The contrast enhancement characterization normalization distillation algorithm of claim 1, wherein: In addition to imitating the teacher network, the student network also interacts with the true labels through prediction results to ensure that its learned knowledge is accurate and reliable.