TSK fuzzy classifier fusing decoupling knowledge distillation and curriculum learning

By integrating the method of decoupling knowledge distillation and curriculum learning, the probability distribution output by the teacher model is decoupled into target class and non-target class probabilities, the non-target class probability weight is increased, and an adaptive weight loss function is constructed. This solves the problem of insufficient generalization and convergence efficiency of the TSK fuzzy classifier in knowledge distillation, and achieves a more efficient training method and accuracy improvement.

CN120671728APending Publication Date: 2025-09-19HUZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510626549.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, the TSK fuzzy classifier has difficulty in efficiently extracting the features and probability distribution output by the teacher model in the field of knowledge distillation, resulting in insufficient generalization and convergence efficiency.

Method used

The TSK fuzzy classifier integrates decoupled knowledge distillation and curriculum learning. It evaluates the difficulty of training samples through a difficulty scale, sorts and filters out noise samples, and uses a training scheduler to arrange the student model for knowledge distillation training from easy to difficult. It also decouples the probability distribution output by the teacher model into target class and non-target class probabilities, increases the probability weight of the non-target class, and constructs a total loss function with adaptive weights.

Benefits of technology

The generalization and convergence efficiency of the TSK fuzzy classifier are improved, which saves training time. The model's accuracy is improved by adaptive weight adjustment to adapt to training changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671728A_ABST
    Figure CN120671728A_ABST
Patent Text Reader

Abstract

The invention discloses a TSK fuzzy classifier fusing decoupling knowledge distillation and curriculum learning, samples are subjected to difficulty ranking through curriculum learning, a training strategy is formulated, a model is made to start learning from easy samples and gradually go to complex samples and knowledge, the TSK fuzzy classifier has higher convergence efficiency, and the classification efficiency is improved. And then knowledge contained in the teacher model is extracted into the TSK fuzzy classifier through decoupling knowledge distillation, and target category loss and non-target category loss can be better coupled through decoupling knowledge distillation, so that the precision and performance of the TSK fuzzy classifier are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field

[0001] The present invention relates to the technical field of knowledge distillation, and in particular to the technical field of a TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning. [Background Technology]

[0002] Knowledge distillation is a method of model compression and improving model accuracy. It uses the probability distribution of the last layer output of the teacher model or the deep features of the middle layer to guide the training of the student model, allowing the student model to imitate the features extracted by the teacher model and the probability distribution of the output to improve its own accuracy. In the current era of large models, knowledge distillation has received increasing attention in the field of deep learning. The TSK fuzzy classifier is an efficient nonlinear approximator that can map input to output through generated fuzzy rules, but there are still problems to be solved. Currently, in the field of knowledge distillation, the effect of feature distillation is generally better than traditional Logits distillation. The TSK fuzzy classifier is not suitable for feature distillation due to its own structure. Therefore, how to more efficiently extract the knowledge in the output distribution of the teacher model into the TSK fuzzy classifier is a problem that needs to be solved urgently. [Summary of the invention]

[0003] The purpose of the present invention is to solve the problems in the prior art and propose a TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning. It can integrate curriculum learning with decoupled knowledge distillation and perform fuzzy modeling on the original data, thereby efficiently extracting the knowledge in the output distribution of the teacher model into the TSK fuzzy classifier, so that the TSK fuzzy classifier has stronger generalization and convergence efficiency.

[0004] To achieve the above objectives, the present invention proposes a TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning, comprising the following steps:

[0005] S1. Use the pre-trained teacher model to calculate the loss value of the training sample Data;

[0006] S2. Use the difficulty scale to evaluate the difficulty of the training sample data according to the loss value, sort the training sample data from easy to difficult, and filter out the noise samples to obtain the sorted training sample Sorted Data;

[0007] S3. Use the training scheduler to arrange the student model to perform knowledge distillation training from easy to difficult. As the training batch increases, the sorted training samples Sorted Data are added to the training through the training scheduler;

[0008] S4. Input the sorted training sample Sorted Data into the teacher model and the student model to obtain the probability distribution of their outputs. Decouple the probability distribution of the teacher model output into the target class label probability and the non-target class label probability.

[0009] S5. Increase the weight of the non-target class label probability, reduce the weight of the target class label probability, and then calculate the loss of the teacher model and the student model;

[0010] S6. Repeat the training until the student model converges.

[0011] Preferably, the student model uses a first-order TSK fuzzy classifier, the difficulty measurer and the teacher model use the same third-order TSK fuzzy classifier, and the third-order TSK fuzzy classifier used by the teacher model has the same architecture as the first-order TSK fuzzy classifier used by the student model.

[0012] Preferably, the training sample Data includes original training data X={x i ,x i =[x1,x2,...,x n ] T ,i=1,2,...,m} and label Y={y i ,y i ∈{1,2,...,c},i=1,2,...,m}, where n, m and c are the number of sample features, the number of samples and the number of categories respectively. Taking X as input, the cross entropy loss L={L1,L2,...,L m}, the smaller the cross entropy loss, the easier the sample training is. The sorted training data is obtained according to the calculated cross entropy loss. and tags

[0013] As a preference, after the training scheduler obtains the sorted training samples Sorted Data, it arranges the student model for training and defines the training data X train =φ and label Y train =φ, the sorted training data is split into and First, X 1 and Y 1 Add new training data X respectively train and Y train Then start training the model, j is the parameter of the number of training rounds E, as j increases, the training data X j and Y j Join X separately train and Y trainand participate in the training. As j increases, difficult samples are continuously added to the training set until X E and Y E Add to X train and Y train In ,the training method is knowledge distillation, using the same third-order TSK fuzzy classifier as the difficulty measure and the teacher model as the student model using the first-order TSK fuzzy classifier for knowledge distillation.

[0014] Preferably, the method for decoupling the probability distribution output by the teacher model includes defining q=[p t ,p \t ]∈ 1×2 is the target class probability (p t ) and non-target class probability (p \t ), the calculation formula is as follows:

[0015]

[0016] Define the probability distribution between non-target classes (excluding the target class t probability) as in The calculation formula is as follows:

[0017]

[0018] According to formulas 4.2 and 4.3, the KL divergence loss is expressed as follows:

[0019]

[0020] where p R and p U are the probability distributions predicted by the teacher model and the student model respectively. According to formulas 4.2, 4.3 and 4.4, L KL It can be expressed as:

[0021]

[0022] According to formula 4.6, the KL divergence loss is decomposed into two coupled parts, namely q R With q U The KL divergence loss between the two parts and the KL divergence loss between the non-target class probabilities of the student and teacher are decoupled and two hyperparameters β and γ are introduced as the weights of the two parts of the loss to increase the loss weight between the non-target class probabilities. The formula is as follows:

[0023]

[0024] Last L DKD Plus the cross entropy loss L between the student model and the true label CE, the total loss of knowledge distillation is obtained as follows:

[0025] L ALL =αL CE +(1-α)L DKD (4.8)

[0026] An adaptive hyperparameter α is introduced into the total loss. At the beginning of training, the value of α is higher to increase the weight of the cross-entropy loss, allowing the model to focus on learning the labels. As the model continues to iterate, the adaptive weight α continues to decrease to increase the weight of the decoupled knowledge distillation loss, allowing the model to focus on learning the teacher model.

[0027] Beneficial effects of the present invention: The present invention further decomposes the knowledge of the teacher model into target class and non-target class knowledge by adopting decoupled knowledge distillation, and further improves the accuracy of the model by increasing the weight of non-target class knowledge. Only one teacher model is required, and this teacher model takes into account the tasks of knowledge distillation and course learning at the same time. While providing supervision information for the TSK fuzzy classifier, it also needs to complete the work of measuring and sorting the difficulty of samples. This not only improves the accuracy performance and convergence speed of the TSK fuzzy classifier, but also saves model training time. Adaptive weight parameters are used in training, and weights can be dynamically adjusted during training. It can flexibly adapt to changes in the TSK fuzzy classifier during training, so as to achieve the effect of improving the model convergence speed and model accuracy. By screening and processing samples, a more efficient training method is obtained, and the accuracy performance of the first-order TSK fuzzy classifier is improved through knowledge distillation. On this basis, a method of decoupling the teacher model knowledge is proposed, which decouples the probability distribution output by the teacher model and decomposes it into target class probability and non-target class probability. At the same time, the weight of the non-target class probability is increased, so that the student model can better learn the hidden knowledge in the output distribution of the teacher model. By further decomposing the loss function, assigning new weights and constructing a new loss function, the accuracy performance of the student model is further improved. The present invention first sorts the samples by difficulty through course learning and formulates a training strategy, so that the model starts learning from easy samples and gradually advances to complex samples and knowledge, so that the TSK fuzzy classifier has stronger convergence efficiency. Then, the knowledge contained in the teacher model is extracted into the TSK fuzzy classifier through decoupling knowledge distillation. Through decoupling knowledge distillation, the target class loss and the non-target class loss can be better coupled, thereby improving the accuracy and performance of the TSK fuzzy classifier.

[0028] The features and advantages of the present invention will be described in detail through embodiments with reference to the accompanying drawings.

Brief Description of the Drawings

[0029] Figure 1 This is the overall framework diagram of a TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning in the present invention;

[0030] Figure 2 This is an overview diagram of knowledge distillation based on curriculum learning of a TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning in the present invention;

[0031] Figure 3 This is the ROC curve of each method on the UCI dataset in a comparative experiment of the TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning in the present invention;

[0032] Figure 4 This is a diagram showing the effect of different distillation temperatures on model performance in a parameter sensitivity experiment of a TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning in the present invention;

[0033] Figure 5 This is a diagram showing the impact of different loss weights on model performance in a parameter sensitivity experiment of a TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning in the present invention;

[0034] Figure 6 This is a convergence analysis diagram in a parameter sensitivity experiment of a TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning in the present invention. [Specific implementation method]

[0035] The overall framework of the TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning is as follows: Figure 1 As shown, the model mainly consists of two parts: the first is the course learning part: first, through pre-training [74-75] The teacher model calculates the loss value for the training sample data and assesses the sample difficulty based on the loss value. It then sorts the training samples from easy to difficult and removes noise samples to obtain sorted data. Finally, as the training batch increases, the sorted data is added to the training through the training scheduler. The second step is the decoupling knowledge distillation: first, the training samples are input into the teacher model and the student model to obtain the probability distribution of their outputs. The probability distribution is then broken down into target class knowledge and non-target class knowledge, and appropriate weights are assigned to each. The loss of the teacher model and the student model is then calculated. Finally, training is repeated until the student model converges.

[0036] This paper designs a knowledge distillation method based on course learning, such as Figure 2 As shown in the figure, the difficulty measurer in course learning is reused as a teacher model in knowledge distillation to distill the TSK fuzzy classifier, which saves training time and improves the performance of the TSK fuzzy classifier. The difficulty measurer and the teacher model use the third-order TSK fuzzy classifier with the same architecture as the first-order TSK fuzzy classifier, because the knowledge distillation efficiency between models with the same architecture is higher. Set the original training data X = {x i ,x i=[x1,x2,...,x n ] T ,i=1,2,...,m} and label Y={y i ,y i ∈{1,2,...,c},i=1,2,...,m}, where n, m and c are the number of sample features, the number of samples and the number of categories respectively. Taking X as input, the cross entropy loss L={L1,L2,...,L m}, the smaller the cross entropy loss, the easier the sample training is. By sorting according to the calculated cross entropy loss, the sorted training data can be obtained. and tags

[0037] After obtaining the sorted training data, the training scheduler arranges the training of the student model. Define new training data X train =φ and label Y train =φ, and then the sorted training data is split into and First, X 1 and Y 1 Add new training data X respectively train and Y train Then start training the model, j is the parameter of the number of training rounds E, as j increases, the training data X j and Y j Join X separately train and Y train and participate in the training. As j increases, difficult samples are continuously added to the training set until X E and Y E Add to X train and Y train In the training method, knowledge distillation is used. The third-order TSK fuzzy classifier as the difficulty measure is also used as the teacher model to perform knowledge distillation on the first-order TSK fuzzy classifier as the student model.

[0038] This invention integrates the design concepts of curriculum learning and knowledge distillation, combining the advantages of the two. By screening and processing samples, a more efficient training method is obtained, and the accuracy performance of the first-order TSK fuzzy classifier is improved through knowledge distillation. On this basis, the invention proposes a method for decoupling the teacher model knowledge, and by further decomposing the loss function, assigning new weights and constructing a new loss function, the accuracy performance of the student model is further improved.

[0039] Traditional knowledge distillation allows the student model to learn the probability distribution p=[p1,p2,...,p c]∈R 1×c , so that the student approaches the teacher's performance, where the probability distribution is processed by the softmax function and p is calculated i The formula is as follows:

[0040]

[0041] Among them, z i is the prediction result of the model on category i, and τ is the temperature parameter in knowledge distillation.

[0042] In traditional knowledge distillation, the probability of the target class is much greater than the probability of the non-target class, which causes the student model to focus on learning the target class knowledge and ignore the non-target class knowledge. In fact, the non-target class probability contains the potential information and hidden knowledge output by the teacher model, which helps the student model learn the probability distribution of the teacher model output, thereby approaching the performance of the teacher model.

[0043] Based on the above idea, this paper attempts to decouple the probability distribution of the teacher model output, decomposing it into the target class probability and the non-target class probability, while increasing the weight of the non-target class probability, so that the student model can better learn the hidden knowledge in the teacher model output distribution, and define q = [p t ,p \t ]∈ 1×2 is the target class probability (p t ) and non-target class probability (p \t ), and decompose Formula 4.1 to obtain:

[0044]

[0045] At the same time, the present invention defines the probability distribution between non-target classes (excluding the target class t probability) as in The calculation formula is as follows:

[0046]

[0047] According to formulas 4.2 and 4.3, the KL divergence loss is expressed as follows:

[0048]

[0049] where p R and p U are the probability distributions predicted by the teacher model and the student model respectively. According to formulas 4.2, 4.3 and 4.4, L KL It can be expressed as:

[0050]

[0051] According to formula 4.6, it can be seen that the KL divergence loss is decomposed into two coupled parts, namely q R With q U The KL divergence loss between the student and the teacher's non-target class probabilities is used to decouple the two coupled losses and introduce two hyperparameters β and γ as weights of the two losses to increase the loss weight between the non-target class probabilities and make the non-target class knowledge effective. The formula is as follows:

[0052]

[0053] Last L DKD Plus the cross entropy loss L between the student model and the true label CE , the total loss of knowledge distillation can be obtained as follows:

[0054] L ALL =αL CE +(1-α)L DKD (4.8)

[0055] In addition, an adaptive hyperparameter α is introduced into the total loss. At the beginning of training, the value of α is higher to increase the weight of the cross-entropy loss, so that the model focuses more on learning the labels. As the model continues to iterate, the adaptive weight α will continue to decrease to increase the weight of the decoupled knowledge distillation loss, so that the model focuses more on learning the teacher model, thereby reducing the difficulty for students to imitate the teacher.

[0056] 4.3.3CDKD-TSK Model Algorithm Flow

[0057] Algorithm 4.1 is the process of the CDKD-TSK algorithm:

[0058]

[0059]

[0060] Experiment and analysis

[0061] This paper first introduces the experimental settings, then demonstrates the performance of the CDKD-TSK model through comparative experiments, and finally analyzes the CDKD-TSK model through parameter sensitivity experiments to study the impact of different strategies on its performance.

[0062] Experimental setup

[0063] The experiment was conducted on a hardware environment equipped with a GPU NVIDIA GeForce RTX4090 24GB, a CPU Intel i9 12900KF, 64GB RAM, and 64-bit Microsoft Windows 10 and Python 3.8.17 with torch. The experiment was carried out in the programming environment of 2.4.1. Seven commonly used datasets and two medical diagnosis datasets in the UCI dataset were selected for experiments and the performance of the model was evaluated. The temperature τ of traditional knowledge distillation was set to 4.0, α was 1.0, and β was 9.0. The temperature τ of decoupled knowledge distillation was set to 4.0, α was 1.0, β was the adaptive weight parameter, and γ was 8.0. The course learning used the loss value to measure the sample difficulty and the default training strategy. The first-order TSK fuzzy classifier with 16 rules was selected as the student model Student, and the third-order TSK fuzzy classifier with 1 rule was selected as the teacher model Teacher. The post-processing part was optimized by gradient descent. The number of model training rounds was 60, the initial value of the learning rate was 1e-1 and it was decreased every 30 rounds. The AdamW optimizer was used to train the model, and other parameters were the default values. To ensure the accuracy of the results, each experiment was repeated five times and the average value was taken as the final result.

[0064] In the experiment, the comparison algorithms selected traditional knowledge distillation (KD) and decoupled knowledge distillation (DKD), which are representative in Logits distillation, were selected. The high-order TSK distillation low-order TSK (HTSK-LLM-DKD) representing the classic fuzzy classifier was also compared. In addition, KD+CL is the method of traditional knowledge distillation plus course learning, DKD+CL is the method of decoupled knowledge distillation plus course learning, and KD+CL+AW is the method of traditional knowledge distillation plus course learning and adaptive weighting. In all the above distillation methods, the student model is a first-order TSK fuzzy classifier with 16 rules, and the teacher model is a third-order TSK fuzzy classifier with 1 rule. The antecedent solution of the TSK fuzzy classifier adopts equal interval division, and the consequent solution adopts gradient descent for parameter optimization.

[0065] In addition to Accuracy and F1-score, the model evaluation indicators in the experiment also use Precision and Recall as evaluation indicators. The specific formulas are as follows:

[0066] Accuracy=(TP+TN) / (TP+FP+FN+TN) (4.9)

[0067] Precision=TP / (TP+FN) (4.10)

[0068] Recall=TP / (TP+FP) (4.11)

[0069] P=TP / (TP+FP) (4.12)

[0070] R=TP / (TP+FN) (4.13)

[0071] F1-score=(2×P×R) / (P+R) (4.14) Among them, TP represents the number of positive samples correctly predicted by the model, FP represents the number of positive samples incorrectly predicted by the model, TN represents the number of negative samples correctly predicted by the model, FN represents the number of negative samples incorrectly predicted by the model, Accuracy refers to the proportion of correct samples in the model prediction results to the total number of samples, Precision refers to the proportion of positive samples in the positive examples in the model prediction results, Recall refers to the proportion of positive examples in the model prediction results to the samples that are actually positive examples, F1-score is the harmonic mean of Precision and Recall, and the higher the F1-score, the more robust the model. The five-fold cross-validation divides the dataset into five subsets, four of which are used for training in turn, and one subset is used to verify the model performance. The training is repeated five times, and the mean and standard deviation of the evaluation indicators are expressed as mean ± standard deviation in the experimental results.

[0072] Comparative test

[0073] Table 4.1 Accuracy comparison of various methods on the UCI dataset

[0074]

[0075]

[0076] Table 4.1 shows the Accuracy and standard deviation of CDKD-TSK and several other comparison methods on the commonly used UCI dataset, where the best results are in bold. According to the data analysis in the table, each knowledge distillation method has a certain improvement on the basis of the student model, among which the improvement of decoupled knowledge distillation is more obvious than that of traditional knowledge distillation. In addition, it can be seen that the introduction of curriculum learning and adaptive parameters in knowledge distillation has a certain degree of improvement on the performance of the student model. Finally, the model CDKD-TSK has better Accuracy and standard deviation on most datasets, especially on datasets with a large number of samples, but it does not perform well in small sample datasets. This is most likely a problem caused by overfitting, which is also the part that this invention will improve later.

[0077] To verify the model's performance on the medical diagnosis datasets, Table 4.2 shows the accuracy and F1-score of each method on the Stroke and Diabetes datasets. As can be seen from the table, CDKD-TSK achieves the best accuracy and the highest F1-score on both datasets. This suggests that CDKD-TSK's prediction results are more balanced and stable, leading to its superior performance on medical diagnosis datasets. Furthermore, a comparison shows that CDKD-TSK's improvement on the Stroke dataset is more significant than on the Diabetes dataset, likely due to the far greater number of samples in the Stroke dataset. Therefore, it can be concluded that CDKD-TSK is more effective with a larger number of samples.

[0078] Table 4.2 Comparison of Accuracy and F1-score of each method on medical diagnosis dataset

[0079]

[0080]

[0081] To further measure the generalization ability of CDKD-TSK in classification problems, the present invention plots the ROC curves of each method on four UCI datasets, as shown in the figure below: Figure 3 As shown in the figure, the model performance is mainly measured by the AUC value, which is the area under the ROC curve. When the AUC value is 0.5, that is, the area is half, it means that the model performance is equivalent to random classification. The larger the AUC value, the better the model performance. It can be concluded from the illustrated curve and the AUC value of the model that CDKD-TSK has the highest AUC value on all four UCI datasets, followed by HTSK-LLM-DKD, and the traditional knowledge distillation effect is the worst. Therefore, the present invention can be considered that CDKD-TSK has better generalization performance in classification problems.

[0082] Parameter sensitivity experiments

[0083] This paper conducts an in-depth study on the hyperparameters in the CDKD-TSK model and explores the impact of hyperparameter settings on the CDKD-TSK model.

[0084] (1) Influence of temperature parameter τ

[0085] The temperature parameter τ is responsible for softening the output of the teacher model and the student model, which plays the role of label smoothing and has a great influence on the distillation effect. For example, Figure 4The following table shows the accuracy of traditional distillation and CDKD-TSK on four UCI datasets. The temperature parameter τ is selected according to an arithmetic progression from 0 to 10. When τ is 1, the student model directly imitates the probability distribution of the teacher model's output. It can be seen that the improvement in student model accuracy is not obvious at this time, and the student model has difficulty imitating the generalization ability of the teacher model. As the temperature parameter τ increases, the probability distribution of the teacher model's output gradually becomes smoother, and the accuracy of the student model also gradually improves. When the temperature parameter τ is between 3 and 4, the student model accuracy reaches the highest and the distillation effect is the best, that is, the temperature most suitable for distillation is reached. Finally, as the temperature τ continues to increase, the accuracy of the student model continues to decline. This is because the output of the teacher model is too smooth at this time, resulting in no room for the student model to learn, so the student model does not improve significantly. On all four datasets, CDKD-TSK shows a certain degree of improvement compared to traditional knowledge distillation. This shows the important influence of temperature τ on the model and also confirms the role of CDKD-TSK in improving model performance.

[0086] (2) Impact of loss weights β and γ

[0087] From formula 4.5, we can see that β and γ control the weights of target class loss and non-target class loss in decoupled knowledge distillation, which has an important impact on distillation efficiency. Therefore, this paper also selects the above four data sets to test the ratio of γ and β γ / β to study its impact on CDKD-TSK, as shown in the following figure. Figure 5 As shown in the figure, the accuracy of CDKD-TSK on the four data sets has a certain upward trend with the increase of the γ / β ratio. It reaches a peak when the γ / β ratio is between 5 and 10, and then shows a downward trend. The present invention usually adopts the setting of γ / β=8 when using the CDKD-TSK model, that is, the proportion of non-target knowledge is 8 times that of target knowledge, which allows the student model to better learn the knowledge of the teacher model. Of course, for different data sets, the optimal γ / β still needs to be determined by specific experiments.

[0088] (3) Convergence analysis

[0089] The number of training rounds E controls the number of training rounds of the model, that is, the number of times the model traverses the training set, which has a great impact on the convergence effect of the model. Too large or too small a number of training rounds is not conducive to the final accuracy of the model. Therefore, the present invention experiments with different numbers of training rounds to find the optimal number of training rounds for the CDKD-TSK model, such as Figure 6As shown in the figure, the accuracy performance of CDKD-TSK on four UCI datasets with different numbers of training rounds. As the number of training rounds increases, the accuracy of the model continues to improve, which means that the model learns more features of the samples in each training. The accuracy of the model reaches the highest when the number of training rounds is around 60 in the Cleve, Sonar and Phoneme datasets, and the accuracy of the model reaches the highest when the number of training rounds is 70 in the QSAR dataset. After the accuracy of the model reaches the highest, as the number of training rounds continues to increase, the accuracy of CDKD-TSK begins to decrease. This is because the model overfits and the accuracy decreases.

[0090] This paper proposes a TSK fuzzy classification model, CDKD-TSK, based on curriculum learning and decoupled knowledge distillation. This model first uses curriculum learning to sort samples by difficulty and develop a training strategy, allowing the model to start learning from easy samples and gradually advance to complex samples and knowledge, giving the TSK fuzzy classifier greater convergence efficiency. Decoupled knowledge distillation is then used to extract the knowledge contained in the teacher model into the TSK fuzzy classifier. Decoupled knowledge distillation allows for better coupling of target category loss and non-target category loss, thereby improving the accuracy and performance of the TSK fuzzy classifier. Experiments have shown that the CDKD-TSK model proposed in this paper has significant effects on most UCI datasets and also performs well on two medical diagnosis datasets.

[0091] The above embodiments are intended to illustrate the present invention, not to limit the present invention. Any solution that is a simple transformation of the present invention falls within the protection scope of the present invention.

Claims

1. A TSK fuzzy classifier that integrates decoupled knowledge distillation and curriculum learning, characterized by: The following steps are involved: S1. Use the pre-trained teacher model to calculate the loss value of the training sample Data; S2. Use the difficulty scale to evaluate the difficulty of the training sample Data according to the loss value, sort the training sample Data from easy to difficult, and filter out the noise samples to obtain the sorted training sample SortedData; S3. Use the training scheduler to arrange the student model to perform knowledge distillation training from easy to difficult. As the training batch increases, the sorted training samples SortedData are added to the training through the training scheduler; S4. Input the sorted training sample SortedData into the teacher model and the student model to obtain the probability distribution of their outputs, and decouple the probability distribution of the teacher model output into the target class label probability and the non-target class label probability; S5. Increase the weight of the non-target class label probability, reduce the weight of the target class label probability, and then calculate the loss of the teacher model and the student model; S6. Repeat the training until the student model converges.

2. The TSK fuzzy classifier integrating decoupled knowledge distillation and curriculum learning according to claim 1, characterized in that: The student model uses a first-order TSK fuzzy classifier, the difficulty measurer and the teacher model use the same third-order TSK fuzzy classifier, and the third-order TSK fuzzy classifier used by the teacher model has the same architecture as the first-order TSK fuzzy classifier used by the student model.

3. The TSK fuzzy classifier integrating decoupled knowledge distillation and curriculum learning according to claim 1, characterized in that: The training sample Data includes original training data X={x i ,x i =[x1,x2,...,x n ] T ,i=1,2,...,m} and label Y={y i ,y i ∈{1,2,...,c},i=1,2,...,m}, where n, m and c are the number of sample features, the number of samples and the number of categories respectively. Taking X as input, the cross entropy loss L={L1,L2,...,L m }, the smaller the cross entropy loss, the easier the sample training is. The sorted training data is obtained according to the calculated cross entropy loss. and tags 4. The TSK fuzzy classifier integrating decoupled knowledge distillation and curriculum learning according to claim 1, characterized in that: After the training scheduler obtains the sorted training samples Sorted Data, it arranges the student model for training and defines the training data X train =φ and label Y train =φ, the sorted training data is split into and First, X 1 and Y 1 Add new training data X respectively train and Y train Then start training the model, j is the parameter of the number of training rounds E, as j increases, the training data X j and Y j Join X separately train and Y train and participate in the training. As j increases, difficult samples are continuously added to the training set until X E and Y E Add to X train and Y train In ,the training method is knowledge distillation, using the same third-order TSK fuzzy classifier as the difficulty measure and the teacher model as the student model using the first-order TSK fuzzy classifier for knowledge distillation.

5. The TSK fuzzy classifier integrating decoupled knowledge distillation and curriculum learning according to claim 1, characterized in that: The method for decoupling the probability distribution output by the teacher model includes defining q=[p t ,p \t ]∈ 1×2 is the target class probability (p t ) and non-target class probability (p \t ), the calculation formula is as follows: Define the probability distribution between non-target classes (excluding the target class t probability) as in The calculation formula is as follows: According to formulas 4.2 and 4.3, the KL divergence loss is expressed as follows: where p R and p U are the probability distributions predicted by the teacher model and the student model respectively. According to formulas 4.2, 4.3 and 4.4, L KL It can be expressed as: According to formula 4.6, the KL divergence loss is decomposed into two coupled parts, namely q R With q U The KL divergence loss between the two parts and the KL divergence loss between the non-target class probabilities of the student and teacher are decoupled and two hyperparameters β and γ are introduced as the weights of the two parts of the loss to increase the loss weight between the non-target class probabilities. The formula is as follows: Last L DKD Plus the cross entropy loss L between the student model and the true label CE , the total loss of knowledge distillation is obtained as follows: L ALL =αL CE +(1-α)L DKD (4.8) An adaptive hyperparameter α is introduced into the total loss. At the beginning of training, the value of α is higher to increase the weight of the cross-entropy loss, allowing the model to focus on learning the labels. As the model continues to iterate, the adaptive weight α continues to decrease to increase the weight of the decoupled knowledge distillation loss, allowing the model to focus on learning the teacher model.