TSK fuzzy classifier fusing reinforcement learning and multi-view knowledge distillation

By integrating reinforcement learning and multi-view knowledge distillation into the TSK fuzzy classifier and dynamically allocating view weights, the problems of low efficiency and insufficient robustness of the TSK fuzzy classifier in high-dimensional data processing are solved, and more efficient knowledge distillation and better model performance are achieved.

CN120671729APending Publication Date: 2025-09-19HUZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510626553.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing TSK fuzzy classifiers are prone to problems such as dimensionality curse, rule explosion, or poor model accuracy when processing high-dimensional data, and the integrity and distribution uniformity of the input data directly affect its accuracy performance.

Method used

A method integrating reinforcement learning and multi-view knowledge distillation is adopted to assign appropriate weights to different perspectives through reinforcement learning. This method is applied to the knowledge distillation process of TSK fuzzy classifier, solving the problem of low efficiency of traditional distillation.

Benefits of technology

The robustness and generalization ability of the TSK fuzzy classifier are improved, overfitting problems are avoided, and more efficient knowledge distillation effects are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671729A_ABST
    Figure CN120671729A_ABST
Patent Text Reader

Abstract

The invention discloses a TSK fuzzy classifier fusing reinforcement learning and multi-view knowledge distillation, appropriate weights are distributed for different views through reinforcement learning, the weights are used for distilling the TSK fuzzy classifier, the problem that traditional distillation is low in efficiency is solved, meanwhile, the robustness of the TSK fuzzy classifier is improved, and the method has the advantages of being high in practicability and the like. According to the method, the output of a teacher model is subjected to knowledge extraction of different visual angles through an integrated visual angle, then appropriate weights are autonomously distributed for the visual angles through reinforcement learning, and finally the knowledge is extracted into a TSK fuzzy classifier through knowledge distillation, so that the TSK fuzzy classifier can autonomously learn the knowledge of different visual angles of the teacher model, and better generalization ability is obtained; and over-fitting is prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge distillation, and in particular to the technical field of a TSK fuzzy classifier that integrates reinforcement learning and multi-view knowledge distillation. Background Art

[0002] Knowledge distillation, as a method for model compression and accuracy improvement, uses the probability distribution of the teacher model's final output layer or the deep features of its intermediate layers to guide the training of the student model, allowing the student model to mimic the features extracted and the probability distribution of the teacher model's output to improve its own accuracy. In the current era of large models, knowledge distillation has attracted increasing attention in the field of deep learning. The TSK fuzzy classifier is an efficient nonlinear approximator that can map input to output through generated fuzzy rules. However, there are still some unresolved issues. For example, the integrity or distribution uniformity of the input data directly affects the accuracy of the TSK fuzzy classifier. When processing high-dimensional data, the TSK fuzzy classifier is prone to the curse of dimensionality, rule explosion, and poor model accuracy. In recent years, many distillation methods have been developed by improving the knowledge transfer method. For example, the multi-view feature extraction method aims to add several MLPs as multiple views after the student model to extract the feature distribution of the student model. This method can keep the student model and the teacher model of the same dimension, avoid overmatching the feature distributions of the student and teacher models, and improve robustness. However, some views promote distillation while others have a negative impact on distillation. How to assign appropriate weights to these views is an urgent problem to be solved. Summary of the Invention

[0003] The purpose of the present invention is to solve the problems in the prior art and propose a TSK fuzzy classifier that integrates reinforcement learning and multi-perspective knowledge distillation. Reinforcement learning is used to assign appropriate weights to different perspectives and use them to distill the TSK fuzzy classifier, which solves the problem of low efficiency of traditional distillation while improving the robustness of the TSK fuzzy classifier.

[0004] To achieve the above objectives, the present invention proposes a TSK fuzzy classifier that integrates reinforcement learning and multi-view knowledge distillation, which includes the following steps: S1, the input sample is ,in m and n The number of samples and the number of features are respectively, and the corresponding output is obtained through the teacher model and TSK fuzzy classifier z R and z U ; S2. Apply reinforcement learning to the knowledge distillation framework. In each round of training, the agentAgent The agent then interacts with the environment and receives a vector representation of the training samples, soft labels of the viewpoints, and the classification loss of the viewpoints. Agent Assign corresponding weights to each perspective through the policy network and participate in knowledge distillation; S3. Distribute the teacher model output z R After inputting multiple perspectives and obtaining the output distribution of each perspective, we assign appropriate weights to each perspective through reinforcement learning and then fuse them to obtain z V , obtained by formula 2.12 , used for training the TSK fuzzy classifier, the formula is as follows: (2.12); in Indicates the i The probability of the class, the temperature is introduced parameter; S4, aggregate the soft labels output by the viewpoints, and use the aggregated features to perform knowledge distillation training on the TSK fuzzy classifier; S5. After completing a round of knowledge distillation training, the performance of the TSK fuzzy classifier is used as a reward Reward To update Agent The policy network parameters are adjusted until the network converges.

[0005] Preferably, the state of the knowledge distillation Design a vector representation of the relevant information for each view, where J is the viewing angle quantity, and the vector consists of the following two parts: a1. Logit obtained after the features extracted by the teacher model on the training sample are input into the perspective , indicating the viewing angle For input samples The classification results on C is the number of data set categories, and the perspective The output Logit is input into the softmax function to calculate the predicted probability distribution. The formula is as follows: (3.1); in Indicates the perspective in the category The predicted probability, is a trainable weight matrix, representing the viewing angle Embeddings in the last layer; a2. Perspective In the sample Cross entropy loss on , that is, the loss of the Logit output of the perspective and the true label, the formula is as follows: (3.2); The above two parts a1 and a2 are connected to obtain the viewing angle Status , the states of all perspectives constitute the sample The state vector S on the agent is provided with the state vector S Agent , through the agent Agent Make decisions.

[0006] As a preference, there are multiple perspectives, and each perspective is assigned a corresponding agent. Agent , the agent Agent After receiving the state vector S from the environment, each viewpoint is selected from several possible actions, where the action represents how much weight is assigned to each viewpoint. V j The weight distribution is designed to be , the actions of all perspectives are expressed as , where the action is selected by random sampling and DQN network To complete, use The strategy is as follows: (3.3); The DQN network selects the appropriate action A for all perspectives from the action state distribution according to the input state vector S, calculates the probability distribution of the weighted output of all perspectives through action A, trains the TSK fuzzy classifier using the obtained probability distribution, and calculates the next state according to the obtained action A. , used to update the DQN network, the loss function for training the DQN network is as follows: (3.4); in For rewards reward , is the attenuation factor, which ranges from [0,1]. The attenuation factor measures the importance the network places on future rewards and determines how the network weighs immediate rewards and future rewards when making decisions. is the target network parameter, and at a certain update step l Later from the DQN network Parameters The loss function is updated and the mean square error (MSE) is used.

[0007] Preferably, the reward Reward Including TSK fuzzy classifier and cross entropy loss of the true value Negative value as reward and and KL divergence loss The negative value of is used as the reward function, as follows: (3.5); The reward Reward Calculated after a batch of training is completed.

[0008] Preferably, there are multiple perspectives, each of which uses a two-layer multi-layer perceptron (MLP) composed of a linear layer and a Relu activation function, and uses true value labels to provide an accurate distribution for each sample, providing accurate supervision information for the TSK fuzzy classifier. The perspective only participates in the training process and does not participate in prediction.

[0009] Preferably, the mapping of the feature perspectives extracted by the teacher model obtains different feature distributions, increasing the weight of the feature distribution that is beneficial to the training of the TSK fuzzy classifier and reducing the weight of the feature distribution that is beneficial to the training of the TSK fuzzy classifier.

[0010] Beneficial effects of the present invention: The present invention adds a set of perspectives after the teacher model to match the dimensions of the student model and the teacher model. At the same time, these perspectives can extract different features of the teacher model, provide more information for knowledge distillation, and thus improve the distillation effect. The method is simple and easy to implement. Reinforcement learning is added to dynamically assign appropriate weights to these perspectives, so that the TSK fuzzy classifier can selectively learn the knowledge of the teacher model from the features extracted from multiple perspectives and obtain better performance. At the same time, RMVD-TSK redefines the environment and reward settings in reinforcement learning, so that reinforcement learning is better suitable for the multi-perspective knowledge distillation framework, which can reveal more hidden knowledge of the samples and better describe the feature distribution of the samples, provide more information for knowledge distillation, improve the performance of the TSK fuzzy classifier while avoiding the overfitting problem, and extract knowledge from different perspectives of the output of the teacher model through integrated perspectives, and then autonomously assign appropriate weights to the perspectives through reinforcement learning. Finally, the knowledge is extracted into the TSK fuzzy classifier through knowledge distillation, which allows the TSK fuzzy classifier to autonomously learn the knowledge of different perspectives of the teacher model, obtain better generalization ability, and prevent overfitting.

[0011] The features and advantages of the present invention will be described in detail through embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is a diagram of the overall model framework of a TSK fuzzy classifier that integrates reinforcement learning and multi-view knowledge distillation in the present invention; Figure 2 This is a multi-perspective knowledge distillation framework diagram of a TSK fuzzy classifier that integrates reinforcement learning and multi-perspective knowledge distillation in the present invention; Figure 3 This is a diagram showing the effect of different distillation temperatures on model performance on PHO in a parameter sensitivity experiment of a TSK fuzzy classifier that integrates reinforcement learning and multi-view knowledge distillation in the present invention; Figure 4 This is a diagram showing the impact of different distillation temperatures on QSA model performance in a parameter sensitivity experiment of a TSK fuzzy classifier that integrates reinforcement learning and multi-view knowledge distillation in the present invention; Figure 5 This is a convergence analysis diagram on the PHO dataset in a parameter sensitivity experiment of a TSK fuzzy classifier that integrates reinforcement learning and multi-view knowledge distillation in the present invention; Figure 6 This is a convergence analysis diagram on the QSA dataset in a parameter sensitivity experiment of a TSK fuzzy classifier that integrates reinforcement learning and multi-view knowledge distillation in the present invention. DETAILED DESCRIPTION

[0013] The overall framework of the TSK fuzzy classifier that integrates reinforcement learning and multi-view knowledge distillation is as follows: Figure 1 As shown, reinforcement learning is applied to the knowledge distillation framework. In each round of training, the agent Agent It interacts with the environment and receives training samples, soft labels of viewpoints and vector representations of classification losses of viewpoints. Agent The policy network assigns corresponding weights to each perspective and participates in knowledge distillation. Then, the soft labels output by the perspectives are aggregated, and the aggregated features are used to perform knowledge distillation on the TSK fuzzy classifier. After completing a round of training, the performance of the TSK fuzzy classifier is used as a reward. Reward To update Agent The policy network parameters are adjusted until the network converges. Because the parameters of this set of perspectives are randomly initialized, the output distribution of the student model from different perspectives can be extracted, which avoids overfitting of the output distribution of the student and teacher models and improves the robustness of the student model. The performance of the teacher model is usually better than that of the student model. The better the model performance of the teacher, the more difficult it is for the student model to imitate the probability distribution of the teacher model. Then, can adding several perspectives after the teacher model reduce the learning difficulty of the student model, while better describing the characteristics of the teacher model and providing more information for the student model?

[0014] Based on the above ideas, the present invention designs a new multi-perspective distillation method, such as Figure 2 As shown, the input sample is ,in m and n The number of samples and the number of features are respectively, and the corresponding output is obtained through the teacher model and TSK fuzzy classifier z R andz U , and then the teacher model output distribution z R Input multiple perspectives, get the output distribution of each perspective, take the average and fuse them to get z V , and finally we get , used for training the TSK fuzzy classifier, the formula is as follows: (2.12); in Indicates the i The probability of the class, the temperature is introduced parameter; The multi-perspective knowledge distillation of the present invention can reveal more hidden knowledge of the samples and better describe the characteristic distribution of the samples, providing more information for knowledge distillation, improving the performance of the TSK fuzzy classifier while avoiding the overfitting problem. Each perspective adopts a two-layer multi-layer perceptron (MLP) composed of a linear layer and a Relu activation function, and uses the true value label to provide a more accurate distribution for each sample, thereby providing better supervision information for the TSK fuzzy classifier. These added perspectives only participate in the training process and do not participate in the prediction, so they do not affect the overall efficiency of the model.

[0015] The effect of directly adding a set of perspectives to the teacher model is not ideal. The possible reason is that the amount of original information from the teacher is large, and some features are easily destroyed when passing through this set of perspectives. Assuming that the mapping of the features extracted by the teacher model to this set of perspectives will obtain different feature distributions, some of these feature distributions are conducive to the training of the TSK fuzzy classifier, while others are not conducive to training. In order to solve this problem, appropriate weights are added to the feature distributions obtained from these perspectives, increasing the weight of the feature distribution that is conducive to training and reducing the weight of the feature distribution that is not conducive to training. In order to adaptively adjust the weight of the perspective, reinforcement learning is introduced to assign appropriate weights to each perspective.

[0016] The method of the present invention is specifically explained from the three basic elements of reinforcement learning: (1) Status: In reinforcement learning, the design of the state is very important for the agent Agent The decision-making of the state of multi-view knowledge distillation based on reinforcement learning has a crucial impact. Design a vector representation of the relevant information for each view, where J is the number of views, this state vector consists of two parts, so that the agent Agent Make a smart choice: The first part is the Logit obtained after the features extracted by the teacher model on the training sample are input into the perspective , indicating the viewing angle For input samples The classification results on C is the number of data set categories, and the perspective The output Logit is input into the softmax function to calculate the predicted probability distribution. The formula is as follows: (3.1); in Indicates the perspective in the category The predicted probability, is a trainable weight matrix, representing the viewing angle Embeddings in the last layer; The second part is perspective In the sample Cross entropy loss on , that is, the loss of the Logit output of the perspective and the true label, the formula is as follows: (3.2); The two parts are connected to get the perspective Status , the states of all perspectives constitute the sample The state vector S on the Agent , allowing them to make better decisions.

[0017] (2) Action: For each perspective in the algorithm V j Each is assigned a corresponding agent Agent , agent Agent After receiving the state S from the environment, each viewpoint is selected from several possible actions. The action here is how much weight is assigned to each viewpoint. V j The value distribution of weights is designed to be , the actions of all perspectives are expressed as , where the action is selected by random sampling and DQN network To complete, use The strategy is as follows: (3.3); The DQN network selects the appropriate action A for all perspectives from the action state distribution according to the input state S. After obtaining the action A, the probability distribution of the weighted output of all perspectives can be calculated. The obtained probability distribution is then used to train the TSK fuzzy classifier. The next state is then calculated based on the obtained action A. , used to update the DQN network, the loss function for training the DQN network is as follows: (3.4); in For rewards reward , is the decay factor, which ranges from [0,1]. The decay factor is used to measure the importance the network places on future rewards. Specifically, it determines how the network weighs immediate rewards and future rewards when making decisions. is the target network parameter, and at a certain update step l Later from the DQN network Parameters The loss function is updated and the mean square error (MSE) is used.

[0018] (3) Rewards: The setting of the reward function is closely related to the performance of the TSK fuzzy classifier, because reward It will directly affect the update of DQN network parameters. Appropriate reward settings can fully reflect the quality of TSK fuzzy classifier training and also help the DQN network make correct decisions.

[0019] In order to fully reflect the quality of action selection, two reward functions are designed in the algorithm for feedback. One is the cross entropy loss between the TSK fuzzy classifier and the true value. Another way is to use a negative value as a reward and KL divergence loss The negative value of is used as the reward function, as follows: (3.5); The reward is not obtained immediately after each iteration, but is calculated after a batch of training is completed.

[0020] Algorithm Flow The following is the process of the RMVD-TSK algorithm:

[0021]

[0022]

[0023]

[0024] Experiment and analysis This paper first introduces the dataset and experimental settings, then demonstrates the performance of RMVD-TSK through comparative experiments, and finally studies the effectiveness of RMVD-TSK and the impact of different strategies on its performance through parameter sensitivity experiments.

[0025] Dataset

[0026] Table 3.1 UCI dataset

[0027]

[0028] Table 3.2 Medical diagnosis dataset

[0029]

[0030] In this experiment, we selected several representative and widely used datasets from the UCI dataset to evaluate RMVD-TSK, as shown in Table 3.1. In addition, we selected two medical diagnosis-related datasets to further verify the model performance, namely the stroke diagnosis dataset Stroke and the diabetes diagnosis dataset Diabetes, as shown in Table 3.2.

[0031] Experimental setup In this section, RMVD-TSK is experimented on seven UCI datasets and two medical diagnosis datasets to demonstrate the effectiveness of its algorithm. The experimental environment is: GPU NVIDIA Geforce RTX4090 24GB and CPU Inteli912900KF, 64GB RAM, programming environment: Python 3.8.17 with troch 2.1.2, a first-order TSK fuzzy classifier with 16 rules is selected as the student model Student, and a third-order TSK fuzzy classifier with 1 rule is selected as the teacher model Teacher. The subsequent parts are optimized using gradient descent, the reward decay factor  in reinforcement learning is set to 0.99, the initial and final values ​​of the random action probability are set to 1.0 and 0.01 respectively, the DQN network learning rate is set to 1e-3, the target network is updated every 10 rounds of training, the number of views in the multi-view distillation method is set to 2 by default, the view defaults to a 2-layer MLP structure, the distillation temperature  is set to 4.0 by default, and the maximum number of iterations during model training is 200 The initial value of the learning rate is 1e-1 and it decreases every 30 rounds. The AdamW optimizer is used to train the model. Other parameters are the default values. To ensure the accuracy of the results, each experiment uses a five-fold cross validation as the final result.

[0032] The comparison algorithms selected are the traditional knowledge distillation (KD), decoupled knowledge distillation (DKD), and relational knowledge distillation (RKD), which are representative in logits distillation. The representative high-order TSK distillation low-order TSK (HTSK-LLM-DKD) of the classic fuzzy classifier is also selected.

[0033] Accuracy and F1-score are selected as the evaluation indicators of the model in the experiment. The specific formulas are as follows: (3.6); (3.7); (3.8); (3.9); Where TP represents the number of positive samples correctly predicted by the model, FP represents the number of positive samples incorrectly predicted by the model, TN represents the number of negative samples correctly predicted by the model, and FN represents the number of negative samples incorrectly predicted by the model.

[0034] Comparative test As shown in Table 3.3 and Table 3.4, the distillation methods are obtained on 7 UCI datasets. Accuracy and F1- score The data shows that RMVD-TSK performs better than other knowledge distillation methods; specifically, RMVD-TSK achieves the best performance in 4 of the 7 benchmark datasets. Accuracy And 6 of them reached the best F1- score; Although RMVD-TSK has a good performance on Titanic and Wine datasets, HTSK-LLM-DKD has better performance on these two datasets. Accuracy And has better performance on the Wine dataset F1-score .

[0035] Table 3.3 Results of various methods on the UCI dataset Accuracy contrast

[0036]

[0037] Table 3.4 F1-score comparison of various methods on the UCI dataset

[0038]

[0039] Table 3.5 Accuracy comparison of various methods on medical diagnosis datasets

[0040]

[0041] Table 3.6 F1-score comparison of various methods on medical diagnosis datasets

[0042]

[0043] As shown in Table 3.5 and Table 3.6, the distillation methods are obtained on the Stroke and Diabetes datasets. Accuracy and F1-score ,The data shows that RMVD-TSK has the same accuracy and F1-score All of them are the best, followed by HTSK-LLM-DKD. After distillation, they all have a significant improvement on the basis of the student model, and even exceed the teacher model. Therefore, the present invention believes that the RMVD-TSK method has a better effect than other methods in extracting teacher model knowledge.

[0044] In summary, it can be concluded that RMVD-TSK has the best or suboptimal performance on all 9 datasets. The present invention believes that RMVD-TSK combines the advantages of reinforcement learning and multi-perspective knowledge distillation. The multi-perspective distillation framework can improve the generalization ability of the student model and avoid overfitting. In addition, RMVD-TSK has a stable improvement on the performance of the student model on all datasets. Therefore, using RMVD-TSK for training can obtain a more stable and more accurate TSK fuzzy classifier.

[0045] Table 3.7 Ablation experiment

[0046]

[0047] In order to prove the necessity of integrating multi-view knowledge distillation with reinforcement learning, ablation experiments were conducted on two medical diagnosis datasets and QSA and PHO datasets, as shown in Table 3.7; where w / o means that a module was removed from the RMVD-TSK model, RL means the reinforcement learning module, and MVKD means multi-view knowledge distillation. According to the data in the table, it can be seen that the accuracy of the model without reinforcement learning is lower than that of the baseline model. The reason is probably because the multi-view disperses or destroys the features of the teacher model. Without the addition of multi-view knowledge distillation, the improvement of model accuracy is not obvious. In summary, combining multi-view knowledge distillation with reinforcement learning can effectively improve the performance of the TSK fuzzy classifier.

[0048] Parameter sensitivity experiments This paper will conduct an in-depth study on the hyperparameters in the RMVD-TSK model and explore the impact of hyperparameter settings on the RMVD-TSK model.

[0049] (1) Influence of the number of viewing angles The number of perspectives indicates that the output of the teacher model is mapped to several feature distributions, which can provide knowledge from different perspectives to the student model, better describe the feature distribution of the teacher model, and have a great impact on the performance of RMVD-TSK. Table 3.8 shows the performance of the model on the Phoneme and QSAR datasets. Accuracy , where the perspectives are all composed of 2-layer MLPs, and the number of perspectives is set in the interval [0,3]. 0 represents the model of traditional knowledge distillation. When the number of perspectives is 1, the distillation improvement effect is negligible and may even reduce the accuracy. It may be that the single perspective destroys the characteristics of the teacher model. When the number of perspectives is 2, the distillation effect is significantly improved. It can be seen that two perspectives can reveal more information and better describe the distribution of teacher model features. When the number of perspectives is 3, the distillation effect is not further improved. It should be that the characteristics of the teacher model are too complex, which makes it difficult for the student model to learn.

[0050] Table 3.8 Impact of the number of viewing angles on model performance

[0051]

[0052] (2) The influence of perspective dimension The perspective dimension indicates that the perspective is composed of several layers of MLP, that is, the output of the teacher model is mapped to the corresponding feature distribution through several layers of MLP. The feature distribution directly affects the quality of the student model's learning. Table 3.9 shows the performance of the model on the Phoneme and QSAR datasets. Accuracy , where the number of perspectives is set to 2, the perspective dimension range is [0,3], and 0 represents the model as traditional knowledge distillation; from the data, it can be seen that perspectives of different dimensions have a certain improvement on the distillation effect, among which the improvement effect is most obvious when the perspective is a two-layer MLP.

[0053] Table 3.9 Impact of viewing angle dimension on model performance

[0054]

[0055] (3) Rewards reward Effect of settings In reinforcement learning, rewards reward The settings directly affect the agent Agent The choice of strategy has a great impact on the performance of the entire model. Table 3.10 shows the performance of the model on the four datasets of Phoneme, QSAR, Adult and Wine.Accuracy , two different reward functions are used, one is the cross entropy loss between TSK fuzzy classifier and the true value Another way is to use a negative value as a reward and KL divergence loss The negative value of is used as the reward function; reward Set to cross entropy loss With KL divergence loss When the value of is negative, the performance of the student model is improved the most.

[0056] Table 3.10 Impact of reward settings on model performance

[0057]

[0058] (4) Distillation temperature Impact Temperature parameters in the knowledge distillation process Responsible for label smoothing of the output distribution of the teacher model and the student model, which has a great impact on the distillation effect. Figure 3 .4 and Figure 3 .5 is the model's performance on the Phoneme and QSAR datasets. Accuracy and F1-score ,temperature The parameter value range is an integer [1,10]. When it is 1, it means that the student model directly imitates the probability distribution of the teacher model output. At this time, the performance of the student model is not significantly improved. As the temperature parameter T increases, the probability distribution of the teacher model output gradually becomes smoother and easier to learn. Therefore, the accuracy of the student model is also slowly improving. When the temperature parameter T increases, the probability distribution of the teacher model output gradually becomes smoother and easier to learn. Model for 4-hour students Accuracy and F1-score The distillation effect is the best, and the present invention believes that this is the most suitable temperature for distillation; Finally, as the temperature As the value of θ continues to increase, the accuracy of the student model will continue to decrease, because the output of the teacher model is too smooth, resulting in no room for the student model to learn, so the student model does not improve much.

[0059] (5) Convergence analysis The number of training rounds E is the number of times the model is trained on the dataset. Too large or too small a value will affect the performance of the model. Figure 3 .6 and Figure 3 .7 is the model on Phoneme and QSAR datasets Accuracy and F1-score; The other parameters of the model are set to default values. According to the data in the figure, as the number of training rounds increases, the model Accuracy andF1-score It also continues to rise. At this time, the model has not converged yet and is in an underfitting state. After the number of training rounds reaches 60, the model begins to converge slowly and achieves a higher classification performance.

[0060] This paper proposes a TSK fuzzy classification model RMVD-TSK based on reinforcement learning and multi-perspective distillation. This framework extracts knowledge from different perspectives of the teacher model's output through integrated perspectives, then autonomously assigns appropriate weights to the perspectives through reinforcement learning. Finally, the knowledge is extracted into the TSK fuzzy classifier through knowledge distillation. This allows the TSK fuzzy classifier to autonomously learn knowledge from different perspectives of the teacher model, obtain better generalization ability, and prevent overfitting. Experiments have shown that RMVD-TSK has good performance on UCI and medical diagnosis datasets.

[0061] The above embodiments are intended to illustrate the present invention, not to limit the present invention. Any solution that is a simple transformation of the present invention falls within the protection scope of the present invention.

Claims

1. A TSK fuzzy classifier that integrates reinforcement learning and multi-view knowledge distillation, characterized by: The following steps are involved: S1, the input sample is ,in m and n The number of samples and the number of features are respectively, and the corresponding output is obtained through the teacher model and TSK fuzzy classifier z R and z U ; S2. Apply reinforcement learning to the knowledge distillation framework. In each round of training, the agent Agent The agent then interacts with the environment and receives a vector representation of the training samples, soft labels of the viewpoints, and the classification loss of the viewpoints. Agent Assign corresponding weights to each perspective through the policy network and participate in knowledge distillation; S3. Distribute the teacher model output z R After inputting multiple perspectives and obtaining the output distribution of each perspective, we assign appropriate weights to each perspective through reinforcement learning and then fuse them to obtain z V , obtained by formula 2.12 , used for training the TSK fuzzy classifier, the formula is as follows: (2.12); in Indicates the i The probability of the class, the temperature is introduced parameter; S4, aggregate the soft labels output by the viewpoints, and use the aggregated features to perform knowledge distillation training on the TSK fuzzy classifier; S5. After completing a round of knowledge distillation training, the performance of the TSK fuzzy classifier is used as a reward Reward To update Agent The policy network parameters are adjusted until the network converges.

2. The TSK fuzzy classifier integrating reinforcement learning and multi-view knowledge distillation according to claim 1, characterized in that: The state of knowledge distillation Design a vector representation of the relevant information for each view, where J is the viewing angle quantity, and the vector consists of the following two parts: a1. Logit obtained after the features extracted by the teacher model on the training sample are input into the perspective , indicating the viewing angle For input samples The classification results on C is the number of data set categories, and the perspective The output Logit is input into the softmax function to calculate the predicted probability distribution. The formula is as follows: (3.1); in Indicates the perspective in the category The predicted probability, is a trainable weight matrix, representing the viewing angle Embeddings in the last layer; a2. Perspective In the sample Cross entropy loss on , that is, the loss of the Logit output of the perspective and the true label, the formula is as follows: (3.2); The above two parts a1 and a2 are connected to obtain the viewing angle Status , the states of all perspectives constitute the sample The state vector S on the agent is provided with the state vector S Agent , through the agent Agent Make decisions.

3. The TSK fuzzy classifier integrating reinforcement learning and multi-view knowledge distillation according to claim 1, characterized in that: There are multiple perspectives, and each perspective is assigned a corresponding agent. Agent , the agent Agent After receiving the state vector S from the environment, each viewpoint is selected from several possible actions, where the action represents how much weight is assigned to each viewpoint. V j The weight distribution is designed to be , the actions of all perspectives are expressed as , where the action is selected by random sampling and DQN network To complete, use The strategy is as follows: (3.3); The DQN network selects the appropriate action A for all perspectives from the action state distribution according to the input state vector S, calculates the probability distribution of the weighted output of all perspectives through action A, trains the TSK fuzzy classifier using the obtained probability distribution, and calculates the next state according to the obtained action A. , used to update the DQN network, the loss function for training the DQN network is as follows: (3.4); in For rewards reward , is the attenuation factor, which ranges from [0,1]. The attenuation factor measures the importance the network places on future rewards and determines how the network weighs immediate rewards and future rewards when making decisions. is the target network parameter, and at a certain update step l Later from the DQN network Parameters The loss function is updated and the mean square error (MSE) is used.

4. The TSK fuzzy classifier integrating reinforcement learning and multi-view knowledge distillation according to claim 1, characterized in that: The reward Reward Including TSK fuzzy classifier and cross entropy loss of the true value Negative value as reward and and KL divergence loss The negative value of is used as the reward function, as follows: (3.5); The reward Reward Calculated after a batch of training is completed.

5. The TSK fuzzy classifier integrating reinforcement learning and multi-view knowledge distillation according to claim 1, characterized in that: There are multiple perspectives, each of which uses a two-layer multi-layer perceptron (MLP) consisting of a linear layer and a Relu activation function. The true value label is used to provide an accurate distribution for each sample, providing accurate supervision information for the TSK fuzzy classifier. The perspective only participates in the training process and does not participate in prediction.

6. The TSK fuzzy classifier integrating reinforcement learning and multi-view knowledge distillation according to claim 1, characterized in that: The mapping of the feature perspectives extracted by the teacher model obtains different feature distributions, increasing the weight of the feature distribution that is beneficial to the training of the TSK fuzzy classifier and reducing the weight of the feature distribution that is beneficial to the training of the TSK fuzzy classifier.