Misleading knowledge diagnosis method based on cross attention mechanism under federated distillation framework
By combining the cross attention mechanism and the Marshall distance scoring function to detect outlier proxy data under the federal distillation framework, the problem of misleading knowledge dissemination is solved, the accuracy of the model and communication efficiency are improved, and risks such as misdiagnosis are avoided.
Patent Information
- Application Number
- CN202510270045.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-11
AI Technical Summary
Under the federal distillation framework, misleading knowledge dissemination of proxy data leads to degradation of model performance, especially in the medical field, which may lead to serious consequences such as misdiagnosis, and existing research lacks effective solutions.
The cross attention mechanism and the Marshall distance scoring function are used to detect outlier agent data. After extracting the agent data and private data features through the federal distillation framework, the cross attention module is used to calculate the attention score, and the outlier data is detected through the Marshall distance scoring function to eliminate misleading knowledge.
It effectively reduces the dissemination of misleading knowledge, improves model performance, enhances the communication efficiency and model accuracy of federal distillation, and avoids problems such as misdiagnosis.
Smart Images

Figure CN120297330A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a misleading knowledge diagnosis method based on cross - attention mechanism under a federated distillation framework, belonging to the technical fields of federated learning, knowledge distillation, and cross - attention. Background Art
[0002] In recent years, the rapid development of deep learning (DL) has been widely applied in various fields, including computer vision, natural language processing, recommendation systems, medical diagnosis, and autonomous driving. Deep learning relies on large amounts of data sets and powerful computing capabilities. Therefore, centralized training has become the main method of deep learning. However, in different fields, data may be distributed in different places, presenting the problem of "data silos". For example, in the medical field, chest X - ray images for pneumonia detection may be stored in different hospitals; in the industrial field, data for mechanical fault diagnosis may be stored in different factories. It is very difficult to collect data stored in different places for centralized training. At the same time, data may contain a large amount of sensitive personal information. Directly collecting data from different places for centralized training may lead to privacy leakage. Countries have implemented regulations such as the General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA) to restrict data collection and storage. Therefore, there are great difficulties for traditional centralized training.
[0003] Federated learning (FL) is a distributed machine learning proposed by the Google team. Federated learning can effectively utilize distributed data without sending the data to the cloud for centralized training, protecting data privacy. Generally, in each round of iteration of federated learning, the central server sends the current global model parameters to a selected group of clients (participants). The clients use these parameters to train the model on their respective private data and send the calculated gradient information back to the central server. The central server then aggregates all the gradients uploaded by the clients and uses a specific method to integrate these gradients. Subsequently, the integrated gradients are used to update the parameters of the global model. The updated model parameters are distributed to the clients for the next round of training. This process is repeated continuously until the model reaches a convergence state. The most commonly used FL algorithm is the parameter averaging - based scheme (such as the FedAvg algorithm, which performs weighted averaging according to the amount of local private data held by the clients). However, the traditional federated averaging algorithm has the following limitations: (i) clients must implement the same model architecture; (ii) transmitting model parameters and model gradients means high communication costs, which increase with the number of model parameters; (iii) in the case of non - independent and identically - distributed data (NIID), the parameter averaging aggregation scheme performs poorly due to client model drift. In summary, federated learning has problems such as model heterogeneity, large communication overhead, and data heterogeneity.
[0004] Regarding the above problems, Knowledge Distillation (KD) is a very effective technique proposed by Geoffrey Hinton et al. Knowledge Distillation draws on the concept of "knowledge transfer" in the field of education, transferring the knowledge in a large and complex model (i.e., the teacher model) to a small and simple-structured model (i.e., the student model). This method can not only reduce the model's requirements for computing resources and storage space but also simplify the model deployment process, making it particularly suitable for use in resource-constrained situations. With the continuous increase in the number of parameters of large models, knowledge distillation is undoubtedly an effective technique for reducing the number of its parameters. Therefore, recently, it has been advocated to combine KD with FL to form Federated Knowledge Distillation (FKD) to solve the above problems. Compared with typical federated learning, the unique feature of FKD is that it conducts global knowledge integration by uploading the average scores (logits) calculated by each client model for each category, rather than uploading the complete local model or gradients. Specifically, compared with traditional FL, FKD has the following advantages: First, since only the logit predictions of each client model are uploaded, compared with traditional FL that uploads the entire model, FKD significantly reduces the communication overhead. Second, FKD does not require the local models to have the same structure, which allows different edge devices to flexibly design personalized model structures according to their own resources and application requirements. Finally, by training multiple models to provide customized services for different data distributions, the problem of data heterogeneity is solved.
[0005] Based on the research of federated knowledge distillation, existing research mainly conducts in two directions, one is the research with proxy datasets, and the other is the research without proxy datasets. However, in the research with proxy datasets, for each client, the proxy samples may not have the same distribution as the local training samples (i.e., they are outlier samples). At this time, the predictions obtained from the proxy data in the locally trained model will be misleading. Thus, misleading knowledge is transmitted, leading to a decline in model performance. The generation of misleading results has a huge impact in various fields, especially in the medical field. For example, if a patient in the early stage of cancer is misdiagnosed as a normal person, this may cause the condition to be delayed until the advanced stage, bringing a huge impact to the patient, and even endangering their life. Therefore, the transmission of misleading knowledge will have a huge impact. However, existing research lacks a federated learning method to solve the above problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a misleading knowledge diagnosis method based on cross-attention mechanism under the federated knowledge distillation framework, aiming to solve the technical problem of the spread of misleading knowledge brought by proxy data under the federated knowledge distillation framework.
[0007] To achieve the above object, the present invention provides a misleading knowledge diagnosis method based on cross-attention mechanism under the federated distillation framework, and the method includes the following steps:
[0008] Step1: Under the federated distillation framework, put the proxy data and the private data of each client into a pre-trained Vision Transform (VIT) model for feature extraction respectively, and use the respectively extracted features as the two inputs of the Cross Attention Transform (CAT) module;
[0009] Step2: Use the output calculated by the Cross Attention Transform (CAT) module as the input of the Mahalanobis Distance (MD) scoring function module;
[0010] Step3: Compare the OOD score output after calculation by the Mahalanobis Distance (MD) scoring function module with a set threshold to detect outlier proxy data;
[0011] Step4: Conduct misleading knowledge diagnosis, remove the outlier proxy data, and use the non-outlier proxy data to transfer knowledge.
[0012] The specific content of Step1 includes:
[0013] Under the federated distillation framework, put the image proxy data into a pre-trained Vision Transform model for feature extraction. The output dimension of the extracted feature vector is (N, Tokens, D), representing the feature vector of the proxy data, where N represents the batch size of the input proxy data, Tokens represents the number of proxy tokens including the classification token CLS, and D represents the feature vector dimension. Then use the feature vector of each proxy data as the Q value queried by the cross attention module for input;
[0014] Put the private image data of each client into a pre-trained Vision Transform model for feature extraction. The output dimension of the extracted feature vector is (M, Tokens, D), representing the feature vector of the private data of each client, where M represents the batch size of the input private data. Then calculate the average to obtain the token proxy of the private data of each client, and use the feature vector of the token proxy as the K value and V value of the cross attention module for input.
[0015] The specific content of the federated distillation framework includes:
[0016] The federated distillation framework includes a server and multiple clients, and a proxy dataset and a private dataset are deployed on the clients.
[0017] The private image data of each client is specifically as follows:
[0018] c = {c1, c2, …, c i , …, c k} (1)
[0019] In formula (1), c represents the set composed of all clients, c i represents the i-th client token proxy, and k represents that there are k clients.
[0020] Step2 specifically includes:
[0021] Perform a scaled dot product multi-head cross-attention calculation on the K value and V value from each client's private data and the Q value from each proxy data. The formula is as follows:
[0022]
[0023] MultiHeadCA(f, c i ) = Concat(CA1(f, c i ), …, CA m (f, c i ))W linear (3)
[0024] In formula (2), CA(f, c i ) represents the output obtained through scaled dot product cross-attention calculation. Softmax represents normalization, mapping the result to a probability value between 0 and 1. represents the Q value of the m-th attention head. represents the K value of the m-th attention head. T represents the transpose matrix. represents the V value of the m-th attention head. represents the dimension;
[0025] In formula (3), MultiHeadCA(f, c i ) represents the output of all cross-attention modules. W linear is the weight matrix of the linear layer. Concat means concatenating multiple tensors (CA1(f, c i ), …, CA m (f, c i )) along a certain dimension.
[0026] Step3 specifically includes:
[0027] Use the classification token T in the feature vector (i, Tokens, D) of the i-th proxy data output by the cross-attention module CLS as an input to the Mahalanobis distance;
[0028] Calculate the mean and covariance of all Tokens of each proxy data as the other two inputs of the Mahalanobis distance scoring function module;
[0029] Calculate the OOD score of each proxy data relative to each client token proxy based on the three inputs, compare the OOD score with a set threshold, and detect outlier proxy data. Among them, the OOD score of each client token proxy is expressed as:
[0030] s = {s1, s2, …, s i , …, s k} (4)
[0031] In formula (4), s represents the set composed of the OOD scores of all client token proxies, and s i represents the OOD score of the i-th client token proxy.
[0032] The specific Mahalanobis distance is as follows:
[0033] The Mahalanobis distance is a measure used to measure the distance between two sample points. It takes into account the covariance structure of the data and can effectively handle the possible correlations and scale differences between different dimensions of the data. Compared with other distance metrics, the Mahalanobis distance is more robust and accurate when dealing with data with complex distributions. The calculation formula of the Mahalanobis distance is as follows:
[0034]
[0035] In formula (5), d M (x) represents the result obtained by calculating the Mahalanobis distance, x represents the sample, μ represents the mean, T represents the transpose, and ∑ -1 represents the inverse of the covariance matrix.
[0036] The specific process of calculating the OOD score of each proxy data relative to each client token proxy based on the three inputs, comparing the OOD score with a set threshold, and detecting outlier proxy data includes:
[0037] Compare the OOD score of each client's token proxy with a given threshold. If the OOD scores of one or more token proxies in each client are less than the given threshold, then this sample is a non-outlier sample;
[0038] Compare the OOD score of each client's token proxy with a given threshold. If the OOD scores of all token proxies in each client are greater than the given threshold, then this sample is an outlier sample.
[0039] The beneficial effects of the present invention are as follows: A misleading knowledge diagnosis method based on cross-attention mechanism under the framework of federated distillation proposed by the present invention combines the idea of cross-attention under the framework of federated distillation, effectively reducing the spread of misleading knowledge while improving communication efficiency and enhancing the performance of the model. At the same time, the present invention also provides new ideas for the further combination and innovation of federated distillation and cross-attention mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is the principle framework diagram of the present invention;
[0041] Figure 2 is the flowchart of the cross-attention mechanism for detecting outlier data in Embodiment 1 of the present invention;
[0042] Figure 3 is the detailed framework diagram of the cross-attention mechanism for detecting outlier data in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0043] In order to more clearly illustrate the uses, solutions and their advantages of the present technology, the following will be explained in detail with reference to the drawings and examples. It should be clear that the examples provided here are only for illustrating the present technology and are not a limitation of the present technology.
[0044] The illustrations provided in the following examples and the setting of specific parameter values in the model are mainly for explaining the basic concept of the present invention and conducting simulation verification on the present invention. In a specific application environment, appropriate adjustments can be made according to the actual scenario and requirements.
[0045] Embodiment 1: A misleading knowledge diagnosis method based on cross-attention mechanism under the framework of federated distillation, using the cross-attention mechanism to combine with the Mahalanobis distance scoring function to detect outlier proxy data. As shown in Figure 1 , Figure 2 and Figure 3 shown, the specific steps are as follows:
[0046] Step1: Under the framework of federated distillation, put the proxy data and the private data of each client into a pre-trained vision Transformer model for feature extraction respectively, and use the respectively extracted features as the two inputs of the cross-attention module.
[0047] Specifically, the present invention collects 2,000 lung X-ray images from different public data sources as proxy data, and sequentially inputs the proxy data of the 2,000 lung X-ray images into a pre-trained Vision Transformer (VIT) model. The output of the VIT model has a feature dimension of (N, Tokens, D), where N represents the number of images input into the model each time (set to 32 in the present invention), Tokens represents the number of image patches (Patches) after image division plus a special classification (CLS) token (determined by the image size and image patch division). The input image is divided into 14×14 image patches, plus the CLS token, and the number of Tokens is 196 + 1 = 197; D is the model feature dimension, set to 768 in the present invention. Each time 32 images are input, a corresponding three-dimensional feature output (32, 197, 768) is obtained, and then the feature vector (1, 197, 768) of each proxy sample is sequentially input as the Q value for querying in the cross-attention module (CAT), as shown in the appendix Figure 3 as shown.
[0048] Furthermore, there are 4 hospitals (each hospital represents a client). Hospital One has 800 lung X-ray images, Hospital Two has 600 lung X-ray images, Hospital Three has 700 lung X-ray images, and Hospital Four has 500 lung X-ray images. Hospital One batches its 800 lung X-ray images and inputs them into the VIT model (also set to 32 for ease of understanding), obtaining a corresponding three-dimensional feature output (also (32, 197, 768)), and then taking the mean of all its samples (its dimension becomes (1, 197, 768)) as the feature vector of Hospital One, and inputting it as the K value and V value of the cross-attention module (CAT).
[0049] Hospitals Two, Three, and Four also input their respective private data into the VIT model to extract features in the same way (also in batches of 32 each time), and take the mean of the extracted feature vectors as the K value and V value of the cross-attention module (CAT) for input.
[0050] The private image data of each client is represented as:
[0051] c = {c1, c2, c3, c4} (1)
[0052] In formula (1), c represents the set composed of all clients, and c1 represents the first client token proxy. The present invention has 4 clients.
[0053] Step2: Use the output calculated by the cross-attention module as the input of the Mahalanobis distance scoring function module.
[0054] Specifically, the K and V values from the 4 client private data are subjected to a scaled dot product multi-head cross-attention calculation with the Q values from each proxy data. In the present invention, the number of attention heads is set to 8, that is:
[0055]
[0056] MultiHeadCA(f,c i ) = Concat(CA1(f,c i ),…,CA m (f,c i ))W linear (3)
[0057] In formula (2), CA(f,c i ) represents the output obtained through the scaled dot product cross-attention calculation. Softmax represents normalization, mapping the result to a probability value between 0 and 1. represents the Q value of the m-th attention head. represents the K value of the m-th attention head, and T represents the transpose matrix. represents the V value of the m-th attention head. represents the dimension.
[0058] In formula (3), MultiHeadCA(f,c i ) represents the output of all cross-attention modules. W linear is the weight matrix of the linear layer. Concat represents concatenating multiple tensors (CA1(f,c i ),…,CA m (f,c i )) along a certain dimension.
[0059] Through the above calculations, for each client, a feature vector (with dimensions (i, 197, 768), i = 1, 2, … 32) for each proxy sample will be obtained, and then it is used as the input to the Mahalanobis distance scoring function module (MD).
[0060] Step 3: Compare the OOD score output after calculation by the Mahalanobis distance scoring function module with the set threshold to detect outlier proxy data.
[0061] Specifically, the classification token T in the feature vector (i, 197, 768) (representing the feature vector of the i-th proxy data) output by the cross-attention module (CAT) CLSA token (whose feature vector becomes (1, d)) is used as an input to the Mahalanobis distance. At the same time, the mean and covariance of all tokens of each proxy data are calculated as the other two inputs to the Mahalanobis distance scoring function module (MD). Then, based on the three inputs, the OOD score of each proxy data relative to each client token proxy is calculated, and the OOD score is compared with a set threshold to detect outlier proxy data. Among them, the OOD score of each client token proxy is expressed as:
[0062] s = {s1, s2, s3, s4} (4)
[0063] In formula (4), s represents the set composed of the OOD scores of all client token proxies, s1 represents the OOD score of the first client token proxy, and there are 4 clients in the present invention.
[0064] The OOD scores of the i-th and j-th proxy data of each client are obtained as: {s 1i = 1.0, s 2i = 0.5, s 3i = 0.6, s 4u = 0.7, u = 1, 2,..., 2000, u ≠ j}, {s 1j = 0.9, s 2j = 1.0, s 3j = 1.2, s 4j = 1.3, j = 1, 2,..., 2000, j ≠ i}. The threshold set in the present invention is 0.8. According to the results, s 1i = 1.0 > 0.8, s 2i = 0.5 < 0.8, s 3i = 0.6 < 0.8, s 4i = 0.7 < 0.8. Since all three clients regard it as non-outlier data, the i-th proxy data is non-outlier data and should be retained; s 1j = 0.9 > 0.8, s 2j = 1.0 > 0.8, s 3j = 1.1 > 0.8, s 4j = 1.2 > 0.8. Since the OOD scores of the j-th proxy data of all four clients are greater than the threshold, that is, each client regards it as outlier data, the j-th proxy data is outlier data and should be excluded.
[0065] Step4: Conduct misleading knowledge diagnosis, exclude outlier proxy data, and use non-outlier proxy data to transfer knowledge.
[0066] The detected outlier proxy data is removed. Among 2,000 lung X-ray images, through detection, 100 lung X-ray images are outlier proxy data, and the remaining 1,900 lung X-ray images are non-outlier proxy data. Then, the 100 lung X-ray images are removed. The remaining 1,900 lung X-ray images are used for subsequent model training and knowledge sharing, thus avoiding misleading information that may be brought by outlier data and resulting in misdiagnosis.
[0067] Embodiment 2: A detection device for outlier proxy data under a federated distillation framework is provided. The device includes:
[0068] A Vision Transformer (VIT) module, which is used to extract the features of proxy data and private data as the input of subsequent modules;
[0069] A cross-attention module, which is used to build the similarity between proxy data and private data, and calculate the attention scores from proxy data and private data through dot product scaling as the input of subsequent modules.
[0070] A Mahalanobis distance scoring function module, which is used to calculate the out-of-distribution (OOD) score of each proxy data, compare it with a given threshold, and detect outlier proxy data.
[0071] Obviously, the foregoing embodiments are only used to illustrate the technical solutions of the present invention and are not a limitation on the application mode of the present invention. Any adjustment, replacement, or optimization carried out under the guidance of the basic principles and spirit of the present invention should be regarded as part of the protection scope of the present invention.
Claims
1. A misleading knowledge diagnosis method based on cross-attention mechanism under the federal distillation framework, characterized in that The method comprises: Step 1: Under the federated distillation framework, the proxy data and each client’s private data are put into the pre-trained visual Transform model for feature extraction, and the extracted features are used as the two inputs of the cross attention module; Step 2: The output calculated by the cross attention module is used as the input of the Mahalanobis distance scoring function module; Step 3: Compare the OOD score calculated by the Mahalanobis distance scoring function module with the set threshold to detect outlier agent data; Step 4: Perform misleading knowledge diagnosis, remove outlier agent data, and use non-outlier agent data to transfer knowledge.
2. A misleading knowledge diagnosis method based on cross-attention mechanism under a federated distillation framework according to claim 1, characterized in that The Step 1 specifically includes: In the federated distillation framework, the image proxy data is put into the pre-trained visual Transform model for feature extraction. The extracted feature vector output dimension is (N, Tokens, D), which represents the feature vector of the proxy data, where N represents the batch size of the input proxy data, Tokens represents the number of proxy tokens including the classification token CLS, and D represents the feature vector dimension. The feature vector of each proxy data is then input as the Q value of the cross-attention module query; The private image data of each client is put into the pre-trained visual Transform model for feature extraction. The output dimension of the extracted feature vector is (M, Tokens, D), which represents the feature vector of each client's private data, where M represents the batch size of the input private data. The token proxy of each client's private data is then calculated and averaged, and the feature vector of the token proxy is input as the K value and V value of the cross-attention module.
3. A misleading knowledge diagnosis method based on cross-attention mechanism under the federal distillation framework according to claim 2, characterized in that The federal distillation framework specifically includes: The federated distillation framework includes a server and multiple clients, and proxy datasets and private datasets are deployed on the clients.
4. A misleading knowledge diagnosis method based on cross-attention mechanism under the federal distillation framework according to claim 2, characterized in that, The private image data of each client is specifically: c = {c1, c2, …, c i , …, c k} (1) In formula (1), c represents the set composed of all clients, and c i represents the i-th client token proxy, and k represents that there are k clients.
5. A misleading knowledge diagnosis method based on cross-attention mechanism under the federal distillation framework according to claim 1, characterized in that The Step 2 specifically includes: The K and V values from each client's private data are scaled dot-product multi-head cross attention calculations with the Q values from each proxy data, as follows: MultiHeadCA(f, c i ) = Concat(CA1(f, c i ), …, CA m (f, c i ))W linear (3) In formula (2), CA(f, c i ) represents the output obtained by scaled dot-product cross-attention calculation. Softmax represents normalization, mapping the result to a probability value between 0 and 1. represents the Q value of the m-th attention head. represents the K value of the m-th attention head. T represents the transposed matrix. represents the V value of the m-th attention head. represents the dimension. In formula (3), MultiHeadCA(f, c i ) represents the output of all cross-attention modules, W linear is the weight matrix of the linear layer, and Concat means to concatenate multiple tensors (CA1(f, c i ), …, CA m (f, c i )) along a certain dimension.
6. A method for diagnosing misleading knowledge based on cross-attention mechanism under a federated distillation framework according to claim 1, characterized in that The Step 3 specifically includes: Take the classification token T in the feature vector (i, Tokens, D) of the i-th proxy data output by the cross-attention module CLS as an input to the Mahalanobis distance; Calculate the mean and covariance of all tokens of each agent data as the other two inputs of the Mahalanobis distance scoring function module; The OOD score of each proxy data relative to each client token proxy is calculated based on the three inputs, and the OOD score is compared with the set threshold to detect outlier proxy data, where the OOD score of each client token proxy is expressed as: s = {s1, s2, …, s i , …, s k} (4) In formula (4), s represents the set composed of the OOD scores of all client token proxies, and s i represents the OOD score of the i-th client token proxy.
7. A misleading knowledge diagnosis method based on cross-attention mechanism under the federal distillation framework according to claim 6, characterized in that, The OOD score of each proxy data relative to each client token proxy is calculated based on the three inputs, the OOD score is compared with a set threshold, and the outlier proxy data is detected, specifically including: Compare the OOD score of the token agent of each client with the given threshold. If the OOD score of one or more token agents in each client is less than the given threshold, the sample is a non-outlier sample. Compare the OOD score of the token proxy for each client with a given threshold. If the OOD score of the token proxy in each client is greater than the given threshold, then the sample is an outlier sample.
Citation Information
Cited By
Image classification method based on personalized federated distillation
CN121708359A