Method, device and equipment for defending federated backdoor attack based on hierarchical knowledge distillation and medium
By singular value decomposition and layered knowledge distillation of model parameters in federated learning, and identifying and stripping backdoor attacks, the problem of poor defense effects in the existing technology is solved, and more efficient defense effects and model robustness are achieved.
Patent Information
- Application Number
- CN202510410481.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-11
AI Technical Summary
When defending against federal backdoor attacks, the existing technology lacks a deep understanding of the nature of backdoor attacks, which makes it difficult to effectively distinguish between benign and malicious behaviors, and may mistakenly regard normal models as suspicious models, resulting in loss of classified information and degraded model performance.
By decomposing the local model parameters uploaded by the client, the cluster is divided into suspicious clients and benign clients, and the student model is improved using layered knowledge distillation technology to generate distilled student models, weakening the impact of backdoor attacks.
Improves defense against federated backdoor attacks, reduces the impact on model performance, maintains the robustness and robustness of the model while protecting data privacy.
Smart Images

Figure CN120301632A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet security technologies, and particularly to a method, device, equipment and medium for defending against federated backdoor attacks based on hierarchical knowledge distillation. Background Art
[0002] In the current field of artificial intelligence, federated learning, as a machine learning method for protecting data privacy, has received extensive attention. In federated learning, multiple participants (such as clients) train models on their local data and then aggregate local model updates through a central server. However, this distributed learning paradigm also introduces some new security challenges, such as federated backdoor attacks. Federated backdoor attacks refer to the situation where attackers embed hidden functions in the model by contaminating local training data or model updates. This attack causes the model to produce incorrect outputs when specific conditions are triggered, while functioning correctly under normal circumstances. Such attacks pose a great threat to the security and fairness of the federated learning system.
[0003] In the prior art, methods for detecting and defending against backdoor attacks mainly rely on analyzing the abnormal performance of the global model or excluding suspicious client updates. These methods include, but are not limited to: screening based on model performance, statistical detection of updates, and model pruning. However, these methods generally have the following defects: (1) Limitation: Lack of a deep understanding of the essence of backdoor attacks, and it is difficult to effectively distinguish between benign and malicious behaviors based solely on the performance of the global model; (2) Overly aggressive: Usually directly discard the client models considered to be suspicious, and may misinterpret some anomalies caused by non-independent and identically distributed data as malicious behaviors, thus losing important classification information; (3) Information loss: Completely removing suspicious client models may lead to the loss of legitimate classification knowledge, especially in a training environment with high diversity, increasing the risk of loss of model performance diversity. Summary of the Invention
[0004] Embodiments of the present invention provide a method, device, equipment and medium for defending against federated backdoor attacks based on hierarchical knowledge distillation, aiming to solve the problem of poor defense effect against federated backdoor attacks in the prior art.
[0005] In a first aspect, embodiments of the present invention provide a method for defending against federated backdoor attacks based on hierarchical knowledge distillation, which includes:
[0006] Receiving local model parameters uploaded by a client, and performing singular value decomposition on the high-level model parameters in the local model parameters to obtain a decomposition matrix;
[0007] Cluster the decomposition matrix through a clustering algorithm to divide the clients into suspicious clients and benign clients, and respectively perform average aggregation on the local model parameters uploaded by the suspicious clients and the benign clients to obtain suspicious model parameters and benign model parameters. Import the suspicious model parameters and the benign model parameters into the global model respectively to obtain a teacher model and a student model;
[0008] Improve the student model through hierarchical knowledge distillation according to the knowledge distillation dataset and the teacher model to obtain a distilled student model, and use the distilled student model as the global model to send it to the client.
[0009] In a second aspect, an embodiment of the present invention further provides a device for defending against federated backdoor attacks based on hierarchical knowledge distillation, which includes:
[0010] A receiving and decomposing unit, configured to receive local model parameters uploaded by a client, and perform singular value decomposition on the high-level model parameters in the local model parameters to obtain a decomposition matrix;
[0011] A clustering and partitioning unit, configured to cluster the decomposition matrix through a clustering algorithm to divide the clients into suspicious clients and benign clients, and respectively perform average aggregation on the local model parameters uploaded by the suspicious clients and the benign clients to obtain suspicious model parameters and benign model parameters. Import the suspicious model parameters and the benign model parameters into the global model respectively to obtain a teacher model and a student model;
[0012] An improving and sending unit, configured to improve the student model through hierarchical knowledge distillation according to the knowledge distillation dataset and the teacher model to obtain a distilled student model, and use the distilled student model as the global model to send it to the client.
[0013] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the above method is implemented.
[0014] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the above method is implemented.
[0015] An embodiment of the present invention provides a method, device, equipment, and medium for defending against federated backdoor attacks based on hierarchical knowledge distillation. Among them, the method includes: receiving local model parameters uploaded by a client, and performing singular value decomposition on the model high-level parameters in the local model parameters to obtain a decomposition matrix; clustering the decomposition matrix through a clustering algorithm to divide the client into a suspicious client and a benign client, and respectively averaging and aggregating the local model parameters uploaded by the suspicious client and the benign client to obtain suspicious model parameters and benign model parameters, and importing the suspicious model parameters and the benign model parameters into a global model to obtain a teacher model and a student model; improving the student model through hierarchical knowledge distillation according to a knowledge distillation dataset and the teacher model to obtain a distilled student model, and sending the distilled student model as the global model to the client. The technical solution of the embodiment of the present invention is to perform singular value decomposition on the model high-level parameters in the local model parameters uploaded by the client to obtain a decomposition matrix, and then cluster the decomposition matrix to divide the client into a suspicious client and a benign client. Because it focuses on the model high-level parameters rather than all model parameters (local model parameters), it can reduce the influence of parameters with low correlation with backdoor features at the bottom layer and can better analyze the properties of the model; because the student model is improved through hierarchical knowledge distillation according to a knowledge distillation dataset and the teacher model to obtain a distilled student model, and the useful classification information in the teacher model is not discarded, that is, the useful classification information existing in the suspicious client model is not removed, making the distilled student model more robust, thereby improving the defense effect against federated backdoor attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is an overall schematic diagram of a method for defending against federated backdoor attacks based on hierarchical knowledge distillation provided by an embodiment of the present invention;
[0018] Figure 2 It is a flowchart of a method for defending against federated backdoor attacks based on hierarchical knowledge distillation provided by an embodiment of the present invention;
[0019] Figure 3 It is a sub-flowchart of a method for defending against federated backdoor attacks based on hierarchical knowledge distillation provided by an embodiment of the present invention;
[0020] Figure 4Another sub - process schematic diagram of a method for defending against federated backdoor attacks based on hierarchical knowledge distillation provided by an embodiment of the present invention;
[0021] Figure 5 Schematic block diagram of a device for defending against federated backdoor attacks based on hierarchical knowledge distillation provided by an embodiment of the present invention;
[0022] Figure 6 Schematic block diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0025] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0026] It should be further understood that the term "and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0027] As used in this specification and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.
[0028] Please refer to Figure 1 , Figure 1It is a schematic diagram of the overall structure of a method for defending against federated backdoor attacks based on hierarchical knowledge distillation provided by an embodiment of the present invention. The method for defending against federated backdoor attacks based on hierarchical knowledge distillation in the embodiment of the present invention is applied to a server. For example, the method for defending against federated backdoor attacks based on hierarchical knowledge distillation can be implemented by configuring corresponding software programs on the server, thereby improving the defense effect against federated backdoor attacks. It should be noted that in this embodiment, as Figure 1 shown, Δw1, Δw2, ……, Δw N-1 , Δw N represent the local model parameters uploaded by the client. The benign model is the student model, the suspected backdoor model is the teacher model, and the aggregated model is the distilled student model.
[0029] Please refer to Figure 2 , Figure 2 It is a schematic diagram of the process of a method for defending against federated backdoor attacks based on hierarchical knowledge distillation provided by an embodiment of the present invention. As Figure 2 shown, the method includes the following steps S110 - S130.
[0030] S110. Receive the local model parameters uploaded by the client, and perform singular value decomposition on the model high - level parameters in the local model parameters to obtain a decomposition matrix.
[0031] In the embodiment of the present invention, in the field of federated learning, the local model and the global model architecture are the same, only the model parameters (such as weights) are different. The client trains the global model sent by the server. After local training is completed, it is called the local model, and then the local model parameters corresponding to the trained local model are uploaded back to the server. The server receives the local model parameters uploaded by the client, and performs singular value decomposition on the model high - level parameters in the local model parameters to obtain a decomposition matrix. It should be noted that in the embodiment of the present invention, since the model consists of many different layers, the model high - level parameters refer to the parameters corresponding to the first few layers of the model.
[0032] It should also be noted that in the embodiments of the present invention, singular value decomposition (SVD) is an important matrix decomposition method in linear algebra, which is widely used in fields such as data analysis, signal processing, image compression, and recommendation systems. Through singular value decomposition, the high-level parameters of the model can be dimensionally reduced. In this embodiment, the reason for performing singular value decomposition on the high-level parameters of the model is that the singular value decomposition technology maps high-dimensional information into a key compact representation through linear transformation, thereby capturing the subtle differences between malicious models and benign models caused by backdoors, which is a commonly used method in current federated learning. However, this method has certain defects. For example, some researchers perform SVD dimensional reduction on the updated gradients of the complete model. The vectors after dimensional reduction will give priority to the underlying gradients with larger numerical fluctuations, and these gradients are often related to the basic feature extraction of the model and have a lower correlation with the backdoor features. This feature selection mechanism based on numerical magnitude will cause the algorithm to overly focus on the natural fluctuations of the underlying gradients and ignore the parameter changes with smaller magnitudes but more sensitive to backdoor attacks in the high-level of the model. This embodiment analyzes from the perspectives of information flow and gradient propagation and proposes to focus the analysis on the high-level parameters of the model. Specifically, first, from the perspective of information flow, the information in the neural network is abstracted layer by layer during the forward propagation process. When a backdoor attack occurs, the attacker must manipulate the high-level features to achieve control of the target output for specific inputs. Therefore, the backdoor attack is likely to leave obvious "traces" in the high-level of the neural network. Second, from the perspective of optimization theory, the gradients of the model parameters actually reflect the influence degree of the parameters on the loss function. In the backdoor attack scenario, the attacker needs to maintain the performance of the model on normal samples at the same time, which means that it must achieve a certain balance between normal tasks and backdoor tasks. Since the high-level parameters of the model directly participate in the final decision-making process, this balance relationship will be most obvious in the high-level parameters of the model. Finally, from the perspective of gradient propagation analysis, since the gradients may decay or explode during the backpropagation process, the underlying gradients are often affected by more noise. In contrast, the high-level gradients have a more direct relationship with the loss function and contain more information related to the behavior of the model.
[0033] S120. Cluster the decomposition matrix through a clustering algorithm to divide the clients into suspicious clients and benign clients, and respectively perform average aggregation on the local model parameters uploaded by the suspicious clients and the benign clients to obtain suspicious model parameters and benign model parameters, and import the suspicious model parameters and the benign model parameters into the global model to obtain a teacher model and a student model.
[0034] In the embodiments of the present invention, the decomposition matrix includes a column vector orthogonal matrix, such as Figure 3As shown, step S120 may specifically include steps S121 - S122: S121. Select the first k principal singular vectors in the column vector orthogonal matrix as the principal component vectors, where k ≥ 2; S122. Cluster the principal component vectors through the K-Means clustering algorithm to divide the clients into the suspicious clients and the benign clients. It can be understood that the decomposition matrix further includes a row vector orthogonal matrix and a diagonal matrix. It should be noted that in the embodiments of the present invention, only the first k principal singular vectors in the column vector orthogonal matrix are selected as the principal component vectors for clustering because the principal component vectors can describe the model features in a lower dimension, facilitating clustering analysis. And dividing the clients into suspicious clients and benign clients aims to quickly distinguish those models that may have been implanted with backdoors by attackers from the mixed local model parameters transmitted from multiple clients, which helps to identify and isolate malicious models in suspicious clients as early as possible, thereby reducing their impact on the subsequent improvement process. It should also be noted that in the embodiments of the present invention, since the local model and the global model architectures are the same, only by importing different model parameters, different models can be obtained. It can be understood that in the server, the local model parameters uploaded by the suspicious clients and the benign clients are respectively averaged and aggregated to obtain suspicious model parameters and benign model parameters, and the suspicious model parameters and the benign model parameters are respectively imported into the global model to obtain a teacher model and a student model. Specifically, the local model parameters uploaded by the suspicious clients are averaged and aggregated to obtain suspicious model parameters, and the suspicious model parameters are imported into the global model to obtain a teacher model; the local model parameters uploaded by the benign clients are averaged and aggregated to obtain benign model parameters, and the benign model parameters are imported into the global model to obtain a student model.
[0035] S130. Improve the student model through hierarchical knowledge distillation according to the knowledge distillation dataset and the teacher model to obtain a distilled student model, and send the distilled student model as the global model to the client.
[0036] In the embodiments of the present invention, after obtaining the teacher model and the student model, the student model is improved through hierarchical knowledge distillation according to the knowledge distillation dataset and the teacher model to obtain a distilled student model, and the distilled student model is sent as the global model to the client, so that the client can use local data to train the global model, realizing the continuous update and enhancement of the global model. The hierarchical knowledge distillation includes output layer distillation, intermediate layer feature distillation, and parameter sparsification, such as Figure 4As shown, step S130 may specifically include steps S131 - S134: S131. Calculate the probability distributions output by the teacher model and the student model based on the knowledge distillation dataset and the preset temperature coefficient to obtain the teacher probability distribution and the student probability distribution, and calculate the output layer distillation loss through the output layer distillation loss function according to the teacher probability distribution, the student probability distribution, and the preset temperature coefficient to perform the output layer distillation; S132. Calculate the intermediate layer distillation loss through the intermediate layer distillation loss function based on the intermediate layer feature maps in the teacher model and the intermediate layer feature maps in the student model to perform the intermediate layer feature distillation; S133. Calculate the parameter sparsification loss through the parameter sparsification loss function according to the weights of the student model and the regularization coefficient to perform the parameter sparsification; S134. Improve the student model according to the output layer distillation loss, the intermediate layer distillation loss, and the sparsification loss to obtain the distilled student model. It should be noted that in the embodiments of the present invention, for the knowledge distillation dataset, taking the image classification task as an example, 500 pictures in the publicly available dataset ImageNet can be selected as the knowledge distillation dataset, and the knowledge distillation dataset is used to extract the potential classification information in the teacher model and transfer it to the student model.
[0037] Further, the calculating the probability distributions output by the teacher model and the student model based on the knowledge distillation dataset and the preset temperature coefficient to obtain the teacher probability distribution and the student probability distribution includes: for each piece of distillation data in the knowledge distillation dataset, input the distillation data into the teacher model and the student model to obtain the output of the teacher model and the output of the student model; calculate the ratio of the output of the teacher model to the preset temperature coefficient to obtain the target teacher model output, and convert the target teacher model output into the teacher probability distribution through an activation function; convert the output of the student model into the student probability distribution through the natural logarithm activation function. The calculating the output layer distillation loss through the output layer distillation loss function according to the teacher probability distribution, the student probability distribution, and the preset temperature coefficient includes: calculating the KL divergence between the teacher probability distribution and the student probability distribution to obtain a divergence value; calculating the output layer distillation loss according to the divergence value and the preset temperature coefficient. The output layer distillation is to soften the backdoor response, specifically as follows: at the output layer, the output probability distribution of the teacher model is softened with the preset temperature coefficient T to make the activation function smoother, while the output of the student model is processed by the natural logarithm activation function to accurately align its response to the data.
[0038] q teacher = softmax(f teacher (x) / T) (1)
[0039] q student= log_softmax(f student (x)) (2)
[0040] Wherein, in Formula (1) and Formula (2), q teacher represents the teacher probability distribution; q student represents the student probability distribution, f teacher (x) is the output of the teacher model, f student (x) is the output of the student model, x is the input of the teacher model and the student model, representing the distilled data in the knowledge distillation dataset, T is the preset temperature coefficient. After dividing the output of the teacher model by the preset temperature coefficient and then performing the activation function Softmax calculation, a softened teacher probability distribution can be obtained, and the value distribution is relatively gentle. Understandably, the larger the value of the preset temperature coefficient T, the more gentle the distribution. Only apply the preset temperature coefficient T to the output of the teacher model, so as to more accurately capture and suppress the backdoor activation and maintain the effectiveness and rationality of the model output. Calculate the KL divergence under this high-temperature state to minimize the difference between the models, thereby weakening the abnormal output response of the model and ensuring the correctness of the output.
[0041]
[0042] In Formula (3), D KL (q teacher ||q student ) is the KL divergence calculation, which is used to measure the difference between the two teacher probability distributions and the student probability distribution. is the output layer distillation loss, and the output layer distillation loss function is Formula (3). The summation in Formula (3) is because there are multiple distilled data in the knowledge distillation dataset. The purpose of designing the output layer distillation is to soften the impact of the backdoor trigger on the model output and remove the abnormal outputs that should not be output.
[0043] Furthermore, calculating the intermediate layer distillation loss according to the intermediate layer feature maps in the teacher model and the intermediate layer feature maps in the student model through the intermediate layer distillation loss function includes: calculating the Gram matrix of the intermediate layer feature maps in the teacher model to obtain the teacher intermediate layer feature map matrix; calculating the Gram matrix of the intermediate layer feature maps in the student model to obtain the student intermediate layer feature map matrix; calculating the intermediate layer distillation loss according to the teacher intermediate layer feature map matrix and the student intermediate layer feature map matrix. It should be noted that the intermediate layer distillation loss function is shown in Formula (4). The intermediate layer feature distillation is to eliminate the backdoor features, specifically as follows:
[0044] Calculate the intermediate layer feature maps in the teacher model and the student model using the Gram matrix to obtain the teacher intermediate layer feature map matrix and the student intermediate layer feature map matrix. The teacher intermediate layer feature map matrix contains the internal structure information of the teacher model, and the student intermediate layer feature map matrix contains the internal structure information of the student model. By aligning the features between the student model and the teacher model, reduce the differences in backdoor features at the intermediate layer and destroy its abnormal propagation chain.
[0045]
[0046] Among them, in formula (4), G teacher is the teacher intermediate layer feature map matrix, G student is the student intermediate layer feature map matrix, is the Euclidean squared norm of G teacher and G student , and is the intermediate layer distillation loss. The purpose of designing the intermediate layer feature distillation is to prevent the propagation of the backdoor malicious feature encoding path by aligning the features at the intermediate layer.
[0047] Furthermore, parameter sparsification is to prune the backdoor parameters, specifically as follows: introduce a regularization loss to make those parameters related to or unnecessary for the backdoor become sparse. This method gradually reduces redundant parameters by imposing penalties on the parameters of each layer during training, retains the important weights in the student model, and thus weakens the role of the hidden backdoor parameters.
[0048]
[0049] Among them, the parameter sparsification loss function is as shown in formula (5). In formula (5), w i is the weight of the student model, and γ is the regularization coefficient. Understandably, too large a γ will cause the student model to underfit, and too small a γ will not be able to effectively prune the backdoor parameters. It should be noted that the purpose of parameter sparsification is to remove unnecessary and potentially malicious parameters in the student model.
[0050] Even further, improving the student model according to the output layer distillation loss, the intermediate layer distillation loss, and the sparsification loss to obtain the distilled student model includes: calculating the product of the preset output layer distillation loss weight and the output layer distillation loss to obtain the target output layer distillation loss; calculating the product of the preset intermediate layer distillation loss weight and the intermediate layer distillation loss to obtain the target intermediate layer distillation loss; calculating the sum of the target output layer distillation loss, the target intermediate layer distillation loss, and the sparsification loss to obtain the total loss, and improving the student model according to the total loss to obtain the distilled student model. The calculation formula of the total loss is as shown in formula (6).
[0051]
[0052] Among them, in formula (6), α is the preset output layer distillation loss weight, and β is the preset intermediate layer distillation loss weight.
[0053] It should be noted that in the embodiments of the present invention, the reason for improving the student model by using hierarchical knowledge distillation is that for backdoor-robust federated learning, in order to ensure that no backdoors are introduced into the global model, relatively extreme discarding strategies are adopted, and the local model parameters in the clients judged to be suspected of having backdoors are directly discarded. The essence of the traditional method is to use some statistical global information to screen out some backdoor models and exclude them from the aggregation process. When the local data among normal participants is heterogeneous, it is very difficult to ensure that the uploaded models of normal participants are not excluded. As a result, the classification information of these normal participants is lost. In addition, assuming that there is an idealized model discrimination method that can accurately classify the models corresponding to the uploaded local model parameters into normal models and backdoor models, then the model aggregated in this round of federated learning only has the information of normal models, and the information of backdoor models is directly discarded. However, according to the definition of the federated backdoor attack, the backdoor model only exhibits specific backdoor behaviors when there are backdoor triggers in the input data. When the input does not contain backdoor triggers, the performance of the backdoor model is the same as that of the normal model. This means that there is also correct model classification knowledge in the backdoor model. If the parameters classified as backdoor models are directly discarded according to the existing method, it will show a phenomenon of loss of classification information globally, resulting in an impact on the global training process of federated learning. Based on the above analysis, knowledge distillation, as a knowledge transfer method based on soft labels, can make good use of the fact that the backdoor model itself performs normally on clean samples. Secondly, as a current mature knowledge transfer technology, knowledge distillation can adjust the distribution of soft labels through temperature parameters, capture the decision boundary information of the teacher model, and effectively transfer this information to the student model, so that the distilled student model has better robustness.
[0054] In summary, in this embodiment, through singular value decomposition of the high-level parameters of the model and hierarchical knowledge distillation technology, the federated backdoor attack in federated learning is effectively defended. By identifying and stripping the backdoor knowledge in the malicious model (teacher model), the security and robustness of the global model are improved. Specifically, by performing singular value decomposition only on the high-level parameters of the model to extract the principal component vectors, the normal model (student model) and the potential backdoor model (teacher model) can be efficiently distinguished, significantly improving the detection accuracy of the potential backdoor model. By using the hierarchical distillation technology to strip the malicious information hidden by the backdoor layer by layer, the ability of the global model to resist the federated backdoor attack is significantly improved, maintaining the output consistency and accuracy under potential attacks. In the whole process, there is no need to access the original client data, and a public benign dataset (knowledge distillation dataset) is used for secure knowledge distillation, perfectly balancing data privacy protection and model performance.
[0055] Figure 5 FIG. 4 is a schematic block diagram of an apparatus 200 for defending against federated backdoor attacks based on hierarchical knowledge distillation provided by an embodiment of the present invention. As Figure 5 shown, corresponding to the above method for defending against federated backdoor attacks based on hierarchical knowledge distillation, the present invention also provides an apparatus 200 for defending against federated backdoor attacks based on hierarchical knowledge distillation. The apparatus 200 for defending against federated backdoor attacks based on hierarchical knowledge distillation includes units for executing the above method for defending against federated backdoor attacks based on hierarchical knowledge distillation, and the apparatus can be configured in a computer device. Specifically, please refer to Figure 5 FIG. 4, the apparatus 200 for defending against federated backdoor attacks based on hierarchical knowledge distillation includes a receiving and decomposing unit 201, a clustering and partitioning unit 202, and an improved distribution unit 203.
[0056] Among them, the receiving and decomposing unit 201 is configured to receive the local model parameters uploaded by the client and perform singular value decomposition on the high-level parameters of the model in the local model parameters to obtain a decomposition matrix; the clustering and partitioning unit 202 is configured to cluster the decomposition matrix through a clustering algorithm to divide the client into a suspicious client and a benign client, and respectively perform average aggregation on the local model parameters uploaded by the suspicious client and the benign client to obtain suspicious model parameters and benign model parameters, and import the suspicious model parameters and the benign model parameters into the global model respectively to obtain a teacher model and a student model; the improved distribution unit 203 is configured to improve the student model through hierarchical knowledge distillation according to the knowledge distillation dataset and the teacher model to obtain a distilled student model, and use the distilled student model as the global model to be distributed to the client.
[0057] In some embodiments, such as this embodiment, the clustering and partitioning unit 202 includes a selection unit and a clustering unit.
[0058] Among them, the selection unit is used to select the first k principal singular vectors in the column vector orthogonal matrix as the principal component vectors, where k≥2; the clustering unit is used to cluster the principal component vectors through the K-Means clustering algorithm to divide the clients into the suspicious clients and the benign clients.
[0059] In some embodiments, such as this embodiment, the improved distribution unit 203 includes a first calculation unit, a second calculation unit, a third calculation unit, and an improvement unit.
[0060] Among them, the first calculation unit is used to calculate the probability distributions output by the teacher model and the student model according to the knowledge distillation data set and the preset temperature coefficient to obtain the teacher probability distribution and the student probability distribution, and calculate the output layer distillation loss through the output layer distillation loss function according to the teacher probability distribution, the student probability distribution, and the preset temperature coefficient to perform the output layer distillation; the second calculation unit is used to calculate the intermediate layer distillation loss through the intermediate layer distillation loss function according to the intermediate layer feature maps in the teacher model and the intermediate layer feature maps in the student model to perform the intermediate layer feature distillation; the third calculation unit is used to calculate the parameter sparsity loss through the parameter sparsity loss function according to the weights of the student model and the regularization coefficient to perform the parameter sparsity; the improvement unit is used to improve the student model according to the output layer distillation loss, the intermediate layer distillation loss, and the sparsity loss to obtain the distilled student model.
[0061] In some embodiments, such as this embodiment, the first calculation unit includes an input unit, a first conversion unit, a second conversion unit, a first calculation subunit, and a second calculation subunit.
[0062] Among them, the input unit is used to input each piece of distillation data in the knowledge distillation data set into the teacher model and the student model to obtain the teacher model output and the student model output; the first conversion unit is used to calculate the ratio of the teacher model output to the preset temperature coefficient to obtain the target teacher model output, and convert the target teacher model output into the teacher probability distribution through the activation function; the second conversion unit is used to convert the student model output into the student probability distribution through the natural logarithm activation function; the first calculation subunit is used to calculate the KL divergence between the teacher probability distribution and the student probability distribution to obtain the divergence value; the second calculation subunit is used to calculate the output layer distillation loss according to the divergence value and the preset temperature coefficient.
[0063] In some embodiments, such as this embodiment, the second computing unit includes a third computing subunit, a fourth computing subunit, and a fifth computing subunit.
[0064] Among them, the third computing subunit is used to calculate the Gram matrix of the intermediate layer feature map in the teacher model to obtain the teacher intermediate layer feature map matrix; the fourth computing subunit is used to calculate the Gram matrix of the intermediate layer feature map in the student model to obtain the student intermediate layer feature map matrix; the fifth computing subunit calculates the intermediate layer distillation loss according to the teacher intermediate layer feature map matrix and the student intermediate layer feature map matrix.
[0065] In some embodiments, such as this embodiment, the improvement unit includes a sixth computing subunit, a seventh computing subunit, and an eighth computing subunit.
[0066] Among them, the sixth computing subunit is used to calculate the product of the preset output layer distillation loss weight and the output layer distillation loss to obtain the target output layer distillation loss; the seventh computing subunit is used to calculate the product of the preset intermediate layer distillation loss weight and the intermediate layer distillation loss to obtain the target intermediate layer distillation loss; the eighth computing subunit is used to calculate the sum of the target output layer distillation loss, the target intermediate layer distillation loss, and the sparsification loss to obtain the total loss, and improve the student model according to the total loss to obtain the distilled student model.
[0067] The specific implementation manner of the device 200 for defending against federated backdoor attacks based on hierarchical knowledge distillation according to the embodiments of the present invention corresponds to the above method for defending against federated backdoor attacks based on hierarchical knowledge distillation, and will not be elaborated here.
[0068] The above device for defending against federated backdoor attacks based on hierarchical knowledge distillation can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 6 .
[0069] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 300 is a device with the function of defending against federated backdoor attacks based on hierarchical knowledge distillation.
[0070] Refer to Figure 6 , the computer device 300 includes a processor 302, a memory, and a network interface 305 connected through a system bus 301. Among them, the memory may include a storage medium 303 and an internal memory 304.
[0071] The storage medium 303 can store an operating system 3031 and a computer program 3032. When the computer program 3032 is executed, it can cause the processor 302 to execute a method for defending against federated backdoor attacks based on hierarchical knowledge distillation.
[0072] The processor 302 is used to provide computing and control capabilities to support the operation of the entire computer device 300.
[0073] The internal memory 304 provides an environment for the operation of the computer program 3032 in the storage medium 303. When the computer program 3032 is executed by the processor 302, it can cause the processor 302 to execute a method for defending against federated backdoor attacks based on hierarchical knowledge distillation.
[0074] The network interface 305 is used for network communication with other devices. Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 300 to which the solution of this application is applied. The specific computer device 300 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0075] Among them, the processor 302 is used to run the computer program 3032 stored in the memory to implement any embodiment of the above method for defending against federated backdoor attacks based on hierarchical knowledge distillation.
[0076] It should be understood that in the embodiments of this application, the processor 302 may be a central processing unit (CPU), and the processor 302 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0077] Those of ordinary skill in the art can understand that all or part of the processes in the methods of implementing the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above method.
[0078] Therefore, the present invention also provides a storage medium. The storage medium can be a computer-readable storage medium. The storage medium stores a computer program. When the computer program is executed by a processor, the processor executes any of the embodiments of the above method for defending against federated backdoor attacks based on hierarchical knowledge distillation.
[0079] The storage medium can be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, an optical disc, or other various computer-readable storage media that can store program codes.
[0080] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0081] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0082] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0083] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.
[0084] In the above embodiments, the descriptions of the respective embodiments each have their own emphasis. For parts not described in detail in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0085] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, provided that these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications therein.
[0086] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for defending against federated backdoor attacks based on hierarchical knowledge distillation, characterized in that, Including: Receiving local model parameters uploaded by a client, and performing singular value decomposition on the model high-level parameters in the local model parameters to obtain a decomposition matrix; Clustering the decomposition matrix through a clustering algorithm to divide the clients into suspicious clients and benign clients, and respectively averaging and aggregating the local model parameters uploaded by the suspicious clients and the benign clients to obtain suspicious model parameters and benign model parameters, and importing the suspicious model parameters and the benign model parameters into a global model respectively to obtain a teacher model and a student model; Improving the student model through hierarchical knowledge distillation according to a knowledge distillation dataset and the teacher model to obtain a distilled student model, and sending the distilled student model as the global model to the client.
2. The method according to claim 1, characterized in that, The decomposition matrix includes a column vector orthogonal matrix, and the clustering the decomposition matrix through a clustering algorithm to divide the clients into suspicious clients and benign clients includes: Selecting the first k principal singular vectors in the column vector orthogonal matrix as principal component vectors, where k≥2; Clustering the principal component vectors through a K-Means clustering algorithm to divide the clients into the suspicious clients and the benign clients.
3. The method according to any one of claims 1-2, characterized in that The hierarchical knowledge distillation includes output layer distillation, intermediate layer feature distillation, and parameter sparsification. The improving the student model through hierarchical knowledge distillation according to a knowledge distillation dataset and the teacher model to obtain a distilled student model includes: Calculating the probability distributions output by the teacher model and the student model according to the knowledge distillation dataset and a preset temperature coefficient to obtain a teacher probability distribution and a student probability distribution, and calculating an output layer distillation loss through an output layer distillation loss function according to the teacher probability distribution, the student probability distribution, and the preset temperature coefficient to perform the output layer distillation; Calculating an intermediate layer distillation loss through an intermediate layer distillation loss function according to the intermediate layer feature maps in the teacher model and the intermediate layer feature maps in the student model to perform the intermediate layer feature distillation; Calculating a parameter sparsification loss through a parameter sparsification loss function according to the weights of the student model and a regularization coefficient to perform the parameter sparsification; Improving the student model according to the output layer distillation loss, the intermediate layer distillation loss, and the sparsification loss to obtain the distilled student model.
4. The method according to claim 3, wherein The calculating the probability distributions output by the teacher model and the student model according to the knowledge distillation dataset and a preset temperature coefficient to obtain a teacher probability distribution and a student probability distribution includes: For each piece of distillation data in the knowledge distillation dataset, inputting the distillation data into the teacher model and the student model to obtain a teacher model output and a student model output; Calculating the ratio of the teacher model output to the preset temperature coefficient to obtain a target teacher model output, and converting the target teacher model output into the teacher probability distribution through an activation function; Converting the student model output into the student probability distribution through a natural logarithm activation function.
5. The method according to claim 3, characterized in that, Calculating the output layer distillation loss through the output layer distillation loss function according to the teacher probability distribution, the student probability distribution, and the preset temperature coefficient includes: Calculating the KL divergence between the teacher probability distribution and the student probability distribution to obtain a divergence value; Calculating the output layer distillation loss according to the divergence value and the preset temperature coefficient.
6. The method according to claim 3, wherein Calculating the intermediate layer distillation loss through the intermediate layer distillation loss function according to the intermediate layer feature maps in the teacher model and the intermediate layer feature maps in the student model includes: Calculating the Gram matrix of the intermediate layer feature maps in the teacher model to obtain the teacher intermediate layer feature map matrix; Calculating the Gram matrix of the intermediate layer feature maps in the student model to obtain the student intermediate layer feature map matrix; Calculating the intermediate layer distillation loss according to the teacher intermediate layer feature map matrix and the student intermediate layer feature map matrix.
7. The method according to claim 3, wherein Improving the student model according to the output layer distillation loss, the intermediate layer distillation loss, and the sparsification loss to obtain the distilled student model includes: Calculating the product of the preset output layer distillation loss weight and the output layer distillation loss to obtain the target output layer distillation loss; Calculating the product of the preset intermediate layer distillation loss weight and the intermediate layer distillation loss to obtain the target intermediate layer distillation loss; Calculating the sum of the target output layer distillation loss, the target intermediate layer distillation loss, and the sparsification loss to obtain the total loss, and improving the student model according to the total loss to obtain the distilled student model.
8. An apparatus for defending against federated backdoor attacks based on hierarchical knowledge distillation, characterized in that, Including: A receiving and decomposing unit, configured to receive the local model parameters uploaded by the client, and perform singular value decomposition on the model high-level parameters in the local model parameters to obtain a decomposition matrix; A clustering and partitioning unit, configured to cluster the decomposition matrix through a clustering algorithm to divide the client into a suspicious client and a benign client, and respectively perform average aggregation on the local model parameters uploaded by the suspicious client and the benign client to obtain suspicious model parameters and benign model parameters, and import the suspicious model parameters and the benign model parameters into the global model respectively to obtain a teacher model and a student model; An improvement and distribution unit, configured to improve the student model through hierarchical knowledge distillation according to the knowledge distillation data set and the teacher model to obtain a distilled student model, and distribute the distilled student model as the global model to the client.
9. A computer device, characterized in that, The computer device includes a memory and a processor, and a computer program is stored on the memory. When the processor executes the computer program, the method according to any one of claims 1-7 is implemented.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-7 is implemented.