A federated learning robust adaptation and privacy protection method and system based on a double-end structure
Patent Information
- Application Number
- CN202610944894.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]第一类是基于梯度的隐私窃取攻击,攻击者或半诚实服务器可能通过深度反演算法,利用梯度匹配重构出客户端本地的原始数据
[0085]1、本发明客户端侧通过多赢者通吃竞争激活实现特征结构化稀疏,根据梯度张量拓扑属性掩码解耦,将携带敏感信息的私有梯度永久留存本地,再通过正交补集投影线性剥离共享梯度中与本地隐私平行的泄露分量;
Smart Images

Figure CN122839422A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of distributed machine learning and information security technology, and in particular to a robust adaptation and privacy protection method and system for federated learning based on a two-end structure. Background Technology
[0002] With the convergence of edge computing and artificial intelligence, traditional centralized machine learning mechanisms face challenges related to data security and communication overhead. Federated learning, through its distributed mechanism of "data remains stationary, model moves," allows participating nodes to train locally and interact only to update parameters, thus mitigating direct data privacy risks. However, existing federated learning architectures still face two main threats:
[0003] The first type is gradient-based privacy-stealing attacks. Attackers or semi-honest servers may use deep inversion algorithms to reconstruct the original local data on the client using gradient matching. Existing defenses mostly employ differential privacy or homomorphic encryption. While differential privacy can obfuscate data, the noise it introduces can corrupt the original gradient structure, causing a significant decrease in the model's performance when handling high-precision tasks. Homomorphic encryption, on the other hand, has the limitation of huge computational and communication overhead, making it difficult to deploy on a large scale on resource-constrained edge nodes.
[0004] The second type of threat is malicious poisoning attacks. In federated networks lacking physical monitoring, anomalous nodes can inject misleading parameters into the global aggregation process by tampering with local data or falsifying model gradients. Existing robust aggregation algorithms are highly susceptible to misinterpreting differences caused by data heterogeneity in real-world scenarios with non-independent and identically distributed data, leading to limited detection accuracy.
[0005] Even more challenging is the mathematical difficulty in synergizing existing privacy protection and robust aggregation mechanisms. For example, the large amount of random noise added to protect privacy can obscure the true distance characteristics of gradients, significantly reducing the effectiveness of algorithms that attempt to identify anomalous nodes through spatial geometric analysis.
[0006] Therefore, there is an urgent need for a lightweight federated computing architecture that can adaptively balance privacy and model robustness. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention provides a robust adaptation and privacy protection method and system for federated learning based on a dual-end structure. The client and server of this invention work together to achieve privacy protection and anti-poisoning compatibility.
[0008] In a first aspect, the present invention provides a robust adaptation and privacy protection method for federated learning based on a dual-end structure, comprising:
[0009] The client, based on a global neural network model containing multi-dimensional feature mapping layers, performs forward propagation computation locally by introducing a winner-take-all competitive activation mechanism to generate hidden layer features with structural sparsity.
[0010] The client performs backpropagation of error locally to generate a set of gradient tensors, and generates a binary mask based on the topological dimension or structural properties of the gradient tensor set; the client decouples the gradient tensor set into private gradients locked locally and shared gradients that are allowed to be uploaded, according to the binary mask.
[0011] A private feature matrix is constructed based on the private gradient, and the sensitive subspace representation of the private feature matrix is extracted. An orthogonal projection matrix is constructed based on the sensitive subspace representation.
[0012] The shared gradient is calculated by orthogonal complement projection using the orthogonal projection matrix, and the associated components parallel to the local sensitive features are stripped to obtain the orthogonal purified gradient, which is then uploaded to the server.
[0013] The server receives the orthogonal purification gradients uploaded by multiple clients, performs momentum tracking over the time dimension, and constructs a client state graph based on the L1 norm momentum distance.
[0014] The server performs dimensionality reduction mapping on the client state graph, uses a clustering algorithm to block updates of abnormal clusters, and only performs momentum updates of normal clusters to iterate the global neural network model.
[0015] Preferably, the client, based on a global neural network model containing a multi-dimensional feature mapping layer, performs forward propagation computation locally by introducing a winner-takes-all competitive activation mechanism to generate structurally sparse hidden layer features, including:
[0016] The client loads the global neural network model issued by the server;
[0017] The client inputs local training samples into the global neural network model, performs forward propagation layer by layer, and calculates the preactivation vectors of each hidden layer in turn.
[0018] Sort the original preactivation output values of all neurons in the preactivation vector of the current hidden layer in descending order to obtain the sorted neuron sequence;
[0019] Before selection The neurons corresponding to the original pre-activation output values are used as candidate neurons. If the original pre-activation output value of a candidate neuron is greater than or equal to the dynamic threshold for competitive screening, the corresponding candidate neuron is determined to be a valid candidate neuron.
[0020] For all valid candidate neurons in the current hidden layer, retain the original pre-activated output values, and assign the remaining neurons the preset inactive values to generate the sparse hidden layer features of the current hidden layer.
[0021] Preferably, the client performs error backpropagation locally to generate a gradient tensor set, and generates a binary mask based on the topological dimension or structural properties of the gradient tensor set, including:
[0022] The client calculates the overall loss function for a single batch of training samples based on the prediction results of the sparse forward propagation of the global neural network model and the true labels of the samples.
[0023] By backpropagating from the output layer to the input layer of the global neural network model, the gradient of the overall loss function with respect to the parameters of each hidden layer is calculated, and the gradient tensor set is obtained by summing the gradients of all hidden layer parameters.
[0024] We classify all gradient tensors in the gradient tensor set, treating the gradient tensors of the bias terms and basic connection weight parameters of the global neural network model as one-dimensional gradient tensors, and the gradient tensors of convolution weights and high-dimensional fully connected weight parameters as multi-dimensional gradient tensors.
[0025] If the gradient tensor in the set of gradient tensors is a one-dimensional gradient tensor, then an exemption mask of 0 is assigned to it;
[0026] If the gradient tensor in the gradient tensor set is determined to be a multidimensional gradient tensor carrying high-dimensional features, the core extreme value component at the front end of the gradient tensor magnitude sorting is extracted according to the sparsity hyperparameter, and the binary mask of the gradient tensor corresponding to the core extreme value component is set to 1, and the rest are 0.
[0027] Traverse all gradient tensors in the gradient tensor set, generate corresponding masks layer by layer, and obtain a mask set with the same structure as the gradient tensor set.
[0028] Preferably, the client decouples the gradient tensor set into locally locked private gradients and shared gradients that are allowed to be uploaded, based on the binary mask, including:
[0029] Based on the mask set and gradient tensor set, gradient decoupling is achieved using the Hadamard product operation, resulting in private and shared gradients, including:
[0030] Multiply the gradient tensor set element by element with the mask set to obtain a private gradient that is locked locally and cannot be uploaded.
[0031] Multiply the gradient tensor set element by element with the inverted mask set to obtain the shared gradient that can be uploaded to the server.
[0032] Preferably, a private feature matrix is constructed based on the private gradient, and a sensitive subspace representation of the private feature matrix is extracted. An orthogonal projection matrix is then constructed based on the sensitive subspace representation, including:
[0033] After flattening the private gradient into a one-dimensional vector, the vectors are concatenated according to the feature dimensions to form a private feature matrix.
[0034] Truncated singular value decomposition is used to reduce the dimensionality of the private feature matrix, and principal component basis vector matrices representing privacy information are extracted, including:
[0035] For private characteristic matrix Perform singular value decomposition, i.e.:
[0036] ;
[0037] In the formula, It is a left singular vector matrix; For a singular value diagonal matrix, the singular values are arranged in descending order of their diagonal elements. ; It is a right singular vector matrix, with each column corresponding to a set of characteristic bases; Indicates the transpose operation;
[0038] Set a truncation hyperparameter s, and extract the first s right singular vectors to form the principal component basis vector matrix of the privacy subspace. ,Right now:
[0039] ;
[0040] in, This represents the basis vector of the s-th principal component;
[0041] Principal component basis vector matrix based on privacy subspace Construct the orthogonal complement projection matrix, i.e.:
[0042] ;
[0043] In the formula, This is the orthogonal complementary projection matrix; It is the identity matrix; This is the projection matrix for the privacy subspace.
[0044] Preferably, the shared gradient is orthogonally complemented by the orthogonal projection matrix, and associated components parallel to local sensitive features are removed to obtain orthogonally purified gradients, which are then uploaded to the server, including:
[0045] The shared gradient is preprocessed into a vectorized form to obtain the vectorized shared gradient vector.
[0046] Multiply the vectorized shared gradient vector by the orthogonal complement projection matrix on the left. The sensitive parallel components are stripped to obtain the orthogonal purification gradient; the orthogonal purification gradient obtained by projection is uploaded to the server.
[0047] Preferably, the server receives the orthogonal purification gradients uploaded by multiple clients, performs momentum tracking over the time dimension, and constructs a client state graph based on the L1 norm momentum distance, including:
[0048] After the t-th round of communication, the server obtains the orthogonal purification gradient from the client. Based on client-side orthogonal purification gradient Create a timing momentum buffer vector for each client. ,Right now:
[0049] ;
[0050] In the formula, For the client The timing momentum buffer vector in round t. ; Indicates the smoothing hyperparameter; For the client The orthogonal purification gradient in round t;
[0051] Calculate the L1 norm momentum distance between the temporal momentum buffer vectors of any two clients, i.e.: ;
[0052] In the formula, For the client and L1 norm momentum distance; The dimension of the timing momentum buffer vector; , For clients and The timing momentum buffer vector is at the 1st Values in each dimension;
[0053] Iterate through all clients participating in the federated training, calculate the L1 norm momentum distance between any two clients, and construct the client state graph. .
[0054] Preferably, the server performs dimensionality reduction mapping on the client state graph, uses a clustering algorithm to block updates of abnormal clusters, and only performs aggregation calculations on momentum updates of normal clusters to iterate the global neural network model, including:
[0055] The similarity between any two clients in each group is calculated based on the L1 norm momentum distance of the client state graph, i.e.:
[0056] ;
[0057] In the formula, For the client and Similarity; The Laplacian kernel-scale hyperparameter;
[0058] Based on the similarity between any two clients Construct a degree matrix, which is a diagonal matrix, where each diagonal element in each row is equal to the sum of the similarities of all clients in the corresponding row, and other elements are 0;
[0059] The graph Laplacian matrix corresponding to the undirected graph is calculated by subtracting the degree matrix from the similarity matrix.
[0060] Perform eigenvalue decomposition on the graph Laplace moments to obtain a diagonal matrix composed of ascending eigenvalues and its corresponding orthogonal eigenvector matrix. Select the eigenvectors matched by the first few smallest non-zero eigenvalues, concatenate them to form a low-dimensional mapping matrix, and then assign the temporal momentum buffer vector corresponding to each client... Transform it into a low-dimensional feature vector; that is:
[0061] ;
[0062] In the formula, It is an orthogonal eigenvector matrix, with each column corresponding to a set of eigenvectors; It is a diagonal eigenvalue matrix, with eigenvalues arranged in ascending order. ;
[0063] Select the eigenvectors corresponding to the k smallest non-zero eigenvalues and concatenate them to form a dimension-reduced mapping matrix. ,Right now: ;
[0064] Each client Mapping to k-dimensional low-dimensional eigenvectors, and taking the dimension reduction mapping matrix. No. Line as a client low-dimensional feature vectors ,Right now:
[0065] ;
[0066] Based on low-dimensional feature vectors The K-means unsupervised clustering algorithm is used to divide the clients into a normal client cluster and a malicious poisoning client cluster, including:
[0067] Randomly initialize the cluster centers of two clusters, calculate the distance between the low-dimensional feature vector of each client and the cluster centers of the two clusters one by one, and assign the client to the cluster with the closer distance;
[0068] After all clients are partitioned, the cluster centers are recalculated based on all clients within the two newly partitioned clusters.
[0069] Then, the distance between the low-dimensional feature vector of each client and the updated cluster center is calculated, and the clients are assigned to clusters that are closer in distance.
[0070] Repeat the iterations until the cluster centers no longer change significantly, thus completing the convergence of client groupings;
[0071] After grouping, the sample size and average momentum distance within the two clusters are compared, and the cluster with smaller sample size and greater average momentum distance within the cluster is identified as the malicious poisoning client cluster, while the other cluster is identified as the normal client cluster.
[0072] The server utilizes the timing momentum buffer vector of all clients within the normal client cluster. Perform aggregate calculations to iterate the global neural network model.
[0073] Secondly, the present invention provides a federated learning robust adaptation and privacy protection system based on a dual-end structure, including multiple clients and servers;
[0074] Each of the aforementioned clients is configured with a feature sparse activation module, a parameter decoupling mask module, a sensitive feature extraction module, and an orthogonal projection calculation module;
[0075] The server is equipped with a momentum tracking module and a clustering secure aggregation module;
[0076] The feature sparse activation module is based on a global neural network model containing a multi-dimensional feature mapping layer. It performs forward propagation calculations locally by introducing a winner-take-all competitive activation mechanism to generate hidden layer features with structural sparsity.
[0077] The parameter decoupling mask module performs error backpropagation locally to generate a gradient tensor set, and generates a binary mask based on the topological dimension or structural properties of the gradient tensor set; the client decouples the gradient tensor set into a private gradient locked locally and a shared gradient that can be uploaded, according to the binary mask.
[0078] The sensitive feature extraction module constructs a private feature matrix based on the private gradient, extracts the sensitive subspace representation of the private feature matrix, and constructs an orthogonal projection matrix based on the sensitive subspace representation.
[0079] The orthogonal projection calculation module uses the orthogonal projection matrix to perform orthogonal complement projection calculation on the shared gradient, strips out the associated components parallel to the local sensitive features, obtains the orthogonal cleaned gradient, and uploads it to the server.
[0080] The momentum tracking module of the server receives the orthogonal purification gradient uploaded by the orthogonal projection calculation modules of multiple clients, performs momentum tracking in the time dimension, and constructs a client state graph based on the L1 norm momentum distance.
[0081] The clustering security aggregation module of the server performs dimensionality reduction mapping on the client state graph, uses a clustering algorithm to block the updates of abnormal clusters, and only performs aggregation calculations on the momentum updates of normal clusters to iterate the global neural network model.
[0082] Thirdly, the present invention provides an electronic device, comprising:
[0083] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program, when executed by the at least one processor, implementing a federated learning robust adaptation and privacy protection method based on a dual-end structure.
[0084] The beneficial effects of this invention are as follows:
[0085] 1. On the client side of this invention, feature structured sparsity is achieved through multi-win competitive activation. Based on the gradient tensor topological attribute mask decoupling, the private gradient carrying sensitive information is permanently stored locally. Then, the leakage component parallel to the local privacy in the shared gradient is linearly stripped away through orthogonal complement projection.
[0086] 2. The orthogonal projection of the present invention only removes linear correlation components in the gradient that can be used for privacy leakage, while retaining the overall gradient distribution, amplitude differences and other geometric features. The orthogonal purified gradient uploaded to the server still has clear node distinguishing features, taking into account both privacy security and anti-poisoning ability.
[0087] 3. The server-side of this invention maintains a temporal momentum buffer for each client, and uses exponential moving average to accumulate gradient update trends over multiple rounds, smoothing out random jitter in a single round of training and amplifying the fixed distortion offset caused by continuous gradient tampering by malicious nodes. At the same time, the L1 norm, which is more sensitive to local amplitude abrupt changes, is used to measure the client momentum vector distance, further widening the feature gap between malicious and benign clients. Based on the distance matrix, a graph Laplacian matrix is constructed and spectral dimensionality reduction and K-means unsupervised clustering are performed, which can divide normal and abnormal clusters without pre-labeling malicious samples, and still maintains a low false negative and false positive rate in scenarios with a high proportion of malicious nodes and strong data heterogeneity.
[0088] 4. The client of this invention blocks the path of privacy data leakage locally through competitive activation, mask decoupling, and orthogonal projection processing; the server avoids malicious gradient pollution of the global model by filtering abnormal node updates through momentum tracking and spectral clustering, and at the same time resists two mainstream attacks: semi-honest server gradient inversion and external malicious client data poisoning. Attached Figure Description
[0089] Figure 1 This is a flowchart illustrating the method of an embodiment of the present invention;
[0090] Figure 2 This is a structural framework diagram of the system according to an embodiment of the present invention;
[0091] In the diagram, 10 is the client; 20 is the server; 101 is the feature sparse activation module; 102 is the parameter decoupling mask module; 103 is the sensitive feature extraction module; 104 is the orthogonal projection calculation module; 201 is the momentum tracking module; and 202 is the clustering secure aggregation module. Detailed Implementation
[0092] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings:
[0093] like Figure 1 As shown, this embodiment provides a robust adaptation and privacy protection method for federated learning based on a dual-end structure, constructing a federated framework including a server and multiple clients, including:
[0094] S1) The client, based on a global neural network model containing a multi-dimensional feature mapping layer, performs forward propagation computation locally by introducing a winner-takes-all competitive activation mechanism to generate hidden layer features with structural sparsity.
[0095] S2) The client performs backpropagation of error locally to generate a set of gradient tensors, and generates a binary mask based on the topological dimension or structural properties of the gradient tensor set; the client decouples the gradient tensor set into private gradients locked locally and shared gradients that are allowed to be uploaded according to the binary mask;
[0096] S3) Construct a private feature matrix based on the private gradient, extract the sensitive subspace representation of the private feature matrix, and construct an orthogonal projection matrix based on the sensitive subspace representation;
[0097] S4) Calculate the orthogonal complement projection of the shared gradient using the orthogonal projection matrix, remove the associated components parallel to the local sensitive features, obtain the orthogonal cleaned gradient, and upload it to the server.
[0098] S5) The server receives the orthogonal purification gradients uploaded by multiple clients, performs momentum tracking in the time dimension, and constructs a client state graph based on the L1 norm momentum distance.
[0099] S6) The server performs dimensionality reduction mapping on the client state graph, uses a clustering algorithm to block the updates of abnormal clusters, and only performs aggregate calculations on the momentum updates of normal clusters to iterate the global neural network model.
[0100] In this embodiment, in step S1), the client, based on a global neural network model containing a multi-dimensional feature mapping layer, performs a forward propagation computation locally that introduces a winner-takes-all competitive activation mechanism to generate structurally sparse hidden layer features, including:
[0101] S11) The client loads the global neural network model issued by the server;
[0102] S12) The client inputs local training samples into the global neural network model, performs forward propagation layer by layer, and calculates the pre-activation vectors of each hidden layer in sequence. The original preactivation vector of the hidden layer for:
[0103] ;
[0104] in, For the first The total number of neurons in the hidden layers; For the first Hidden layer The original pre-activation output value of each neuron.
[0105] S13) Pre-activation vector of the current hidden layer The original pre-activation output values of all neurons are sorted in descending order from largest to smallest to obtain the sorted neuron sequence;
[0106] S14) Before selection The neurons corresponding to the original pre-activation output values are used as candidate neurons. If the original pre-activation output value of a candidate neuron is greater than or equal to the dynamic threshold for competitive selection... If the corresponding candidate neuron is selected, then the candidate neuron is determined to be a valid candidate neuron.
[0107] In some embodiments, the competitive screening dynamic threshold ,in, This is the scaling factor; This represents the average value of the pre-activation vectors in the current feature mapping layer.
[0108] S15) Retain the original pre-activation output values for all valid candidate neurons in the current hidden layer, and uniformly assign the preset inactive values to the remaining neurons. This generates sparse hidden layer features for the current hidden layer.
[0109] S16) Traverse all hidden layers to obtain the hidden layer features of the structural sparsity of all hidden layers.
[0110] In this embodiment, in step S2), the client performs backpropagation of the error locally to generate a set of gradient tensors, and generates a binary mask based on the topological dimension or structural properties of the gradient tensor set; the client decouples the gradient tensor set into a private gradient locked locally and a shared gradient that is allowed to be uploaded, according to the binary mask, including:
[0111] S21) The client calculates the overall loss function for a single batch of training samples based on the prediction results of the sparse forward propagation of the global neural network model and the true labels of the samples. ,Right now:
[0112] ;
[0113] In the formula, For sample batch size, For mission losses; For the sample The prediction results; For the sample The true label;
[0114] S22) The overall loss function is calculated by taking the derivative layer by layer from the output layer to the input layer of the global neural network model through backpropagation. Gradient of parameters for each hidden layer The gradient tensor set is obtained by summing the gradients of all hidden layer parameters. ,Right now:
[0115] ;
[0116] ;
[0117] In the formula, For the first The full gradient set of each client in round t; Represents the overall loss function For the Hidden layer parameters Find the partial derivative; This represents the number of hidden layers. For the first The client number Round The gradient tensor corresponding to the layer parameters.
[0118] S23), for the gradient tensor set All gradient tensors are classified as follows: the gradient tensors of the bias terms and basic connection weight parameters of the global neural network model are treated as one-dimensional gradient tensors; the gradient tensors of the convolution weights and high-dimensional fully connected weight parameters are treated as multi-dimensional gradient tensors.
[0119] S24) Generate masks corresponding to the one-dimensional gradient tensor and multi-dimensional gradient tensor for classification, resulting in binary masks, including:
[0120] If the gradient tensor set contains gradient tensors If it is a one-dimensional gradient tensor, then an exemption mask is assigned to it; a binary mask is also assigned. ;
[0121] If the gradient tensors in the set of gradient tensors The multidimensional gradient tensor, which is determined to carry high-dimensional features, is based on the sparsity hyperparameter. Extract the core extremum components from the front of the gradient tensor magnitude sorting, and then sort the gradient tensors corresponding to the core extremum components. The binary mask is set to 1. The remaining positions are 0, that is:
[0122] ;
[0123] In the formula, The proportion of the absolute value of the gradient tensor before the magnitude. A collection of elements; Represents a multidimensional gradient tensor;
[0124] Traversing the gradient tensor set For all gradient tensors, generate corresponding masks layer by layer to obtain a mask set with the same structure as the gradient tensor set. ;
[0125] S25), based on mask set With gradient tensor set Gradient decoupling is achieved using the Hadamard product operation, resulting in private and shared gradients, including:
[0126] set of gradient tensors With mask set Element-wise multiplication yields a private gradient that is locked locally and prohibited from being uploaded. ;Right now:
[0127] ;
[0128] set of gradient tensors The set of masks with their inverses Element-wise multiplication yields the shared gradient that can be uploaded to the server. ,Right now:
[0129] ;
[0130] In the formula, It is the product of Hadama.
[0131] In this embodiment, step S3) involves constructing a private feature matrix based on the private gradient, extracting the sensitive subspace representation of the private feature matrix, and constructing an orthogonal projection matrix based on the sensitive subspace representation, including:
[0132] S31) Private gradient After being flattened into a one-dimensional vector, the vectors are concatenated according to their feature dimensions to form a private feature matrix. ,Right now:
[0133] ;
[0134] In the formula, For reshaping operation;
[0135] S32) The private feature matrix is reduced in dimensionality using truncated singular value decomposition to extract the principal component basis vector matrix representing the privacy information, including:
[0136] For private characteristic matrix Perform singular value decomposition, i.e.:
[0137] ;
[0138] In the formula, It is a left singular vector matrix; For a singular value diagonal matrix, the singular values are arranged in descending order of their diagonal elements. ; It is a right singular vector matrix, with each column corresponding to a set of characteristic bases; Indicates the transpose operation;
[0139] Set a truncation hyperparameter s, and extract the first s right singular vectors to form the principal component basis vector matrix of the privacy subspace. ,Right now:
[0140] ;
[0141] in, This represents the basis vector of the s-th principal component;
[0142] S33), Principal Component Basis Vector Matrix Based on Privacy Subspace Construct the orthogonal complement projection matrix, i.e.:
[0143] ;
[0144] In the formula, This is the orthogonal complementary projection matrix; It is the identity matrix; This is the projection matrix for the privacy subspace.
[0145] In this embodiment, step S4) involves using the orthogonal projection matrix to perform orthogonal complement projection calculation on the shared gradient, stripping away the associated components parallel to the local sensitive features, obtaining the orthogonal cleaned gradient, and uploading it to the server. This includes:
[0146] S41) Share gradient Perform vectorization preprocessing to obtain the vectorized shared gradient vector; that is:
[0147] ;
[0148] In the formula, This is the vectorized shared gradient vector; As a vectorization operator, it stretches a tensor of arbitrary size into a one-dimensional column vector by columns;
[0149] S42) The vectorized shared gradient vector Left multiplication by orthogonal complement projection matrix By stripping away the sensitive parallel components, orthogonal purification gradients are obtained. ,Right now:
[0150] ;
[0151] S43) Project the orthogonal purification gradient obtained from the projection. Uploaded to the server.
[0152] In this embodiment, in step S5), the server receives the orthogonal purification gradients uploaded by multiple clients, performs momentum tracking over the time dimension, and constructs a client state graph based on the L1 norm momentum distance, including:
[0153] S51) After the t-th round of communication, the server obtains the client's orthogonal purification gradient. Based on client-side orthogonal purification gradient Create a timing momentum buffer vector for each client. ,Right now:
[0154] ;
[0155] In the formula, For the client The timing momentum buffer vector in round t. ; Indicates the smoothing hyperparameter; For the client The orthogonal purification gradient in round t;
[0156] S52), calculate the L1 norm momentum distance between the temporal momentum buffer vectors of any two clients, i.e.:
[0157] ;
[0158] In the formula, For the client and L1 norm momentum distance; The dimension of the timing momentum buffer vector; , For clients and The timing momentum buffer vector is at the 1st The values in each dimension.
[0159] S53) Traverse all clients participating in federated training, calculate the L1 norm momentum distance between any two clients, and construct the client state graph. ,Right now:
[0160] ;
[0161] In the formula For the client and The L1 norm momentum distance ; The total number of clients participating in the federal training.
[0162] In this embodiment, in step S6), the server performs dimensionality reduction mapping on the client state graph, uses a clustering algorithm to block updates of abnormal clusters, and only performs aggregation calculations on the momentum updates of normal clusters to iterate the global neural network model, including:
[0163] S61) Calculate the similarity between any two clients in each group based on the L1 norm momentum distance of the client state graph, i.e.:
[0164] ;
[0165] In the formula, For the client and Similarity; The Laplacian kernel-scale hyperparameter;
[0166] S62), based on the similarity between any two clients Construct a degree matrix, which is a diagonal matrix, where each diagonal element of a row equals the sum of the similarities of all clients in that row, and all other elements are 0.
[0167]
[0168] In the formula, It is a degree matrix;
[0169] S63) Calculate the graph Laplacian matrix L corresponding to the undirected graph by subtracting the degree matrix and the similarity matrix, i.e.:
[0170] ;
[0171] In the formula, This is a similarity matrix;
[0172] S64) Perform eigenvalue decomposition on the graph Laplacian matrix L to obtain a diagonal matrix composed of ascending eigenvalues and its corresponding orthogonal eigenvector matrix. Select the eigenvectors matched by the first few smallest non-zero eigenvalues, concatenate them to form a low-dimensional mapping matrix, and then assign the timing momentum buffer vector corresponding to each client to the matrix. Transform it into a low-dimensional feature vector; that is:
[0173] ;
[0174] In the formula, It is an orthogonal eigenvector matrix, with each column corresponding to a set of eigenvectors; It is a diagonal eigenvalue matrix, with eigenvalues arranged in ascending order. ;
[0175] Select the eigenvectors corresponding to the k smallest non-zero eigenvalues and concatenate them to form a dimension-reduced mapping matrix. ,Right now: ;
[0176] Each client Mapping to k-dimensional low-dimensional eigenvectors, and taking the dimension reduction mapping matrix. No. Line as a client low-dimensional feature vectors ,Right now:
[0177] ;
[0178] S65), based on low-dimensional feature vectors The K-means unsupervised clustering algorithm is used to divide the clients into a normal client cluster and a malicious poisoning client cluster, including:
[0179] Randomly initialize the cluster centers of two clusters, calculate the distance between the low-dimensional feature vector of each client and the cluster centers of the two clusters one by one, and assign the client to the cluster with the closer distance;
[0180] After all clients are partitioned, the cluster centers are recalculated based on all clients within the two newly partitioned clusters.
[0181] Then, the distance between the low-dimensional feature vector of each client and the updated cluster center is calculated, and the clients are assigned to clusters that are closer in distance.
[0182] Repeat the iterations until the cluster centers no longer change significantly, thus completing the convergence of client groupings;
[0183] After grouping, the sample size and average momentum distance within the two clusters are compared, and the cluster with smaller sample size and greater average momentum distance within the cluster is identified as the malicious poisoning client cluster, while the other cluster is identified as the normal client cluster.
[0184] S66), the server utilizes the timing momentum buffer vector of all clients within the normal client cluster. Perform aggregate calculations to iterate the global neural network model.
[0185] like Figure 2 As shown, the second embodiment of this application provides a federated learning robust adaptation and privacy protection system based on a dual-end structure, including multiple clients 10 and a server 20;
[0186] Each of the client 10 is configured with a feature sparse activation module 101, a parameter decoupling mask module 102, a sensitive feature extraction module 103, and an orthogonal projection calculation module 104.
[0187] The server 20 is configured with a momentum tracking module 201 and a clustering security aggregation module 202;
[0188] The feature sparse activation module 101 is based on a global neural network model containing a multi-dimensional feature mapping layer. It performs forward propagation calculations locally by introducing a winner-take-all competitive activation mechanism to generate hidden layer features with structural sparsity.
[0189] The parameter decoupling mask module 102 performs error backpropagation locally to generate a gradient tensor set, and generates a binary mask based on the topological dimension or structural properties of the gradient tensor set; the client decouples the gradient tensor set into a private gradient locked locally and a shared gradient that is allowed to be uploaded, according to the binary mask.
[0190] The sensitive feature extraction module 103 constructs a private feature matrix based on the private gradient, extracts the sensitive subspace representation of the private feature matrix, and constructs an orthogonal projection matrix based on the sensitive subspace representation.
[0191] The orthogonal projection calculation module 104 uses the orthogonal projection matrix to perform orthogonal complement projection calculation on the shared gradient, strips out the associated components parallel to the local sensitive features, obtains the orthogonal cleaned gradient, and uploads it to the server.
[0192] The momentum tracking module 201 of the server 20 receives the orthogonal purification gradient uploaded by the orthogonal projection calculation modules of multiple clients, performs momentum tracking in the time dimension, and constructs a client state graph based on the L1 norm momentum distance.
[0193] The clustering security aggregation module 202 of the server 20 performs dimensionality reduction mapping on the client state graph, uses a clustering algorithm to block the updates of abnormal clusters, and only performs aggregation calculations on the momentum updates of normal clusters to iterate the global neural network model.
[0194] In this embodiment, the feature sparse activation module 101, based on a global neural network model containing a multi-dimensional feature mapping layer, performs forward propagation computation locally by introducing a winner-takes-all competitive activation mechanism to generate hidden layer features with structural sparsity, including:
[0195] The feature sparse activation module 101 of the client 10 loads the global neural network model issued by the server 20;
[0196] The feature sparse activation module 101 of the client 10 inputs local training samples into the global neural network model, performs forward propagation layer by layer, and calculates the pre-activation vectors of each hidden layer in sequence. The original preactivation vector of the hidden layer for:
[0197] ;
[0198] in, For the first The total number of neurons in the hidden layers; For the first Hidden layer The original pre-activation output value of each neuron.
[0199] Preactivation vector of the current hidden layer The original pre-activation output values of all neurons are sorted in descending order from largest to smallest to obtain the sorted neuron sequence;
[0200] Before selection The neurons corresponding to the original pre-activation output values are used as candidate neurons. If the original pre-activation output value of a candidate neuron is greater than or equal to the dynamic threshold for competitive selection... If the corresponding candidate neuron is selected, then the candidate neuron is determined to be a valid candidate neuron.
[0201] In some embodiments, the competitive screening dynamic threshold ,in, This is the scaling factor; This represents the average value of the pre-activation vectors in the current feature mapping layer.
[0202] For all valid candidate neurons in the current hidden layer, retain the original pre-activation output values, and assign the remaining neurons the preset inactive values. This generates sparse hidden layer features for the current hidden layer.
[0203] Traverse all hidden layers to obtain the hidden layer features of the structural sparsity of all hidden layers.
[0204] In this embodiment, the parameter decoupling mask module 102 of the client 10 performs backpropagation of error locally to generate a gradient tensor set, and generates a binary mask based on the topological dimension or structural properties of the gradient tensor set; the client decouples the gradient tensor set into a private gradient locked locally and a shared gradient that is allowed to be uploaded, according to the binary mask, including:
[0205] The parameter decoupling mask module 102 of the client 10 calculates the overall loss function of a single batch of training samples based on the prediction results of the sparse forward propagation of the global neural network model and the true labels of the samples. ,Right now:
[0206] ;
[0207] In the formula, For sample batch size, For mission losses; For the sample The prediction results; For the sample The true label;
[0208] The overall loss function is calculated by taking the derivative layer by layer from the output layer to the input layer of the global neural network model through backpropagation. Gradient of parameters for each hidden layer The gradient tensor set is obtained by summing the gradients of all hidden layer parameters. ,Right now:
[0209] ;
[0210] In the formula, For the first The full gradient set of each client in round t; Represents the overall loss function For the Hidden layer parameters Find the partial derivative; This represents the number of hidden layers. For the first The client number Round The gradient tensor corresponding to the layer parameters.
[0211] For gradient tensor set All gradient tensors are classified as follows: the gradient tensors of the bias terms and basic connection weight parameters of the global neural network model are treated as one-dimensional gradient tensors; the gradient tensors of the convolution weights and high-dimensional fully connected weight parameters are treated as multi-dimensional gradient tensors.
[0212] Generate masks corresponding to the one-dimensional and multi-dimensional gradient tensors for classification, resulting in binary masks, including:
[0213] If the gradient tensor set contains gradient tensors If it is a one-dimensional gradient tensor, then an exemption mask is assigned to it; a binary mask is also assigned. ;
[0214] If the gradient tensors in the set of gradient tensors The multidimensional gradient tensor, which is determined to carry high-dimensional features, is based on the sparsity hyperparameter. Extract the core extremum components from the front of the gradient tensor magnitude sorting, and then sort the gradient tensors corresponding to the core extremum components. The binary mask is set to 1. The remaining positions are 0, that is:
[0215] ;
[0216] In the formula, The proportion of the absolute value of the gradient tensor before the magnitude. A collection of elements; Represents a multidimensional gradient tensor;
[0217] Traversing the gradient tensor set For all gradient tensors, generate corresponding masks layer by layer to obtain a mask set with the same structure as the gradient tensor set. ;
[0218] Based on mask set With gradient tensor set Gradient decoupling is achieved using the Hadamard product operation, resulting in private and shared gradients, including:
[0219] set of gradient tensors With mask set Element-wise multiplication yields a private gradient that is locked locally and prohibited from being uploaded. ;Right now:
[0220] ;
[0221] set of gradient tensors The set of masks with their inverses Element-wise multiplication yields the shared gradient that can be uploaded to the server. ,Right now:
[0222] ;
[0223] In the formula, It is the product of Hadama.
[0224] In this embodiment, the sensitive feature extraction module 103 constructs a private feature matrix based on the private gradient, extracts the sensitive subspace representation of the private feature matrix, and constructs an orthogonal projection matrix based on the sensitive subspace representation, including:
[0225] Private gradient After being flattened into a one-dimensional vector, the vectors are concatenated according to their feature dimensions to form a private feature matrix. ;
[0226] Truncated singular value decomposition is used to reduce the dimensionality of the private feature matrix, and principal component basis vector matrices representing privacy information are extracted, including:
[0227] For private characteristic matrix Perform singular value decomposition, i.e.: ;
[0228] In the formula, It is a left singular vector matrix; For a singular value diagonal matrix, the singular values are arranged in descending order of their diagonal elements. ; It is a right singular vector matrix, with each column corresponding to a set of characteristic bases; Indicates the transpose operation;
[0229] Set a truncation hyperparameter s, and extract the first s right singular vectors to form the principal component basis vector matrix of the privacy subspace. ,Right now: ;
[0230] in, This represents the basis vector of the s-th principal component;
[0231] Principal component basis vector matrix based on privacy subspace Construct the orthogonal complement projection matrix, i.e.: ;
[0232] In the formula, This is the orthogonal complementary projection matrix; It is the identity matrix; This is the projection matrix for the privacy subspace.
[0233] In this embodiment, the orthogonal projection calculation module 104 uses the orthogonal projection matrix to perform orthogonal complement projection calculation on the shared gradient, strips off the associated components parallel to the local sensitive features, obtains the orthogonal purified gradient, and uploads it to the server 20, including:
[0234] The orthogonal projection calculation module 104 will share gradients. Perform vectorization preprocessing to obtain the vectorized shared gradient vector;
[0235] The orthogonal projection calculation module 104 will vectorize the shared gradient vector Left multiplication by orthogonal complement projection matrix By stripping away the sensitive parallel components, an orthogonal purification gradient is obtained.
[0236] The orthogonal projection calculation module 104 uploads the orthogonal purification gradient obtained from the projection to the server 20.
[0237] In this embodiment, the momentum tracking module 201 of the server 20 receives the orthogonal purification gradients uploaded by multiple clients, performs momentum tracking along the time dimension, and constructs a client state graph based on the L1 norm momentum distance, including:
[0238] After the t-th round of communication, the momentum tracking module 201 of the server 20 obtains the orthogonal purification gradient from the orthogonal projection calculation module 104 of the client 10. Orthogonal purification gradient based on client 10 Create a timing momentum buffer vector for each client. ,Right now:
[0239] ;
[0240] In the formula, For the client The timing momentum buffer vector in round t. ; Indicates the smoothing hyperparameter; For the client The orthogonal purification gradient in round t;
[0241] The momentum tracking module 201 calculates the L1 norm momentum distance between the temporal momentum buffer vectors of any two clients, i.e.:
[0242] ;
[0243] In the formula, For the client and L1 norm momentum distance; The dimension of the timing momentum buffer vector; , For clients and The timing momentum buffer vector is at the 1st The values in each dimension.
[0244] Iterate through all clients participating in the federated training, calculate the L1 norm momentum distance between any two clients, and construct the client state graph. .
[0245] In this embodiment, the clustering security aggregation module 202 of the server 20 performs dimensionality reduction mapping on the client state graph, uses a clustering algorithm to block updates of abnormal clusters, and only performs aggregation calculations on momentum updates of normal clusters to iterate the global neural network model, including:
[0246] The clustering security aggregation module 202 calculates the similarity between any two clients in each group based on the L1 norm momentum distance of the client state graph. ;in, For the client and Similarity; The Laplacian kernel-scale hyperparameter;
[0247] The clustering security aggregation module 202 is based on the similarity between any two clients. Construct a degree matrix, which is a diagonal matrix, where each diagonal element in each row is equal to the sum of the similarities of all clients in the corresponding row, and other elements are 0;
[0248] The clustering security aggregation module 202 utilization matrix and similarity matrix The difference operation calculates the graph Laplacian matrix of an undirected graph. ;
[0249] The clustering security aggregation module 202 performs eigenvalue decomposition on the graph Laplacian matrix L, obtaining a diagonal matrix composed of ascending eigenvalues and its corresponding orthogonal eigenvector matrix. It selects the eigenvectors matched by the first few smallest non-zero eigenvalues, concatenates them to form a low-dimensional mapping matrix, and then assigns the timing momentum buffer vector corresponding to each client to the matrix. Transform it into a low-dimensional feature vector; that is:
[0250] ;
[0251] In the formula, It is an orthogonal eigenvector matrix, with each column corresponding to a set of eigenvectors; It is a diagonal eigenvalue matrix, with eigenvalues arranged in ascending order. ;
[0252] Select the eigenvectors corresponding to the k smallest non-zero eigenvalues and concatenate them to form a dimension-reduced mapping matrix. ,Right now: ;
[0253] Each client Mapping to k-dimensional low-dimensional eigenvectors, and taking the dimension reduction mapping matrix. No. Line as a client low-dimensional feature vectors ;
[0254] The clustering security aggregation module 202 is based on low-dimensional feature vectors. The K-means unsupervised clustering algorithm is used to divide the clients into a normal client cluster and a malicious poisoning client cluster, including:
[0255] Randomly initialize the cluster centers of two clusters, calculate the distance between the low-dimensional feature vector of each client and the cluster centers of the two clusters one by one, and assign the client to the cluster with the closer distance;
[0256] After all clients are partitioned, the cluster centers are recalculated based on all clients within the two newly partitioned clusters.
[0257] Then, the distance between the low-dimensional feature vector of each client and the updated cluster center is calculated, and the clients are assigned to clusters that are closer in distance.
[0258] Repeat the iterations until the cluster centers no longer change significantly, thus completing the convergence of client groupings;
[0259] After grouping, the sample size and average momentum distance within the two clusters are compared, and the cluster with smaller sample size and greater average momentum distance within the cluster is identified as the malicious poisoning client cluster, while the other cluster is identified as the normal client cluster.
[0260] The clustering security aggregation module 202 of the server 20 utilizes the time momentum buffer vector of all clients within the normal client cluster. Perform aggregate calculations to iterate the global neural network model.
[0261] Embodiments of this application provide an electronic device, including:
[0262] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program, when executed by the at least one processor, implementing a federated learning robust adaptation and privacy protection method based on a dual-end structure.
[0263] Embodiments of this application provide a computer storage medium storing a computer program, which, when executed by a processor, implements the described robust adaptation and privacy protection method based on a dual-end structure federated learning.
[0264] The embodiments and descriptions above are merely illustrative of the principles and preferred embodiments of the present invention. Various changes and modifications may be made to the present invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed.
Claims
1. A robust adaptation and privacy protection method for federated learning based on a two-terminal structure, characterized in that, include: The client, based on a global neural network model containing multi-dimensional feature mapping layers, performs forward propagation computation locally by introducing a winner-take-all competitive activation mechanism to generate hidden layer features with structural sparsity. The client performs backpropagation of error locally to generate a set of gradient tensors, and generates a binary mask based on the topological dimension or structural properties of the gradient tensor set. The client decouples the gradient tensor set into a private gradient locked locally and a shared gradient that can be uploaded, based on the binary mask. A private feature matrix is constructed based on the private gradient, and the sensitive subspace representation of the private feature matrix is extracted. An orthogonal projection matrix is constructed based on the sensitive subspace representation. The shared gradient is calculated by orthogonal complement projection using the orthogonal projection matrix, and the associated components parallel to the local sensitive features are stripped to obtain the orthogonal purified gradient, which is then uploaded to the server. The server receives the orthogonal purification gradients uploaded by multiple clients, performs momentum tracking over the time dimension, and constructs a client state graph based on the L1 norm momentum distance. The server performs dimensionality reduction mapping on the client state graph, uses a clustering algorithm to block updates of abnormal clusters, and only performs aggregate calculations on the momentum updates of normal clusters to iterate the global neural network model.
2. The robust adaptation and privacy protection method for federated learning based on a dual-end structure according to claim 1, characterized in that: The client, based on a global neural network model containing multi-dimensional feature mapping layers, performs forward propagation computation locally by introducing a winner-takes-all competitive activation mechanism to generate structurally sparse hidden layer features, including: The client loads the global neural network model issued by the server; The client inputs local training samples into the global neural network model, performs forward propagation layer by layer, and calculates the preactivation vectors of each hidden layer in turn. Sort the original preactivation output values of all neurons in the preactivation vector of the current hidden layer in descending order to obtain the sorted neuron sequence; Before selection The neurons corresponding to the original pre-activation output values are used as candidate neurons. If the original pre-activation output value of a candidate neuron is greater than or equal to the dynamic threshold for competitive screening, the corresponding candidate neuron is determined to be a valid candidate neuron. For all valid candidate neurons in the current hidden layer, retain the original pre-activated output values, and assign the remaining neurons the preset inactive values to generate the sparse hidden layer features of the current hidden layer.
3. The robust adaptation and privacy protection method for federated learning based on a dual-end structure according to claim 2, characterized in that: The client performs backpropagation of the error locally to generate a set of gradient tensors, and generates a binary mask based on the topological dimension or structural properties of the gradient tensor set, including: The client calculates the overall loss function for a single batch of training samples based on the prediction results of the sparse forward propagation of the global neural network model and the true labels of the samples. By backpropagating from the output layer to the input layer of the global neural network model, the gradient of the overall loss function with respect to the parameters of each hidden layer is calculated, and the gradient tensor set is obtained by summing the gradients of all hidden layer parameters. We classify all gradient tensors in the gradient tensor set, treating the gradient tensors of the bias terms and basic connection weight parameters of the global neural network model as one-dimensional gradient tensors, and the gradient tensors of convolution weights and high-dimensional fully connected weight parameters as multi-dimensional gradient tensors. If the gradient tensor in the set of gradient tensors is a one-dimensional gradient tensor, then an exemption mask of 0 is assigned to it; If the gradient tensor in the gradient tensor set is determined to be a multidimensional gradient tensor carrying high-dimensional features, the core extreme value component at the front end of the gradient tensor magnitude sorting is extracted according to the sparsity hyperparameter, and the binary mask of the gradient tensor corresponding to the core extreme value component is set to 1, and the rest are 0. Traverse all gradient tensors in the gradient tensor set, generate corresponding masks layer by layer, and obtain a mask set with the same structure as the gradient tensor set.
4. The robust adaptation and privacy protection method for federated learning based on a dual-end structure according to claim 3, characterized in that: The client decouples the gradient tensor set into private gradients locked locally and shared gradients that are allowed to be uploaded, based on the binary mask, including: Based on the mask set and gradient tensor set, gradient decoupling is achieved using the Hadamard product operation, resulting in private and shared gradients, including: Multiply the gradient tensor set element by element with the mask set to obtain a private gradient that is locked locally and cannot be uploaded. Multiply the gradient tensor set element by element with the inverted mask set to obtain the shared gradient that can be uploaded to the server.
5. The robust adaptation and privacy protection method for federated learning based on a dual-end structure according to claim 4, characterized in that: A private feature matrix is constructed based on the private gradient, and a sensitive subspace representation of the private feature matrix is extracted. An orthogonal projection matrix is then constructed based on the sensitive subspace representation, including: After flattening the private gradient into a one-dimensional vector, the vectors are concatenated according to the feature dimensions to form a private feature matrix. Truncated singular value decomposition is used to reduce the dimensionality of the private feature matrix, and principal component basis vector matrices representing privacy information are extracted, including: For private characteristic matrix Perform singular value decomposition, i.e.: ; In the formula, It is a left singular vector matrix; For a singular value diagonal matrix, the singular values are arranged in descending order of their diagonal elements. ; It is a right singular vector matrix, with each column corresponding to a set of characteristic bases; Indicates the transpose operation; Set a truncation hyperparameter s, and extract the first s right singular vectors to form the principal component basis vector matrix of the privacy subspace. ,Right now: ; in, This represents the basis vector of the s-th principal component; Principal component basis vector matrix based on privacy subspace Construct the orthogonal complement projection matrix, i.e.: ; In the formula, This is the orthogonal complementary projection matrix; It is the identity matrix; This is the projection matrix for the privacy subspace.
6. The robust adaptation and privacy protection method for federated learning based on a dual-end structure according to claim 5, characterized in that: The shared gradient is orthogonally complemented by the orthogonal projection matrix, and associated components parallel to local sensitive features are removed to obtain the orthogonally purified gradient, which is then uploaded to the server. This process includes: The shared gradient is preprocessed into a vectorized form to obtain the vectorized shared gradient vector. Multiply the vectorized shared gradient vector by the orthogonal complementary projection matrix on the left, remove the sensitive parallel components, and obtain the orthogonal purified gradient; upload the orthogonal purified gradient obtained by projection to the server.
7. The robust adaptation and privacy protection method for federated learning based on a dual-end structure according to claim 6, characterized in that: The server receives orthogonal purification gradients uploaded by multiple clients, performs momentum tracking over the time dimension, and constructs a client state graph based on the L1 norm momentum distance, including: After the t-th round of communication, the server obtains the orthogonal purification gradient from the client. Based on client-side orthogonal purification gradient Create a timing momentum buffer vector for each client. ,Right now: ; In the formula, For the client The timing momentum buffer vector in round t. ; Indicates the smoothing hyperparameter; For the client The orthogonal purification gradient in round t; Calculate the L1 norm momentum distance between the temporal momentum buffer vectors of any two clients, i.e.: ; In the formula, For the client and L1 norm momentum distance; The dimension of the timing momentum buffer vector; , Clients respectively and The timing momentum buffer vector is at the 1st Values in each dimension; Iterate through all clients participating in the federated training, calculate the L1 norm momentum distance between any two clients, and construct the client state graph. ,Right now: ; In the formula For the client and The L1 norm momentum distance ; The total number of clients participating in the federal training.
8. The robust adaptation and privacy protection method for federated learning based on a dual-end structure according to claim 7, characterized in that: The server performs dimensionality reduction mapping on the client state graph, uses a clustering algorithm to block updates of abnormal clusters, and only performs aggregation calculations on momentum updates of normal clusters to iterate the global neural network model, including: The similarity between any two clients in each group is calculated based on the L1 norm momentum distance of the client state graph, i.e.: ; In the formula, For the client and Similarity; The Laplacian kernel-scale hyperparameter; Based on the similarity between any two clients Construct a degree matrix, which is a diagonal matrix, where each diagonal element in each row is equal to the sum of the similarities of all clients in the corresponding row, and other elements are 0; The graph Laplacian matrix corresponding to the undirected graph is calculated by subtracting the degree matrix from the similarity matrix. Perform eigenvalue decomposition on the graph Laplace moments to obtain a diagonal matrix composed of ascending eigenvalues and its corresponding orthogonal eigenvector matrix. Select the eigenvectors matched by the first few smallest non-zero eigenvalues, concatenate them to form a low-dimensional mapping matrix, and then assign the temporal momentum buffer vector corresponding to each client... Transform it into a low-dimensional feature vector; that is: ; In the formula, It is an orthogonal eigenvector matrix, with each column corresponding to a set of eigenvectors; It is a diagonal eigenvalue matrix, with eigenvalues arranged in ascending order. ; Select the eigenvectors corresponding to the k smallest non-zero eigenvalues and concatenate them to form a dimension-reduced mapping matrix. ,Right now: ; Each client Mapping to k-dimensional low-dimensional eigenvectors, and taking the dimension reduction mapping matrix. No. Line as a client low-dimensional feature vectors ,Right now: ; Based on low-dimensional feature vectors The K-means unsupervised clustering algorithm is used to divide the clients into a normal client cluster and a malicious poisoning client cluster, including: Randomly initialize the cluster centers of two clusters, calculate the distance between the low-dimensional feature vector of each client and the cluster centers of the two clusters one by one, and assign the client to the cluster with the closer distance; After all clients are partitioned, the cluster centers are recalculated based on all clients within the two newly partitioned clusters. Then, the distance between the low-dimensional feature vector of each client and the updated cluster center is calculated, and the clients are assigned to clusters that are closer in distance. Repeat the iterations until the cluster centers no longer change significantly, thus completing the convergence of client groupings; After grouping, the sample size of the two clusters and the average L1 norm momentum distance within the clusters are compared. The cluster with the smaller sample size and the greater average L1 norm momentum distance within the cluster is designated as the malicious poisoning client cluster, and the other cluster is designated as the normal client cluster. The server utilizes the timing momentum buffer vector of all clients within the normal client cluster. Perform aggregate calculations to iterate the global neural network model.
9. A robust adaptation and privacy protection system for federated learning based on a dual-end architecture, comprising multiple clients and servers; characterized in that: Each of the aforementioned clients is configured with a feature sparse activation module, a parameter decoupling mask module, a sensitive feature extraction module, and an orthogonal projection calculation module; The server is equipped with a momentum tracking module and a clustering secure aggregation module; The feature sparse activation module is based on a global neural network model containing a multi-dimensional feature mapping layer. It performs forward propagation calculations locally by introducing a winner-take-all competitive activation mechanism to generate hidden layer features with structural sparsity. The parameter decoupling mask module performs error backpropagation locally to generate a gradient tensor set, and generates a binary mask based on the topological dimension or structural properties of the gradient tensor set. The client decouples the gradient tensor set into a private gradient locked locally and a shared gradient that can be uploaded, based on the binary mask. The sensitive feature extraction module constructs a private feature matrix based on the private gradient, extracts the sensitive subspace representation of the private feature matrix, and constructs an orthogonal projection matrix based on the sensitive subspace representation. The orthogonal projection calculation module uses the orthogonal projection matrix to perform orthogonal complement projection calculation on the shared gradient, strips out the associated components parallel to the local sensitive features, obtains the orthogonal cleaned gradient, and uploads it to the server. The momentum tracking module of the server receives the orthogonal purification gradient uploaded by the orthogonal projection calculation modules of multiple clients, performs momentum tracking in the time dimension, and constructs a client state graph based on the L1 norm momentum distance. The clustering security aggregation module of the server performs dimensionality reduction mapping on the client state graph, uses a clustering algorithm to block the updates of abnormal clusters, and only performs aggregation calculations on the momentum updates of normal clusters to iterate the global neural network model.
10. An electronic device, comprising: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, wherein the computer program, when executed by the at least one processor, implements the federated learning robust adaptation and privacy protection method based on a dual-end structure as described in any one of claims 1-8.