A ubiquitous intelligent federated learning privacy protection system and method with cloud-edge-device collaboration

The ubiquitous intelligent federated learning system, which utilizes cloud-edge-device collaboration, employs masking and adaptive differential perturbation to address the balance between privacy protection and training efficiency in cloud-based and edge-based federated learning. This approach achieves high-efficiency privacy protection and model accuracy, making it suitable for ubiquitous intelligent systems.

CN115017541BActive Publication Date: 2025-12-02UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210634337.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-06
Publication Date
2025-12-02
Estimated Expiration
2042-06-06

AI Technical Summary

Technical Problem

Existing cloud and edge federated learning frameworks struggle to balance privacy protection and training efficiency, and are subject to gradient leakage attack risks and high communication costs.

Method used

A ubiquitous intelligent federated learning system with cloud-edge-device collaboration is adopted. By adding masks on terminal devices and performing adaptive differential perturbation on edge servers, combined with aggregation on the central cloud server, a lightweight mask transmission and adaptive differential privacy scheme is designed to protect model and data privacy.

Benefits of technology

While ensuring model accuracy and training efficiency, it achieves privacy protection for data and model parameters, improves the privacy and communication efficiency of ubiquitous intelligent systems, and is suitable for large-scale applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115017541B_ABST
    Figure CN115017541B_ABST
Patent Text Reader

Abstract

This invention discloses a cloud-edge-device collaborative ubiquitous intelligent federated learning privacy protection system and method. The system includes terminal devices, parameter servers, edge servers, and a central cloud server. The method includes: S1, setting a mask-based terminal training output protection mechanism; S2, adding adaptive differential perturbations to local models; S3, global model aggregation and adding adaptive differential perturbations. This invention proposes a lightweight privacy protection scheme. Partial model training is performed on the terminal device with matrix masking to ensure secure transmission between the terminal and the edge server. Furthermore, the remaining model training is performed on the edge server with differential perturbations added. After aggregation in the cloud, noise is added before feeding back to the edge server. Experimental results show that this scheme achieves 86% accuracy on the CIFAR10 dataset while ensuring privacy, effectively meeting the needs of ubiquitous intelligence. Therefore, this invention is highly suitable for large-scale application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a ubiquitous intelligent federated learning privacy protection system and method for cloud-edge-device collaboration. Background Technology

[0002] With the rapid development of artificial intelligence (AI), machine learning, supported by big data, has become a major force in ubiquitous intelligence. Machine learning trains large amounts of data through complex neural networks to obtain better classification and decision-making models. Among these, deep learning, combined with various fields, has promoted the development of ubiquitous intelligence. However, deep learning requires the support of big data, and the confidentiality of data makes data collection difficult, resulting in a lack of big data support for model training.

[0003] To address the data silo problem, some scholars have applied Google's distributed federated learning framework to model training, enabling model training while ensuring that user data remains on the user's local machine (Kairouz P, McMahan HB, Avent B, et al. Advances and open problems in federated learning[J]. Foundations and (in Machine Learning, 2021, 14(1–2): 1-210). Federated learning, coordinated by servers and terminal devices, can guarantee data privacy and improve the reliability of training.

[0004] However, existing federated learning frameworks also have some problems. In cloud-based federated learning, terminal equipment (TE) lacks sufficient memory and computing power (Khan LU, Saad W, Han Z, et al. Federated learning for internet of things: Recent advances, taxonomy, and open challenges[J]. IEEE Communications Surveys & Tutorials, 2021), which cannot support the training of complex neural network models. At the same time, a large amount of communication brings huge latency. To address this, some scholars have proposed applying the federated learning framework to mobile edge scenarios (Lim WYB, Luong NC, Hoang DT, et al. Federated learning in mobile edge networks: A comprehensive survey[J]. IEEE Communications Surveys & Tutorials, 2020, 22(3): 2031-2063), where the terminal equipment aggregates the trained model on an edge server (ES). While this approach can alleviate the pressure of insufficient bandwidth and reduce some transmission latency, it increases computational latency. To balance communication and computation latency, Wang R et al. proposed offloading model training tasks to edge servers closer to mobile devices (Wang R, Lai J, Zhang Z, et al. Privacy-Preserving Federated Learning for Internet of Medical Things under Edge Computing[J]. IEEE Journal of Biomedical and Health Informatics, 2022). Placing data closer to the user at the edge balances data localization and the limited computing resources of terminal devices. On the one hand, data is not transmitted through the core network, reducing the risk of attacks during transmission; on the other hand, the computing power of the edge server can support rapid model training and iterative updates.

[0005] Furthermore, existing research indicates that edge-based federated learning can achieve similar accuracy to cloud-based federated learning while offering lower latency, thus addressing bandwidth limitations inherent in cloud-based federated learning. However, gradient leakage attacks targeting edge-based federated learning can exploit gradients to deduce user data information. Additionally, training data on edge servers presents vulnerabilities during data transmission between terminal devices and edge servers, potentially leading to user privacy breaches.

[0006] To address the privacy breach issue, Zhang C et al. proposed a hierarchical federated learning homomorphic encryption method (Zhang C, Li S, Xia J, et al. {BatchCrypt}: Efficient Homomorphic Encryption for {Cross-Silo} Federated Learning[C] / / 2020USENIX Annual Technical Conference(USENIX ATC20).2020:493-506). This method homomorphically encrypts the data before training the federated learning. However, homomorphic encryption requires a significant amount of time and space, which is not feasible for ubiquitous intelligent systems that require high real-time performance. Kanagavelu R et al. proposed cryptographically-based multi-party computation (Kanagavelu R, Li Z, Samsudin J, et al. Two-phase multi-party computation enabled privacy-preserving federated learning[C] / / 2020 20th IEEE / ACM International Symposium on Cluster, Cloud and Internet Computing(CCGRID).IEEE,2020:410-419), but attackers can obtain data or gradients by acquiring the key, and multi-party computation is not suitable for distributed scenarios. Wei K et al. proposed a differential privacy scheme (Wei K, Li J, Ding M, et al. Federated learning with differential privacy: Algorithms and performance analysis[J].IEEE Transactions on Information Forensics and Security,2020,15:3454-3469), which adds noise to the model gradient or adds noise to the data locally. Although this method can protect the data well, the addition of noise will distort the data to some extent, which will reduce the accuracy of the model to a certain extent.

[0007] To balance security and accuracy, Wu X et al. proposed an adaptive differential privacy scheme (Wu X, Zhang Y, Shi M, et al. An adaptive federated learning scheme with differential privacy preserving[J]. Future Generation Computer Systems, 2022, 127: 362-372), which improves accuracy by adaptively pruning gradients while reducing throughput and latency. Yu R et al. proposed using the sensitivity of different layers of the network to compress the model and solve the problem of redundant weight parameters, so as to achieve a balance between model training efficiency and model complexity (Yu R, Li A, Chen CF, et al. Nisp: Pruning networks using neuron importance score propagation[C] / / Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018: 9194-9203). Li T et al. proposed using parameter sparsity to transmit parameters that are not zero after being ANDed with the mask, which can prevent model parameter leakage (Li T, Sahu AK, Talwalkar A, et al. Federated learning: Challenges, methods, and future directions[J].IEEE Signal Processing Magazine, 2020, 37(3):50-60).

[0008] While the above methods can protect user privacy to some extent, cloud-based federated learning incurs significant communication costs, and edge-based federated learning generally suffers from low training efficiency. Therefore, improved solutions are needed to reduce communication costs and enhance edge-based federated learning training efficiency while protecting user privacy. Summary of the Invention

[0009] To address the shortcomings of the existing technologies, this invention provides a ubiquitous intelligent federated learning privacy protection system and method with cloud-edge-device collaboration.

[0010] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0011] A ubiquitous intelligent federated learning privacy protection system with cloud-edge-device collaboration includes:

[0012] Terminal devices are used to generate personal mask vectors and components of the coefficient matrix, collect datasets and perform partial model training, and add masks to the model output.

[0013] The parameter server is responsible for key negotiation and distribution, collecting personal masks from terminal devices within a predetermined area, generating a public mask and coefficient matrix, and sending the public mask and coefficient matrix to the terminal devices and edge servers participating in the training.

[0014] Edge servers are used to obtain the model framework to be trained from the central cloud server, collect the output after adding the mask and remove the mask, then use the output of the terminal device as input to train the neural network, and differentially perturb the local model parameters obtained by training before sending them to the central cloud server for aggregation.

[0015] The central cloud server is used to aggregate the local model parameters trained by different edge servers, and then send the aggregated global model parameters to the edge servers for the next round of training.

[0016] Based on the above system, this invention also provides a privacy protection method for ubiquitous intelligent federated learning with cloud-edge-device collaboration. This method involves adding masks and adaptive differential perturbations to the cloud-edge-device collaborative federated learning process, thereby protecting the privacy of data and model parameters while ensuring model accuracy and training efficiency. The specific process is as follows:

[0017] S1. Set up a mask-based terminal training output protection mechanism:

[0018] S101 and the parameter server respectively negotiate symmetric key S with the terminal devices and edge servers participating in the training to obtain the symmetric key S. s ;

[0019] S102. Each terminal device generates a personal random mask and a personal coefficient vector, and sends them to the edge server after encrypting them with a symmetric key.

[0020] S103. After receiving the symmetric key, the edge server uses the coefficient vector group to generate a full-rank matrix A as the coefficient matrix.

[0021] S104, Calculate the inverse A of the coefficient matrix of the edge server. - The individual random masks of the corresponding terminal devices are added together to obtain the global public mask R, which is then symmetrically encrypted with the coefficient matrix and sent to each terminal device participating in the training.

[0022] S105. After receiving the ciphertext, the terminal device decrypts it to obtain the common mask and coefficient matrix, and then adds a mask to part of the model output during the training process.

[0023] S2. Add adaptive difference perturbation to the local model:

[0024] S201. The edge server downloads and initializes the model parameters from the central cloud server. The terminal device sends the masked partial model output to the edge server. After the edge server removes the mask, it uses the components of the partial model output as input to train the neural network and obtain a local model.

[0025] S202. Calculate the local model gradient values ​​and perform clipping;

[0026] S203. Update the local model parameters using the clipped local model gradient values;

[0027] S204. Add Gaussian noise to the updated local model parameters;

[0028] S205. Each edge server sends the local model parameters with added noise to the central cloud server.

[0029] S3, Global Model Aggregation and Adding Adaptive Differential Perturbation:

[0030] S301. After receiving the local model parameters from each edge server, the central cloud server aggregates them to obtain the global model parameters.

[0031] S302, Trim the global model parameters;

[0032] S303. After clipping, add Gaussian noise to the global model parameters and update the global model parameters.

[0033] S304. The central cloud server will distribute the updated global model parameters to the edge servers for the next round of training, until the model converges or reaches the required number of iterations.

[0034] Further, the specific process of step S102 is as follows: each terminal device d i Use the random number generator random.uniform() to generate a personal random mask r i , that r i It is an n-dimensional column vector with values ​​in [0,1], and r i Normalization yields the coefficient vector a i After encryption, r i and a i Send to the edge server; for r i The j-th component in the matrix; the normalization formula is as follows:

[0035]

[0036] In the formula, ||r i||2 is the 2-norm of the vector, and the traversal of r i For each component element in the vector, perform the above operations to obtain the normalized vector a. i .

[0037] Specifically, step S103 involves the following process: the edge server collects the personal mask r from the terminal device. i and coefficient vector a i And determine the coefficient vector a i If the coefficient vectors are linearly independent of the vector group in the coefficient matrix A, then use the append() function to add the coefficient vectors to the vector group, and simultaneously R = R + r. i Continue until rank(A) reaches N, obtaining a full-rank matrix A; otherwise, change the coefficient vector a. i Remove and re-collect personal mask r i and coefficient vector a i .

[0038] Specifically, in step S105, the partial model output Y is performed according to the following formula. i c By adding the mask, we obtain an n-order mask matrix M. i :

[0039] M i =A(Y i c +R)+r i

[0040] Where R and r i It will automatically expand into an n-dimensional vector.

[0041] Furthermore, in step S201, the mask is removed according to the following formula:

[0042] Y i c =A - (M i -r i )–R.

[0043] Specifically, in step S202, for the j-th edge server, the local model gradient value obtained in the t-th round of training is:

[0044]

[0045] In the formula, ω t is the weight parameter obtained after the j-th edge server in the t-th round of training; N is the total number of terminals uploading to the edge server; Loss is the loss function, and Among them, D i It is the training set of the i-th terminal device, and Di ={x1, x2, x3…x n}, The input data is x n The resulting model output; L i It is the tag set of the i-th terminal device;

[0046] Furthermore, the local model gradient values ​​are obtained. Then, calculate the Loss value pair. derivative And feed it back to the i-th terminal device, which then processes the local neural network parameters ω according to the following formula. c Update:

[0047]

[0048] In the formula, η is the learning rate.

[0049] Specifically, in step S202, the local gradient values ​​are clipped using the following formula:

[0050]

[0051] In the formula, C is the cropping threshold, which is a set constant;

[0052] In step S203, the local model parameters are updated using the following formula.

[0053]

[0054] In step S204, Gaussian noise is added using the following formula:

[0055]

[0056] In the formula, N(0,σ) 2 C 2 I) is a normalized distribution with a mean of 0; σ is the standard deviation of Gaussian noise, and Where ∈ represents the privacy budget, δ represents the confidence parameter, and Δf represents the global sensitivity.

[0057] Specifically, in step S301, the following formula is used to perform aggregation and obtain the global model parameter ω. t+1 :

[0058]

[0059] In the formula, is the weight of the j-th edge server; W is the sum of the weights of all edge servers participating in the training, i.e.:

[0060] Specifically, in step S302, the global model parameter ω is calculated using the following formula. t+1 Perform cropping and updating:

[0061]

[0062] In step S303, Gaussian noise is added and the global model parameters are updated using the following formula:

[0063]

[0064] Where σ1 is the standard deviation of Gaussian noise, and

[0065] Compared with the prior art, the present invention has the following beneficial effects:

[0066] (1) This invention designs a ubiquitous intelligent federated learning privacy protection system and method for cloud-edge-device collaboration. It adopts a lightweight masking scheme. After partial model training on the terminal device, the model is masked and transmitted to the edge server. This ensures the integrity and accuracy of the model while achieving a good protection effect.

[0067] (2) In addition to mask addition, the present invention also designs an adaptive differential privacy scheme for the edge server. The edge server performs adaptive pruning in each round of training using the pruning threshold and prior gradient, which reduces the amount of transmission and the error caused by noise.

[0068] (3) After collecting the local model parameters of each edge server, the present invention aggregates them through the central cloud server and then adds Gaussian noise. In this way, the central cloud server adjusts the variance of the noise using the pruning threshold, and calculates the parameter change value according to the weight value of each edge server, so as to protect privacy during feedback while ensuring that the model parameters are as accurate as possible.

[0069] (4) The cloud-edge-device collaborative ubiquitous intelligent federated learning privacy protection scheme designed in this invention has been shown by experimental results to achieve an accuracy of over 86% on the CIFAR10 dataset while ensuring privacy. This scheme effectively meets and promotes the development of ubiquitous intelligent systems. Therefore, this invention is very suitable for large-scale application. Attached Figure Description

[0070] Figure 1 This is a schematic diagram of the system framework of an embodiment of the present invention.

[0071] Figure 2 This is a graph showing the relationship between the accuracy of the model and the number of training rounds in an embodiment of the present invention.

[0072] Figure 3 This is a comparison chart of the loss function values ​​during model training in an embodiment of the present invention.

[0073] Figure 4 This is a comparison chart of model training accuracy on different devices in an embodiment of the present invention.

[0074] Figure 5 This is a comparison chart of model training time on different devices in an embodiment of the present invention.

[0075] Figure 6 This is a comparison chart of model training energy consumption on different devices in an embodiment of the present invention. Detailed Implementation

[0076] The present invention will be further described below with reference to the accompanying drawings and embodiments. The implementation of the present invention includes, but is not limited to, the following embodiments.

[0077] Example

[0078] This embodiment provides a cloud-edge-device collaborative ubiquitous intelligent federated learning privacy protection system, designed to meet the needs of ubiquitous intelligence. For example... Figure 1 As shown, this ubiquitous intelligent federated learning privacy protection system framework, which integrates cloud, edge, and device collaboration, mainly comprises four parts:

[0079] Terminal devices are used to generate personal mask vectors and components of coefficient matrices, collect datasets and perform partial model training, and add masks to the model output.

[0080] The parameter server is responsible for key negotiation and distribution, collecting personal masks from terminal devices within a predetermined area, generating a public mask and coefficient matrix, and sending the public mask and coefficient matrix to the terminal devices and edge servers participating in the training.

[0081] Edge servers are used to obtain the model framework to be trained from the central cloud server, collect the output after adding the mask and remove the mask, then use the output of the terminal device as input to train the neural network, and differentially perturb the local model parameters obtained by training before sending them to the central cloud server for aggregation.

[0082] The central cloud server is used to aggregate the local model parameters trained by different edge servers, and then send the aggregated global model parameters to the edge servers for the next round of training.

[0083] The implementation process of this privacy protection system will be described in detail below.

[0084] Problem definition:

[0085] First, when evaluating the model, the following relevant definitions are determined. By defining positive and negative classes and comparing them with basic facts, Table 1 can be obtained, and the accuracy can then be calculated.

[0086] Positive Class negative class True positive (TP) False positive (FP) False negative (FN) True negative (TN)

[0087] Table 1. Determination Table for Positive and Negative Class Results

[0088] Accuracy is defined as: Accuracy = (TP + TN) / (TP + FT + FN + TN) * 100%

[0089] For each edge server, let the total energy consumption be P, and the energy consumption when no training task is being performed be Pi. static If the CPU utilization is u, then the relationship between energy consumption and CPU utilization during model training is as follows:

[0090] P train =(PP static )(2u-u k )

[0091] Among them, through regression analysis, the energy consumption prediction accuracy is highest when k is 1.3.

[0092] Differential privacy: Given a dataset D and neighboring datasets D′, for a query function f, if the following expression is satisfied, then f satisfies differential privacy:

[0093] Pr[f(D)∈R]≤exp(∈)*Pr[f(D)∈R]+δ

[0094] Wherein, ∈ represents the privacy budget. The larger ∈ is, the higher the data availability. The smaller ∈ is, the higher the degree of privacy protection. Correspondingly, the more noise is added. When ∈ is 0, it means that no differential privacy is added. δ represents the confidence parameter. In strict differential privacy, δ is 0. When δ>0, it is approximate differential privacy.

[0095] Privacy budget ∈ controls the privacy level of function f for a query set S = {f1, f2, f3, ..., f n If each query function provides ∈ for disjoint subsets i Differential privacy means that for any subset, the query set S provides {max(∈1,∈2,...,∈...} i Differential privacy of n*∈. For a dataset D, if the query function in the query set S satisfies the ∈-differential privacy of sequential query, then the query set S provides differential privacy of n*∈

[14] .

[0096] Laplace mechanism: Given a dataset D and a query function f, the mechanism M providing differential privacy satisfies:

[0097]

[0098] Where Δf represents global sensitivity, and the calculation formula is as follows:

[0099] Δf=max||f(D)-f(D′)||

[0100] The maximum value of the norm of the difference between the query results of two adjacent datasets.

[0101] Exponential Mechanism: The exponential mechanism is mostly used for non-numerical queries. Given a dataset D, its output is R, and given a scoring function q(D,R), its sensitivity is calculated as follows:

[0102] Δq=||q(D,R)-q(D′,R)||

[0103] Index mechanism The probability output result R.

[0104] Gaussian mechanism: Compared to the Laplace mechanism, the Gaussian mechanism is more widely used because it achieves approximate differential privacy. The sensitivity calculation under the Gaussian mechanism is as follows:

[0105] Δf = max||f(D) - f(D′)||2

[0106] In contrast to the Laplace mechanism, the Gaussian mechanism uses the L2 norm as the sensitivity, and its implementation mechanism is as follows:

[0107] M(D)=f(D)+Y

[0108] Where Y is the added noise, and for any confidence parameter δ, the variance is... Y follows a sequence N ~ (0, σ). 2 ).

[0109] Based on the above problem definition, the entire privacy protection scheme is described below.

[0110] I. Setting up a mask-based terminal training output protection mechanism:

[0111] (1) The parameter server negotiates symmetric key S with the terminal devices and edge servers participating in the training, respectively. s .

[0112] (2) Each terminal device d i Use the random number generator random.uniform() to generate a personal random mask r i , that r i It is an n-dimensional column vector with values ​​in [0,1], and r iNormalization yields the coefficient vector a i After encryption, r i and a i Send to the edge server; for r i The j-th component in the matrix; the normalization formula is as follows:

[0113]

[0114] In the formula, ||r i ||2 is the 2-norm of the vector, and the traversal of r i For each component element in the vector, perform the above operations to obtain the normalized vector a. i .

[0115] (3) After receiving the symmetric key, the edge server determines the coefficient vector a. i If the coefficient vectors are linearly independent of the vector group in the coefficient matrix A, then use the append() function to add the coefficient vectors to the vector group, and simultaneously R = R + r. i Continue until rank(A) reaches N, obtaining a full-rank matrix A; otherwise, change the coefficient vector a. i Remove and re-collect personal mask r i and coefficient vector a i .

[0116] (4) The edge server calculates the inverse A of the coefficient matrix. - The global public mask R is obtained by adding the personal random masks of the corresponding terminal devices, and then symmetrically encrypting it with the coefficient matrix before sending it to each terminal device participating in the training.

[0117] (5) After receiving the ciphertext, the terminal device decrypts it to obtain the common mask and coefficient matrix, and then adds a mask to part of the model output during training. In this embodiment, the mask for part of the model output Y is added according to the following formula. i c By adding the mask, we obtain an n-order mask matrix M. i :

[0118] M i =A(Y i c +R)+r i

[0119] Where R and r i It will automatically expand into an n-dimensional vector. Since asymmetric encryption of the model requires significant memory and computing power, and symmetric encryption is easily cracked, this embodiment performs a masking operation on the model. Each terminal device outputs a portion of its model, Y. c After adding a public mask R, multiply the result by A on the left, and then add a personal mask r to the result.i This yields an n-order mask matrix M. i Therefore, only a lightweight symmetric key is needed to encrypt the mask and coefficient matrix, because even if an attacker obtains a portion of the mask or coefficient matrix, they will not be able to reverse engineer the model and data information.

[0120] II. Add adaptive difference perturbation to the local model:

[0121] The edge server downloads and initializes model parameters from the central cloud server. The terminal device sends a masked portion of the model output to the edge server, which then removes the mask and uses this portion of the model output as input to train the neural network, resulting in a local model. After the edge server obtains the local model parameters, they need to be sent to the central cloud server for aggregation. To ensure the privacy of the model parameters, this embodiment adds adaptive differential perturbations.

[0122] In this embodiment, the partial output of the input model is: The edge server removes the mask using the following formula:

[0123] Y i c =A - (M i -r i )–R.

[0124] First, calculate the local model gradient value. For the j-th edge server, the local model gradient value obtained in the t-th round of training is:

[0125]

[0126] In the formula, ω t is the weight parameter obtained after the j-th edge server in the t-th round of training; N is the total number of terminals uploading to the edge server; Loss is the loss function, and Among them, D i It is the training set of the i-th terminal device, and D i ={x1, x2, x3…x n}, The input data is x n The resulting model output; L i It is the tag set of the i-th terminal device.

[0127] Then, the gradient values ​​of the local model are clipped:

[0128]

[0129] In the formula, the cropping threshold C>1.

[0130] After cutting, use Update local model parameters:

[0131]

[0132] In the formula, η is the learning rate. In this embodiment, the learning rate η is set to 0.01 at the beginning of training, and the learning rate decreases by more than 100 times near the end of training.

[0133] At the same time, calculate the Loss value for Y i c derivative

[0134]

[0135] Then Feedback is sent to the i-th terminal device, which then processes the local neural network parameters ω according to the following formula. c Update:

[0136]

[0137] Add Gaussian noise to the updated local model parameters:

[0138]

[0139] In the formula, N(0,σ) 2 C 2 I) is a normalized distribution with a mean of 0; σ is the standard deviation of Gaussian noise, and

[0140] Finally, each edge server sends the local model parameters with added noise to the central cloud server.

[0141] III. Global model aggregation and adding adaptive difference perturbations:

[0142] (1) After receiving the local model parameters from each edge server, the central cloud server aggregates them to obtain the global model parameters. In this embodiment, the following formula is used to aggregate and obtain the global model parameters ω. t+1 :

[0143]

[0144] In the formula, It is the weight of the j-th edge server.

[0145] (2) Pruning global model parameters:

[0146]

[0147] (3) After clipping, add Gaussian noise to the global model parameters and update the global model parameters. The update formula is as follows:

[0148]

[0149] Where σ1 is the standard deviation of Gaussian noise, and

[0150] (4) Send the updated global model parameters to the edge server for the next round of training until the model converges or the number of iterations is reached.

[0151] Thus, this embodiment has achieved privacy protection for data and model parameters while ensuring model accuracy and training efficiency.

[0152] The following is a theoretical analysis of the design of the above process.

[0153] I. Privacy Protection of Terminal Devices

[0154] Based on the above mask addition, it can be seen that when the attacker obtains mask M... i To obtain the model from the mask, a personal mask ri, a public mask R, and a coefficient matrix A are needed. However, all three parameters are transmitted after being symmetrically encrypted, so even if an attacker decrypts and obtains one of the parameters, they will not be able to successfully decode the model.

[0155] During transmission, based on the number of transmissions, the public mask R and the coefficient matrix A are the easiest to decrypt, but the personal mask is not easy to decrypt. When an attacker obtains R and A, he can calculate A. - Simultaneously, the spurious model output matrix Y is obtained by decoding according to the following formula. i c :

[0156] Y i c =AM i -R

[0157] From the decoding formula:

[0158] Y i c =A - (M i -r i )-R

[0159] The true Y can be obtained i c By comparing the two datasets, we can see that ΔY i c =A - r iBecause of the randomness of the coefficient matrix and the personal mask, this error is random and cannot be compensated for by fixed values. Therefore, even if an attacker obtains the mask and some parameters, they cannot deduce the accurate model parameters.

[0160] II. Privacy Protection of Edge Server Model

[0161] (1) Model privacy analysis

[0162] Due to the heterogeneity of the data, this invention adds Gaussian noise to the model gradient to achieve a relaxed differential privacy mechanism. By the definition of relaxed differential privacy, Therefore, on the edge server side, Simultaneously, the pruning threshold C is set to be greater than 1, and Cσ is used as the Gaussian noise value for the edge server. This satisfies relaxed differential privacy and allows for the addition of different Gaussian noise by changing the pruning threshold. For each data point, after calculating the gradient, the edge server prunes the data according to the pruning threshold and prior gradient values. The existence of prior knowledge and the limitation of the pruning threshold enable adaptive gradient pruning. Furthermore, for each training round, the edge server iterates locally, adding Gaussian noise to the local model after multiple iterations to achieve differential privacy. This improves the efficiency of model training and reduces the communication overhead between the edge server and the central cloud server.

[0163] (2) Privacy proof

[0164] Prove: Pr[f(x)+x∈R]≤exp(∈)*Pr[f(x)+x∈R]+δ, where x~N(0,σ 2 ),

[0165] Pr[f(x)+x∈R]=Pr[f(x)+x∈R1]+Pr[f(x)+x∈R2]

[0166] Divide Pr[f(x)+x∈R] into two parts. The first part strictly follows ∈-pure differential privacy, and the second part violates differential privacy. The second part should be less than δ.

[0167] So there is: Pr[f(x)+x∈R1]+Pr[f(x)+x∈R2]≤Pr[f(x)+x∈R1]+δ≤exp(∈)*Pr[f(y)+x∈R1]+δ.

[0168] For the first part, by

[0169] Since the probability is always greater than 0,

[0170]

[0171] Solving for:

[0172]

[0173] Will Substituting, we get:

[0174]

[0175] For the second part, the probability of violating differential privacy is less than δ, which can be derived from symmetry:

[0176]

[0177] Since x follows a Gaussian distribution, we have:

[0178]

[0179]

[0180] make Substituting, we get:

[0181]

[0182] Since the privacy budget ∈ <1 and c>1, we get

[0183]

[0184] Therefore:

[0185]

[0186] Substituting back to σ, we get:

[0187] If the clipping threshold C is greater than 1, then:

[0188]

[0189] Satisfy the relaxed differential privacy mechanism.

[0190] (3) Privacy protection of the central cloud server

[0191] Set Gaussian noise on the central cloud server. Where C is the pruning threshold, N is the number of edge servers participating in training, and W is the sum of the weights of all edge servers participating in training. The noise condition for relaxed differential privacy satisfies...

[0192] In this way, by adding Gaussian noise to the edge server and the central cloud server respectively to achieve differential privacy for the cloud and the edge, the leakage of gradients can be well prevented.

[0193] Proof:

[0194]

[0195] Since δ is a very small value, usually take So there is:

[0196]

[0197] Also because lnx < x, so there is:

[0198] Because C > 1, W ≤ 1, only need to set at initialization to satisfy differential privacy.

[0199] To sum up, in theoretical analysis, the scheme designed by the present invention is feasible.

[0200] Next, the performance of the scheme of the present invention in terms of model accuracy, loss function value, training efficiency and energy consumption is experimentally verified.

[0201] I. Experimental configuration

[0202] In terms of hardware, a Raspberry Pi is used as the terminal device and a PC is used as the edge server. In terms of software, each module is written in the Python language, and the accuracy, training efficiency and other performances of the model are obtained through training and evaluation using the resnet18 residual network on the CIFAR10 dataset. At the same time, three schemes, namely FedAvg and Transfer by Layer Sensitivity (TLS) and Parameter Sparse (PS) based on FedAvg, are used for comparison. In addition, in the experiment, different clipping thresholds C are also set to compare the accuracy and efficiency.

[0203] II. Accuracy analysis <​​​In UIFLPP, by adding a mask, partial model output can be accurately transmitted from the terminal device to the edge server. After model training on the edge server, the model is adaptively pruned and noise is added. When adding noise, the pruning threshold is added to the standard deviation to fine-tune the noise. When noise has an excessive impact on model accuracy, the pruning threshold C is reduced to increase the accuracy of model training. At the same time, as C decreases, the pruning degree increases, resulting in increased pruning degree in local models. Due to the existence of prior gradients, the pruning degree will adaptively decrease in the next round of pruning. By adjusting the pruning threshold C, a balance can be achieved between accuracy, noise, and pruning degree.

[0206] Model training and evaluation were performed on the CIFAR10 dataset, and the relationship between accuracy and training epochs is as follows: Figure 2 As shown.

[0207] from Figure 2 As can be seen, at the beginning of training, the accuracy of UIFLPP, PS, and TLS was not high, all below 20%, with FedAvg having the highest accuracy. After 15 rounds of training, the accuracy of UIFLPP, PS, and TLS increased rapidly, approaching 80%, all higher than FedAvg, and then increased slowly.

[0208] It can be seen that PS exhibits good stability with almost no jitter, while TLS has more inflection points. Although the model achieves good accuracy with increasing training epochs, it still suffers from significant jitter. Comparing the two, UIFLPP is very unstable in the early stages of training. Due to noise, the model's accuracy is inconsistent. However, the jitter gradually decreases with increasing training epochs, reaching 86% accuracy by the 60th epoch, surpassing TLS's 85% and PS's 84.5%. Furthermore, while TLS and PS can hide model gradients, their accuracy at convergence is not high.

[0209] Figure 3 It is the loss function value during model training, from Figure 3 As can be seen, after the first 10 training rounds, the loss values ​​of UIFLPP, PS, and FedAvg all drop below 1, while TLS remains around 1. From the 10th to the 50th training round, TLS has a higher loss value, while PS and UIFLPP have lower values, although UIFLPP's loss is slightly higher than PS's due to noise. With increasing training rounds, PS and UIFLPP converge around the 50th round, while TLS requires nearly 60 rounds to converge. This demonstrates that UIFLPP can achieve high accuracy with increasing training rounds while effectively protecting gradients. Furthermore, with the masking of the dataset, UIFLPP can achieve a good model while maintaining privacy.

[0210] To verify the stability of this scheme, 10 PCs were used as edge servers for model training, and the resulting accuracy was compared to... Figure 4 As shown, from Figure 4 As can be seen, on devices 3, 4, 5, and 8, the accuracy of UIFLPP is comparable to that of FedAvg. On other devices, the accuracy at convergence is higher than that of PS, TLS, and FedAvg. The accuracy achieved during device training can reach over 85%. This shows that UIFLPP achieves good model accuracy while also exhibiting good stability.

[0211] III. Efficiency Analysis

[0212] In terms of communication, UIFLPP uses symmetric encryption and matrix transmission between terminal devices and edge servers. It does not require the transmission of large datasets, requires very little bandwidth, and has low throughput. Therefore, UIFLPP can guarantee real-time performance and is suitable for scenarios where ubiquitous intelligence requires a large number of terminal devices.

[0213] To accurately assess efficiency, the time from the first to the 20th training round was used as the efficiency evaluation metric. At the end of the 20th training round, all four methods achieved an accuracy rate of 80%. We recorded the training time. Specific times are shown in Table 2.

[0214] Time UIFLPP FedAvg PS TLS Total time 26304 27720 17520 17580 Average round time 430 462 290 292

[0215] Table 2

[0216] As shown in Table 2, the time taken by UIFLPP is roughly the same as that of FedAvg, and slightly shorter overall, but significantly longer than that taken by TLS and PS. This is because TLS only uploads a portion of the model, and PS uploads a sparse parameter matrix, thus reducing both model training and transmission times. However, UIFLPP still requires model processing, increasing the time per iteration. Nevertheless, UIFLPP features adaptive pruning, which reduces model transmission time, resulting in a slightly shorter total time compared to FedAvg.

[0217] Generally, PS has the shortest training time, but TLS achieves time compression by reducing model accuracy. UIFLPP has a shorter training time than FedAvg and offers privacy protection. Furthermore, UIFLPP has higher accuracy than PS, making it the best overall performer.

[0218] In addition, we compared model training times on different devices. From Figure 5 As can be seen, the convergence time to obtain a good model ranges from 21,000 to 27,500 seconds. This indicates that the convergence time of the model does not differ significantly across different devices, and UIFLPP can adapt to various devices.

[0219] IV. Energy Consumption Analysis

[0220] Calculate the energy consumption during model training based on the computational expression. Figure 6 As can be seen, on devices 8, 9, and 10, UIFLPP consumes less power than FedAvg, with training power consumption around 0.65P. On other devices, while UIFLPP consumes more power than FedAvg, for example, training power consumption on devices 1, 2, and 3 is approximately 0.75P, it is mostly lower than TLS except for devices 4 and 7. Furthermore, UIFLPP consumes slightly more power than PS on devices 1, 2, and 3, while its power consumption is lower than PS on devices 8, 9, and 10. Therefore, UIFLPP consumes less power than TLS, and compared to PS and FedAvg, it depends on the specific device. The average power consumption of UIFLPP is approximately 0.69P, and the average power consumption of PS is also approximately 0.69P. Additionally, FedAvg consumes approximately 0.67P, and TLS consumes 0.7P. However, FedAvg lacks a privacy protection scheme, and compared to PS and TLS, the power consumption of UIFLPP is acceptable.

[0221] However, the four schemes do not differ significantly in energy consumption. When the terminal device or edge server is under sufficient load, UIFLPP is the preferred choice to achieve a balance between model accuracy and privacy protection.

[0222] In summary, compared with existing model parameter pruning privacy protection schemes such as PS, FedAvg, and TLS, UIFLPP achieves higher accuracy and is more suitable for ubiquitous intelligent applications. Therefore, compared with existing technologies, this invention has outstanding substantive features and significant progress.

[0223] The above embodiments are merely preferred embodiments of the present invention and should not be used to limit the scope of protection of the present invention. Any modifications or refinements made to the main design concept and spirit of the present invention that are not of substantial significance, but which still solve the same technical problem as the present invention, should be included within the scope of protection of the present invention.

Claims

1. A privacy-preserving method for ubiquitous intelligent federated learning with cloud-edge-device collaboration, characterized in that, It utilizes cloud-edge-device collaborative federated learning, employing masking and adaptive differential perturbation to protect the privacy of data and model parameters while ensuring model accuracy and training efficiency. The specific process is as follows: S1. Set up a mask-based terminal training output protection mechanism: S101 and the parameter server respectively negotiate symmetric key S with the terminal devices and edge servers participating in the training to obtain the symmetric key S. s ; S102. Each terminal device generates a personal random mask and a personal coefficient vector, and sends them to the edge server after encrypting them with a symmetric key. S103. After receiving the symmetric key, the edge server uses the coefficient vector group to generate a full-rank matrix A as the coefficient matrix. S104, Calculate the inverse A of the coefficient matrix of the edge server. - The individual random masks of the corresponding terminal devices are added together to obtain the global public mask R, which is then symmetrically encrypted with the coefficient matrix and sent to each terminal device participating in the training. S105. After receiving the ciphertext, the terminal device decrypts it to obtain the common mask and coefficient matrix, and then adds a mask to part of the model output during the training process. S2. Add adaptive difference perturbation to the local model: S201. The edge server downloads and initializes the model parameters from the central cloud server. The terminal device sends the masked partial model output to the edge server. After the edge server removes the mask, it uses the components of the partial model output as input to train the neural network and obtain a local model. S202. Calculate the local model gradient values ​​and perform clipping; S203. Update the local model parameters using the clipped local model gradient values; S204. Add Gaussian noise to the updated local model parameters; S205. Each edge server sends the local model parameters with added noise to the central cloud server. S3, Global Model Aggregation and Adding Adaptive Differential Perturbation: S301. After receiving the local model parameters from each edge server, the central cloud server aggregates them to obtain the global model parameters. S302, Trim the global model parameters; S303. After clipping, add Gaussian noise to the global model parameters and update the global model parameters. S304. The central cloud server will distribute the updated global model parameters to the edge servers for the next round of training, until the model converges or reaches the required number of iterations.

2. The privacy protection method for ubiquitous intelligent federated learning with cloud-edge-device collaboration according to claim 1, characterized in that, The specific process of S102 is as follows: Each terminal device d i Use the random number generator random.uniform() to generate a personal random mask r i , that r i It is an n-dimensional column vector with values ​​in [0,1], and r i Normalization yields the coefficient vector a i After encryption, r i and a i Send to the edge server; for r The j-th component in i; The normalization formula is as follows: In the formula, ||r i ||2 is the 2-norm of the vector, and the traversal of r i For each component element in the vector, perform the above operations to obtain the normalized vector a. i .

3. The privacy protection method for ubiquitous intelligent federated learning with cloud-edge-device collaboration according to claim 2, characterized in that, The specific process of S103 is as follows: The edge server collects the personal mask r from the terminal device. i and coefficient vector a i And determine the coefficient vector a i If the coefficient vectors are linearly independent of the vector group in the coefficient matrix A, then use the append() function to add the coefficient vectors to the vector group, and simultaneously R = R + r. i Continue until rank(A) reaches N, at which point a full-rank matrix A is obtained; No, then the coefficient vector a i Remove and re-collect personal mask r i and coefficient vector a i .

4. The privacy protection method for ubiquitous intelligent federated learning with cloud-edge-device collaboration according to claim 3, characterized in that, In step S105, the partial model output Y is performed according to the following formula. i c By adding the mask, we obtain an n-order mask matrix M. i : M i =A(Y i c +R)+r i Where R and r i It will automatically expand into an n-dimensional vector.

5. A privacy protection method for ubiquitous intelligent federated learning with cloud-edge-device collaboration according to claim 4, characterized in that, In step S201, the mask is removed according to the following formula: Y i c =A - (M i -r i )–R。 6. A privacy protection method for ubiquitous intelligent federated learning with cloud-edge-device collaboration according to any one of claims 1 to 5, characterized in that, In step S202, for the j-th edge server, the local model gradient value obtained in the t-th round of training is: In the formula, ω represents the weight parameter, ω t is the weight parameter obtained after the j-th edge server in the t-th round of training; N is the total number of terminals uploading to the edge server; Loss is the loss function, and Among them, D i It is the training set of the i-th terminal device, and D i ={x1, x2, x3…x n }, The input data is x n The resulting model output; L i It is the tag set of the i-th terminal device; Furthermore, the local model gradient values ​​are obtained. Then, calculate the Loss value against Y. i c derivative And feed it back to the i-th terminal device, which then processes the local neural network parameters ω according to the following formula. c Update: In the formula, η is the learning rate.

7. A privacy protection method for ubiquitous intelligent federated learning with cloud-edge-device collaboration according to claim 6, characterized in that, In step S202, the local gradient values ​​are clipped using the following formula: In the formula, C is the cropping threshold, which is a set constant; In step S203, the local model parameters are updated using the following formula. In step S204, Gaussian noise is added using the following formula: In the formula, N(0,σ) 2 C 2 I) is a normalized distribution with a mean of 0; σ is the standard deviation of Gaussian noise, and Where ∈ represents the privacy budget, δ represents the confidence parameter, Δf represents the global sensitivity, and I represents a unit vector with the same dimension as the model parameters.

8. A privacy protection method for ubiquitous intelligent federated learning with cloud-edge-device collaboration according to claim 7, characterized in that, In step S301, the following formula is used to perform aggregation and obtain the global model parameter ω. t+1 : In the formula, It is the weight of the j-th edge server; W is the sum of the weights of all edge servers participating in the training, that is:

9. A privacy protection method for ubiquitous intelligent federated learning with cloud-edge-device collaboration according to claim 8, characterized in that, In step S302, the following formula is used to calculate the global model parameter ω. t+1 Perform cropping and updating: In step S303, Gaussian noise is added and the global model parameters are updated using the following formula: Where σ1 is the standard deviation of Gaussian noise, and 10. A system for implementing the cloud-edge-device collaborative ubiquitous intelligent federated learning privacy protection method according to any one of claims 1 to 9, characterized in that, include: Terminal devices are used to generate personal mask vectors and components of coefficient matrices, collect datasets and perform partial model training, and add masks to the model output. The parameter server is responsible for key negotiation and distribution, collecting personal masks from terminal devices within a predetermined area, generating a public mask and coefficient matrix, and sending the public mask and coefficient matrix to the terminal devices and edge servers participating in the training. Edge servers are used to obtain the model framework to be trained from the central cloud server, collect the output after adding the mask and remove the mask, then use the output of the terminal device as input to train the neural network, and differentially perturb the local model parameters obtained by training before sending them to the central cloud server for aggregation. The central cloud server is used to aggregate the local model parameters trained by different edge servers, and then send the aggregated global model parameters to the edge servers for the next round of training.