Federal learning model training method, system and device and storage medium

By sub-data division and sensitivity level determination of local data in federated learning, and adding hierarchical noise to gradient information on the client, the problem of data privacy in federated learning cannot be effectively guaranteed, achieving stronger privacy protection and data security.

CN120146157AInactive Publication Date: 2025-06-13GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510140399.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In federated learning, although local training and gradient calculations reduce the risk of data breaches, attackers can still infer sensitive information by analyzing gradients, resulting in data privacy being unable to be effectively protected.

Method used

By dividing the local data into several sub-data and determining their sensitivity level, the local model parameters are initialized on the client and trained in batches. After obtaining the gradient information, hierarchical noise is added to the gradient information according to the sensitivity level and model level, model parameters are updated, and the adjusted parameters are uploaded to the central server for aggregation.

Benefits of technology

This method effectively improves data security, reduces the risk of privacy leakage, and reduces communication costs, ensuring that users' sensitive information will not be leaked during the data sharing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146157A_ABST
    Figure CN120146157A_ABST
Patent Text Reader

Abstract

The invention discloses a federal learning model training method, system and device and a storage medium. Local data D is divided into a plurality of sub-data di, and a corresponding sensitivity level is determined; training a local model by using the sub-data di in batches to obtain gradient information of the local model in the current batch, the gradient information comprising a gradient value of each layer of the local model; according to the sensitivity grade of the sub-data di and the layer pole of the local model, layered noise is added to gradient information of the local model, and adjustment gradient information of the current batch is obtained; the client updates the model parameters of the local model by using the adjustment gradient information to obtain the model update parameters of the current batch; the client uploads the model updating parameters of the current batch to the central server; and the central server forms global model parameters of the current batch based on the model updating parameters of all the clients and distributes the global model parameters to each client for local training of the next batch. And the data security can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly to a method, system, device, and storage medium for training a federated learning model. Background Art

[0002] In the current era of rapid development of information technology, the network security of colleges and universities has encountered huge challenges. Nowadays, online learning and digital management are widely carried out, and a large amount of sensitive data, including students' personal information, academic achievements, and financial information, is frequently transmitted within the campus network. This undoubtedly gives hackers an opportunity, resulting in a sharp increase in the network security risks of colleges and universities. To protect data privacy, each college and university has started to use federated learning to analyze data and train models locally to ensure the protection of data privacy. The personal information of students and faculty does not need to be centrally stored, which can reduce the risk of data leakage. Although the updates sent after the client trains the model locally and calculates the gradients do not contain the original data, attackers can still infer sensitive information by analyzing the gradients. Therefore, there is an urgent need for a method that can better protect data privacy. Summary of the Invention

[0003] For this reason, the embodiments of the present application provide a method, system, device, and storage medium for training a federated learning model, which can effectively improve the security of data.

[0004] In a first aspect, the present application provides a method for training a federated learning model.

[0005] The present application is achieved through the following technical solutions:

[0006] A method for training a federated learning model includes:

[0007] Dividing the local data D into several sub-data d i , and determining the sensitivity level of each sub-data d i ;

[0008] Initializing the model parameters of the local model on multiple clients, where the model parameters include the model parameters of each layer of the local model;

[0009] Training the local model in batches using the sub-data d i to obtain the gradient information of the local model in the current batch, where the gradient information includes the gradient values of each layer of the local model;

[0010] According to the sensitivity level of the sub-data d i and the layer level of the local model, adding hierarchical noise to the gradient information of the local model to obtain the adjusted gradient information of the current batch, where the adjusted gradient information includes the adjusted gradient values of each layer of the local model;

[0011] The client updates the model parameters of the local model using the adjusted gradient information to obtain the model update parameters for the current batch;

[0012] Each client uploads the model update parameters for the current batch to the central server;

[0013] The central server aggregates the model update parameters of all clients to form the global model parameters for the current batch;

[0014] The central server distributes the global model parameters for the current batch to each client for each client to use the global model parameters for the next batch of local training.

[0015] In a preferred example of the present application, it can be further set that,

[0016] Divide the local data D into several sub - data d i , and determine the sensitivity level of each sub - data d i , including:

[0017] Analyze each sub - data d i to determine the variance of each sub - data d i ;

[0018] Based on the variance of each sub - data d i to determine the sensitivity level of all sub - data d i .

[0019] In a preferred example of the present application, it can be further set that, use the sub - data d i to train the local model batch by batch to obtain the gradient information of the local model in the current batch. The gradient information includes the gradient value of each layer of the local model, and further includes:

[0020] Judge whether the gradient value of each layer is greater than a preset threshold. If it is greater than the preset threshold, clip the gradient value greater than the preset threshold, and determine the gradient information of the current batch based on the clipped gradient value.

[0021] In a preferred example of the present application, it can be further set that the clipping process for the gradient value greater than the preset threshold is specifically expressed by the formula:

[0022]

[0023] In the formula, g l represents the gradient value of the l - th layer, C represents the clipping threshold, and g l-clipped represents the gradient value of the l - th layer after clipping.

[0024] In a preferred example of the present application, it can be further set that, according to the sub - data d iThe sensitivity level and the layer of the local model are used to add hierarchical noise to the gradient information of the local model to obtain the adjusted gradient information of the current batch. The adjusted gradient information includes the adjusted gradient value of each layer of the local model, including:

[0025] Based on the base standard deviation and the layer of the local model, determine the noise standard deviation of each layer:

[0026] σ l = σ base ×(1 + α × l);

[0027] According to the sensitivity level of the sub-data d i adjust the noise standard deviation of each layer to obtain the adjusted gradient information of each layer:

[0028] σ final = σ l ×(1 + β·A(x));

[0029] Where, σ base represents the base standard deviation, l represents the layer of the local model, σ l represents the noise standard deviation, α is the coefficient controlling the increase of noise, A(x) represents the data analysis index, β represents the coefficient of the control factor, x represents the input data, and σ final represents the adjusted gradient information.

[0030] In a preferred example of the present application, it can be further set that the central server aggregates the model update parameters of all clients to form the global model parameters of the current batch, including:

[0031] The central server calculates the global model parameters of the current batch by using the method of weighted average.

[0032] In a preferred example of the present application, it can be further set that the sensitivity level of all sub-data d i is determined based on the variance of each sub-data d i , including:

[0033] When the variance σ i of the sub-data d 2 belongs to σ 2 > 0.5, the sub-data d i is determined as the first sensitivity;

[0034] When the variance σ i of the sub-data d 2 belongs to 0.2 < σ 2 ≤ 0.5, the sub-data d i is determined as the second sensitivity;

[0035] When the sub-data d iThe variance σ 2 belongs to 0 ≤ σ 2 ≤ 0.2, the sub - data d i is determined as the third sensitivity.

[0036] In a second aspect, the present application provides a training system for a federated learning model.

[0037] The present application is achieved through the following technical solutions:

[0038] A training system for a federated learning model, used to execute the training method of the federated learning model described in the first aspect above, includes:

[0039] Multiple clients and a central server;

[0040] Each client includes a data analysis module, a gradient confirmation module, a gradient adjustment module, and a model update module;

[0041] The data analysis module is used to divide the local data D into several sub - data d i , and determine the sensitivity level of each sub - data d i ;

[0042] The gradient confirmation module is used to initialize the model parameters of the local model in multiple clients, and the model parameters include the model parameters of each layer of the local model; use the sub - data d i to train the local model batch by batch, and obtain the gradient information of the local model in the current batch, where the gradient information includes the gradient values of each layer of the local model;

[0043] The gradient adjustment module is used to add hierarchical noise to the gradient information of the local model according to the sensitivity level of the sub - data d i and the layer level of the local model, and obtain the adjusted gradient information of the current batch, where the adjusted gradient information includes the adjusted gradient values of each layer of the local model;

[0044] The model update module is used to update the model parameters of the local model using the adjusted gradient information, obtain the model update parameters of the current batch, and upload the model update parameters of the current batch to the central server;

[0045] The central server is used to aggregate the model update parameters of all clients to form the global model parameters of the current batch, and distribute the global model parameters of the current batch to each client for each client to use the global model parameters for the next batch of local training.

[0046] In a third aspect, the present application is achieved through the following technical solutions:

[0047] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the training method of any one of the above-mentioned federated learning models are implemented.

[0048] In a fourth aspect, the present application provides a computer-readable storage medium.

[0049] The present application is achieved through the following technical solutions:

[0050] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the training method of any one of the above-mentioned federated learning models are implemented.

[0051] In summary, compared with the prior art, the beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0052] In the present application, the local data D is divided into several sub-data d i , and the sensitivity level of each sub-data d i is determined; the model parameters of the local model are initialized on multiple clients, and the model parameters include the model parameters of each layer of the local model; the local model is trained in batches using the sub-data d i to obtain the gradient information of the local model in the current batch, and the gradient information includes the gradient values of each layer of the local model; according to the sensitivity level of the sub-data d i and the layer level of the local model, hierarchical noise is added to the gradient information of the local model to obtain the adjusted gradient information of the current batch, and the adjusted gradient information includes the adjusted gradient values of each layer of the local model; the client uses the adjusted gradient information to update the model parameters of the local model to obtain the model update parameters of the current batch; each client uploads the model update parameters of the current batch to the central server; the central server aggregates the model update parameters of all clients to form the global model parameters of the current batch; the central server distributes the global model parameters of the current batch to each client for each client to perform the next batch of local training using the global model parameters. The local data of each client is fully utilized to train the local model, and hierarchical noise is added to the model parameters of the local model layer by layer to provide stronger privacy protection; each client performs model training locally, avoiding centralized storage of data, thereby reducing the risk of privacy leakage; the client only uploads the model parameters protected by noise to the central server to ensure that the sensitive information of the user will not be leaked during the data sharing process; not only improves the security of the data, but also reduces the communication cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1Schematic flowchart of a method for training a federated learning model provided by an embodiment of the present application;

[0054] Figure 2 Schematic structural diagram of a training system for a federated learning model provided by an embodiment of the present application;

[0055] Explanation of reference numerals:

[0056] Client 1, central server 2, data analysis module 101, gradient confirmation module 102, gradient adjustment module 103, model update module 104. Detailed implementation manners

[0057] This specific embodiment is only an interpretation of the present application and does not limit the present application. After reading this specification, those skilled in the art can make modifications to this embodiment without creative contributions as needed, but as long as they are within the scope of the claims of the present application, they are protected by the patent law.

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0059] In addition, the term "and / or" in the present application is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present application generally represents an "or" relationship between the associated objects before and after, unless otherwise specified.

[0060] The terms "first", "second", etc. in the present application are used to distinguish the same items or similar items with basically the same functions and effects. It should be understood that there is no logical or chronological dependence between "first", "second", and "nth", nor are the quantity and execution order limited.

[0061] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.

[0062] Federated learning is an innovative distributed machine learning method aimed at protecting data privacy while training models. Different from traditional centralized learning, federated learning allows multiple clients to train models locally without uploading the original data to a central server. The data always remains local, significantly reducing the risk of data leakage.

[0063] The following further describes the embodiments of the present application in conjunction with the accompanying drawings of the specification.

[0064] As Figure 1 shown, it is a training method for a federated learning model provided by the first exemplary embodiment of the present application. The training method of the present application is applicable to data of educational institutions or research institutions, including but not limited to student grades, personal information, health records, etc. The client can be a smartphone, an Internet of Things device, or other edge devices.

[0065] Among them, dedicated local models are designed for different types of data such as text, voice, and image data in the university scenario to meet diverse task requirements. For text data, such as course materials and academic papers, the Bidirectional Encoder Representations from Transformers (BERT), or the Generative Pretrained Transformer (GPT) is used as the local model. These models can perform text classification and sentiment analysis for course classification, keyword extraction, etc.; for voice data, such as classroom recordings and meeting records, the Long Short-Term Memory (LSTM) or the Convolutional Neural Network-Long Short-Term Memory combined model (CNN-LSTM) can be used as the local model to accurately process speech-to-text, sentiment classification, and keyword recognition; while for image data, such as courseware charts and experimental image scenarios, the Convolutional Neural Network (CNN) can be used as the local model for experimental result analysis, image classification, and object detection.

[0066] The training method proposed in the present application includes:

[0067] S1: Divide the local data D into several sub-data d i , and determine the sensitivity level of each sub-data d i .

[0068] Specifically, before each client loads the local data D to train the local model, the client needs to preprocess the data to ensure data quality and consistency.

[0069] First, divide the local data D into several sub-data d i , and the sub-data d i is used for model training in different batches. For example, n clients are set in n different educational institutions. The i-th client loads the corresponding local data D i , and divides it into:

[0070] D i = {(d i 1 : (x 1 , y 1 ), (d i 2 : (x 2 , y 2 ),, (d i n : (x n , y n )},

[0071] where d i 1 represents the first sub - data of the i - th client for the local model training of the first batch. x 1 represents the input sample, and y 1 represents the corresponding label of the input sample; d i 2 represents the second sub - data of the i - th client for the local model training of the second batch. x 2 represents the input sample, and y 2 represents the corresponding label of the input sample; d i n represents the n - th sub - data of the i - th client for the local model training of the n - th batch. x n represents the input sample, and y n represents the corresponding label of the input sample.

[0072] Each client analyzes its local sub - data d i respectively to determine the sensitivity of the local data included in each sub - data d i . The sensitivity is used to reflect the privacy of the sub - data d i . Among them, the sensitivity includes high sensitivity, medium sensitivity, and low sensitivity.

[0073] S2: Initialize the model parameters of the local model on multiple clients. The model parameters include the model parameters of each layer of the local model.

[0074] In this embodiment, the local model M on each client includes a multi - layer structure, denoted as {L 1 , L 2 ,..., L n}, where L 1 represents the first - layer structure of the local model M, L 2 represents the second - layer structure of the local model M, and L n represents the n - th layer structure of the local model M. At the beginning of training, first set the initial model parameters {θ1 ,θ 2 ,...,θ n}, where θ 1 Represents the model parameters of the first layer structure of the local model, θ 2 Represents the model parameters of the second layer structure of the local model, θ n Represents the model parameters of the nth level structure of the local model.

[0075] Taking the local model M as a convolutional neural network (CNN) as an example, the local model M includes a 6-layer structure, represented as {L 1 ,L 2 ,L 3 ,L 4 ,L 5 ,L 6}, L 1 Represents the first layer structure (convolutional layer) of the local model M, L 2 represents the second layer structure (pooling layer) of the local model M, L 3 represents the third layer structure (convolutional layer) of the local model M, L 4 Represents the fourth layer structure (pooling layer) of the local model M, L 5 represents the fifth layer structure (fully connected layer) of the local model M, L 6 Represents the sixth layer structure (output layer) of the local model M. At the beginning of training, the initial model parameters {θ 1 ,θ 2 ,θ 3 ,θ 4 ,θ 5 ,θ 6}, where θ 1 represents the model parameters of the first layer of the local model, θ 2 represents the model parameters of the second layer of the local model, θ 3 represents the model parameters of the third layer of the local model, θ 4 represents the model parameters of the fourth layer of the local model, θ 5 represents the model parameters of the fifth layer of the local model, θ 6 Represents the model parameters of the sixth layer of the local model.

[0076] S3: Utilize sub-data in batches i Train the local model to obtain the gradient information of the local model in the current batch. The gradient information includes the gradient value of each layer of the local model.

[0077] In this embodiment, local data is used on each client to perform model training and calculate loss and gradient. iInput into the local model, and the local model outputs the corresponding predicted value. Based on the predicted value and the data label of the sub-data d i to construct a loss function, calculate the loss based on the loss function, and then determine the corresponding gradient value g i :

[0078]

[0079] where f(x,θ) is the local model, x is the input parameter, y is the corresponding label, θ is the model parameter, and L is the loss function;

[0080] After calculating the corresponding gradient value g i use the gradient descent method to update the model parameter:

[0081]

[0082] In the formula, is the model parameter of the t-th iteration, is the model parameter of the (t - 1)-th iteration, and η represents the learning rate.

[0083] S4: According to the sensitivity level of the sub-data d i and the layer level of the local model, add hierarchical noise to the gradient information of the local model to obtain the adjusted gradient information of the current batch. The adjusted gradient information includes the adjusted gradient value of each layer of the local model.

[0084] In this embodiment, according to the sensitivity level of the sub-data d i and the layer level of the local model, add hierarchical noise σ 1 ,σ 2 ,...,σ n to the gradient value of each layer of the gradient information θ = {θ i ,θ 1 ,...,θ 2} of the local model to obtain the adjusted gradient information θ' = {θ n ',θ 1 ',...,θ 2 '} of the current batch, where θ n ' represents the adjusted gradient value of the first layer structure after adding hierarchical noise, θ 2 ' represents the adjusted gradient value of the second layer structure after adding hierarchical noise, and θ n ' represents the adjusted gradient value of the n-th layer structure after adding hierarchical noise, and n represents the number of layers of the local model.

[0085] It should be noted that the hierarchical noise σ i added to each layer is jointly determined by the sensitivity level of the currently processed sub-data d i and the layer level of each layer.

[0086] S5: The client updates the model parameters of the local model using the adjusted gradient information to obtain the model update parameters for the current batch.

[0087] According to the adjusted gradient value calculated in step S4, the client uses the gradient descent method to update the model parameters (w, b) of the local model. The update formula is expressed as:

[0088]

[0089]

[0090] where w t represents the weight of the t-th iteration, b t represents the bias of the t-th iteration, w t+1 represents the weight of the (t + 1)-th iteration, b t+1 represents the weight of the (t + 1)-th iteration, and η represents the learning rate, which is used to control the step size of parameter update.

[0091] Repeat the above update iteration. When the preset number of iterations is reached, determine the model parameters of the local model at the last iteration as the model update parameters.

[0092] S6: Each client uploads the model update parameters for the current batch to the central server.

[0093] All clients upload the model update parameters for the current batch to the central server, which is expressed as:

[0094] upload: g i,noisy → server.

[0095] In specific implementation, to further protect privacy, the model update parameters can be encrypted and then sent to the central server, and the central server performs decoding processing after receiving them.

[0096] S7: The central server aggregates the model update parameters of all clients to form the global model parameters for the current batch.

[0097] After receiving the model update parameters uploaded by all clients, the central server fuses the model update parameters of all clients according to certain rules to form the global model parameters for the current batch.

[0098] Federated learning realizes the balance between effective data utilization and privacy protection by aggregating the model update parameters of each client to the central server and then forming a global model.

[0099] S8: The central server distributes the global model parameters for the current batch to each client for each client to use the global model parameters for the next batch of local training.

[0100] The server distributes the updated global model parameters to each client to prepare for the next round of training. The clients will continue local training using the new global model parameters, forming a loop to gradually improve the model performance:

[0101] distribute: θ t+1 → client i 。

[0102] In some embodiments, the local data D is divided into several sub - data d i , and the sensitivity level of each sub - data d i is determined, including:

[0103] Analyze each sub - data d i to determine the variance of each sub - data d i ; determine the sensitivity level of all sub - data d i based on the variance of each sub - data d i .

[0104] Among them, variance is mainly used to evaluate the dispersion degree of numerical features, thus reflecting the sensitivity of the data. For example: 1. Numerical features:

[0105] For numerical data such as student grades and tuition fees, calculate the variance to measure the fluctuations of these features in the sub - data. It can be calculated in the following way:

[0106]

[0107] In the formula, σ 2 represents the variance, x i represents each data point, μ represents the overall average value, and N is the total amount of data.

[0108] 2. Time - series data:

[0109] For time - series data such as students' learning progress, variance is used to measure the volatility of the data. It can be calculated in the following way:

[0110] First, calculate the mean of the time series Calculate the autocorrelation coefficient at zero lag. Use the autocorrelation coefficient and the volatility of the series to estimate the variance, and the variance can be estimated as the square of the standard deviation of the series, which can be expressed as:

[0111] σ 2 =Var(X)=(std(X)) 2 。

[0112] 3. Numerical features of categorical data:

[0113] For categorical data that undergoes encoding conversion, such as subject classification, variance is used to measure the distribution differences between categories. The calculation method of grouped variance can be used:

[0114] First, divide the data into different categories, calculate the mean and variance of each category, and calculate the variance for each classification i

[0115]

[0116]

[0117] In the formula, n i is the amount of data in category i, xij represents the jth data in category i, represents the mean of category i;

[0118] Calculate the total variance σ 2 :

[0119]

[0120] where N represents the total amount of data, k represents the total number of categories, and x is the overall mean.

[0121] Specifically, when the variance σ i of the sub-data d 2 belongs to σ 2 > 0.5, the sub-data d i is determined as the first sensitivity, and the first sensitivity is used to indicate that the sub-data d i is highly sensitive data;

[0122] When the variance σ i of the sub-data d 2 belongs to 0.2 < σ 2 ≤ 0.5, the sub-data d i is determined as the second sensitivity, and the second sensitivity is used to indicate that the sub-data d i is medium-sensitive data;

[0123] When the variance σ i of the sub-data d 2 belongs to 0 ≤ σ 2 ≤ 0.2, the sub-data d i is determined as the third sensitivity, and the third sensitivity is used to indicate that the sub-data d i is low-sensitive data.

[0124] Different sensitivities can be represented by different numerical values, and the numerical values are added as labels to the sub-data d i above.

[0125] In some embodiments, when the gradient value is too large, it will cause the model parameters to be updated too violently, thus affecting the convergence of the model. To avoid this situation, this embodiment controls the gradient value through gradient clipping to ensure that the update amplitude each time is not too large, thereby improving the stability of the model.

[0126] The specific steps include: in each training batch, train the local model with sub-data and obtain gradient information. Then, determine whether the gradient value of each layer is greater than a preset threshold. If the gradient value is greater than the preset threshold, clip it, and determine the gradient information of the current batch based on the clipped gradient value to ensure that the gradient value is within a reasonable range. In this way, the parameter update of the model will not be too violent, avoiding instability caused by too large gradient values during the training process.

[0127] Considering that gradient value clipping requires a reasonable threshold, this embodiment adopts a dynamically changing clipping threshold. Specifically, the threshold can be dynamically adjusted according to the gradient value calculated in each iteration. An adaptive threshold can be set by calculating the standard deviation of the current gradient. Such a method ensures that the threshold can be flexibly adjusted according to the change of the gradient distribution during the training process, thereby improving the effectiveness and flexibility of gradient clipping.

[0128] Specifically, perform clipping processing on the gradient value greater than the preset threshold, and the specific formula is expressed as:

[0129]

[0130] In the formula, g i represents the gradient value of the i-th layer, C represents the clipping threshold, and g i-clipped represents the gradient value of the i-th layer after clipping.

[0131] In this embodiment, the basic standard deviation is determined through statistical analysis of the training data. First, collect a certain amount of training data and standardize it in the preprocessing stage to ensure that the mean is 0 and the variance is 1. Then, calculate the standard deviation of the preprocessed data, and the formula is:

[0132]

[0133] where μ is the data mean, N is the number of samples, and x i is the sample point.

[0134] The value range of the basic standard deviation is usually from 0.1 to 1.0, depending on the data characteristics. For data with a larger range of eigenvalue, the basic standard deviation may increase; otherwise, it may decrease. Through these steps, the basic standard deviation can be effectively determined, providing a stable basis for subsequent training.

[0135] In some embodiments, according to the sub-data di Based on the sensitivity level and the layer level of the local model, hierarchical noise is added to the gradient information of the local model to obtain the adjusted gradient information for the current batch. The adjusted gradient information includes the adjusted gradient value for each layer of the local model, including:

[0136] Based on the base standard deviation and the layer level of the local model, determine the noise standard deviation for each layer:

[0137] σ l = σ base ×(1 + α×l);

[0138] According to the sensitivity level of the sub-data d i adjust the noise standard deviation for each layer to obtain the adjusted gradient information for each layer:

[0139] σ final = σ l ×(1 + β·A(x));

[0140] where, σ base represents the base standard deviation, l represents the layer level of the local model, σ l represents the noise standard deviation, α is the coefficient controlling the noise increase, A(x) represents the data analysis index, β represents the coefficient of the control factor, x represents the input data, and σ final represents the adjusted gradient information. It should be noted that when the sensitivity level of the data is higher, the value of A(x) is larger.

[0141] In some embodiments, the central server aggregates the model update parameters of all clients to form the global model parameters for the current batch. Specifically, the central server uses the weighted average method to calculate the global model parameters. To ensure the effectiveness of the aggregation, the server first evaluates the performance of each local model, usually choosing accuracy as the main metric.

[0142] The formula for the server to calculate the new global model parameter θ is:

[0143]

[0144] where, θ i represents the local model parameter of the i-th client, and w i is the weight assigned based on the local model performance. Specifically, the weight w i can be calculated in the following way:

[0145]

[0146] In this formula, Accuracy iis the accuracy of the local model of the i-th client on the validation set. Only clients with an accuracy higher than the preset accuracy threshold will be included in the weight calculation, thus ensuring that the global model is mainly influenced by local models with better performance. Specifically, the central server receives the local model parameters and the corresponding accuracy, compares the accuracy of each local model with the preset accuracy threshold. When the accuracy of the local model is greater than the preset accuracy threshold, the local model is included in the weight calculation, and the central server assigns different weight values based on the accuracy of the local model. In actual implementation, the preset accuracy threshold can be set according to requirements, for example, set to 70%.

[0147] The dynamic hierarchical differential privacy federated learning proposed in this application can effectively protect data privacy. This mechanism improves the intensity and flexibility of data privacy protection by dynamically adjusting the noise intensity according to the importance of the data and the hierarchical structure of the model. Specifically, the dynamic hierarchical differential privacy federated learning divides the data into different levels, adds stronger noise to highly sensitive data (such as students' grades and personal information), and adds weaker noise to low-sensitivity data (such as course evaluations and browsing habits). This hierarchical method ensures the accuracy of privacy protection, enables privacy protection measures to be optimized according to the specific characteristics of the data, and thus minimizes the risk of privacy leakage without compromising the performance of the model.

[0148] The dynamic hierarchical differential privacy federated learning strictly follows the basic principles of differential privacy to ensure that when a single record is added or deleted, the change in the query result is imperceptible to the observer. This feature effectively prevents the leakage of sensitive information and maintains the privacy of user participation. With the increasing strictness of data privacy regulations, the dynamic hierarchical differential privacy federated learning provides an innovative solution for the compliant use of data. This not only helps educational institutions comply with laws and regulations when using data for analysis and decision-making, but also enhances users' trust in data use.

[0149] To verify the effectiveness of the dynamic hierarchical differential privacy federated learning method proposed in this application, this application conducts experiments on dynamic hierarchical differential privacy federated learning using the MNIST and FashionMNIST datasets.

[0150] The MNIST (Modified National Institute of Standards and Technology) dataset is a classic handwritten digit recognition dataset. This dataset contains a large number of grayscale images of 28x28 pixels, representing handwritten digits from 0 to 9. Grayscale images of preset data are obtained from this dataset and further divided into a training set and a test set. The training set contains 60,000 images, and the test set contains 10,000 images. The FashionMNIST dataset contains 70,000 grayscale images of 28x28 pixels, representing different fashion items. Different from MNIST, Fashion-MNIST contains 10 categories, including T-shirts, pants, shoes, skirts, etc. The distribution of the training set and the test set is the same as that of MNIST, with 60,000 and 10,000 images respectively.

[0151] On the MNIST dataset, the accuracy of the model using dynamic hierarchical differential privacy on the test set reached 97.6%. In contrast, the accuracy of the differential privacy federated learning method based on hierarchical noise was 94.7%, and that based on Laplace noise for federated learning was 91.6%. On the FashionMNIST dataset, the accuracy of the model using dynamic hierarchical differential privacy on the test set reached 95.2%, while that of the differential privacy federated learning based on Laplace noise was 89.5%, and that of the differential privacy federated learning method based on hierarchical noise was 91.2%. These results indicate that dynamic hierarchical differential privacy can significantly improve the accuracy of the model while protecting privacy.

[0152] In terms of running time, 20 clients were selected for testing to demonstrate the running speed, and then the running time of each round was recorded. Taking the MNIST dataset as an example, the time of the model using dynamic hierarchical differential privacy on the test set was 31.64% faster than that of federated learning based on Laplace noise. This shows that dynamic hierarchical differential privacy not only improves the performance of the model but also has obvious advantages in efficiency.

[0153] Table 1

[0154]

[0155] Another exemplary embodiment of the present application provides a training system for a federated learning model, as Figure 2 shown, including: a plurality of clients 1 and a central server 2;

[0156] Each client 1 includes a data analysis module 101, a gradient confirmation module 102, a gradient adjustment module 103, and a model update module 104;

[0157] The data analysis module 101 is used to divide the local data D into several sub - data d i , and determine the sensitivity level of each sub - data d i .

[0158] The gradient confirmation module 102 is used to initialize the model parameters of the local model on multiple clients. The model parameters include the model parameters of each layer of the local model; use the sub - data d in batches i to train the local model, and obtain the gradient information of the local model in the current batch. The gradient information includes the gradient values of each layer of the local model;

[0159] The gradient adjustment module 103 is used to add hierarchical noise to the gradient information of the local model according to the sensitivity level of the sub - data d i and the layer level of the local model to obtain the adjusted gradient information of the current batch. The adjusted gradient information includes the adjusted gradient values of each layer of the local model;

[0160] The model update module 104 is used to update the model parameters of the local model using the adjusted gradient information to obtain the model update parameters of the current batch, and upload the model update parameters of the current batch to the central server;

[0161] The central server 2 is used to aggregate the model update parameters of all clients to form the global model parameters of the current batch, and distribute the global model parameters of the current batch to each client for each client to use the global model parameters for the next batch of local training.

[0162] For the specific limitations of the training system of the federated learning model provided in this embodiment, reference can be made to the embodiments of the training method of the federated learning model in the above text, which will not be elaborated here. Each module in the above - mentioned client can be implemented in whole or in part by software, hardware, and their combination. The above - mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above - mentioned modules.

[0163] An embodiment of the present application provides a computer device, which may include a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the processor executes the steps of the training method of the federated learning model in any of the above embodiments.

[0164] For the working process, working details, and technical effects of the computer device provided in this embodiment, reference may be made to the embodiments of the training method of the federated learning model in the foregoing text, and details are not described herein again.

[0165] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the training method of the federated learning model in any of the above embodiments are implemented. Among them, the computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, a floppy disk, an optical disc, a hard disk, a flash memory, a USB flash drive, and / or a Memory Stick, etc. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices.

[0166] For the working process, working details, and technical effects of the computer-readable storage medium provided in this embodiment, reference may be made to the embodiments of the training method of the federated learning model in the foregoing text, and details are not described herein again.

[0167] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0168] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0169] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system described in the present application is divided into different functional units or modules to complete all or part of the functions described above.

Claims

1. A training method for a federated learning model, characterized in that: include: Divide the local data D into several sub-data d i , determine each sub-data d i sensitivity level; Initializing model parameters of local models on multiple clients, wherein the model parameters include model parameters of each layer of the local model; Utilize subdata in batches i Train the local model to obtain gradient information of the local model in the current batch, where the gradient information includes the gradient value of each layer of the local model; According to the sub-data d i The sensitivity level and the layer extreme of the local model are used to add hierarchical noise to the gradient information of the local model to obtain the adjusted gradient information of the current batch, wherein the adjusted gradient information includes the adjusted gradient value of each layer of the local model; The client updates the model parameters of the local model using the adjustment gradient information to obtain the model update parameters of the current batch; Each client uploads the model update parameters of the current batch to the central server; The central server aggregates the model update parameters of all clients to form the global model parameters of the current batch; The central server distributes the global model parameters of the current batch to each client, so that each client can use the global model parameters to perform local training of the next batch.

2. The method for training a federated learning model according to claim 1, characterized in that: Divide the local data D into several sub-data d i , determine each sub-data d i The sensitivity levels include: For each sub-data d i Analyze and determine each sub-data d i The variance of Based on each sub-data d i The variance of all sub-data d i sensitivity level.

3. The training method of the federated learning model according to claim 1, characterized in that: Utilize subdata in batches i The local model is trained to obtain the gradient information of the local model in the current batch. The gradient information includes the gradient value of each layer of the local model and also includes: It is determined whether the gradient value of each layer is greater than a preset threshold value. If so, the gradient values ​​greater than the preset threshold value are clipped, and the gradient information of the current batch is determined based on the clipped gradient values.

4. The method for training a federated learning model according to claim 3, characterized in that: The gradient values ​​greater than the preset threshold are clipped. The specific formula is expressed as: In the formula, g i represents the gradient value of the i-th layer, C represents the clipping threshold, g i-clipped Represents the gradient value after clipping of the i-th layer.

5. The method for training a federated learning model according to claim 3, characterized in that: According to the sub-data d i The sensitivity level and the layer extreme of the local model are used to add layered noise to the gradient information of the local model to obtain the adjusted gradient information of the current batch, where the adjusted gradient information includes the adjusted gradient value of each layer of the local model, including: Based on the base standard deviation and the level of the local model, determine the noise standard deviation of each layer: s l =s base ×(1+α×l); According to the sub-data d i The sensitivity level of each layer is adjusted to adjust the noise standard deviation of each layer to obtain the adjustment gradient information of each layer: s final =s l ×(1+β·A(x)); Among them, σ base represents the basic standard deviation, l represents the level of the local model, σ l represents the standard deviation of noise, α is the coefficient for controlling the increase of noise, A(x) represents the data analysis index, β represents the coefficient of the control factor, x represents the input data, σ final Indicates adjustment gradient information.

6. The method for training a federated learning model according to claim 1, characterized in that: The central server aggregates the model update parameters of all clients to form the global model parameters of the current batch, including: The central server uses the weighted average method to calculate the global model parameters of the current batch.

7. The method for training a federated learning model according to claim 2, characterized in that: Based on each sub-data d i The variance of all sub-data d i The sensitivity levels include: When the sub-data d i The variance σ 2 Belong to σ 2 >0.5, the sub-data d i Determined as the first sensitivity; When the sub-data d i The variance σ 2 Belongs to 0.2<σ 2 ≤0.5, the sub-data d i Determined as the second sensitivity; When the sub-data d i The variance σ 2 Belongs to 0≤σ 2 ≤0.2, the sub-data d i Determined as the third sensitivity.

8. A training system for a federated learning model, characterized in that: The training system includes: a plurality of clients and a central server; Each client includes a data analysis module, a gradient confirmation module, a gradient adjustment module, and a model update module; The data analysis module is used to divide the local data D into several sub-data d i , determine each sub-data d i sensitivity level; The gradient confirmation module is used to initialize the model parameters of the local model on multiple clients, wherein the model parameters include the model parameters of each layer of the local model; and to use the sub-data d i Train the local model to obtain gradient information of the local model in the current batch, where the gradient information includes the gradient value of each layer of the local model; Gradient adjustment module, used to adjust the gradient according to the sub-data d i The sensitivity level and the layer extreme of the local model are used to add hierarchical noise to the gradient information of the local model to obtain the adjusted gradient information of the current batch, wherein the adjusted gradient information includes the adjusted gradient value of each layer of the local model; A model updating module, used to update the model parameters of the local model using the adjustment gradient information, obtain the model update parameters of the current batch, and upload the model update parameters of the current batch to the central server; The central server is used to aggregate the model update parameters of all clients to form the global model parameters of the current batch, and distribute the global model parameters of the current batch to each client so that each client can use the global model parameters to perform local training for the next batch.

9. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.