Distributed learning method and device of linear regression model, equipment and storage medium
By introducing a mechanism of proof key and honest proof information between participant nodes and central nodes of federated learning, the honesty of the gradient data feedback from participant nodes is solved, and the problem of model poisoning attacks in federated learning is ensured, ensuring the correctness of the model and the security of data privacy.
Patent Information
- Application Number
- CN202510476487.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In federated learning, participants may conduct model poisoning attacks, provide incorrect calculation results or use incorrect calculation processes, resulting in degradation of model performance, increasing model deviation, and even causing the model to make incorrect information predictions, endangering data security.
By introducing a mechanism of proof keys and honest proof information between the participant nodes and the central node, we ensure that the central node can verify the honesty of the gradient data feedback from the participant nodes and prevent poisoning attacks. The specific steps include: the participant node calculates the gradient data of the loss function based on the local training sample, and uses the proof key to generate honest proof information; after the central node receives this information, it verifys its authenticity, and after the verification is passed, the gradient data of each participant node is integrated to update the model parameters.
It effectively prevents model poisoning attacks, ensures the correctness of the model and the security of data privacy during the federated learning process, and prevents the problems of degradation in model performance and increase in deviations.
Smart Images

Figure CN120012042A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning, and in particular to a distributed learning method, apparatus, device and storage medium for a linear regression model. Background Art
[0002] In the traditional linear regression modeling process, data can be cleaned and verified before training to reduce the mixing of malicious data.
[0003] Federated Learning (FL) is a distributed machine learning paradigm with decentralized computing characteristics and privacy protection requirements. When training a linear regression model under the federated learning framework, federated learning allows multiple participating nodes to collaborate on training the model without sharing the original data, thereby protecting data privacy. Its basic model is that data is scattered among multiple parties, and multiple parties conduct multiple rounds of data interaction under the organization of the center to complete the model training.
[0004] One problem is that each party participates in part of the calculation and data provision, so it will affect the final result to a certain extent. In some cases, some participants may launch model poisoning attacks, provide incorrect calculation results or use completely wrong calculation processes. In this case, federated learning cannot guarantee the correctness of the final result.
[0005] These attacks can lead to a decline in model performance and an increase in model deviation, causing the global model to converge to the undesirable state expected by the attacker, resulting in damage to model training during the federated learning process, and even causing the trained model to make incorrect information predictions, endangering data security and affecting the model's data processing performance. Summary of the invention
[0006] The embodiments of the present application provide a distributed learning method, apparatus, device and storage medium for a linear regression model to solve the problem of damage to model training in the federated learning process caused by poisoning attacks in the prior art.
[0007] A first aspect of an embodiment of the present application provides a distributed learning method for a linear regression model, which is applied to a participant node, and the method includes: Obtain the model parameters of the target model and the certification keys of each model parameter sent by the central node; Calculate the gradient data of the loss function of the target model to each model parameter based on the local training sample, and calculate the honesty proof information based on the proof key and the local training sample; Sending the honesty proof information and the gradient data to the central node; Obtain the updated model parameters and the proof keys of each updated model parameter sent by the central node, and return to execute the step of calculating the gradient data of each model parameter of the loss function of the target model based on the local training samples until the target model training is completed; wherein the updated model parameters are obtained by updating the model parameters after the central node verifies that the gradient data is honest data based on the honesty proof information.
[0008] A second aspect of an embodiment of the present application provides a distributed learning method for a linear regression model, which is applied to a central node. The method includes: Send the model parameters of the target model and the certification keys of each model parameter to each participating node; Receive the honesty proof information sent by each of the participant nodes and the gradient data of the loss function of the target model to each model parameter; the gradient data is calculated by each of the participant nodes based on local training samples; the honesty proof information is calculated by the participant node based on the proof key and the local training samples; Verifying the corresponding gradient data based on the honesty proof information, and after verifying that the gradient data is honest data, integrating the gradient data of each of the participating nodes to obtain model correction data; The model parameters are updated based on the model correction data to obtain updated model parameters, and the proof keys of each updated model parameter are calculated, and the step of sending the model parameters of the target model and the proof keys of each model parameter to each participating node is returned to execute until the target model training is completed.
[0009] A third aspect of an embodiment of the present application provides a distributed learning device for a linear regression model, comprising: An acquisition module, used to acquire the model parameters of the target model and the certification keys of each model parameter sent by the central node; A calculation module, used to calculate the gradient data of the loss function of the target model to each model parameter based on the local training sample, and to calculate the honesty proof information based on the proof key and the local training sample; A sending module, used for sending the honesty proof information and the gradient data to the central node; An update module is used to obtain the updated model parameters sent by the central node and the proof keys of each updated model parameter, and return to execute the step of calculating the gradient data of each model parameter of the loss function of the target model based on the local training samples until the training of the target model is completed; wherein the updated model parameters are obtained by updating the model parameters after the central node verifies that the gradient is honest data based on the honest proof information.
[0010] A fourth aspect of an embodiment of the present application provides a distributed learning device for a linear regression model, comprising: A sending module, used for sending the model parameters of the target model and the certification keys of each model parameter to each participating node; A receiving module, used to receive the honesty proof information sent by each of the participating nodes and the gradient data of the loss function of the target model to each model parameter; the gradient data is calculated by each of the participating nodes based on the local training samples; the honesty proof information is calculated by the participating node based on the proof key and the local training samples; A verification module, used to verify the corresponding gradient data based on the honesty proof information, and after verifying that the gradient data is honest data, integrate the gradient data of each of the participating nodes to obtain model correction data; An update module is used to update the model parameters based on the model correction data to obtain updated model parameters, calculate the proof key of each updated model parameter, and return to execute the step of sending the model parameters of the target model and the proof key of each model parameter to each participating node until the target model training is completed.
[0011] The fifth aspect of an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect or the second aspect when executing the computer program.
[0012] A sixth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.
[0013] The seventh aspect of the present application provides a computer program product. When the computer program product is run on a computer device, the computer device executes the steps of the method described in the first aspect or the second aspect.
[0014] In an embodiment of the present application, the participating nodes generate honesty proof information for the data objects and data calculation processes involved in the model training, so that the central node can verify the gradient data fed back by the participating parties based on the honesty proof information to prevent poisoning attacks. The participating nodes need to perform some additional operations and use the results of these calculations as evidence to prove their innocence and the correctness of their own operations, and ensure that the data privacy of the participants is not leaked during the federated learning process, ensure data security, and ensure that the model is effectively trained. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 This is the process of the distributed learning method of the linear regression model provided in the embodiment of the present application Figure 1 ; Figure 2 This is the process of the distributed learning method of the linear regression model provided in the embodiment of the present application Figure 2 ; Figure 3 It is a structural diagram of a distributed learning device for a first linear regression model provided in an embodiment of the present application; Figure 4 is a structural diagram of a distributed learning device for a second linear regression model provided in an embodiment of the present application; Figure 5 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0018] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0019] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0020] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0021] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0022] It should be understood that the size of the serial numbers of the steps in this embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0023] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.
[0024] Linear regression modeling is a statistical method used to describe the linear relationship between variables. It is used to determine whether there is a mutual dependence between two or more variables and try to establish a mathematical model to describe this relationship.
[0025] Specifically, a linear regression model assumes that there is a linear relationship between a target variable (usually denoted as Y) and one or more independent variables (usually denoted as X). This relationship is expressed in an exemplary mathematical formula as: Y = β0 + β1X1+ β2X2.
[0026] Among them, the independent variable X corresponds to the characteristic data in the training sample of the linear regression model. β0 is the intercept term, which corresponds to the bias term in the linear regression model. β1, β2,... are slope coefficients, which indicate the influence of the independent variable on the dependent variable, and correspond to the characteristic weights of each characteristic data in the linear regression model. β0,β1, β2,... form the model parameters.
[0027] In some cases, the training of the model in machine learning can be carried out based on the local data of multiple participating nodes, and the central node aggregates the trained data of each participating node to implement the update of the model parameters. During the implementation process, the distributed learning and gradient descent algorithms in machine learning can be used to achieve it.
[0028] In a distributed learning scenario, a central node (or server) is usually responsible for coordinating the entire learning process. Federated learning is a distributed machine learning framework that allows multiple participating nodes to collaborate on training models without sharing original data.
[0029] Each participating node trains the model on its own local data and only uploads the gradient information to the central node for aggregation, thereby effectively protecting data privacy.
[0030] Federated learning is applied to linear regression model training, so the federated linear regression operation method is executed, which requires collaboration among all parties.
[0031] Before starting to train the model, this central node randomly selects or initializes the parameters of the model. In one example, these model parameters include w0, w1, w2, and d. These parameters are the feature weights (w0, w1, w2) and possible bias terms (d) learned by the model, which are used to make subsequent predictions on the input data.
[0032] Once the initial model parameters are determined, the central node will distribute these parameters to all participating nodes (or clients) participating in the training. These participating nodes have their own sample data sets locally for training the model.
[0033] The participating nodes use local sample data to calculate the gradient of the model's loss function with respect to these initial model parameters. After receiving the initial model parameters, each participating node will use its own local sample data set to calculate the gradient of the loss function (a function that measures the difference between the model's predicted data and the actual data) relative to these model parameters, and feed these gradients back to the central node. The central node adjusts the model parameters based on these gradients and feeds them back to each participating node, repeating the above process until the model training is completed.
[0034] Among them, the gradient is the derivative of the loss function with respect to the model parameters, which indicates in which direction and how much the parameters should be adjusted in order to minimize the loss function.
[0035] The above process is part of the gradient descent algorithm in distributed learning, where the model parameters are updated iteratively between the central node and the participating nodes to minimize the global loss function. Each participating node independently performs gradient calculations on its local sample data without sharing data with other participants. In this way, even if the sample data is distributed on different participating nodes, a global model can be trained together while protecting the privacy of each participant's local data.
[0036] In the above process, since the participating nodes need to use the model parameters of the current round obtained from the central node to calculate the gradient data of this round, malicious participating nodes may not use the correct model parameters, nor calculate according to the agreed algorithm formula, or even fabricate data out of thin air and upload it to the central node.
[0037] The behavior of a participant node in federated learning intentionally providing false calculations for the purpose of interfering with the final result can be defined as a poisoning attack. In some cases, a model poisoning attack means that a participant node intentionally provides incorrect calculation results during data calculation or uses a completely wrong calculation process to obtain incorrect technical results. These attacks can lead to a decrease in model performance, an increase in model bias, and even cause the model to make incorrect predictions, so federated learning cannot guarantee the correctness of the final result.
[0038] For example, in a text classification task, an attacker can upload a modified model update, causing the model to produce incorrect classification results for certain keywords.
[0039] These malicious updates will affect the aggregation process of the global model, causing the global model to converge to the bad state expected by the attacker, resulting in damage to the model training during the federated learning process. It will be difficult for the central node to identify the honesty of the calculation process and feedback data, making it difficult to resist poisoning attacks during the distributed training of the model.
[0040] See also Figure 1 , Figure 1 This is a process of a distributed learning method for a linear regression model provided in an embodiment of the present application. Figure 1 .like Figure 1 As shown, a distributed learning method of a linear regression model is applied to a participant node, and the method includes the following steps: Step 101, obtaining the model parameters of the target model and the certification keys of the various model parameters sent by the central node.
[0041] Optionally, the target model here is a linear regression model. Model parameters are specific parameters included in different types of models. When the target model is a linear regression model, the model parameters may include feature weights and bias items of feature variables in different dimensions.
[0042] The proof key is the data sent by the central node, which is used to enable the participating nodes to generate honesty proof information based on the proof key to prove the honesty of the data calculation of the participating nodes during the model training process based on local sample data.
[0043] The proof key may be a selected random number, or data calculated in other ways, such as the proof key obtained based on cryptographic calculations.
[0044] Step 102, calculate the gradient data of the loss function of the target model with respect to each model parameter based on the local training sample, and calculate the honest proof information based on the proof key and the local training sample.
[0045] Each participating node uses local data to train the target model and calculates the gradient data of the target model's loss function for each model parameter.
[0046] The calculation of gradient data can be done by taking the derivative of the loss function with respect to each model parameter.
[0047] In an optional embodiment, calculating the gradient data of the loss function of the target model to each model parameter based on the local training sample includes: The feature data in the local training samples are input into the target model to obtain the output prediction results; the sample labels and prediction results in the local training samples are input into the derivative functions of each model parameter under the loss function of the target model, and the gradient data of the loss function for each model parameter is calculated.
[0048] The local training samples include feature data of different dimensions and sample labels corresponding to the samples to which these feature data belong. For example, the local training samples include feature data such as user consumption amount, user investment amount, user savings amount, user debt amount, etc. The type of population to which the user belongs can be used as the label of the user sample.
[0049] Based on the loss function of the target model, the derivative functions of the loss function for different model parameters can be determined, and the gradient data of the corresponding model parameters can be calculated through these derivative functions. The result data predicted by the target model based on the input feature data based on the model parameters and the sample labels corresponding to the input feature data are input into the derivative function, and the gradient data of the loss function for each model parameter can be obtained.
[0050] In some optional implementations, the sample labels and prediction results in the local training samples are input into the derivative functions of each model parameter under the loss function of the target model, and the gradient data of the loss function to each model parameter are calculated, including: Input the sample labels and prediction results in the local training samples into the following derivative function to calculate the gradient data of the loss function for each model parameter:
[0051]
[0052] in, is the feature weight corresponding to the feature data of the jth dimension in the model parameters; is the derivative function of the loss function; is the gradient data of the feature weight corresponding to the feature data of the j-th dimension of the loss function; ∈[0, m], m is the total number of dimensions of the feature data in the local training sample; n is the total number of local training samples; is the feature data of the jth dimension in the i-th training sample, i∈[1,n].
[0053] is the bias term in the model parameters; is the gradient data of the loss function with respect to the bias term; is the prediction result output by the target model based on the feature data in the i-th training sample in the local training sample, is the sample label of the i-th training sample in the local training sample.
[0054] In this process, the training samples include multiple sample data. Each feature data in each sample data corresponds to a feature dimension. The feature parameters include a feature weight and a bias term corresponding to each feature dimension. The loss function of the linear regression model is used to effectively calculate the gradient data of each model parameter on the participating node side to achieve distributed model training.
[0055] The method of calculating the honest proof information based on the proof key and the local training sample can be a calculation method determined in advance by the participating nodes and the central node.
[0056] The calculation method may be a randomly determined calculation formula, such as addition, multiplication, dot multiplication, etc., or a proof calculation method determined based on cryptography. The calculated honest proof information may include one or more proof data.
[0057] In an optional implementation, the honest proof information is calculated based on the proof key and the local training sample, including: Based on the feature data of each dimension in the local training sample and the proof key, the first proof information corresponding to each of the model parameters is calculated respectively; based on the feature data of each dimension in the local training sample and the sample label in the local training sample, the second proof information corresponding to the local training sample is calculated; and the honest proof information including the first proof information and the second proof information is obtained.
[0058] The proof key can be multiplied by the feature data of each feature dimension to obtain the proof information, which is used as the proof information corresponding to the model parameters in each feature dimension. The feature data of each feature dimension can be multiplied by the sample label to obtain the proof information, which is used as the proof information corresponding to the local training sample. This realizes the proof processing of different objects involved in model training.
[0059] In an optional implementation, based on the feature data of each dimension in the local training sample and the proof key, the first proof information corresponding to each model parameter is calculated respectively, including: The first proof information corresponding to each model parameter is calculated based on the following formula:
[0060]
[0061] Among them, the model parameters include feature weights and bias items corresponding to the feature data of each dimension; is the first certification information of the feature weight corresponding to the feature data of the j-th dimension; ∈[0, m], m is the total number of dimensions of the feature data in the local training sample; is the feature weight corresponding to the feature data of the mth dimension; n is the total number of local training samples; is the feature data of the jth dimension in the i-th training sample, i∈[1,n]; D is the first proof information of the bias item; is the bias term; Q is a point on the elliptic curve, [ ] represents elliptic curve multiplication, where, for Q points added together.
[0062] In the above formula, the formula is constructed by imitating the derivative function of the loss function, based on which the corresponding proof information of the model parameters and gradient calculation rules used by the participating nodes when calculating the gradient data is generated.
[0063] This process uses the elliptic curve in cryptography, selects points from it and uses the linear properties of the elliptic curve, which only has addition, subtraction and multiplication, to perform linear combinations, combining the points of the elliptic curve with the linear characteristics of linear regression to achieve the generation of proof information, effectively improving the validity of data proof in linear regression model training, and using cryptographic principles combined with federated linear regression technology to prevent poisoning attacks.
[0064] In an optional implementation, based on the feature data of each dimension in the local training sample and the sample label in the local training sample, calculating the second certification information corresponding to the local training sample includes: The second proof information corresponding to the local training sample is calculated based on the following formula:
[0065] in, is the second certification information; n is the total number of the local training samples; m is the total number of dimensions of the feature data in the local training samples; is the feature data of the mth dimension in the i-th training sample; i∈[1,n]; is the sample label of the i-th training sample in the local training sample.
[0066] In this implementation, the second proof information is formed as the balancing data required to verify the first proof information, ensuring the effective implementation of the honest proof data in subsequent data verification and improving the effectiveness of data proof in linear regression model training.
[0067] Step 103, sending the honesty proof information and gradient data to the central node.
[0068] Step 104: Obtain the updated model parameters and the certification keys of the updated model parameters sent by the central node.
[0069] Among them, the updated model parameters are obtained by the central node updating the model parameters after verifying that the gradient data is honest data based on the honest proof information.
[0070] After receiving the gradient data, the central node first verifies whether the data is honest. Only after the verification is passed will it use the data feedback from the participating nodes to calculate the results.
[0071] Here, when the central node updates the model parameters, it may select another set of random data and superimpose the integrated gradient data to obtain new model parameters. Alternatively, it may superimpose the previous set of model parameters on the integrated gradient data to obtain new model parameters. Alternatively, it may determine the learning coefficient based on the integrated gradient data, and multiply the learning coefficient by the previous set of model parameters to obtain new model parameters.
[0072] Based on the updated model parameters and the updated certification keys of each model parameter, return to executing step 102 to calculate the gradient data of the loss function of the target model for each model parameter based on the local training samples until the training of the target model is completed.
[0073] Each time the execution returns to step 102, the number of loop iterations of the model training increases by one.
[0074] The cutoff condition for completing the training of the target model may be when the value of the loss function of the target model is less than a set value, or when the number of iterations of the model training cycle reaches a threshold.
[0075] In the above implementation process, since it is impossible to determine whether any participating node deliberately fails to perform data calculations in the model training process according to the agreed rules, the participating nodes generate honesty proof information for the data objects and data calculation processes involved in the model training, so that the central node can verify the gradient data fed back by the participating nodes based on the honesty proof information to prevent poisoning attacks. The participating nodes need to perform some additional operations and use the results of these calculations as evidence to prove their innocence, prove the correctness of their own operations, and ensure that the data privacy of the participants is not leaked during the federated learning process.
[0076] See also Figure 2 , Figure 2 This is a process of a distributed learning method for a linear regression model provided in an embodiment of the present application. Figure 2 .like Figure 2 As shown, a distributed learning method of a linear regression model is applied to a central node, and the method includes the following steps: Step 201, sending the model parameters of the target model and the certification key of each model parameter to each participant node.
[0077] There are multiple participating nodes.
[0078] Optionally, the target model here is a linear regression model. Model parameters are specific parameters included in different types of models. When the target model is a linear regression model, the model parameters may include feature weights and bias items of feature variables in different dimensions.
[0079] The proof key is the data sent by the central node, which is used to enable the participating nodes to generate honesty proof information based on the proof key to prove the honesty of the data calculation of the participating nodes during the model training process based on local sample data.
[0080] The proof key may be a selected random number, or data calculated in other ways, such as the proof key obtained based on cryptographic calculations.
[0081] In an optional embodiment, sending the model parameters of the target model and the certification key of each model parameter to each participant node includes: Select initial parameters as model parameters of the target model; select random numbers from the domain parameters of the elliptic curve, and use the point product of the random number and the base point in the elliptic curve as the verification key according to the elliptic curve multiplication; use the point product of each model parameter and the verification key as the proof key of each model parameter according to the elliptic curve multiplication; send the model parameters of the target model and the proof keys of each model parameter to each participating node.
[0082] Among them, the verification key is used by the central node to verify the gradient data of each model parameter based on the honesty proof information sent by the participating nodes, and the gradient data is verified together with the honesty proof information.
[0083] This process uses the elliptic curve in cryptography, selects points from it and uses the linear property of the elliptic curve that only addition, subtraction and multiplication can be used. The points of the elliptic curve are combined with the linear characteristics of linear regression to generate the verification key and then obtain the proof key.
[0084] Step 202, receiving the honesty proof information sent by each participating node and the gradient data of the loss function of the target model to each model parameter.
[0085] Among them, the gradient data is calculated by each of the participating nodes based on local training samples; the honesty proof information is calculated by the participating nodes based on the proof key and local training samples.
[0086] The method for generating gradient data and honest proof information can be found in the description of the aforementioned implementation method, which will not be repeated here.
[0087] Step 203, verify the corresponding gradient data based on the honesty proof information, and after verifying that the gradient data is honest data, integrate the gradient data of each participating node to obtain model correction data.
[0088] Model calibration data is used to adjust model parameters.
[0089] Optionally, the gradient data of each participant node is integrated, and the gradient data of each model parameter is summed to obtain the model correction data. Alternatively, the gradient data of each model parameter is summed and divided by the number of gradient data from different participant nodes under the same model parameter to obtain the model correction data.
[0090] The corresponding gradient data can be verified based on the honest proof information, which can be reverse calculated according to the generation method of the honest proof information by the participating node to determine whether the relationship between the calculation result and the gradient data meets the expectations. If it meets the expectations, the gradient data is determined to be honest data, otherwise it is not. Alternatively, the corresponding gradient data can be verified in other ways based on the honest proof information.
[0091] In some optional implementations, the honest proof information includes first proof information and second proof information. The first proof information is calculated by the participant node based on the feature data of each dimension in the local training sample and the proof key; the second proof information is calculated by the participant node based on the feature data of each dimension in the local training sample and the sample label in the local training sample.
[0092] When verifying the gradient based on the honest proof information, it can be: According to elliptic curve multiplication, the point product of the second proof information and the aforementioned verification key is calculated as the target verification parameter; based on the difference between the sum of all the first proof information and the target verification parameter, the first verification value is calculated; based on the gradient of each model parameter and the point product of the aforementioned verification key, the second verification value is summed up; when the first verification value is equal to the second verification value, the gradient data is determined to be honest data.
[0093] Optionally, in one embodiment, verifying the corresponding gradient data based on the honest proof information includes: The first verification information and the second verification information are calculated based on the following formula:
[0094]
[0095] The model parameters include feature weights and bias items corresponding to feature data of each dimension; The first verification information; is the first certification information of the bias item; is the first certification information of the feature weight corresponding to the feature data of the j-th dimension in the model parameters; ∈[0, m], m is the total number of dimensions of the feature data in the local training sample; ; is the second proof information; Q is a point on the elliptic curve, [ ] represents elliptic curve multiplication, where, for Q points added together; The second verification information; is the feature weight corresponding to the feature data of the mth dimension in the local training sample; The gradient data of the feature weight corresponding to the feature data of the mth dimension of the loss function; is the bias term; is the gradient data of the loss function with respect to the bias term;
[0096]
[0097]
[0098] Wherein, n is the total number of local training samples; is the feature data of the jth dimension in the i-th training sample, i∈[1,n]; for Q points added together; is the sample label of the i-th training sample in the local training sample; When the first verification information is equal to the second verification information, the gradient data is determined to be honest data.
[0099] This process uses the elliptic curve in cryptography and the linear property of the elliptic curve that only addition, subtraction and multiplication can be used to perform linear combination. The elliptic curve is combined with the linear characteristics of linear regression to generate proof information, effectively improve the validity of data proof in linear regression model training, and use cryptographic principles combined with federated linear regression technology to prevent poisoning attacks.
[0100] Step 204: Update the model parameters based on the model correction data to obtain updated model parameters, and calculate the certification key of each updated model parameter.
[0101] Here, when the central node updates the model parameters based on the model correction data, it can select another set of random data and superimpose the integrated gradient data to obtain new model parameters. Alternatively, the new model parameters can be obtained by superimposing the previous set of model parameters on the integrated gradient data. Alternatively, the learning coefficient is determined based on the integrated gradient data, and the learning coefficient is multiplied by the previous set of model parameters to obtain new model parameters.
[0102] Return to step 201 to send the model parameters of the target model and the certification key of each model parameter to each participant node until the training of the target model is completed.
[0103] Each time step 201 is returned to be executed, the number of loop iterations of the model training increases by one.
[0104] The cutoff condition for completing the training of the target model may be when the value of the loss function of the target model is less than a set value, or when the number of iterations of the model training cycle reaches a threshold.
[0105] In an optional implementation, updating the model parameters based on the model correction data to obtain updated model parameters includes:
[0106]
[0107] in, is the updated feature weight corresponding to the feature data of the j-th dimension in the model parameters; is the learning rate; is the feature weight before updating corresponding to the feature data of the j-th dimension in the model parameters; is the derivative function of the loss function; the model calibration data includes and ; , is the gradient of the feature weight before updating corresponding to the feature data of the jth dimension, The total gradient after aggregating the gradients of the feature weights before updating corresponding to the feature data of the j-th dimension sent by all the participating nodes; is the updated bias term in the model parameters, is the bias term before updating in the model parameters; , is the gradient of the bias term before updating; It is the total gradient after aggregating the gradients of the bias items before the update sent by all the participating nodes.
[0108] In this process, the feature weights in the corresponding dimension are updated based on the gradient data of the integrated feature weights in each feature dimension, and the original bias items are updated based on the gradient data of the integrated bias items, so as to update different model parameters and obtain the updated model parameters, thereby ensuring the effectiveness of the model's iterative training.
[0109] In the above implementation process, since it is impossible to determine whether any participating node deliberately fails to perform data calculations in the model training process according to the agreed rules, the central node sends the proof key of each model parameter to the participating node during each round of model training, and then verifies the gradient data fed back by the participating node with the proof key and the honest proof information generated by the participating node to prevent poisoning attacks. The central node will only update the model parameters and feed back to each participating node after the gradient data verification is passed, ensuring that the data privacy of the participants is not leaked during the federated learning process while ensuring that the model is trained correctly and effectively.
[0110] In the following, the above implementation process is described in an example, where the model parameters of a linear regression model are 4 (w0, w1, w2, d), a central node, and two participant nodes. The data used for model training includes feature data of 3 dimensions (x0, x1, x2) and a sample label (y). The model parameters can actually be any number greater than 2, and the number of participant nodes can be any number greater than 1.
[0111] Negotiate and prepare relevant data before model training: The central node and each participating node agree in advance on the elliptic curve, including the curve equation, domain parameter F, base point G and other parameters. For example, these parameters can specifically adopt the recommended parameters in the fifth part of the SM2 national cryptographic standard (GM / T 0003.5-2012). In order to ensure the correctness of subsequent cryptographic operations and avoid calculation errors caused by the natural precision loss of floating-point numbers, the floating-point data is converted into fractions according to a certain precision, and all subsequent calculations can be selected as accurate calculations based on integers or integer ratios.
[0112] Linear regression assumes that there is a certain linear relationship between y and x0, x1, x2. The corresponding target model can be expressed as y=w0*x0+w1*x1+w2*x2+d. The purpose of training is to find the ideal values of these feature weights (w0, w1, w2) and the intercept d.
[0113] During the implementation, the central node determines the initial model parameters (initial w0, w1, w2, d). The initial model parameters can be randomly selected by the central node, or all set to a certain agreed value. The central node selects a random number r on F, calculates the verification key: Q=[r] and the proof key The central node will prove the key and initial model parameters (w0 w1 w2 d) Send to each participating node. Q is a point on the elliptic curve, for Q points are added together; when Q is the base point on the elliptic curve, it can be abbreviated as [ ].
[0114] The participating nodes use local sample data to calculate the gradient of the loss function on these model parameters. The commonly selected loss function is the mean square error (formula is , which can be regarded as a multivariate function of model parameters w0,w1,w2,d).
[0115] When the mean square error is specifically selected as the loss function, the derivative of the loss function with respect to w0 is: , its derivatives with respect to w0, w1, w2 and d are also similar formulas, and the vector composed of all these derivatives is the gradient data.
[0116] Assuming that a participant's local training sample has a total of n sample data, the calculations required are as follows:
[0117]
[0118]
[0119]
[0120] For the i-th data:
[0121]
[0122]
[0123]
[0124]
[0125]
[0126] All participating nodes will transfer gradient data ( , , , ) and honest proof information ( , , ,D,s) is sent to the central node.
[0127] After the central node receives data from a participant node, it first verifies the honest proof information. The central node performs the following calculations:
[0128]
[0129]
[0130] If T1 is equal to T2, then the validation is passed, otherwise the batch of gradient data is rejected.
[0131] After the gradient data is verified, the central node verifies the gradient data of each participant node ( , , , ) to get the total gradient. The aggregation method can be selected to sum up the gradients of all parties:
[0132]
[0133]
[0134]
[0135] The total gradient is then used to update the model parameters (e.g. ), the other model parameter update formulas are the same. Among them, is the learning rate, and the prime in the formula represents the derivative of the loss function with respect to w0. The specific updating model parameters are:
[0136]
[0137]
[0138]
[0139] The updated new model parameters are then sent to each participating node.
[0140] The above steps are repeated until the model training is completed.
[0141] The following example uses actual data and a possible scenario. Suppose there are two participant nodes A and B, each of which has a portion of customer sales data. Participant node A shows customer sales data for user IDs 1 and 2, and participant node B shows customer sales data for user IDs 3 and 4 (this is just an example and is only used to demonstrate the algorithm content).
[0142]
[0143]
[0144] Now we want to train a model through federated linear regression to try to provide credit scores for new users.
[0145] The central node selects a random number r=3, the initial model parameters are (3,3,100), the elliptic curve base point is G, and sends the following data to the participating nodes A and B: (3,3,100) and (9G,9G,300G). Among them, (9G, 9G, 300G) can also be expressed as ([9],[9],
[300] ).
[0146] Participant node A calculates the gradient:
[0147]
[0148]
[0149]
[0150]
[0151]
[0152] s = (620* 10000+620*3000+620)+(330* 5000+330*2000+330)= 10370950; Participant node A sends (488650000, 156980000, 59250), ([1489500000], [478500000], [180600]) and 10370950 to the central node.
[0153] Participant node B processes the data using the same principle and sends it to the central node.
[0154] The central node verifies the data fed back by participant node A: ;
[0155]
[0156] If the two results are consistent, the verification is confirmed to be successful.
[0157] The central node verifies the data fed back by the participant node B in the same way.
[0158] The central node updates the model parameters based on the gradient data fed back by the participating nodes A and B, and then sends the new model parameters to the two participating nodes, starting a new round of processing operations until the model training is completed.
[0159] The above scheme of the embodiment of the present application requires the participating nodes to prove that their calculation process is correct and honest through some cryptographic calculations by means of active defense. The central node receives the data and verifies it. Only when the verification is passed will the relevant calculation results be used. During the implementation process, the calculation mainly occurs at the participating nodes, and the central node mainly performs verification operations, so there is no need to consume too much computing power of the central node. The overall amount of calculation is small, and the calculation process is relatively simple, ensuring the efficiency of model training processing. And the possibility of malicious people wanting to pretend to verify it through the central node decreases exponentially with the setting of the verification key. For example, when a 256-bit verification key is used, the probability is 1 in 2 to the power of 256. If necessary, the security parameters can be set to 512, 1024 until any computing power under a traditional computer cannot be fooled, ensuring the reliability of verifying the honesty of the feedback information of the participating nodes.
[0160] See also Figure 3 , Figure 3 This is a structural diagram of a distributed learning device for the first linear regression model provided in an embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0161] The first distributed learning device 300 of the linear regression model includes: The acquisition module 301 is used to acquire the model parameters of the target model and the certification key of each model parameter sent by the central node; A calculation module 302 is used to calculate the gradient data of the loss function of the target model to each model parameter based on the local training sample, and to calculate the honesty proof information based on the proof key and the local training sample; A first sending module 303, used to send the honesty proof information and the gradient data to the central node; The first update module 304 is used to obtain the updated model parameters sent by the central node and the proof key of each updated model parameter, and return to execute the step of calculating the gradient data of each model parameter of the loss function of the target model based on the local training sample until the training of the target model is completed; wherein the updated model parameters are obtained by updating the model parameters after the central node verifies that the gradient is honest data based on the honest proof information.
[0162] See also Figure 4 , Figure 4 This is a structural diagram of a distributed learning device for a second linear regression model provided in an embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0163] The second distributed learning device 400 of the linear regression model includes: The second sending module 401 is used to send the model parameters of the target model and the certification key of each model parameter to each participant node; The receiving module 402 is used to receive the honesty proof information sent by each of the participating nodes and the gradient data of the loss function of the target model to each model parameter; the gradient data is calculated by each of the participating nodes based on the local training samples; the honesty proof information is calculated by the participating node based on the proof key and the local training samples; A verification module 403 is used to verify the corresponding gradient data based on the honesty proof information, and after verifying that the gradient data is honest data, integrate the gradient data of each of the participating nodes to obtain model correction data; The second update module 404 is used to update the model parameters based on the model correction data to obtain updated model parameters, and calculate the proof key of each updated model parameter, and return to execute the step of sending the model parameters of the target model and the proof key of each model parameter to each participating node until the target model training is completed.
[0164] The distributed learning device for the two linear regression models provided in the embodiment of the present application can correspond to the various processes of the embodiment of the distributed learning method for the above-mentioned two linear regression models, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0165] Figure 5 is a structural diagram of a computer device provided in an embodiment of the present application. As shown in the figure, the computer device 5 of this embodiment includes: at least one processor 50 ( Figure 5 Only one is shown in the figure), a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, wherein the processor 50 implements the steps of any of the above-mentioned method embodiments when executing the computer program 52.
[0166] The computer device 5 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The computer device 5 may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will appreciate that Figure 5 It is only an example of the computer device 5 and does not constitute a limitation of the computer device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, etc.
[0167] The processor 50 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0168] The memory 51 may be an internal storage unit of the computer device 5, such as a hard disk or memory of the computer device 5. The memory 51 may also be an external storage device of the computer device 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 5. Further, the memory 51 may also include both an internal storage unit and an external storage device of the computer device 5. The memory 51 is used to store the computer program and other programs and data required by the computer device. The memory 51 may also be used to temporarily store data that has been output or is to be output.
[0169] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0170] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0171] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0172] In the embodiments provided in the present application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which can be electrical, mechanical or other forms.
[0173] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0174] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0175] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, electric signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0176] The present application implements all or part of the processes in the above-mentioned embodiment methods, and may also be implemented through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0177] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A distributed learning method for a linear regression model, characterized in that: Applied to a participant node, the method comprises: Obtain the model parameters of the target model and the certification keys of each model parameter sent by the central node; Calculate the gradient data of the loss function of the target model to each model parameter based on the local training sample, and calculate the honesty proof information based on the proof key and the local training sample; Sending the honesty proof information and the gradient data to the central node; Obtain the updated model parameters and the proof keys of each updated model parameter sent by the central node, and return to execute the step of calculating the gradient data of each model parameter of the loss function of the target model based on the local training samples until the target model training is completed; wherein the updated model parameters are obtained by updating the model parameters after the central node verifies that the gradient data is honest data based on the honesty proof information.
2. The method according to claim 1, characterized in that The step of calculating the gradient data of the loss function of the target model to each model parameter based on the local training sample includes: Inputting the feature data in the local training sample into the target model to obtain an output prediction result; The sample labels in the local training samples and the prediction results are input into the derivative functions of each model parameter under the loss function of the target model, and the gradient data of the loss function for each model parameter is calculated.
3. The method according to claim 2, characterized in that The step of inputting the sample labels in the local training samples and the prediction results into the derivative functions of each model parameter under the loss function of the target model, and calculating the gradient data of the loss function with respect to each model parameter, comprises: The sample labels in the local training samples and the prediction results are input into the following derivative function to calculate the gradient data of the loss function for each model parameter: in, is the feature weight corresponding to the feature data of the jth dimension in the model parameters; is the derivative function of the loss function; is the gradient data of the feature weight corresponding to the feature data of the j-th dimension of the loss function; ∈[0, m], m is the total number of dimensions of the feature data in the local training samples; n is the total number of local training samples; is the feature data of the jth dimension in the i-th training sample, i∈[1,n]; is the bias term in the model parameters; is the gradient data of the loss function with respect to the bias term; is the prediction result output by the target model based on the feature data in the i-th training sample in the local training sample, is the sample label of the i-th training sample in the local training sample.
4. The method according to claim 1, characterized in that: The calculating and obtaining the honesty proof information based on the proof key and the local training sample includes: Based on the feature data of each dimension in the local training sample and the proof key, respectively calculate the first proof information corresponding to each of the model parameters; Calculate second certification information corresponding to the local training sample based on the feature data of each dimension in the local training sample and the sample label in the local training sample; The honesty certification information including the first certification information and the second certification information is obtained.
5. The method according to claim 4, characterized in that The calculating, based on the feature data of each dimension in the local training sample and the proof key, the first proof information corresponding to each model parameter includes: The first proof information corresponding to each of the model parameters is calculated based on the following formula: The model parameters include feature weights and bias items corresponding to feature data of each dimension; is the first certification information of the feature weight corresponding to the feature data of the j-th dimension; ∈[0, m], m is the total number of dimensions of the feature data in the local training sample; is the feature weight corresponding to the feature data of the mth dimension; n is the total number of local training samples; is the feature data of the jth dimension in the i-th training sample, i∈[1,n]; D is the first proof information of the bias item; is the bias term; Q is a point on the elliptic curve, [ ] represents elliptic curve multiplication, where, for Q points added together.
6. The method according to claim 4, characterized in that The calculating, based on the feature data of each dimension in the local training sample and the sample label in the local training sample, the second certification information corresponding to the local training sample includes: The second proof information corresponding to the local training sample is calculated based on the following formula: in, is the second certification information; n is the total number of the local training samples; m is the total number of dimensions of the feature data in the local training samples; is the feature data of the mth dimension in the i-th training sample; i∈[1,n]; is the sample label of the i-th training sample in the local training sample.
7. A distributed learning device for a linear regression model, characterized in that: include: An acquisition module, used to acquire the model parameters of the target model and the certification keys of each model parameter sent by the central node; A calculation module, used to calculate the gradient data of the loss function of the target model to each model parameter based on the local training sample, and to calculate the honesty proof information based on the proof key and the local training sample; A first sending module, used for sending the honesty proof information and the gradient data to the central node; The first update module is used to obtain the updated model parameters sent by the central node and the proof key of each updated model parameter, and return to execute the step of calculating the gradient data of each model parameter of the loss function of the target model based on the local training sample until the training of the target model is completed; wherein the updated model parameters are obtained by updating the model parameters after the central node verifies that the gradient is honest data based on the honesty proof information.
8. A distributed learning method for a linear regression model, characterized in that: Applied to the central node, the method includes: Send the model parameters of the target model and the certification keys of each model parameter to each participating node; Receive the honesty proof information sent by each of the participant nodes and the gradient data of the loss function of the target model to each model parameter; the gradient data is calculated by each of the participant nodes based on local training samples; the honesty proof information is calculated by the participant node based on the proof key and the local training samples; Verifying the corresponding gradient data based on the honesty proof information, and after verifying that the gradient data is honest data, integrating the gradient data of each of the participating nodes to obtain model correction data; The model parameters are updated based on the model correction data to obtain updated model parameters, and the proof keys of each updated model parameter are calculated, and the step of sending the model parameters of the target model and the proof keys of each model parameter to each participating node is returned to execute until the target model training is completed.
9. The method according to claim 8, characterized in that The sending of the model parameters of the target model and the certification keys of each model parameter to each participant node includes: Selecting initial parameters as model parameters of the target model; Selecting a random number from a domain parameter of an elliptic curve, and using the multiplication result of the random number and a base point in the elliptic curve as a verification key according to elliptic curve multiplication; Using the elliptic curve multiplication method to calculate the point multiplication result of each model parameter and the verification key as the proof key of each model parameter; The model parameters of the target model and the certification key of each model parameter are sent to each of the participant nodes.
10. The method according to claim 8, characterized in that The honest proof information includes first proof information and second proof information; and the verifying the corresponding gradient data based on the honest proof information includes: The first verification information and the second verification information are calculated based on the following formula: The model parameters include feature weights and bias items corresponding to feature data of each dimension; The first verification information; is the first certification information of the bias item; is the first certification information of the feature weight corresponding to the feature data of the j-th dimension in the model parameters; ∈[0, m], m is the total number of dimensions of the feature data in the local training sample; ; is the second proof information; Q is a point on the elliptic curve, [ ] represents elliptic curve multiplication, where, for Q points added together; The second verification information; is the feature weight corresponding to the feature data of the mth dimension in the local training sample; The gradient data of the feature weight corresponding to the feature data of the mth dimension of the loss function; is the bias term; is the gradient data of the loss function with respect to the bias term; Wherein, n is the total number of local training samples; is the feature data of the jth dimension in the i-th training sample, i∈[1,n]; for Q points added together; is the sample label of the i-th training sample in the local training sample; When the first verification information is equal to the second verification information, the gradient data is determined to be honest data.
11. The method according to claim 8, characterized in that The updating of the model parameters based on the model correction data to obtain updated model parameters includes: in, is the updated feature weight corresponding to the feature data of the j-th dimension in the model parameters; is the learning rate; is the feature weight before updating corresponding to the feature data of the j-th dimension in the model parameters; is the derivative function of the loss function; , is the gradient of the feature weight before updating corresponding to the feature data of the jth dimension, The total gradient after aggregating the gradients of the feature weights before updating corresponding to the feature data of the j-th dimension sent by all the participating nodes; is the updated bias term in the model parameters, is the bias term before updating in the model parameters; , is the gradient of the bias term before update; It is the total gradient after aggregating the gradients of the bias items before the update sent by all the participating nodes.
12. A distributed learning device for a linear regression model, characterized in that: include: The second sending module is used to send the model parameters of the target model and the certification key of each model parameter to each participant node; A receiving module, used to receive the honesty proof information sent by each of the participating nodes and the gradient data of the loss function of the target model to each model parameter; the gradient data is calculated by each of the participating nodes based on the local training samples; the honesty proof information is calculated by the participating node based on the proof key and the local training samples; A verification module, used to verify the corresponding gradient data based on the honesty proof information, and after verifying that the gradient data is honest data, integrate the gradient data of each of the participating nodes to obtain model correction data; The second updating module is used to update the model parameters based on the model correction data to obtain updated model parameters, calculate the proof keys of each updated model parameter, and return to execute the step of sending the model parameters of the target model and the proof keys of each model parameter to each participating node until the training of the target model is completed.
13. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 or 8 to 11 are implemented.
14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 or 8 to 11 are implemented.
Citation Information
Patent Citations
Federal learning processing method, device and system based on block chain and medium
CN114443754A
Multi-party federal learning method, system and device, storage medium and program product
CN119443317A
Joint learning model iterative update method, apparatus, system, and storage medium
WO2023124219A1