Joint Training Method, Device, Equipment, Medium and Product of Model
By splitting and sending the training data matrix of the first training participant in joint training of the model, and updating the training model with the gradient feedback of the second training participant and the computing service provider, the problem of low privacy and security of the model joint training data in the prior art is solved, and data security and system performance are improved.
Patent Information
- Application Number
- CN202411783738.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-06
AI Technical Summary
The existing joint training method of model ensures data privacy and security, while also leading to low system performance.
By splitting the training data matrix held by the first training participant, the local matrix, the first shared matrix and the second shared matrix are obtained, and the first shared matrix is sent to the second training participant, and the second shared matrix is sent to the computing service provider. Finally, the training model is updated using the gradient shard value feedbacked by the second training participant and the computing service provider and the gradient shard value generated by the first training participant.
On the premise of ensuring the joint training effect, the privacy of the training data of the first training participants is protected and the security of the training data is improved.
Smart Images

Figure CN119249158B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment, medium and product for jointly training models. Background Art
[0002] Model joint training is a technology that improves the overall performance by simultaneously training multiple models and utilizing the interaction and information sharing between them. Through the collaborative work of multiple models, this method can solve challenges that a single model may face, such as insufficient data, overfitting, etc., thereby improving the accuracy and stability of prediction.
[0003] However, to ensure data privacy and security, current model joint training methods are usually implemented based on homomorphic encryption and fully secret sharing technologies, which results in low performance of the entire system. Summary of the Invention
[0004] The present invention provides a method, device, equipment, medium and product for jointly training models to solve the problem of low security of training data in existing model joint training.
[0005] According to one aspect of the present invention, there is provided a method for jointly training models, including:
[0006] Splitting a first training data matrix held by a first training participant to obtain a target number of first split matrices, and determining a local matrix, a first shared matrix, and a second shared matrix from the first split matrices, where the local matrix is different from the first shared matrix and the second shared matrix;
[0007] Sending the first shared matrix to a second training participant and sending the second shared matrix to a computing service provider;
[0008] Updating first to-be-updated model parameters of a first to-be-trained model of the first training participant in a current iteration round according to a first gradient first shard value sent by the second training participant, a first gradient second shard value sent by the computing service provider, and a first gradient third shard value generated by the first training participant;
[0009] Wherein, the first gradient first shard value is determined by the second training participant according to the first shared matrix, the first gradient second shard value is determined by the computing service provider according to the second shared matrix, and the first gradient third shard value is determined by the first training participant according to the local matrix.
[0010] According to another aspect of the present invention, there is provided a device for jointly training models, including:
[0011] A matrix splitting module, which is used to split the first training data matrix held by the first training participant to obtain a target number of first split matrices, and determine a local matrix, a first shared matrix, and a second shared matrix from the first split matrices, where the local matrix is different from the first shared matrix and the second shared matrix;
[0012] A matrix sending module, which is used to send the first shared matrix to the second training participant and send the second shared matrix to the computing service party;
[0013] A model parameter updating module, which is used to update the first model parameter to be updated of the first training participant's first model to be trained in the current iteration round according to the first shard value of the first gradient sent by the second training participant, the second shard value of the first gradient sent by the computing service party, and the third shard value of the first gradient generated by the first training participant;
[0014] Wherein, the first shard value of the first gradient is determined by the second training participant according to the first shared matrix, the second shard value of the first gradient is determined by the computing service party according to the second shared matrix, and the third shard value of the first gradient is determined by the first training participant according to the local matrix.
[0015] According to another aspect of the present invention, an electronic device is provided, and the electronic device includes:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the joint training method of the model according to any one of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the joint training method of the model according to any one of the present invention when executed by a processor.
[0020] According to another aspect of the present invention, a computer program product is provided, including a computer program, and the computer program implements the joint training method of the model according to any one of the present invention when executed by a processor.
[0021] In the present invention, by introducing a computing service provider during the joint training process and splitting the first training data matrix held by the first training participant to obtain a local matrix, a first shared matrix, and a second shared matrix, then sending the first shared matrix to the second training participant and sending the second shared matrix to the computing service provider, and finally using the gradient shard values respectively fed back by the second training participant and the computing service provider and the gradient shard values generated by the first training participant to train the first model to be trained of the first training participant, the second training participant cannot independently restore the complete first training data matrix, which protects the privacy of the training data of the first training participant on the premise of ensuring the joint training effect and further improves the security of the training data.
[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0024] Figure 1 It is a flowchart of a method for joint training of a model provided in Embodiment 1 of the present invention;
[0025] Figure 2 It is a flowchart of a method for joint training of a model provided in Embodiment 2 of the present invention;
[0026] Figure 3 It is a schematic structural diagram of a device for joint training of a model provided in Embodiment 3 of the present invention;
[0027] Figure 4 It is a schematic structural diagram of an electronic device for implementing the method for joint training of the model in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first", "second", "local", "shared", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0030] Embodiment 1
[0031] Figure 1 FIG. is a flowchart of a method for joint training of a model provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of joint training of a model between a first training participant and a second training participant. This method can be executed by a joint training device of the model, and the joint training device of the model can be implemented in the form of hardware and / or software. As Figure 1 shown, the method includes:
[0032] S101. Split the first training data matrix held by the first training participant to obtain a target number of first split matrices, and determine a local matrix, a first shared matrix, and a second shared matrix from the first split matrices.
[0033] Among them, the first training participant refers to any one of at least two training participants participating in the joint training of the model. For example, assuming that the training participants participating in the joint training of the model include Training Participant 1 and Training Participant 2, then the first training participant can be either Training Participant 1 or Training Participant 2. In other words, any training participant participating in the joint training of the model can be used as the first training participant in this embodiment.
[0034] The first training data matrix refers to the training data matrix held by the first training participant, and the training data matrix refers to the matrix representation of the training data.
[0035] Taking the model application scenario as the loan amount prediction scenario as an example, the training data matrix is illustrated as follows: In the loan amount prediction scenario, assuming that the training participants include a certain management institution and a certain lending bank, the training data matrix held by the certain management institution can be expressed as:
[0036] ;
[0037] The training data matrix held by a lending bank can be represented as:
[0038] ;
[0039] In this embodiment, only the above examples are used to illustrate the training data matrix, rather than limiting the specific content of the training data matrix. Moreover, in this embodiment, only the above examples are used to illustrate the model application scenarios, rather than limiting the model application scenarios. The model application scenarios in this embodiment include, but are not limited to, loan amount prediction, housing price prediction, investment risk analysis, crop fertilization amount analysis, student learning ability assessment, patient organ function prediction, etc.
[0040] In one implementation, the first training participant splits the first training data matrix to obtain a target number of first split matrices. Among them, the target number is at least two and can be set and adjusted according to actual business needs. Taking the target number as three as an example, how to split the first training data matrix is illustrated as follows:
[0041] Suppose the first training data matrix , then the first training data matrix is split into three first split matrices: , and , and it satisfies . Among them, , and can be represented in matrix form as:
[0042]
[0043]
[0044]
[0045] In this embodiment, only the above examples are used to illustrate how to split the first training data matrix, rather than limiting the target number corresponding to the first split matrix.
[0046] After obtaining the target number of first split matrices, the first training participant determines the first split matrices that are not sent to the second training participant and the computing service party from the first split matrices as local matrices, the first split matrices used to be sent to the second training participant as first shared matrices, and the first split matrices used to be sent to the computing service party as second shared matrices. Among them, the local matrices are different from the first shared matrices and the second shared matrices. The first shared matrices and the second shared matrices are only used to distinguish different first split matrices, rather than referring to a specific first split matrix.
[0047] It is understandable that, in order to ensure the security of the training data, the first shared matrix and the second shared matrix are respectively only a part of the first split matrix, rather than the entire first split matrix. For example, assume that the first split matrix includes , and , then the first shared matrix / second shared matrix can be { , }, { , } or { , }.
[0048] Preferably, in order to ensure the training effect of the joint training of the model, the number of matrices of the first shared matrix and the second shared matrix is set to be the same, and the combination between the first shared matrix and the second shared matrix can restore the first training data matrix. For example, assume that the first split matrix includes , and , when the first shared matrix is { , }, the second shared matrix can be { , } or { , }; when the first shared matrix is { , }, the second shared matrix can be { , } or { , }; when the first shared matrix is { , }, the second shared matrix can be { , } or { , }.
[0049] This embodiment only gives examples of the first shared matrix and the second shared matrix with the above examples, rather than limiting the first shared matrix and the second shared matrix.
[0050] S102. Send the first shared matrix to the second training participant, and send the second shared matrix to the computing service provider.
[0051] Among them, the second training participant refers to the training participants other than the first training participant among the training participants participating in the joint training of the model. For example, assuming that the training participants participating in the joint training of the model include training participant 1 and training participant 2, then if training participant 1 is the first training participant, training participant 2 is the second training participant; if training participant 2 is the first training participant, training participant 1 is the second training participant. The computing service provider is a third-party service provider that can provide data computing services for any training participant, such as a data computing server, etc.
[0052] In one implementation, the first training participant sends the first shared matrix to the second training participant through the data channel between the first training participant and the second training participant, and sends the second shared matrix to the computing service provider through the data channel between the first training participant and the computing service provider.
[0053] Exemplarily, assume that the first shared matrix is { 、 }, the second shared matrix is { 、 }, the first training participant then sends { 、 } to the second training participant, and sends { 、 } to the computing service provider.
[0054] S103. Update the first model parameters to be updated of the first model to be trained by the first training participant in the current iteration round according to the first shard value of the first gradient sent by the second training participant, the second shard value of the first gradient sent by the computing service provider, and the third shard value of the first gradient generated by the first training participant.
[0055] Among them, the first shard value of the first gradient is determined by the second training participant according to the first shared matrix, the second shard value of the first gradient is determined by the computing service provider according to the second shared matrix, and the third shard value of the first gradient is determined by the first training participant according to the local matrix.
[0056] In one implementation, the second training participant determines the second error vector according to the second model parameters to be updated of the corresponding second model to be trained in the current iteration round and the second training data matrix held by the second training participant. The second training participant splits the second error vector to obtain a target number of first split vectors, and determines a local vector, a first shared vector, and a second shared vector from the first split vectors. After receiving the first shared matrix, the second training participant determines the first shard value of the first gradient according to the local vector and the first shared matrix, and sends the first shard value of the first gradient to the first training participant.
[0057] Meanwhile, the second training participant sends the second shared vector to the computing service party. The computing service party determines the second shard value of the first gradient based on the second shared vector sent by the second training participant and the second shared matrix sent by the first training participant, and then the computing service party sends the second shard value of the first gradient to the first training participant.
[0058] Meanwhile, the first training participant calculates the third shard value of the first gradient according to the local matrix and the first shared vector sent by the second training participant.
[0059] Furthermore, the first training participant determines the first gradient value based on the first shard value of the first gradient, the second shard value of the first gradient, and the third shard value of the first gradient, and determines the first error vector according to the first model parameter to be updated and the first training data matrix. Further, the first training participant determines the second gradient value according to the first error vector, and determines the total first iterative gradient value of the first model to be trained in the current iteration round according to the first gradient value and the second gradient value. Finally, the first training participant performs model training on the first model to be trained according to the total first iterative gradient value, that is, updates the first model parameter to be updated of the first model to be trained by the first training participant in the current iteration round.
[0060] In the embodiment of the present invention, by introducing a computing service party during the joint training process and splitting the first training data matrix held by the first training participant into a local matrix, a first shared matrix, and a second shared matrix, then sending the first shared matrix to the second training participant and sending the second shared matrix to the computing service party, and finally using the gradient shard values respectively fed back by the second training participant and the computing service party and the gradient shard value generated by the first training participant to train the first model to be trained by the first training participant, the second training participant cannot independently restore the complete first training data matrix, which protects the privacy of the training data of the first training participant while ensuring the joint training effect, and further improves the security of the training data.
[0061] Embodiment 2
[0062] Figure 2 The figure is a flowchart of a method for joint training of a model provided by Embodiment 2 of the present invention. This embodiment further optimizes and expands the above embodiment and can be combined with the above various optional implementation manners. As Figure 2 shown, the method includes:
[0063] S201. Split the first training data matrix held by the first training participant to obtain a target number of first split matrices, and determine a local matrix, a first shared matrix, and a second shared matrix from the first split matrices.
[0064] S202. Send the first shared matrix to the second training participant and send the second shared matrix to the computing service provider.
[0065] S203. Determine the first gradient value according to the first shard value of the first gradient, the second shard value of the first gradient, and the third shard value of the first gradient, and determine the first prediction vector according to the first model parameter to be updated of the first model to be trained and the first training data matrix.
[0066] In one implementation, the first training participant determines the first gradient value according to the sum of the first shard value of the first gradient, the second shard value of the first gradient, and the third shard value of the first gradient. Moreover, the first training participant determines the first prediction vector according to the product result between the first model parameter to be updated and the first training data matrix.
[0067] Exemplarily, assume that the first model parameter to be updated is , and the first training data matrix is , then the first prediction vector .
[0068] Optionally, the determining the first gradient value according to the first shard value of the first gradient, the second shard value of the first gradient, and the third shard value of the first gradient includes:
[0069] S2031. Obtain the first shared vector sent by the second training participant, and calculate the third shard value of the first gradient according to the first shared vector and the local matrix.
[0070] Wherein,
[0071] the first shared vector is determined in the following manner:
[0072] The second training participant respectively determines the second error vector according to the second model parameter to be updated of the corresponding second model to be trained in the current iteration round and the second training data matrix held by the second training participant; the second training participant splits the second error vector to obtain a target number of first split vectors, and determines the first shared vector from the first split vectors.
[0073] Wherein, the second model to be trained refers to the model to be trained corresponding to each second training participant.
[0074] In one embodiment, the second training participant determines a second error vector based on the second model parameters to be updated and the second training data matrix, and splits the second error vector to obtain a target number of first split vectors. After obtaining the target number of first split vectors, the first training participant determines, from the first split vectors, the first split vectors to be sent to the first training participant as first shared vectors. Further, the second training participant sends the first shared vectors to the first training participant.
[0075] The first training participant obtains the first shared vectors sent by the second training participant, and determines a first gradient third shard value based on the product result between the first shared vectors and the local matrix.
[0076] Exemplarily, assume that the first shared vectors include and . The local matrix includes .
[0077] Then, according to the first gradient third gradient shard value is:
[0078] .
[0079] S2032. Determine a first gradient value based on the sum of the first gradient first shard value, the first gradient second shard value, and the first gradient third shard value.
[0080] Exemplarily, assume that the first gradient first shard value is , the first gradient second shard value is the first gradient third shard value is , then the first gradient value is + +
[0081] By obtaining the first shared vectors sent by the second training participant and calculating the first gradient third shard value based on the first shared vectors and the local matrix; and determining the first gradient value based on the sum of the first gradient first shard value, the first gradient second shard value, and the first gradient third shard value, the training of the first model to be trained not only depends on the gradient shard values generated by the first model to be trained itself, but also depends on the gradient shard values generated by the second model to be trained, thus ensuring the training effect of the joint training of the models.
[0082] S204. Determine a first error vector based on the first standard vector and the first prediction vector, and determine a second gradient value based on the first error vector and the first training data matrix.
[0083] In one implementation, the first training participant determines a first error vector based on the difference between the first standard vector and the first prediction vector, and determines a second gradient value based on the product result of the first error vector and the first training data matrix.
[0084] Exemplarily, assume that the first error vector is , and the first training data matrix is The second gradient value .
[0085] S205. Determine a first total iterative gradient of the first model to be trained in the current iteration round according to the first gradient value and the second gradient value, and update the first model parameters to be updated according to the first total iterative gradient.
[0086] In one implementation, the first training participant determines a first total iterative gradient of the first model to be trained in the current iteration round according to the sum of the first gradient value and the second gradient value. Exemplarily, assume that the first gradient value is , and the second gradient value is , then the first total iterative gradient = .
[0087] After determining the first total iterative gradient of the current iteration round, the first training participant updates the first model to be trained of the first training participant by using the first total iterative gradient, that is, updates the first model parameters to be updated of the first model to be trained of the first training participant in the current iteration round.
[0088] By determining the first gradient value according to the first gradient first shard value, the first gradient second shard value, and the first gradient third shard value, and determining the first prediction vector according to the first model parameters to be updated of the first model to be trained and the first training data matrix; determining the first error vector according to the first standard vector and the first prediction vector, and determining the second gradient value according to the first error vector and the first training data matrix; determining the first total iterative gradient of the first model to be trained in the current iteration round according to the first gradient value and the second gradient value, and updating the first model parameters to be updated according to the first total iterative gradient, the joint training of the first model to be trained is enabled, thereby ensuring the training effect of the model joint training.
[0089] Optionally, updating the first model parameters to be updated according to the first total iterative gradient includes:
[0090] S2051. Obtain the current learning rate of the first model to be trained in the current iteration round, and obtain the current sample size corresponding to the first training data matrix.
[0091] Among them, the current learning rate refers to the speed at which the first model to be trained updates its knowledge or parameters in the current iteration round. The current sample size reflects the scale of the training data of the current training data.
[0092] S2052. Update the first model parameters to be updated according to the current learning rate, the current sample size, and the total value of the first iteration gradients, so as to obtain the first updated model parameters of the first model to be trained in the current iteration round.
[0093] In one implementation, the first training participant determines the model change parameters generated in the current iteration round according to the current learning rate, the current sample size, and the total value of the first iteration gradients, updates the first model parameters to be updated according to the model change parameters, and uses the updated model parameters as the first updated model parameters.
[0094] By obtaining the current learning rate of the first model to be trained in the current iteration round, and obtaining the current sample size corresponding to the first training data matrix; updating the first model parameters to be updated according to the current learning rate, the current sample size, and the total value of the first iteration gradients, so as to obtain the first updated model parameters of the first model to be trained in the current iteration round, the effect of training the first model to be trained by the first training participant in the model joint training scenario is achieved, and the business requirements of the first training participant for model training are met.
[0095] Optionally, updating the first model parameters to be updated according to the current learning rate, the current sample size, and the total value of the first iteration gradients, so as to obtain the first updated model parameters of the first model to be trained in the current iteration round includes:
[0096] The first updated model parameters are obtained through the following formula:
[0097] ;
[0098] Among them, is the first updated model parameter, is the first model parameter to be updated, is the current learning rate, is the total value of the first iteration gradients, is the current sample size.
[0099] Determining the first updated model parameters through the formula can ensure the effect and accuracy of parameter update.
[0100] Optionally, the second training participant determines the first shard value of the first gradient in the following manner:
[0101] A1. Determine the second prediction vector according to the second model parameters to be updated corresponding to the second model to be trained and the corresponding second training data matrix.
[0102] In one implementation, the second training participant determines a second prediction vector based on the product result between the second model parameters to be updated and the second training data matrix.
[0103] Exemplarily, assume that the second model parameters to be updated are , and the second training data matrix is , then the second prediction vector .
[0104] B1. Determine a second error vector based on the second standard vector and the second prediction vector, split the second error vector to obtain a target number of first split vectors, and determine a local vector, a second shared vector, and a first shared vector from the first split vectors.
[0105] In one implementation, the second training participant determines the second error vector based on the difference between the second standard vector and the second prediction vector. Further, the second training participant splits the second error vector to obtain a target number of first split vectors. Wherein, the target number of the first split vectors is the same as the target number of the first split matrix. Taking the target number being three as an example for illustration:
[0106] For example, assume the second error vector , then split the second error vector into three first split vectors: , and , and satisfy . Wherein, , and in matrix form can be expressed as:
[0107]
[0108] This embodiment only uses the above example to illustrate how to split the second error vector, rather than limiting the target number corresponding to the first split vector.
[0109] After obtaining the target number of first split vectors, the second training participant determines the first split vectors that are not sent to the first training participant and the computing service party as local vectors, the first split vectors used to be sent to the first training participant as the first shared vectors, and the first split vectors used to be sent to the computing service party as the second shared vectors. Wherein, the first shared vectors and the second shared vectors are only used to distinguish different first split vectors, rather than referring to a specific first split vector.
[0110] The local vector refers to the first split vector selected by the second training participant from the first split vector and not sent to the first training participant and the computing service party. The local vector is used to combine the first shared matrix sent by the first training participant to calculate the first shard value of the first gradient. The local vector is different from the first shared vector and the second shared vector.
[0111] C1. Determine the first shard value of the first gradient according to the first shared matrix and the local vector.
[0112] In one implementation, the second training participant determines the first shard value of the first gradient according to the product result between the first shared matrix and the local vector.
[0113] Exemplarily, assume the local vector includes and E 23 =[ e2 11 3 … e2 n1 3 ] . The first shared matrix includes .
[0114] Then according to the first shard value of the first gradient is:
[0115] .
[0116] The second training participant determines the second prediction vector according to the second model parameter to be updated corresponding to the second model to be trained and the corresponding second training data matrix; determines the second error vector according to the second standard vector and the second prediction vector, splits the second error vector to obtain the target number of first split vectors, and determines the local vector, the second shared vector and the first shared vector from the first split vectors; determines the first shard value of the first gradient according to the first shared matrix and the local vector, achieving the effect of calculating the gradient value required for the first training participant to perform model training by using the error vector of the second training participant, so that the first model to be trained learns the data values of both itself and the second model to be trained, further ensuring the effect of joint model training.
[0117] Optionally, the computing service party determines the second shard value of the first gradient in the following manner:
[0118] Obtain the second shared vector sent by the second training participant; determine the second shard value of the first gradient according to the second shared matrix and the second shared vector.
[0119] In one implementation, the second training participant sends the obtained second shared vector to the computing service party. After the computing service party obtains the second shared vector sent by the second training participant, it determines the second shard value of the first gradient according to the product result between the second shared matrix and the second shared vector.
[0120] Exemplarily, assume that the second shared vector includes and E 23 =[ e2 11 3 … e2 n1 3 ] . The second shared matrix includes .
[0121] Then, according to , , the first gradient second shard value is:
[0122] .
[0123] The computing party obtains the second shared vector sent by the second training party; determines the first gradient second shard value according to the second shared matrix and the second shared vector, so as to jointly calculate the gradient shard value by the computing party and the second training party, protecting the privacy of the training data of the first training party while ensuring the joint training effect, and further improving the security of the training data.
[0124] Optionally, it further includes:
[0125] A2. Determine the second total iteration gradient of the second model to be trained in the current iteration round.
[0126] Among them, regarding the second training party as the first training party and the second model to be trained as the first model to be trained, the second total iteration gradient of the second model to be trained in the current iteration round can be determined by referring to the determination method of the first total iteration gradient of the first training party in the current iteration round, which will not be elaborated here.
[0127] B2. Determine the first gradient difference according to the first historical total gradient and the first total iteration gradient, and determine the second gradient difference according to the second historical total gradient and the second total iteration gradient.
[0128] Among them, the first historical total gradient is the total gradient of the first model to be trained in the previous iteration round. The second historical total gradient is the total gradient of the second model to be trained in the previous iteration round. For example, assume that the current iteration round is N, then the previous iteration round is N - 1.
[0129] In one implementation, determine the first gradient difference according to the absolute value of the difference between the first historical total gradient and the first total iteration gradient, and determine the second gradient difference according to the absolute value of the difference between the second historical total gradient and the second total iteration gradient.
[0130] Exemplarily, assume that the first total iteration gradient is , the total value of the first historical gradient is , then the first gradient difference . Assume that the total value of the second iterative gradient is , and the total value of the second historical gradient is , then the second gradient difference .
[0131] C2. When both the first gradient difference and the second gradient difference are less than or equal to the gradient difference threshold, generate a first trained model according to the first updated model parameters, and generate a second trained model according to the second updated model parameters corresponding to the current iteration round of the second model to be trained.
[0132] Among them, the gradient difference threshold is used to identify whether the model to be trained has completed iterative convergence. That is, if the gradient difference between the current iteration round and the previous iteration round of any model to be trained is less than or equal to the gradient difference threshold, it means that the model to be trained has completed iterative convergence. It can be understood that for the scenario of joint model training, it is necessary that the gradient differences corresponding to the second model to be trained and the first model to be trained are both less than or equal to the gradient difference threshold to indicate that the joint model training has completed iterative convergence. Further, generate a first trained model according to the first updated model parameters, and generate a second trained model according to the second updated model parameters.
[0133] Correspondingly, if the first gradient difference or the second gradient difference is greater than the gradient difference threshold, it means that at least one model to be trained has not completed iterative convergence. Therefore, continue to execute the next iteration round to continue model training until both the first gradient difference and the second gradient difference are less than or equal to the gradient difference threshold.
[0134] Exemplarily, assume that the first gradient difference is , and the second gradient difference is , and the gradient difference threshold is . If , then end the iteration, the model training ends, generate a first trained model according to the first updated model parameters, and generate a second trained model according to the second updated model parameters. If , then continue the iteration, continue to execute the joint training method of the model provided in this embodiment until .
[0135] By determining the total second iteration gradient of the second model to be trained in the current iteration round; determining a first gradient difference according to the total first historical gradient and the total first iteration gradient, and determining a second gradient difference according to the total second historical gradient and the total second iteration gradient; wherein, the total first historical gradient is the total gradient of the first model to be trained in the previous iteration round, and the total second historical gradient is the total gradient of the second model to be trained in the previous iteration round; when both the first gradient difference and the second gradient difference are less than or equal to the gradient difference threshold, generating a first trained model according to the first updated model parameters, and generating a second trained model according to the second updated model parameters corresponding to the second model to be trained in the current iteration round, so as to determine that the joint training of the models has iteratively converged only after both the second model to be trained and the first model to be trained have iteratively converged, ensuring the training effect of the joint training of the models.
[0136] Embodiment III
[0137] Figure 3 The following is a schematic structural diagram of a joint training device for a model provided in Embodiment III of the present invention. As Figure 3 shown, the device includes:
[0138] A matrix splitting module 31, configured to split the first training data matrix held by the first training participant to obtain a target number of first split matrices, and determine a local matrix, a first shared matrix, and a second shared matrix from the first split matrices, wherein the local matrix is different from the first shared matrix and the second shared matrix;
[0139] A matrix sending module 32, configured to send the first shared matrix to the second training participant, and send the second shared matrix to the computing service party;
[0140] A model parameter updating module 33, configured to update the first model parameter to be updated of the first model to be trained by the first training participant in the current iteration round according to the first gradient first shard value sent by the second training participant, the first gradient second shard value sent by the computing service party, and the first gradient third shard value generated by the first training participant;
[0141] wherein, the first gradient first shard value is determined by the second training participant according to the first shared matrix, the first gradient second shard value is determined by the computing service party according to the second shared matrix, and the first gradient third shard value is determined by the first training participant according to the local matrix.
[0142] Optionally, the model parameter updating module 33 is specifically configured to:
[0143] Determine the first gradient value based on the first shard value of the first gradient, the second shard value of the first gradient, and the third shard value of the first gradient, and determine the first prediction vector based on the first model parameters to be updated of the first model to be trained and the first training data matrix;
[0144] Determine the first error vector based on the first standard vector and the first prediction vector, and determine the second gradient value based on the first error vector and the first training data matrix;
[0145] Determine the total first iteration gradient of the first model to be trained in the current iteration round based on the first gradient value and the second gradient value, and update the first model parameters to be updated according to the total first iteration gradient.
[0146] Optionally, the model parameter update module 33 is further specifically configured to:
[0147] Obtain the first shared vector sent by the second training party, and calculate the third shard value of the first gradient according to the first shared vector and the local matrix;
[0148] Determine the first gradient value according to the sum of the first shard value of the first gradient, the second shard value of the first gradient, and the third shard value of the first gradient;
[0149] Wherein, the first shared vector is determined by the following method:
[0150] The second training party respectively determines the second error vector according to the second model parameters to be updated of the corresponding second model to be trained in the current iteration round, and the second training data matrix held by the second training party respectively;
[0151] The second training party splits the second error vector to obtain a target number of first split vectors, and determines the first shared vector from the first split vectors.
[0152] Optionally, the model parameter update module 33 is further specifically configured to:
[0153] Obtain the current learning rate of the first model to be trained in the current iteration round, and obtain the current sample size corresponding to the first training data matrix;
[0154] Update the first model parameters to be updated according to the current learning rate, the current sample size, and the total first iteration gradient, to obtain the first updated model parameters of the first model to be trained in the current iteration round.
[0155] Optionally, the model parameter update module 33 is further specifically configured to:
[0156] The first updated model parameter is obtained through the following formula:
[0157] ;
[0158] where is the first updated model parameter, is the first model parameter to be updated, is the current learning rate, is the total value of the first iterative gradient, is the current sample size.
[0159] Optionally, the second training participant determines the first shard value of the first gradient in the following manner:
[0160] Determine a second prediction vector according to the second model parameter to be trained corresponding to the second model to be updated and the corresponding second training data matrix;
[0161] Determine the second error vector according to the second standard vector and the second prediction vector, split the second error vector to obtain a target number of first split vectors, and determine a local vector, a second shared vector, and the first shared vector from the first split vectors;
[0162] Determine the first shard value of the first gradient according to the first shared matrix and the local vector.
[0163] Optionally, the computing service party determines the second shard value of the first gradient in the following manner:
[0164] Obtain the second shared vector sent by the second training participant;
[0165] Determine the second shard value of the first gradient according to the second shared matrix and the second shared vector.
[0166] Optionally, the device further includes a gradient value comparison module, which is specifically used for:
[0167] Determine the total value of the second iterative gradient of the second model to be trained in the current iteration round;
[0168] Determine a first gradient difference according to the total value of the first historical gradient and the total value of the first iterative gradient, and determine a second gradient difference according to the total value of the second historical gradient and the total value of the second iterative gradient; where the total value of the first historical gradient is the total value of the gradient of the first model to be trained in the previous iteration round, and the total value of the second historical gradient is the total value of the gradient of the second model to be trained in the previous iteration round;
[0169] When both the first gradient difference and the second gradient difference are less than or equal to the gradient difference threshold, a first trained model is generated according to the first updated model parameters, and a second trained model is generated according to the second updated model parameters corresponding to the second model to be trained in the current iteration round.
[0170] The model joint training device provided by the embodiments of the present invention can execute the model joint training method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0171] Embodiment 4
[0172] Figure 4 FIG. shows a schematic structural diagram of an electronic device 40 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0173] As Figure 4 shown, the electronic device 40 includes at least one processor 41, and a memory communicatively connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. The memory stores a computer program executable by the at least one processor. The processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other through a bus 44. The input / output (I / O) interface 45 is also connected to the bus 44.
[0174] A plurality of components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0175] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the method for jointly training the model.
[0176] In some embodiments, the method for jointly training the model can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the method for jointly training the model described above can be executed. Alternatively, in other embodiments, the processor 41 can be configured to execute the method for jointly training the model by any other suitable means (e.g., by means of firmware).
[0177] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0178] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0179] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0180] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0181] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0182] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0183] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0184] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A joint training method of a model, characterized in that: include: Splitting a first training data matrix held by a first training participant to obtain a target number of first split matrices, and determining a local matrix, a first shared matrix, and a second shared matrix from the first split matrix, wherein the local matrix is different from the first shared matrix and the second shared matrix; wherein the first training data matrix consists of salary, accumulated provident fund, job stability coefficient, age, accumulated assets, credit score, or loan amount; The first shared matrix is sent to the second training participant by using the data channel between the first training participant and the second shared matrix is sent to the computing service provider by using the data channel between the first training participant and the computing service provider; wherein the computing service provider is a third-party service provider that can provide data computing services to any training participant, and the computing service provider is a data computing server; According to the first gradient first sharding value sent by the second training participant, the first gradient second sharding value sent by the computing service provider, and the first gradient third sharding value generated by the first training participant, the first model to be trained of the first training participant in the current iteration round is updated with the first model parameter to be updated; wherein the first model to be trained is used for loan amount prediction; Among them, the first gradient first slice value is determined by the second training participant according to the first shared matrix, the first gradient second slice value is determined by the computing service provider according to the second shared matrix, and the first gradient third slice value is determined by the first training participant according to the local matrix.
2. The method according to claim 1, characterized in that The updating of the first model parameter to be updated of the first model to be trained of the first training participant in the current iteration round according to the first gradient first slice value sent by the second training participant, the first gradient second slice value sent by the computing service provider, and the first gradient third slice value generated by the first training participant includes: Determine a first gradient value according to the first gradient first slice value, the first gradient second slice value, and the first gradient third slice value, and determine a first prediction vector according to the first to-be-trained model parameter to be updated and the first training data matrix of the first to-be-trained model; Determine a first error vector according to a first standard vector and the first prediction vector, and determine a second gradient value according to the first error vector and the first training data matrix; According to the first gradient value and the second gradient value, a first iteration gradient total value of the first model to be trained in the current iteration round is determined, and the first model parameter to be updated is updated according to the first iteration gradient total value.
3. The method according to claim 2, characterized in that The determining the first gradient value according to the first gradient first slice value, the first gradient second slice value and the first gradient third slice value includes: Obtaining a first shared vector sent by the second training participant, and calculating a first gradient third slice value according to the first shared vector and the local matrix; Determine the first gradient value according to the sum of the first gradient first slice value, the first gradient second slice value and the first gradient third slice value; The first shared vector is determined in the following manner: The second training participants determine the second error vector according to the second model parameters to be updated of the corresponding second model to be trained in the current iteration round and the second training data matrices respectively held by the second training participants; The second training participant splits the second error vector to obtain a target number of first split vectors, and determines the first shared vector from the first split vectors.
4. The method according to claim 3, characterized in that The updating of the first to-be-updated model parameter according to the first iterative gradient total value includes: Obtaining a current learning rate of the first model to be trained in a current iteration round, and obtaining a current sample size corresponding to the first training data matrix; According to the current learning rate, the current sample size and the first iterative gradient total value, the first model parameters to be updated are updated to obtain first updated model parameters of the first model to be trained in the current iterative round.
5. The method according to claim 4, characterized in that The updating of the first to-be-updated model parameters according to the current learning rate, the current sample size, and the first iterative gradient total value to obtain first updated model parameters of the first to-be-trained model in the current iterative round includes: The first updated model parameter is obtained by the following formula: ; in, Update the model parameters for the first step, is the first model parameter to be updated, The current learning rate, is the total value of the first iteration gradient, is the current sample size.
6. The method according to claim 3, characterized in that The second training participant determines the first gradient first slice value in the following manner: Determine a second prediction vector according to the second to-be-updated model parameters corresponding to the second to-be-trained model and the corresponding second training data matrix; Determine the second error vector according to the second standard vector and the second prediction vector, split the second error vector to obtain a target number of first split vectors, and determine a local vector, a second shared vector, and the first shared vector from the first split vectors; The first gradient first slice value is determined according to the first shared matrix and the local vector.
7. The method according to claim 6, characterized in that The computing service provider determines the second slice value of the first gradient in the following manner: Acquire the second shared vector sent by the second training participant; The first gradient second slice value is determined according to the second shared matrix and the second shared vector.
8. The method according to claim 4, characterized in that Also includes: Determine a second iteration gradient total value of the second to-be-trained model in the current iteration round; Determine a first gradient difference value according to a first historical gradient total value and the first iteration gradient total value, and determine a second gradient difference value according to a second historical gradient total value and the second iteration gradient total value; wherein the first historical gradient total value is the gradient total value of the first to-be-trained model in the previous iteration round, and the second historical gradient total value is the gradient total value of the second to-be-trained model in the previous iteration round; When the first gradient difference and the second gradient difference are both less than or equal to the gradient difference threshold, a first training completion model is generated according to the first updated model parameters, and a second training completion model is generated according to the second updated model parameters corresponding to the second model to be trained in the current iteration round.
9. A joint training device for a model, characterized in that: include: A matrix splitting module is used to split the first training data matrix held by the first training participant to obtain a target number of first split matrices, and determine a local matrix, a first shared matrix, and a second shared matrix from the first split matrix, wherein the local matrix is different from the first shared matrix and the second shared matrix; wherein the first training data matrix is composed of salary, accumulated provident fund, job stability coefficient, age, accumulated assets, credit score or loan amount; a matrix sending module, configured to send the first shared matrix to the second training participant by using a data channel with the second training participant, and to send the second shared matrix to the computing service provider by using a data channel with the computing service provider; wherein the computing service provider is a third-party service provider that can provide data computing services to any training participant, and the computing service provider is a data computing server; A model parameter updating module, configured to update the first model parameter to be updated of the first model to be trained of the first training participant in the current iteration round according to the first gradient first sharding value sent by the second training participant, the first gradient second sharding value sent by the computing service provider, and the first gradient third sharding value generated by the first training participant; wherein the first model to be trained is used for loan amount prediction; Among them, the first gradient first slice value is determined by the second training participant according to the first shared matrix, the first gradient second slice value is determined by the computing service provider according to the second shared matrix, and the first gradient third slice value is determined by the first training participant according to the local matrix.
10. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the joint training method of the model described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the joint training method of the model described in any one of claims 1-8.
12. A computer program product, comprising a computer program, which, when executed by a processor, implements the joint training method of the model according to any one of claims 1-8.
Citation Information
Patent Citations
Method and device for jointly training service prediction model by two parties for protecting data privacy
CN111178549A
Model training method, device and system
CN111523673A
Federal learning-based model training method and device, medium and electronic equipment
CN116432040A