Data Processing Method, Apparatus, Computer Device, and Storage Medium
By mapping and encrypting gradient values in federated learning to avoid overflow, the method enhances model training efficiency and accuracy in vertical federated scenarios, addressing the issue of gradient value overflow during transmission.
Patent Information
- Application Number
- CN202210089132.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-01-25
AI Technical Summary
In the vertical federal scenario, the sample size is large, and the gradient value is prone to overflow during the federal modeling process, affecting the modeling effect.
By obtaining the initial gradient value, performing mapping processing to obtain a positive mapping gradient value, then performing encoding processing to obtain the encoded gradient value, and homomorphic encryption process is performed to generate an encrypted gradient value and send it to the second device.
It effectively avoids the overflow of gradient values in the federated modeling process, improves the modeling effect, optimizes the training process, and reduces the model training time.
Smart Images

Figure CN114448597B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of Internet technologies, and in particular, to a data processing method, apparatus, computer device, and storage medium. Background Art
[0002] With the rise of big data, the machine learning field in the artificial intelligence industry has started to grow "explosively". It has been widely applied and practiced in both the financial and Internet fields. Data is the basic content for the implementation of machine learning algorithm models, so data in all walks of life is very important. However, data contains customer privacy, so preventing data leakage and strengthening data protection have become very necessary tasks. With the introduction of various laws and regulations, while the data of all parties is protected, the data between them has gradually become isolated from each other, resulting in the phenomenon of "data islands". In this scenario, in order to improve the accuracy of modeling, joint modeling between different parties without mutual disclosure of data has become a new development trend, and thus federated learning has emerged.
[0003] Currently, federated learning can achieve federated modeling and federated training without data sharing, with relatively high data security, and can also solve the data island problem, and has been widely applied. However, in related technologies, in the vertical federated scenario, the sample size is very large, and the gradient values transmitted during the federated modeling process are prone to overflow, thus affecting the effect of federated modeling. Summary of the Invention
[0004] The present disclosure provides a data processing method, apparatus, computer device, and storage medium, aiming to solve at least one of the technical problems in related technologies to a certain extent.
[0005] According to a first aspect, there is provided a data processing method, which is applied to a first device that provides first data. The method includes: obtaining an initial gradient value corresponding to the first data, where the initial gradient value is used to determine the timing of splitting the current node, and the current node belongs to the current federated model to be trained; performing a mapping process on the initial gradient value to obtain a mapped gradient value, where the mapped gradient value is a positive value; performing an encoding process on the mapped gradient value to obtain an encoded gradient value; and performing a homomorphic encryption process on the encoded gradient value to obtain an encrypted gradient value, and sending the encrypted gradient value to a second device.
[0006] According to a second aspect, a data processing method is provided, which is applied to a second device. The second device provides second data, and the second data corresponds to second data features. The method includes: receiving an encrypted gradient value sent by a first device, where the encrypted gradient value is generated by the first device. The first device provides first data, and the first data is different from the second data; generating an encrypted tag and a value according to the second data features and the encrypted gradient value; wherein the first device determines the timing of splitting the current node based on the encrypted tag and the value, and the current node belongs to the current federated model to be trained.
[0007] According to a third aspect, a data processing device is provided. The device provides first data. The device includes: an acquisition module, configured to acquire an initial gradient value corresponding to the first data, where the initial gradient value is used to determine the timing of splitting the current node, and the current node belongs to the current federated model to be trained; a mapping module, configured to perform a mapping process on the initial gradient value to obtain a mapped gradient value, where the mapped gradient value is a positive value; an encoding module, configured to perform an encoding process on the mapped gradient value to obtain an encoded gradient value; and an encryption module, configured to perform a homomorphic encryption process on the encoded gradient value to obtain an encrypted gradient value, and send the encrypted gradient value to the second device.
[0008] According to a fourth aspect, a data processing device is provided. The device provides second data, and the second data corresponds to second data features. The device includes: a second receiving module, configured to receive an encrypted gradient value sent by a first device, where the encrypted gradient value is generated by the first device. The first device provides first data, and the first data is different from the second data; a generating module, configured to generate an encrypted tag and a value according to the second data features and the encrypted gradient value; wherein the first device determines the timing of splitting the current node based on the encrypted tag and the value, and the current node belongs to the current federated model to be trained.
[0009] According to a fifth aspect, a computer device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the data processing method provided by the embodiments of the present disclosure.
[0010] According to a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is proposed, where the computer instructions are used to cause the computer to execute the data processing method provided by the embodiments of the present disclosure.
[0011] In the embodiments of the present disclosure, by obtaining an initial gradient value corresponding to first data, the initial gradient value is used to determine the timing of splitting the current node. The current node belongs to the current federated model to be trained, and the initial gradient value is mapped to obtain a mapped gradient value. The mapped gradient value is a positive value, and the mapped gradient value is encoded to obtain an encoded gradient value. Then, the encoded gradient value is homomorphically encrypted to obtain an encrypted gradient value, and the encrypted gradient value is sent to a second device. It is possible to perform mapping and encoding processing on the gradient value transmitted during the federated modeling process, effectively avoiding the phenomenon of gradient value overflow, thereby improving the effect of federated modeling.
[0012] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and / or additional aspects and advantages of the present disclosure will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0014] Figure 1 is a schematic flowchart of a data processing method according to an embodiment of the present disclosure;
[0015] Figure 2 is a schematic diagram of a mapping and encoding process according to an embodiment of the present disclosure;
[0016] Figure 3 is a schematic flowchart of a data processing method according to another embodiment of the present disclosure;
[0017] Figure 4 is a schematic diagram of a gradient value splicing process according to an embodiment of the present disclosure;
[0018] Figure 5 is a schematic flowchart of a data processing method according to another embodiment of the present disclosure;
[0019] Figure 6 is a schematic flowchart of a data processing method according to another embodiment of the present disclosure;
[0020] Figure 7 is a schematic diagram of an overall federated training process according to an embodiment of the present disclosure;
[0021] Figure 8 is a schematic diagram of a data processing device according to another embodiment of the present disclosure;
[0022] Figure 9 is a schematic diagram of a data processing device according to another embodiment of the present disclosure;
[0023] Figure 10Schematic diagram of a data processing device according to another embodiment of the present disclosure; and
[0024] Figure 11 It is a block diagram of a computer device for implementing the data processing method of the embodiment of the present disclosure. Detailed implementation manners
[0025] The embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present disclosure and should not be construed as a limitation of the present disclosure. On the contrary, the embodiments of the present disclosure include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.
[0026] In related technologies, in the vertical federated scenario, the sample size is very large, and the gradient values transmitted during the federated modeling process are prone to overflow, thus affecting the effect of federated modeling. The technical solution of this embodiment provides a data processing method, which will be described below in combination with specific embodiments.
[0027] It should be noted that the execution subject of the data processing method disclosed in this embodiment may be a data processing device, which may be implemented in software and / or hardware, and this device may be configured in an electronic device, and the electronic device may include but is not limited to a terminal, a server, etc.
[0028] Figure 1 It is a flowchart of a data processing method proposed in an embodiment of the present disclosure. This data processing method can be applied to a first device, and the first device provides first data.
[0029] As Figure 1 shown, this data processing method includes:
[0030] S101: Obtain an initial gradient value corresponding to the first data. The initial gradient value is used to determine the timing of splitting the current node, and the current node belongs to the currently trained federated model.
[0031] In the embodiment of the present disclosure, the first device first obtains an initial gradient value corresponding to the first data.
[0032] Among them, the first device is a device of a participant (for example: business party Guest) participating in the federated model training. This first device may be a server of the business party Guest, a terminal device, and any other possible electronic device, and there is no limitation thereto.
[0033] The first data is provided by the first device. That is to say, the data provided by the business party Guest can be called the first data. This first data can be used as training samples during the federated model training process, and the quantity of this first data can be multiple. For example, it is expressed as label y ∈ R n .
[0034] The federated model in the embodiments of the present disclosure can be based on the SecureBoost model. The key to training the SecureBoost model lies in the establishment of each gradient boosting decision tree, and the key to the establishment of the gradient boosting decision tree lies in: the prediction result of the m-th tree is determined by the prediction residual of the (m - 1)-th tree; when each node of the decision tree is split, each first data will generate a binning situation and its corresponding encrypted gradient value. Through this gradient value, it can be determined whether to perform a splitting process on the leaf node of the decision tree. Among them, the gradient value can be calculated according to the prediction value of the previous tree (for example: the (m - 1)-th tree) and the loss function (Loss function).
[0035] The gradient value generated by the first data at the current node can be called the initial gradient value. This initial gradient value can be used to determine the timing of performing a splitting process on the current node. That is to say, according to the initial gradient value, it can be determined whether to perform a splitting process on the current node.
[0036] In some embodiments, the initial gradient value can include the first-order gradient value and the second-order gradient value. That is to say, multiple first data can correspond to multiple first-order gradient values and multiple second-order gradient values. Among them, the first-order gradient value can be expressed as The second-order gradient value can be expressed as i is the sample serial number, and n is the quantity of multiple first data.
[0037] S102: Perform a mapping process on the initial gradient value to obtain a mapped gradient value, and the mapped gradient value is a positive value.
[0038] After obtaining the initial gradient value corresponding to the first data as described above, further, the first device can perform a mapping process on the initial gradient value. The initial gradient value after the mapping process can be called the mapped gradient value, and the mapped gradient value is a positive value.
[0039] In some embodiments, the first-order gradient value can be subjected to a mapping process to obtain a first-order mapped gradient value, and the second-order gradient value can be subjected to a mapping process to obtain a second-order mapped gradient value. The first-order mapped gradient value and the second-order mapped gradient value are jointly used as the mapped gradient value. By performing mapping processes on the first-order gradient value and the second-order gradient value respectively, the accuracy of the mapping operation can be improved.
[0040] Among them, in the process of mapping a first-order gradient value, a first reference gradient value can be determined first, for example, represented by c, and this first reference gradient value is used to map the first-order gradient value to a positive value.
[0041] Among them, the first reference gradient value c is less than or equal to the smallest first-order gradient value among multiple first-order gradient values. The smallest first-order gradient value can be represented as min(x), then c <= min(x).
[0042] Furthermore, the difference is processed between multiple first-order gradient values and the first reference gradient value respectively to obtain multiple first-order mapped gradient values corresponding to the multiple first-order gradient values respectively.
[0043] For example, Figure 2 is a schematic diagram of the mapping and encoding process provided according to an embodiment of the present disclosure. As Figure 2 shown, the actual value of the first-order gradient value can be represented by x i The mapping process (i.e., the first-order mapped gradient value) can be represented as: x' i = x i - c, where i is the serial number of multiple first data, and x' i is the first-order mapped gradient value. In addition, other arbitrary possible ways can also be used for mapping, and no limitation is imposed on this.
[0044] Thus, in the mapping process, the difference operation can be performed between each first-order gradient value and the first reference gradient value. Since the first reference gradient value is less than or equal to the smallest first-order gradient value among multiple first-order gradient values, it can be ensured that each first-order mapped gradient value is a positive value, which is beneficial to subsequent encoding operations and prevents encoding overflow.
[0045] It can be understood that in the above embodiment, only the smallest first-order gradient value is used as the first-order target gradient value for exemplary illustration. In actual applications, other arbitrary possible values can also be used as the first-order target gradient value under the condition that the first-order mapped gradient value after mapping is a positive value, and no limitation is imposed on this.
[0046] In other embodiments, in the operation of mapping the second-order gradient value, a second reference gradient value is determined first. Among them, the second reference gradient value is less than or equal to the smallest second-order gradient value among multiple second-order gradient values; the difference is processed between multiple second-order gradient values and the second reference gradient value respectively to obtain multiple second-order mapped gradient values corresponding to the multiple second-order gradient values respectively. The mapping process of the second-order gradient value can be the same as the process of mapping the first-order gradient value described above, and will not be elaborated here.
[0047] S103: Perform encoding processing on the mapped gradient value to obtain an encoded gradient value.
[0048] After obtaining the mapped gradient values as described above, further, the first device may perform encoding processing on the mapped gradient values (the first-order mapped gradient value and the second-order mapped gradient value) to obtain encoded gradient values, and the encoded gradient values may be represented by z i (m-1) For representation.
[0049] In some embodiments, the first-order mapped gradient value and the second-order mapped gradient value may be encoded between a specified negative quadratic to positive quadratic power number to obtain corresponding encoded gradient values. Among them, the encoded gradient value corresponding to the first-order mapped gradient value may be represented as g i (m-1) , and the encoded gradient value corresponding to the second-order mapped gradient value may be represented as h i (m-1) , that is to say, the values of g i (m-1) and h i (m-1) are in the range between the specified negative quadratic to positive quadratic power number.
[0050] It can be understood that the above examples are only for illustrative purposes of the encoding process. In actual applications, encoding can also be performed within any other possible numerical range. And any encoding method can be adopted for encoding, which is not limited herein.
[0051] S104: Perform homomorphic encryption processing on the encoded gradient values to obtain encrypted gradient values, and send the encrypted gradient values to the second device.
[0052] After obtaining the encoded gradient value z i (m-1) , the first device may perform homomorphic encryption processing on the encoded gradient value of each first data to obtain encrypted gradient values, and the encrypted gradient values may be represented as Further, the first device sends the encrypted gradient value to the second device.
[0053] Among them, the second device is a device of a participant (for example: data party Host) participating in the federated model training. The second device may be a server, a terminal device, or any other possible electronic device of the data party Host, which is not limited herein. The second device of the Host party may perform binning calculations based on the homomorphically encrypted encrypted gradient values sent by the first device to implement federated training.
[0054] In an embodiment of the present disclosure, by obtaining an initial gradient value corresponding to first data, the initial gradient value is used to determine the timing of splitting the current node. The current node belongs to the current federated model to be trained, and the initial gradient value is subjected to a mapping process to obtain a mapped gradient value. The mapped gradient value is a positive value, and the mapped gradient value is encoded to obtain an encoded gradient value, and the encoded gradient value is subjected to a homomorphic encryption process to obtain an encrypted gradient value, and the encrypted gradient value is sent to a second device. It is possible to perform mapping and encoding on the gradient value transmitted during the federated modeling process, effectively avoiding the phenomenon of gradient value overflow, thereby improving the effect of federated modeling.
[0055] Figure 3 It is a schematic flowchart of a data processing method proposed in another embodiment of the present disclosure.
[0056] As Figure 3 shown, the data processing method includes:
[0057] S301: Obtain an initial gradient value corresponding to first data. The initial gradient value is used to determine the timing of splitting the current node. The current node belongs to the current federated model to be trained.
[0058] S302: Perform a mapping process on the initial gradient value to obtain a mapped gradient value. The mapped gradient value is a positive value.
[0059] For the specific descriptions of S301 - S302, reference can be made to the above - mentioned embodiments and will not be elaborated here.
[0060] S303: Encode the first - order mapped gradient value to obtain a first - order encoded gradient value.
[0061] In an embodiment of the present disclosure, during the process of encoding the mapped gradient value, the first - order mapped gradient value x′ i can be encoded to obtain a first - order encoded gradient value.
[0062] In some embodiments, the number of encoding bits and the first encoding interval can be determined first. Among them, the number of encoding bits is the number of bits in a specified binary form code, which can be represented by r. The first encoding interval, which can also be called the quantization interval, can be represented by [min, max], and is used to encode the first - order mapped gradient value. The first encoding interval can be set according to the actual application scenario and is not limited thereto.
[0063] In some embodiments, before encoding the mapped gradient value, the mapped gradient value can also be clipped using a gradient hyperparameter a (a > 0). The mapped gradient values exceeding a are uniformly clipped to α, and the mapped gradient values exceeding - a are uniformly clipped to - a.
[0064] Further, a first coding range value and a second coding range value are determined according to the first coding interval. In some embodiments, the first coding range value may be, for example, the minimum value min in the first coding interval, and the second coding range value may be, for example, the maximum value max in the first coding interval. There is no limitation on the selection of the first coding range value and the second coding range value.
[0065] Further, according to the first coding range value min and the second coding range value max, combined with the number of coding bits r, the first-order mapping gradient value x' i is coded to obtain a first-order coding gradient value.
[0066] For example, as shown in Figure 2, the first-order coding gradient value can be represented by g i (m-1) and the specific value can be represented by x encode . The process of coding the first-order mapping gradient value can be expressed as:
[0067] S304: Code the second-order mapping gradient value to obtain a second-order coding gradient value. The first-order coding gradient value and the second-order coding gradient value are jointly used as the coding gradient value.
[0068] Embodiments of the present disclosure can also code the second-order mapping gradient value to obtain a second-order coding gradient value, which can be represented by h i (m-1) . This first-order coding gradient value g i (m-1) and the second-order coding gradient value h i (m-1) are jointly used as the coding gradient value.
[0069] In some embodiments, during the process of coding the second-order mapping gradient value, a second coding interval can be first determined; further, a third coding range value and a fourth coding range value are determined according to the second coding interval; and according to the third coding range value and the fourth coding range value, combined with the number of coding bits, the second-order mapping gradient value is coded to obtain a second-order coding gradient value.
[0070] The process of coding the second-order mapping gradient value can be the same as the process of coding the first-order mapping gradient value described above and will not be elaborated here.
[0071] In this embodiment, by setting the reference gradient values (the first reference gradient value and the second reference gradient value) c to be less than or equal to the minimum of the first-order gradient value and the second-order gradient value, the clipping result is mapped to a positive value, solving the overflow problem caused by adding negative samples; in addition, by reasonably setting the max and min parameters and expanding the quantization interval, the encoding can represent a larger range of positive numbers, thereby ensuring that the aggregation result will not overflow.
[0072] S305: Determine the first starting position corresponding to the first-order encoded gradient value.
[0073] Among them, the position before the first-order encoded gradient value can be called the first starting position.
[0074] S306: Determine the second starting position corresponding to the second-order encoded gradient value.
[0075] Similarly, the position before the second-order encoded gradient value can be called the second starting position.
[0076] S307: Configure the overflow bits with a set number of bits at the first starting position and the second starting position respectively.
[0077] After determining the first starting position and the second starting position as described above, further, configure the overflow bits with a set number of bits at the first starting position and the second starting position respectively.
[0078] Figure 4 It is a schematic diagram of the gradient value splicing process provided by an embodiment of the present disclosure. As Figure 4 shown, for example: the first-order encoded gradient value in the form of r-bit (8-bit) binary after encoding is: 00001110, and the second-order encoded gradient value is: 11011111. In this embodiment, the overflow bits with a set number of bits (for example, 2 bits) can be configured at the first starting position and the second starting position. That is to say, 2 bits of overflow bits (for example, 00) can be set at the first starting position, and 2 bits of overflow bits (for example, 00) can be set at the second starting position. The first-order encoded gradient value with overflow bits can be expressed as The second-order encoded gradient value with overflow bits can be expressed as By setting the overflow bits, the overflow space can be increased. In the case of encoding overflow, the original number can be restored according to the overflow bits, and setting the overflow bits can avoid continuously causing errors in the previous encoded numbers.
[0079] S308: Perform splicing processing on the first-order encoded gradient value and the second-order encoded gradient value, and use the gradient value obtained by the splicing processing as the encoded gradient value.
[0080] Further, as Figure 4 shown, for the first-order encoded gradient value and the second-order encoded gradient value Perform splicing processing to obtain the encoded gradient value z i (m-1) . Through splicing processing, the number of transmissions of encrypted data can be reduced, and the training speed of the federated model can be improved.
[0081] S309: Perform homomorphic encryption processing on the encoded gradient value to obtain the encrypted gradient value, and send the encrypted gradient value to the second device.
[0082] For the specific description of S309, reference can be made to the above embodiments, and details are not described here again.
[0083] In the embodiments of the present disclosure, by obtaining the initial gradient value corresponding to the first data, the initial gradient value is used to determine the timing of splitting the current node. The current node belongs to the currently to-be-trained federated model, and perform mapping processing on the initial gradient value to obtain the mapped gradient value. The mapped gradient value is a positive value, and perform encoding processing on the mapped gradient value to obtain the encoded gradient value, and perform homomorphic encryption processing on the encoded gradient value to obtain the encrypted gradient value, and send the encrypted gradient value to the second device. It is possible to perform mapping and encoding processing on the gradient value transmitted during the federated modeling process, effectively avoiding the phenomenon of gradient value overflow, thereby improving the effect of federated modeling. In addition, gradient hyperparameters can be introduced during the encoding process to limit the gradient range, which is beneficial to encoding the first-order mapped gradient value. And by setting the overflow bit, it is beneficial to perform splicing processing on the first-order encoded gradient value and the second-order encoded gradient value, and can prevent gradient value overflow during the splicing process. In addition, through splicing processing, the number of transmissions of encrypted data can be reduced, and the training speed of the federated model can be improved.
[0084] Figure 5 It is a schematic flowchart of a data processing method proposed in another embodiment of the present disclosure.
[0085] As Figure 5 shown, the data processing method includes:
[0086] S501: Obtain the initial gradient value corresponding to the first data. The initial gradient value is used to determine the timing of splitting the current node. The current node belongs to the currently to-be-trained federated model.
[0087] S502: Perform mapping processing on the initial gradient value to obtain the mapped gradient value. The mapped gradient value is a positive value.
[0088] S503: Perform encoding processing on the mapped gradient value to obtain the encoded gradient value.
[0089] S504: Perform homomorphic encryption processing on the encoded gradient value to obtain the encrypted gradient value, and send the encrypted gradient value to the second device.
[0090] For the specific descriptions of S501 - S504, reference can be made to the above - mentioned embodiments, and details are not repeated here.
[0091] S505: Receive the encrypted tag and value sent by the second device.
[0092] In the embodiments of the present disclosure, the first data corresponds to the first data feature, the second device provides the second data, and the second data corresponds to the second data feature.
[0093] Among them, the feature carried by the first data of the business party Guest can be referred to as the first data feature.
[0094] The second data is provided by the second device. That is to say, the data provided by the data party (Host) can be referred to as the second data, and the second data corresponds to the second data feature. This second data can be used as a training sample during the federated model training process.
[0095] Among them, there can be multiple data parties Host participating in the federated model training (for example: p), and these p Host parties can be represented as X 1 ,..., X p . Among them, the data of the i - th host participating party is d i , and the entire data set (i.e., multiple second data) is X = [X 0 , X 1 ,..., X p ∈ R n×d ,
[0096] After sending the encrypted gradient value to the second device, the second device can calculate the encrypted tag and value of the encrypted gradient value and send them to the first device. In this case, the first device can receive the encrypted tag and value sent by the second device.
[0097] Among them, the encrypted tag and value is the sum value of multiple encrypted tag values. The encrypted tag value is parsed from the encrypted gradient value by the corresponding second device according to the second data feature.
[0098] In practical applications, after the second device of the Host party receives the encrypted gradient value sent by the first device, the second device can perform binning calculations on the encrypted gradient value according to its own second data feature to obtain multiple encrypted tag values, and the encrypted tag values can be expressed as And the sum value of multiple encrypted tag values can be expressed as: After all the features (i.e., the second data features) are binned for each Host party, each Host participating party transmits the calculated encrypted tag and value to the first device of the Guest party.
[0099] S506: Decrypt the encrypted tag and value to obtain the decrypted tag and value.
[0100] After receiving the encrypted tag and value sent by the second device as described above further, the first device decrypts the encrypted tag and value to obtain the decrypted tag and value.
[0101] S507: Decompose the decrypted tag and value according to the first data feature to obtain the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped corresponding to the first-order gradient value.
[0102] Further, the first device may decompose the decrypted tag and value according to the first data feature to obtain the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped corresponding to the first-order gradient value.
[0103] Among them, the first-order gradient value-to-be-mapped can be expressed as The second-order gradient value-to-be-mapped can be expressed as
[0104] S508: Perform inverse mapping processing on the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped respectively to obtain the first-order gradient initial value and the second-order gradient initial value.
[0105] In some embodiments, during the inverse mapping process, the sample size and the mapping translation amount may be obtained first. Among them, the sample size is the sample size in each bin of the Host side, which can be represented by n kv The mapping translation amount is the translation amount between the initial gradient value and the mapped gradient value, that is, the translation amount x of the mapping process in the above embodiments i_new = x i - min(x).
[0106] Further, the first device of the Guest side, according to the sample size n kv and the mapping translation amount, performs decoding and inverse mapping processing on the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped respectively to obtain the first-order gradient initial value and the second-order gradient initial value. Among them, the first-order gradient initial value can be represented as g kv (m-1) , and the second-order gradient initial value can be represented as h kv (m-1) .
[0107] S509: Aggregate the first-order gradient initial value and the second-order gradient initial value to obtain the reference gradient value.
[0108] After obtaining the initial first-order gradient value and the initial second-order gradient value, further, the first device performs an aggregation process on the initial first-order gradient value and the initial second-order gradient value to obtain a reference gradient value.
[0109] Among them, the reference gradient value is used to determine the gain score value of the first data feature corresponding to the current node, and the gain score value is used to determine the timing of splitting the current node.
[0110] In practical applications, the Guest party obtains the gain score values of various binning methods of the first data feature at the splitting node according to the information gain calculation formula, and takes the maximum value max of all results. If the score value max is greater than the threshold γ, continue to split; otherwise, stop splitting. After continuing to split, obtain the sample ID sets IL and IR of the left and right child nodes respectively. If the maximum value corresponds to a local feature, the Guest party directly records the feature information and the threshold; if it corresponds to a Host party feature, record the Host party where the feature comes from and its corresponding index value F, and then synchronize the splitting result (IL, IR) to each Host party. When the leaf node l stops splitting, the Guest party calculates the weights w of each child node at this time. The current m-th tree is constructed, and the next step is to start constructing the m-th tree. Repeat step 2 to step 11 until m reaches T. When T trees are constructed, the federated learning model obtained is where x is the sample feature information.
[0111] In this way, the federated learning model training of multiple participants can be realized, and the training process of the model can be optimized, the training time of the model can be reduced, it is ensured that the sum of the aggregated gradient values will not overflow, and the model is more concise.
[0112] In the embodiments of the present disclosure, by obtaining the initial gradient value corresponding to the first data, the initial gradient value is used to determine the timing of splitting the current node. The current node belongs to the current federated model to be trained, and the initial gradient value is mapped to obtain a mapped gradient value. The mapped gradient value is a positive value, and the mapped gradient value is encoded to obtain an encoded gradient value, and the encoded gradient value is homomorphically encrypted to obtain an encrypted gradient value, and the encrypted gradient value is sent to the second device. It is possible to perform mapping and encoding processing on the gradient values transmitted during the federated modeling process, effectively avoiding the phenomenon of gradient value overflow, thereby improving the effect of federated modeling. In addition, the federated learning model training of multiple participants can be realized, and the training process of the model can be optimized, the training time of the model can be reduced, it is ensured that the sum of the aggregated gradient values will not overflow, and the model is more concise.
[0113] Figure 6It is a schematic flowchart of a data processing method proposed in another embodiment of the present disclosure, which is applied to a second device. The second device provides second data, and the second data corresponds to second data features.
[0114] As Figure 6 shown, the data processing method includes:
[0115] S601: Receive the encrypted gradient value sent by the first device. The encrypted gradient value is generated by the first device. The first device provides first data, and the first data is different from the second data.
[0116] In the embodiment of the present disclosure, the second device first receives the encrypted gradient value sent by the first device. The encrypted gradient value can be expressed as
[0117] Among them, the second device and the first device can be parties participating in the federated model training respectively. The first device can be represented as the device of the business party Guest, and the second device can be represented as the device of the data party Host. Moreover, the second device and the first device can be servers, terminal devices, and any other possible electronic devices, and there is no limitation thereto. Also, the first device and the second device can provide one or more first data and second data respectively. The first data can have corresponding first data features, and the second data can have corresponding second data features for federated model training.
[0118] Among them, the first device can generate a corresponding encrypted gradient value according to the first data and send it to the second device. The generation process of the encrypted gradient value can refer to the above embodiment and will not be elaborated here. In this case, the second device can receive the encrypted gradient value sent by the first device.
[0119] S602: Generate an encrypted tag and a value according to the second data feature and the encrypted gradient value.
[0120] Furthermore, the second device generates an encrypted tag and a value according to the second data feature and the encrypted gradient value.
[0121] In some embodiments, in the process of generating the encrypted tag and the value, the second device first parses the encrypted tag value from the encrypted gradient value according to the second data feature of the Host party. The encrypted tag value can be expressed as
[0122] Furthermore, the second device processes the encrypted tag value according to the binning information corresponding to the second data feature to obtain the encrypted tag and the value. For example: calculate the sum of the encrypted tag values corresponding to q + 1 bins of its own included feature (the second data feature) as the encrypted tag and the value, which can be expressed as Among them, the calculation process is as follows
[0123] S603: Send the encrypted tag and value to the first device.
[0124] Among them, the first device can determine the timing of splitting the current node based on the encrypted tag and value, and the current node belongs to the current federated model to be trained. In this way, the first device and the second device can complete the training task of the federated model.
[0125] In the embodiment of the present disclosure, by receiving the encrypted gradient value sent by the first device, the encrypted gradient value is generated by the first device, the first device provides the first data, the first data is different from the second data, and according to the second data feature and the encrypted gradient value, an encrypted tag and value are generated; among them, the first device determines the timing of splitting the current node based on the encrypted tag and value, and the current node belongs to the current federated model to be trained, which can realize the joint training of the federated model by the first device and the second device.
[0126] Figure 7 It is a schematic diagram of the overall federated training process provided by the embodiment of the present disclosure, as Figure 7 shown. The training process of the federated learning model includes the following steps:
[0127] Assume that the data of the guest party participating in the federated model training is represented as label y∈R n ; there can be multiple data parties Host participating in the federated model training (for example: p), and these p Host parties can be represented as X 1 ,... X p . Among them, the data of the i-th host participating party is d i , and the entire data set (that is, multiple second data) is X = [X 0 , X 1 ,..., X p ∈R n×d , In this scenario, the federated random forest belongs to vertical federated learning, and the whole stage is mainly divided into the following steps:
[0128] Step 1: The guest party performs sample encryption alignment with each host party to unify the common sample ids.
[0129] Step 2: The guest party generates a set of public and private key pairs and sends the public key to each host party.
[0130] By performing iterative operations, construct the t = 1, 2,..., T tree models f t (here the decision tree uniformly considers the binary case):
[0131] Step 3: The decision tree adopts a pre-pruning strategy to split the leaf nodes from top to bottom, where l = 1, 2,..., n t :
[0132] Step 4: The Guest also generates a public key Public Key and a private key Private Key and transfers the public key to the Host
[0133] Step 5: For all Hosts and Guests, after determining all the sample IDs of the node, binning is performed on the feature k (k = 1,..., N) to obtain the thresholds at the specified q quantiles, denoted as S k ={S k1 ,...S kq}, which is used as the screening for tree splitting nodes
[0134] Step 6: The Guest calculates the first derivative (corresponding to the above first-order gradient value) and the second derivative (corresponding to the above second-order gradient value) according to the prediction value of the (m - 1)-th tree and the calculation formula of the loss function (Loss function) and the second derivative (corresponding to the above second-order gradient value)
[0135] Step 7: The Guest first maps g io (m-1) and h io (m-1) to positive values and performs clipping. The clipping process is as follows: Given a hyperparameter a (a > 0) representing the gradient range, the gradient values exceeding a are uniformly clipped to a, and the gradient values exceeding -a are uniformly clipped to -a. Subsequently, the Guest encodes the obtained positive values between the specified negative quadratic to positive quadratic powers to obtain g i (m-1) and h i (m -1) . The mapping method is: x′ i =x i -c, where c <= min(x), and min(x) is the minimum value of the first-order gradient value x i in. The encoding method is: where r is the specified number of bits in the binary form code, i represents the sample subscript, max and min respectively represent the maximum and minimum parameters of the set quantization interval [min, max]; finally, x_{encode} is quantized and rounded. We map the clipped result to a positive value by setting c less than or equal to the minimum value of x i
[0136] Step 8: The Guest transfers g i (m-1) and hi (m-1) Perform splicing. When splicing, introduce a binary form code with the first 2 overflow bits at the starting positions of the first-order and second-order gradients (r is the power value of the above-mentioned power of 2) to encode them, and obtain and The data after splicing is denoted as z i (m-1) . (The splicing binary form code when r = 8 is as shown in Figure 4 ).
[0137] Step Nine: Subsequently, the Guest party homomorphically encrypts the z i (m-1). value obtained for each sample and then becomes and transmits it to the Host party. The Host party calculates the sum of the encrypted tag values according to the q + 1 bins corresponding to the features it incorporates After all the bins corresponding to the features incorporated by each Host party are completed, it transmits all to the Guest party. The features incorporated by the Guest party itself are directly binned to obtain the sum of the g values and h values in each Part.
[0138] Step Ten: The Guest party decrypts and disassembles the transmitted by the Host party to obtain and At the same time, the Host party also needs to transmit the sample size n kv in this bin to the Guest party. Subsequently, the Guest party decodes and again and inverse maps them to the initial values g kv and h kv (m-1) according to n kv (m -1) , to obtain the aggregated values of the first-order and second-order gradients.
[0139] Step Eleven: The Guest party obtains the gain score values of various binning methods for each feature at this split node according to the information gain calculation formula and extracts the maximum value max among all the results. If the score value max is greater than the threshold γ, continue to split; otherwise, stop splitting. After continuing to split, obtain the sample ID sets IL and IR of the left and right child nodes respectively. If the maximum value corresponds to a local feature, the Guest party directly records the feature information and the threshold; if it corresponds to a Host party feature, record the Host party where the feature comes from and its corresponding index value F, and then synchronize the split result (IL, IR) to each Host party.
[0140] Step Twelve: When the leaf node l stops splitting, the Guest party calculates the weights w of each child node at this time. The construction of the current m-th tree is completed. Next, start constructing the m-th tree, and repeat Steps Three to Twelve until m reaches T.
[0141] Step Thirteen: When the construction of T trees is completed, the obtained ScureBoost model is where x is the sample feature information.
[0142] Compared with the existing secureboost model training scheme, the technical solution implemented in this disclosure only requires the Guest party to perform encryption once (i.e., encrypt the binary code after aggregating the first and second order gradients), and the number of transmissions from the Guest party to the Host party and from the Host party to the Guest party is reduced by half. Moreover, it is ensured that the sum of the aggregated gradient values will not overflow, and the model is more concise. In addition, since the gradient values are shifted to positive numbers during the gradient binary encoding in this scheme, there is no need to introduce a sign bit.
[0143] Figure 8 It is a schematic diagram of a data processing device according to another embodiment of the present disclosure.
[0144] As Figure 8 shown, the data processing device 80 includes:
[0145] An acquisition module 801, configured to acquire an initial gradient value corresponding to the first data. The initial gradient value is used to determine the timing of splitting the current node, and the current node belongs to the current federated model to be trained;
[0146] A mapping module 802, configured to perform mapping processing on the initial gradient value to obtain a mapped gradient value, and the mapped gradient value is a positive value;
[0147] An encoding module 803, configured to perform encoding processing on the mapped gradient value to obtain an encoded gradient value; and
[0148] An encryption module 804, configured to perform homomorphic encryption processing on the encoded gradient value to obtain an encrypted gradient value, and send the encrypted gradient value to a second device.
[0149] Optionally, Figure 9 It is a schematic diagram of a data processing device according to another embodiment of the present disclosure. As Figure 9 shown, in some embodiments, the initial gradient value includes: a first-order gradient value and a second-order gradient value. Among them, the mapping module 802 includes:
[0150] A first mapping sub-module 8021, configured to perform mapping processing on the first-order gradient value to obtain a first-order mapped gradient value;
[0151] The second mapping sub-module 8022 is used to perform mapping processing on the second-order gradient value to obtain a second-order mapped gradient value, and the first-order mapped gradient value and the second-order mapped gradient value are jointly used as the mapped gradient value.
[0152] Optionally, in some embodiments, the number of the first data is multiple, and the multiple first data respectively correspond to multiple first-order gradient values. The first mapping sub-module 8021 is specifically configured to: determine a first reference gradient value, where the first reference gradient value is less than or equal to the smallest first-order gradient value among the multiple first-order gradient values; respectively perform difference processing on the multiple first-order gradient values and the first reference gradient value to obtain multiple first-order mapped gradient values corresponding to the multiple first-order gradient values respectively.
[0153] Optionally, in some embodiments, the multiple first data respectively correspond to multiple second-order gradient values. The second mapping sub-module 8022 is specifically configured to: determine a second reference gradient value, where the second reference gradient value is less than or equal to the smallest second-order gradient value among the multiple second-order gradient values; respectively perform difference processing on the multiple second-order gradient values and the second reference gradient value to obtain multiple second-order mapped gradient values corresponding to the multiple second-order gradient values respectively.
[0154] Optionally, in some embodiments, as Figure 9 shown, the encoding module 803 includes:
[0155] The first encoding sub-module 8031 is used to perform encoding processing on the first-order mapped gradient value to obtain a first-order encoded gradient value;
[0156] The second encoding sub-module 8032 is used to perform encoding processing on the second-order mapped gradient value to obtain a second-order encoded gradient value, and the first-order encoded gradient value and the second-order encoded gradient value are jointly used as the encoded gradient value.
[0157] Optionally, in some embodiments, the first encoding sub-module 8031 is specifically configured to: determine the number of encoding bits and a first encoding interval; determine a first encoding range value and a second encoding range value according to the first encoding interval; and perform encoding processing on the first-order mapped gradient value in combination with the number of encoding bits according to the first encoding range value and the second encoding range value to obtain a first-order encoded gradient value.
[0158] Optionally, in some embodiments, the second encoding sub-module 8032 is specifically configured to: determine a second encoding interval; determine a third encoding range value and a fourth encoding range value according to the second encoding interval; and perform encoding processing on the second-order mapped gradient value in combination with the number of encoding bits according to the third encoding range value and the fourth encoding range value to obtain a second-order encoded gradient value.
[0159] Optionally, in some embodiments, as Figure 9As shown, the apparatus 80 further includes: a splicing module 805, configured to splice the first-order encoded gradient value and the second-order encoded gradient value, and use the gradient value obtained by the splicing process as the encoded gradient value.
[0160] Optionally, in some embodiments, as Figure 9 shown, the apparatus 80 further includes: an overflow bit configuration module 806, specifically configured to: determine a first starting position corresponding to the first-order encoded gradient value; determine a second starting position corresponding to the second-order encoded gradient value; and configure a set number of overflow bits at the first starting position and the second starting position, respectively.
[0161] Optionally, in some embodiments, the first data corresponds to a first data feature, the second apparatus provides second data, and the second data corresponds to a second data feature. As Figure 9 shown, the apparatus 80 further includes:
[0162] a first receiving module 807, configured to receive an encrypted tag and a value sent by the second apparatus;
[0163] a decryption module 808, configured to decrypt the encrypted tag and the value to obtain a decrypted tag and a value;
[0164] a disassembling module 809, configured to disassemble the decrypted tag and the value according to the first data feature to obtain a first-order gradient value-to-be-mapped and a second-order gradient value-to-be-mapped corresponding to the first-order gradient value;
[0165] an inverse mapping module 810, configured to perform inverse mapping processing on the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped respectively to obtain a first-order gradient initial value and a second-order gradient initial value;
[0166] an aggregation module 811, configured to aggregate the first-order gradient initial value and the second-order gradient initial value to obtain a reference gradient value;
[0167] wherein, the encrypted tag and the value are the sum of multiple encrypted tag values, and the encrypted tag value is parsed from the encrypted gradient values by the corresponding second apparatus according to the second data feature. The reference gradient value is used to determine the gain score value corresponding to the first data feature for the current node, and the gain score value is used to determine the timing of splitting the current node.
[0168] Optionally, in some embodiments, the inverse mapping module 810 is specifically configured to: obtain the sample size and the mapping translation amount, where the mapping translation amount is the translation amount between the initial gradient value and the mapped gradient value; and perform inverse mapping processing on the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped respectively according to the sample size and the mapping translation amount.
[0169] Figure 10It is a schematic diagram of a data processing device according to another embodiment of the present disclosure. The device provides second data, and the second data corresponds to second data characteristics.
[0170] As Figure 10 shown, the data processing device 100 includes:
[0171] A second receiving module 1001, configured to receive an encrypted gradient value sent by a first device. The encrypted gradient value is generated by the first device, and the first device provides first data, and the first data and the second data are different;
[0172] A generating module 1002, configured to generate an encrypted tag and a value according to the second data characteristics and the encrypted gradient value;
[0173] Wherein, the first device determines the timing of splitting the current node based on the encrypted tag and the value, and the current node belongs to the current federated model to be trained.
[0174] Optionally, in some embodiments, the generating module 1002 is specifically configured to: parse an encrypted tag value from the encrypted gradient values according to the second data characteristics; process the encrypted tag value according to the binning information corresponding to the second data characteristics to obtain the encrypted tag and the value.
[0175] Optionally, in some embodiments, as Figure 10 shown, the data processing device 100 further includes: a sending module 1003, configured to send the encrypted tag and the value to the first device.
[0176] In the embodiments of the present disclosure, by obtaining an initial gradient value corresponding to the first data, the initial gradient value is used to determine the timing of splitting the current node, the current node belongs to the current federated model to be trained, and the initial gradient value is mapped to obtain a mapped gradient value. The mapped gradient value is a positive value, and the mapped gradient value is encoded to obtain an encoded gradient value, and the encoded gradient value is homomorphically encrypted to obtain an encrypted gradient value, and the encrypted gradient value is sent to the second device, which can map and encode the gradient value transmitted during the federated modeling process. Therefore, the phenomenon of gradient value overflow can be effectively avoided, thereby improving the effect of federated modeling.
[0177] To implement the above embodiments, the present disclosure also proposes a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the data processing method proposed in the foregoing embodiments of the present disclosure.
[0178] To implement the above embodiments, the present disclosure also proposes a non-transitory computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the data processing method proposed in the foregoing embodiments of the present disclosure.
[0179] To implement the above embodiments, the present disclosure also provides a computer program product. When the instruction processor in the computer program product executes, it performs the data processing method proposed in the foregoing embodiments of the present disclosure.
[0180] Figure 11 The block diagram of an exemplary computer device suitable for implementing the embodiments of the present disclosure is shown. Figure 11 The displayed computer device 12 is only an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0181] As Figure 11 shown, the computer device 12 is presented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0182] The bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnection (PCI) bus.
[0183] The computer device 12 typically includes a variety of computer system-readable media. These media can be any available media accessible by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0184] The memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 11not shown and is commonly referred to as a "hard disk drive").
[0185] Although Figure 11 not shown in Figure 11 , a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as: Compact Disc Read Only Memory (hereinafter referred to as: CD-ROM), Digital Video Disc Read Only Memory (hereinafter referred to as: DVD-ROM) or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 through one or more data medium interfaces. The memory 28 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.
[0186] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 42 generally perform the functions and / or methods in the embodiments described in the present disclosure.
[0187] The computer device 12 may also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and may also communicate with one or more devices that enable a user to interact with the computer device 12, and / or communicate with any device that enables the computer device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 22. Moreover, the computer device 12 may also communicate with one or more networks (such as a Local Area Network (hereinafter referred to as: LAN), a Wide Area Network (hereinafter referred to as: WAN) and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the computer device 12 through the bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in combination with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0188] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing the data processing method mentioned in the foregoing embodiments.
[0189] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the application disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0190] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
[0191] It should be noted that in the description of the present disclosure, the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present disclosure, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0192] Any process or method description in the flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present disclosure pertain.
[0193] It should be understood that the various parts of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0194] Those of ordinary skill in the art can understand that all or part of the steps carried out in implementing the above-described embodiment method can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0195] In addition, in each of the various embodiments of the present disclosure, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0196] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, or the like.
[0197] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0198] Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A data processing method, characterized in that Applied to a first device, the first device providing first data, the method comprising: Determining a gradient value of the first data at a current node as an initial gradient value, the initial gradient value being initial data for determining an opportunity to perform a splitting process on the current node, the current node belonging to a currently to-be-trained federated model; Performing a mapping process on the initial gradient value to obtain a mapped gradient value, the mapped gradient value being a positive value; Performing an encoding process on the mapped gradient value to obtain an encoded gradient value; and Performing a homomorphic encryption process on the encoded gradient value to obtain an encrypted gradient value, and sending the encrypted gradient value to a second device.
2. The method according to claim 1, wherein The initial gradient value includes: a first-order gradient value and a second-order gradient value, wherein, performing the mapping process on the initial gradient value to obtain a mapped gradient value includes: Performing the mapping process on the first-order gradient value to obtain a first-order mapped gradient value; Performing the mapping process on the second-order gradient value to obtain a second-order mapped gradient value, the first-order mapped gradient value and the second-order mapped gradient value being jointly used as the mapped gradient value.
3. The method according to claim 2, wherein The number of the first data is multiple, and the multiple first data respectively correspond to multiple first-order gradient values. Performing the mapping process on the first-order gradient value to obtain a first-order mapped gradient value includes: Determining a first reference gradient value, wherein the first reference gradient value is less than or equal to the smallest first-order gradient value among the multiple first-order gradient values; Respectively performing a difference process on the multiple first-order gradient values and the first reference gradient value to obtain multiple first-order mapped gradient values respectively corresponding to the multiple first-order gradient values.
4. The method according to claim 3, wherein The multiple first data respectively correspond to multiple second-order gradient values. Performing the mapping process on the second-order gradient value to obtain a second-order mapped gradient value includes: Determining a second reference gradient value, wherein the second reference gradient value is less than or equal to the smallest second-order gradient value among the multiple second-order gradient values; Respectively performing a difference process on the multiple second-order gradient values and the second reference gradient value to obtain multiple second-order mapped gradient values respectively corresponding to the multiple second-order gradient values.
5. The method according to claim 3, wherein Performing the encoding process on the mapped gradient value to obtain an encoded gradient value includes: Performing the encoding process on the first-order mapped gradient value to obtain a first-order encoded gradient value; Performing the encoding process on the second-order mapped gradient value to obtain a second-order encoded gradient value, the first-order encoded gradient value and the second-order encoded gradient value being jointly used as the encoded gradient value.
6. The method according to claim 5, characterized in that, Performing the encoding process on the first-order mapped gradient value to obtain a first-order encoded gradient value includes: Determining the number of encoding bits and a first encoding interval; Determining a first encoding range value and a second encoding range value according to the first encoding interval; and Performing an encoding process on the first-order mapped gradient value according to the first encoding range value and the second encoding range value in combination with the number of encoding bits to obtain the first-order encoded gradient value.
7. The method according to claim 6, wherein Performing the encoding process on the second-order mapped gradient value to obtain a second-order encoded gradient value includes: Determining a second encoding interval; Determine a third coding range value and a fourth coding range value according to the second coding interval; and According to the third coding range value and the fourth coding range value, combine the number of coding bits to perform coding processing on the second-order mapping gradient value to obtain the second-order coding gradient value.
8. The method according to claim 5, characterized in that Before performing homomorphic encryption processing on the coding gradient value to obtain an encrypted gradient value, it further includes:[[]] Perform splicing processing on the first-order coding gradient value and the second-order coding gradient value, and use the gradient value obtained by the splicing processing as the coding gradient value.
9. The method according to claim 8, wherein, Before performing splicing processing on the first-order coding gradient value and the second-order coding gradient value, it further includes:[[]] Determine a first starting position corresponding to the first-order coding gradient value; Determine a second starting position corresponding to the second-order coding gradient value; Configure overflow bits with a set number of bits at the first starting position and the second starting position respectively.
10. The method according to claim 2, characterized in that, The first data corresponds to a first data feature, the second device provides second data, and the second data corresponds to a second data feature. The method further includes:[[]] Receive the encrypted tag and value sent by the second device; Perform decryption processing on the encrypted tag and value to obtain a decrypted tag and value; According to the first data feature, perform disassembling processing on the decrypted tag and value to obtain a first-order gradient value-to-be-mapped corresponding to the first-order gradient value and a second-order gradient value-to-be-mapped; Perform inverse mapping processing on the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped respectively to obtain a first-order gradient initial value and a second-order gradient initial value; Perform aggregation processing on the first-order gradient initial value and the second-order gradient initial value to obtain a reference gradient value; Wherein, the encrypted tag and value is the sum value of multiple encrypted tag values, and the encrypted tag value is parsed from the encrypted gradient value by the corresponding second device according to the second data feature. The reference gradient value is used to determine the gain score value corresponding to the first data feature for the current node, and the gain score value is used to determine the timing of splitting the current node.
11. The method according to claim 10, characterized in that, The performing inverse mapping processing on the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped respectively includes:[[]] Obtain the sample size and the mapping translation amount, where the mapping translation amount is the translation amount between the initial gradient value and the mapping gradient value; According to the sample size and the mapping translation amount, perform inverse mapping processing on the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped respectively.
12. A data processing method, characterized in that, Applied to the second device, the second device provides second data, and the second data corresponds to a second data feature. The method includes:[[]] Receive the encrypted gradient value sent by the first device, and the encrypted gradient value is generated by the first device; Parse the encrypted tag value from the encrypted gradient value according to the second data feature; Process the encrypted tag value according to the binning information corresponding to the second data feature to obtain an encrypted tag and value; Send the encrypted tag and value to the first device, where the first device determines the timing of splitting the current node based on the encrypted tag and value, and the current node belongs to the federated model to be currently trained.
13. A data processing device, characterized in that, The device provides first data, and the device includes: An acquisition module, configured to determine the gradient value of the first data at the current node as an initial gradient value, where the initial gradient value is initial data for determining the timing of splitting the current node, and the current node belongs to the federated model to be currently trained; A mapping module, configured to perform a mapping process on the initial gradient value to obtain a mapped gradient value, where the mapped gradient value is a positive value; An encoding module, configured to perform an encoding process on the mapped gradient value to obtain an encoded gradient value; and An encryption module, configured to perform a homomorphic encryption process on the encoded gradient value to obtain an encrypted gradient value, and send the encrypted gradient value to a second device.
14. The device according to claim 13, characterized in that, The initial gradient value includes: a first-order gradient value and a second-order gradient value, where the mapping module includes: A first mapping sub-module, configured to perform the mapping process on the first-order gradient value to obtain a first-order mapped gradient value; A second mapping sub-module, configured to perform the mapping process on the second-order gradient value to obtain a second-order mapped gradient value, and the first-order mapped gradient value and the second-order mapped gradient value are jointly used as the mapped gradient value.
15. The device according to claim 14, characterized in that, The number of the first data is multiple, and the multiple first data respectively correspond to multiple first-order gradient values. The first mapping sub-module is specifically configured to: Determine a first reference gradient value, where the first reference gradient value is less than or equal to the smallest first-order gradient value among the multiple first-order gradient values; Respectively perform a difference process on the multiple first-order gradient values and the first reference gradient value to obtain multiple first-order mapped gradient values respectively corresponding to the multiple first-order gradient values.
16. The device according to claim 15, characterized in that, The multiple first data respectively correspond to multiple second-order gradient values. The second mapping sub-module is specifically configured to: Determine a second reference gradient value, where the second reference gradient value is less than or equal to the smallest second-order gradient value among the multiple second-order gradient values; Respectively perform a difference process on the multiple second-order gradient values and the second reference gradient value to obtain multiple second-order mapped gradient values respectively corresponding to the multiple second-order gradient values.
17. The device according to claim 15, wherein The encoding module includes: A first encoding sub-module, configured to perform the encoding process on the first-order mapped gradient value to obtain a first-order encoded gradient value; A second encoding sub-module, configured to perform the encoding process on the second-order mapped gradient value to obtain a second-order encoded gradient value, and the first-order encoded gradient value and the second-order encoded gradient value are jointly used as the encoded gradient value.
18. The device according to claim 17, characterized in that, The first encoding sub-module is specifically configured to: Determine the number of encoding bits and a first encoding interval; Determine a first encoding range value and a second encoding range value according to the first encoding interval; and According to the first encoding range value and the second encoding range value, and in combination with the number of encoding bits, perform an encoding process on the first-order mapped gradient value to obtain the first-order encoded gradient value.
19. The device according to claim 18, characterized in that, The second encoding sub-module is specifically configured to: Determine a second encoding interval; Determine a third encoding range value and a fourth encoding range value according to the second encoding interval; and Perform encoding processing on the second-order mapping gradient value according to the third encoding range value and the fourth encoding range value in combination with the number of encoding bits to obtain the second-order encoded gradient value.
20. The device according to claim 17, characterized in that, The device further includes: A splicing module, configured to perform splicing processing on the first-order encoded gradient value and the second-order encoded gradient value, and use the gradient value obtained by the splicing processing as the encoded gradient value.
21. The device according to claim 20, characterized in that, The device further includes an overflow bit configuration module, specifically configured to: Determine a first starting position corresponding to the first-order encoded gradient value; Determine a second starting position corresponding to the second-order encoded gradient value; Configure a set number of overflow bits at the first starting position and the second starting position respectively.
22. The device according to claim 14, characterized in that, The first data corresponds to a first data feature, the second device provides second data, the second data corresponds to a second data feature, and the device further includes: A first receiving module, configured to receive an encrypted tag and a value sent by the second device; A decryption module, configured to perform decryption processing on the encrypted tag and the value to obtain a decrypted tag and a value; A disassembling module, configured to perform disassembling processing on the decrypted tag and the value according to the first data feature to obtain a first-order gradient value-to-be-mapped and a second-order gradient value-to-be-mapped corresponding to the first-order gradient value; An inverse mapping module, configured to perform inverse mapping processing on the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped respectively to obtain a first-order gradient initial value and a second-order gradient initial value; An aggregation module, configured to perform aggregation processing on the first-order gradient initial value and the second-order gradient initial value to obtain a reference gradient value; Wherein, the encrypted tag and the value are a sum value of multiple encrypted tag values, and the encrypted tag value is parsed from the encrypted gradient values by the corresponding second device according to the second data feature. The reference gradient value is used to determine a gain score value corresponding to the first data feature for the current node, and the gain score value is used to determine the timing of splitting the current node.
23. The device according to claim 22, wherein The inverse mapping module is specifically configured to: Obtain a sample size and a mapping translation amount, where the mapping translation amount is the translation amount between the initial gradient value and the mapping gradient value; Perform inverse mapping processing on the first-order gradient value-to-be-mapped and the second-order gradient value-to-be-mapped respectively according to the sample size and the mapping translation amount.
24. A data processing device, characterized in that, The device provides second data, the second data corresponds to a second data feature, and the device includes: A second receiving module, configured to receive an encrypted gradient value sent by the first device, and the encrypted gradient value is generated by the first device; An analysis module, configured to parse an encrypted tag value from the encrypted gradient values according to the second data feature; A generation module, configured to process the encrypted tag value according to binning information corresponding to the second data feature to obtain an encrypted tag and a value; The first sending module is configured to send the encrypted tag and value to a first device, where the first device determines, based on the encrypted tag and value, the timing for splitting the current node, and the current node belongs to the federated model to be currently trained.
25. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method described in any one of claims 1-11, or implements the method described in claim 12.
26. A storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method described in any one of claims 1-11, or execute the method described in claim 12.
Citation Information
Patent Citations
Sample prediction method and device based on federation training and storage medium
CN109165683A
Data protection method based on WOE mask in multi-party longitudinal federated learning LightGBM training
CN113779608A