Data Processing Method and System
By constructing and splitting the processing of full binary trees in multi-party security computing, the problem of inefficient data processing of decision tree models and Bagging models is solved, and more efficient data processing is achieved. Especially in the privacy computing scenario, the computing speed is significantly improved and resource consumption is reduced.
Patent Information
- Application Number
- CN202310210538.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-02-28
AI Technical Summary
In multi-party security computing, the data processing efficiency of the decision tree model and the Bagging model is low, mainly due to the large number of comparison operations interacting, resulting in excessive computing resources and transmission consumption.
In the computing node, the initial node full binary tree and the initial leaf full binary tree are built, the model processing parameters and pending data are added to these trees respectively, and the target node full binary tree and the target leaf full binary tree are generated through split processing, and these trees are used to determine the target prediction results of the pending task.
Through dimensionality reduction processing, the communication interaction volume and memory usage between computing nodes are reduced, and data processing efficiency is improved. Especially in multi-party security computing scenarios, faster data processing speed and lower resource consumption are achieved.
Smart Images

Figure CN116227591B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a data processing method. This application also relates to a data processing system, a data processing device, a computing device, and a computer-readable storage medium. Background Art
[0002] Multi-party secure computation is a cryptographic technology. Multiple parties participating in the computation can jointly perform secure computation based on their respective input data by executing a multi-party secure computation algorithm to obtain a computation result without revealing the input data they hold to other parties.
[0003] When the computing party uses decision tree models and Bagging model algorithms to classify or regress data, due to a large number of comparison operation operations involved in the algorithms themselves, that is, the number of comparison operation interactions is relatively large, the data processing efficiency is low. Summary of the Invention
[0004] In view of this, embodiments of this application provide a data processing method. This application also relates to a data processing device, a computing device, and a computer-readable storage medium to solve the problem of low data processing efficiency in the prior art.
[0005] According to a first aspect of the embodiments of this application, a data processing method is provided, which is applied to each computing node in a set of computing nodes and includes:
[0006] Receiving model processing parameters corresponding to a task to be processed and the data to be processed, where both the model processing parameters and the data to be processed are encrypted;
[0007] Determining target decision tree structure information corresponding to a preset decision tree model set, and generating an initial node full binary tree and an initial leaf full binary tree according to the target decision tree structure information;
[0008] Adding the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively to generate a target node full binary tree and a target leaf full binary tree;
[0009] Performing splitting processing on the target node full binary tree, and determining a target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree.
[0010] According to a second aspect of the embodiments of this application, a data processing system is provided. The system includes a service initiator, at least one data provider, and at least one computing node;
[0011] The service initiator is configured to determine a task to be processed and send the model processing parameters corresponding to the task to be processed to each computing node, where the model processing parameters are encrypted data;
[0012] The at least one data provider is configured to determine the original data corresponding to the task to be processed, perform encryption processing on the original data, obtain the data to be processed, and send it to each computing node;
[0013] The at least one computing node is configured to receive the model processing parameters and the data to be processed, and determine the data prediction result corresponding to the task to be processed.
[0014] According to a third aspect of the embodiments of the present application, there is provided a data processing device, which is applied to each computing node in a set of computing nodes, and includes:
[0015] A data receiving module, which is configured to receive the model processing parameters and the data to be processed corresponding to the task to be processed, where both the model processing parameters and the data to be processed are encrypted;
[0016] A first structure tree generation module, which is configured to determine the target decision tree structure information corresponding to a preset decision tree model set, and generate an initial node full binary tree and an initial leaf full binary tree according to the target decision tree structure information;
[0017] A second structure tree generation module, which is configured to add the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively, and generate a target node full binary tree and a target leaf full binary tree;
[0018] A prediction result determination module, which is configured to perform splitting processing on the target node full binary tree, and determine the target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree.
[0019] According to a fourth aspect of the embodiments of the present application, there is provided a computing device, which includes a memory, a processor, and computer instructions stored on the memory and executable on the processor. When the processor executes the computer instructions, the steps of the data processing method are implemented.
[0020] According to a fifth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a processor, the steps of the data processing method are implemented.
[0021] The data processing method provided by this application is applied to each computing node in a set of computing nodes, and includes: receiving model processing parameters corresponding to a task to be processed and the data to be processed, where both the model processing parameters and the data to be processed are encrypted; determining target decision tree structure information corresponding to a preset decision tree model set, and generating an initial node full binary tree and an initial leaf full binary tree according to the target decision tree structure information; adding the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively to generate a target node full binary tree and a target leaf full binary tree; performing splitting processing on the target node full binary tree, and determining a target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree.
[0022] In an embodiment of this application, in each computing node, according to the target decision tree structure information in a preset decision tree model set, an initial node full binary tree and an initial leaf full binary tree are constructed, the model processing parameters and the data to be processed are added to these two full binary trees, and the target node full binary tree is split in a splitting manner, and then the split result is restored by using the target leaf full binary tree. This method slightly increases the amount of computation while greatly reducing the amount of communication interaction and memory occupation. At the same time, it can be applied to the privacy computing scenario of multi-party secure computing to improve data processing efficiency. Brief Description of the Drawings
[0023] Figure 1 is a schematic structural diagram of a data processing system provided by an embodiment of this application;
[0024] Figure 2 is a flowchart of a data processing method provided by an embodiment of this application;
[0025] Figure 3 is a flowchart of the model preprocessing stage of a data processing method provided by an embodiment of this application;
[0026] Figure 4 is a flowchart of the actual prediction stage of a data processing method provided by an embodiment of this application;
[0027] Figure 5 is a schematic structural diagram of a data processing device provided by an embodiment of this application;
[0028] Figure 6 is a block diagram of the structure of a computing device provided by an embodiment of this application. Detailed Description of the Embodiment
[0029] In the following description, numerous specific details are set forth to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the spirit of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.
[0030] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "said", and "the" used in one or more embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0031] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0032] First, the noun terms related to one or more embodiments of the present application are explained.
[0033] Privacy computing: Privacy computing is a new branch of data application technology formed in recent years driven by data security. Its main purpose is to protect data privacy while performing calculations on the data and then mining the value of the data. Privacy computing is divided into several main technical routes, namely multi-party secure computation, homomorphic encryption, federated learning, trusted hardware TEE, differential privacy. In this patent, the multi-party secure computation technical route is mainly used.
[0034] Multi-party secure computation: Abbreviated as MPC (multi-party secure computation), multi-party secure computation is an important branch of privacy computing. Its main purpose is to obtain the calculation result without exposing any input data and intermediate results. Since multi-party secure computation needs to protect the input data and intermediate results, a variety of cryptographic means are required to encrypt or obfuscate the data. These means include secret sharing of floating-point numbers, secret sharing on binary fields, etc. The introduction of these means has slowed down the calculation efficiency to a great extent, especially for comparison operations.
[0035] Tree Model: The tree model generally refers to the decision tree model, which is a commonly used method in traditional machine learning. Its generality and simple interpretability make it a very commonly used algorithm. There are generally two structures in the tree model, one is the NODE node, and the other is the LEAF leaf. The essential condition in the NODE is a judgment condition, and then the samples are divided into two parts according to certain conditions. When the samples are continuously split and finally stop according to certain conditions, these samples form a LEAF leaf. It can be regarded as the bifurcation of a tree, continuously bifurcating to the leaves. For example: among a group of people, there are men, women, the elderly, young and middle-aged people, and children. First, the first NODE divides the population into two parts, men and women, according to whether they are male. Then the judgment condition in the NODE is whether they are male. Then, among the men, they are further divided according to age into whether they are the elderly. If they are the elderly, the splitting stops, and a leaf LEAF is formed, which contains men and the elderly.
[0036] Ensemble Learning: The purpose of ensemble learning is that when the effect of traditional machine learning models is not good, multiple basic machine learning models can be combined in a certain way to form a larger and more complex machine learning model to improve accuracy and precision.
[0037] Bagging Model: The Bagging model is a type of ensemble learning. The Bagging method is a relatively simple ensemble learning method. Its main method is to form multiple basic learners through random sampling. Then the input data of these learners is often random, and the learners are independent of each other. When using the Bagging model for prediction, after obtaining the prediction results of multiple learners, a weighted average is performed.
[0038] Addition Circuit: The addition circuit is an important part of multi-party secure computing for computing comparison operations. Its main function is to achieve a secure addition through bit-level circuit operations, so that the input of each computing party is secure. Since in multi-party secure computing, a normal floating-point number is often divided into multiple seemingly random numbers and then given to multiple computing parties for custody. When these computing parties compare two numbers, they need to do subtraction first, and then add up the obtained several random numbers in the form of a secure addition circuit to get the final highest bit. If it is 0, then this number is greater than or equal to 0. If it is 1, then this number is less than 0. Then the result of the comparison operation is obtained by recursion.
[0039] Secret Sharing: It is to split the secret in an appropriate way. Each share after splitting is managed by different participants. A single participant cannot recover the secret information, and only several participants can cooperate to recover the secret message. More importantly, when any corresponding range of participants has problems, the secret can still be completely recovered.
[0040] Privacy computing is a new data application technology driven by data security needs in recent years. Its primary purpose is to unlock the value of data while protecting its privacy, making it "available but invisible." Generally speaking, privacy computing involves an initiator, multiple data providers, and multiple computing parties. The initiator wants to use the data provider's data for computation to extract value, while at the same time preventing the data provider from exposing its original data.
[0041] Among the many areas of privacy-preserving computing applications, machine learning is the most widespread and challenging. Machine learning in privacy-preserving computing is generally divided into two main scenarios: training algorithms and prediction algorithms. Decision tree and bagging models are very commonly used in prediction algorithms. This is because many traditional machine learning models are not very complex and require high interpretability and high model stability. Therefore, long-term plaintext models maintain long-term usability. However, since a good model often has many input variables, which may be distributed across multiple data sources, driven by data security requirements, the traditional plaintext prediction method needs to be replaced by the new secure prediction method based on privacy-preserving computing.
[0042] Among the various technical approaches in privacy-focused computing, multi-party secure computation (MPCC) is a key technology and branch. It utilizes rigorous cryptographic assumptions and encryption methods to ensure that each data party's original data remains private and that any intermediate results are not accessible to any participating party. However, while offering excellent confidentiality, MPC also carries significant computational and transmission overhead, especially for logical operations and some nonlinear operations.
[0043] At the same time, compared with other solutions, such as homomorphic encryption, multi-party secure computing is more universal and has lower computational complexity. Compared with differential privacy, it has higher computational accuracy. Differential privacy cannot handle some high-precision calculations, and decision trees often require high-precision logical operations. Federated learning often has certain defects in security.
[0044] Decision tree models and Bagging models often involve a large number of comparison operations, and comparison operations need to use addition circuits in MPC multi-party secure computing. We will slightly extend this method and the reasons for its poor performance here. Since all numbers running in MPC are either in integer form or in pure binary form, while most data in daily life is actually in floating-point (decimal) form, this makes it necessary to convert floating-point numbers to integers when performing operations in MPC. Generally, the method is to multiply by a large integer to ensure accuracy. Generally, 2^32 is selected as the large integer, and the general limit of floating-point numbers is within 2^32. In this way, the range of the multiplied large integer will not exceed 2^64, that is, it can be represented by 64 bits. Then, in order to ensure security, the obtained large integer needs to be divided into multiple parts and given to multiple computing parties. When the computing parties want to determine whether the integer is greater than 0, they need to add up the multiple parts. Since the addition circuit runs on multiple computing parties, each bit requires one round of interaction, which makes the overall interaction rounds of the addition circuit reach as many as 64 rounds. That is to say, for one logical operation, multiple participating parties need to interact 64 times.
[0045] During the process of decision tree splitting, each NODE needs to perform one logical operation and then perform an AND operation with the result of the previous NODE to obtain the final splitting result of this NODE. Generally, one AND operation requires one round of interaction. Therefore, each NODE split finally requires 64 + 1 = 65 interactions. The above methods are huge in terms of computational volume and operation volume because a tree usually has many NODES, and there are usually more than 100 decision trees in the Bagging method. Based on this, an efficient Bagging decision tree prediction method based on multi-party secure computing is required in the embodiments of this application.
[0046] In this application, a data processing method is provided. This application is also related to a data processing system, a data processing device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0047] Figure 1 It is a schematic diagram of the system structure of a data processing system provided by an embodiment of this application.
[0048] Figure 1As shown, it includes a service initiator 102, at least one data provider 104, and at least one computing node 106. It should be noted that the service initiator can be understood as the party that needs to perform multi-party secure computing, and it has model processing parameters and tasks to be executed for computing. The data provider can be understood as multiple entities that provide encrypted data, and the data they provide can be used for multi-party secure computing tasks. The computing node can be understood as the node entity that performs multi-party secure computing on the provided encrypted data. In this embodiment, the forms of the above-mentioned entities are not limited in any way.
[0049] The service initiator 102 is configured to determine the task to be processed and send the model processing parameters corresponding to the task to be processed to each computing node, where the model processing parameters are encrypted data.
[0050] The at least one data provider 104 is configured to determine the original data corresponding to the task to be processed, perform encryption processing on the original data, obtain the data to be processed, and send it to each computing node.
[0051] The at least one computing node 106 is configured to receive the model processing parameters and the data to be processed and determine the data prediction result corresponding to the task to be processed.
[0052] Among them, the task to be processed can be understood as the task that needs to be processed by multi-party secure computing. The model processing parameters can be understood as the parameters related to the decision tree model, such as node judgment operation conditions, etc. The data to be processed can be understood as the data obtained by encrypting the original data provided by the data provider. The data prediction result can be understood as the task result of the task to be processed predicted by the computing node according to each input condition and data.
[0053] In practical applications, during the process of multi-party secure computation, each computing node can use the model processing parameters provided by the service initiator and the data to be processed provided by each data provider to obtain the data prediction results of the task to be processed. It should be noted that both the model processing parameters and the data to be processed are encrypted data, and the initiator and data providers can send them to each computing node using the secret sharing method. One implementation method is as follows: The initiator, i.e., the model owner, divides the model parameters into multiple parts in the form of secret sharing and then sends them to each computing party. Without collusion with each other, the computing parties cannot obtain the model parameters. Generally speaking, the initiator also acts as a computing party. Therefore, as long as the initiator does not collude with other parties, the model parameters are secure. All data parties divide the original data into multiple parts in the form of secret sharing and then send them to each computing party. At the same time, here the data parties generally also act as each computing party. At this time, the data distribution stage is completed. This stage ensures the security of the data. Secret sharing is an important means in MPC and can well protect the data and intermediate results.
[0054] Further, the at least one computing node 106 is further configured to determine an initial node full binary tree and an initial leaf full binary tree based on a preset decision tree model set; add the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively to generate a target node full binary tree and a target leaf full binary tree; perform splitting processing on the target node full binary tree, and determine the data prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree; and send the data prediction result to the service initiator.
[0055] Among them, the initial node full binary tree can be understood as a full tree of nodes containing all possible bifurcations of a tree with the maximum depth, and the initial leaf full binary tree can be understood as a full tree of leaves containing all possible bifurcations of a tree with the maximum depth. At the same time, the depth of the initial leaf full binary tree is 1 more than that of the initial node full binary tree.
[0056] In practical applications, each computing node can extract the decision tree model set, determine the initial node full binary tree and the initial leaf full binary tree according to the decision tree structure in the decision tree model set, and add the received model processing parameters and data to be processed to the two full trees respectively to generate a target node full binary tree and a target leaf full binary tree. Then, split the target node full binary tree, and determine the data prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree. Each computing node sends the calculated data prediction result to the service initiator.
[0057] This method can reduce the dimensionality of multiple decision tree models into two full trees. One of these two full trees is used for splitting, and the other is used for recovering the result after splitting to achieve acceleration in data processing performance.
[0058] Further, the service initiator 102 is further configured to receive the data prediction results sent by each computing node; and process the data prediction results sent by each computing node based on the random forest algorithm to obtain the target prediction result corresponding to the task to be processed.
[0059] In practical applications, in multi-party secure computing, the random forest algorithm can be used to process the data prediction results reported by each computing node, so that the service initiator can obtain the target prediction result corresponding to the task to be processed. It should be noted that for tasks to be processed in different service scenarios, the calculated target prediction results will vary, and no limitations are made in this embodiment. For example, after the initiator decrypts, it obtains the prediction results of each tree, and then the initiator can perform Bagging processing on the prediction results of each tree to obtain the Bagging prediction result.
[0060] Figure 2 The flowchart shows a data processing method according to an embodiment of the present application, which specifically includes the following steps:
[0061] It should be noted that the data processing method provided in this embodiment is applied to each computing node in the computing node set. That is to say, each computing node can calculate and process the data to be processed using the same steps or methods, but their specific implementation methods will also be different, and no limitations are made in this embodiment. For the sake of understanding, only the processing method of one computing node is used as an example for description.
[0062] Step 202: Receive the model processing parameters corresponding to the task to be processed and the data to be processed, where both the model processing parameters and the data to be processed are encrypted.
[0063] Among them, the task to be processed can be understood as a task that needs to calculate a certain data set in multi-party secure computing. Different application scenarios and requirements can correspond to different tasks to be processed; the model processing parameters can be understood as the model parameters of the model used when calculating the data corresponding to the task to be processed. The models here include but are not limited to decision tree models, Bagging models, etc.; the data to be processed can be understood as the data content processed by multi-party secure computing, and both the data to be processed and the model processing parameters are encrypted data.
[0064] In practical applications, in order to ensure data security, the data can be shared using the secret sharing method. The computing node can receive the model processing parameters corresponding to the model calculation sent by the model initiator, and can also receive the data to be processed sent by the data provider, thus completing the data reception step.
[0065] Step 204: Determine the target decision tree structure information corresponding to the preset decision tree model set, and generate an initial node full binary tree and an initial leaf full binary tree according to the target decision tree structure information.
[0066] Among them, the preset decision tree model set can be understood as a set composed of multiple preset trained decision tree models, and there is no data in this decision tree; the target decision tree structure information can be understood as the structure information such as the depth of each NODE node and LEAF in each decision tree and their positions in this layer.
[0067] In practical applications, the computing node can also preprocess the preset decision tree model set. In order to improve the data processing efficiency and reduce the number and time of mutual decision interactions of decision trees, in this embodiment, after extracting each decision tree model in all preset decision tree model sets, two large trees can be generated, one is a NODE tree and the other is a LEAF tree, that is, an initial node full binary tree and an initial leaf full binary tree; it should be noted that the NODE tree and the LEAF tree will be full trees, and at the same time, the depth of the LEAF tree is the depth of the NODE tree + 1. The meaning of a full tree is that these two trees will contain all possible bifurcations of the tree with the maximum depth. For example: in all decision tree models, the depth of the deepest tree is 3, then the generated NODE and LEAF trees will contain all bifurcations of the tree with a depth of 3.
[0068] Further, the determining the target decision tree structure information corresponding to the preset decision tree model set includes:
[0069] Determine at least one decision tree model in the preset decision tree model set;
[0070] Obtain the initial decision tree structure information corresponding to each decision tree model, where the initial decision tree structure information includes node position information and leaf position information;
[0071] Based on the initial decision tree structure information corresponding to each decision tree, determine the target decision tree structure information.
[0072] In practical applications, the preset decision tree model set contains at least one decision tree model. Among them, the structures of each decision tree model, such as depth, bifurcation relationship, etc., may not be the same. Furthermore, the initial decision tree result information corresponding to each decision tree model can be obtained, including but not limited to node position information and leaf position information; then, according to the node position information, leaf position information, etc. corresponding to each decision tree, the target decision tree structure information is determined to facilitate the subsequent construction of a full binary tree according to the target decision tree structure information.
[0073] Even further, the generating an initial node full binary tree and an initial leaf full binary tree according to the target decision tree structure information includes:
[0074] Process the at least one decision tree model according to the target decision tree structure information to generate an initial node full binary tree and an initial leaf full binary tree, where the depth of the initial leaf full binary tree is 1 greater than the depth of the initial node full binary tree.
[0075] In practical applications, after the computing node determines the target decision tree structure information, it can perform a merging process on each decision tree model to generate two trees that contain all possible bifurcations of the tree with the maximum depth. One is the initial node full binary tree, and the other is the initial leaf full binary tree. At the same time, the depth of the LEAF tree is the depth of the NODE tree + 1.
[0076] Step 206: Add the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively to generate a target node full binary tree and a target leaf full binary tree.
[0077] Furthermore, the computing node can also add the received model processing parameters and the data to be processed to these two full binary trees respectively. It should be noted that the specific data mounting method is not limited in this embodiment.
[0078] In addition, in these two full binary trees, the computing party can also store all the NODE position information extracted previously in the NODE tree and store all the LEAF position information in the LEAF tree. Specifically, before adding the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively, it further includes:
[0079] Traverse all nodes of the initial node full binary tree and store the decision tree number corresponding to the node position information that is consistent with the current node position information in each node.
[0080] Traverse all leaf nodes of the initial leaf full binary tree and store the decision tree number corresponding to the leaf position information that is consistent with the current leaf position information in each leaf node.
[0081] In practical applications, the computing node needs to traverse these two full binary trees respectively. In each node of the NODE tree, the decision tree numbers corresponding to the node position information in each decision tree model consistent with the current node position information are stored. In each leaf of the LEAF tree, the decision tree numbers corresponding to the leaf position information in each decision tree consistent with the current leaf node position information are stored. For example, at depth 3 and position 3, it is the NODE position of the four trees [1, 5, 7, 100] and the LEAF position of the two trees [2, 88], that is, the large NODE tree: at depth 3 and position 3 in the large NODE tree, store the four numbers [1, 5, 7, 100]; the large LEAF tree: at depth 3 and position 3 in the large LEAF tree, store the two numbers [2, 88].
[0082] Furthermore, the computing node adds the model parameters distributed by the initiator before in the large NODE tree and the large LEAF tree. The model parameters mainly include the splitting information of the NODE and the predicted values of the LEAF. Specifically, the model processing parameters include the splitting information of each tree node and the predicted values of each leaf node;
[0083] Correspondingly, adding the model processing parameters to the initial node full binary tree and the initial leaf full binary tree includes:
[0084] Determine the splitting information of each tree node at the first storage position corresponding to the initial node full binary tree, and add the splitting information of each tree node to the initial node full binary tree;
[0085] Determine the predicted values of each leaf node at the second storage position corresponding to the initial leaf full binary tree, and add the predicted values of each leaf node to the initial leaf full binary tree.
[0086] In practical applications, determine the splitting information corresponding to each NODE, store the splitting information in the first storage position, and then, the splitting information can be added to the initial node full binary tree in each NODE node; correspondingly, the predicted value of each leaf node can also be determined, store the predicted value in the second storage position, and then, the predicted values of each leaf node are added to the initial leaf full binary tree.
[0087] Based on this, the preprocessing stage of the model is completed at this time. Multiple decision trees are successfully reduced to two trees. One of these two trees will be used for splitting and the other for restoring the results after splitting, achieving performance acceleration.
[0088] Step 208: Perform splitting processing on the target node full binary tree, and determine the target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree.
[0089] In practical applications, the computing node can split the target node full binary tree, and determine the target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree.
[0090] Specifically, it can be performed by traversing and executing the splitting layer by layer for each node, and this embodiment does not limit this; the splitting process of the target node full binary tree includes:
[0091] Traverse all nodes in the target node full binary tree based on a preset traversal rule, and extract the decision tree number and the stored splitting information corresponding to each node;
[0092] Determine the variable value of the data to be processed stored in each node, and based on the variable value, the decision tree number and the stored splitting information, split the target node full binary tree.
[0093] In practical applications, an optional method is to first use the NODE tree, start traversing from depth 1 downward, and the traversal method is to traverse each NODE in the current depth first, and then traverse to the next depth. Every time a NODE is reached, extract the tree number and the corresponding splitting value in the NODE, and then, according to the splitting value, variable and tree number, pack the information to be split into a long vector, use this long vector for logical operations, and then complete the splitting to form a long vector X. Then, extract the information of these trees in the previous splitting from the splitting information of the previous parent NODE, and then pack it to form another long vector Y. Then calculate X and Y to get the final result Z of this splitting. Then, split the information of Z into the form of each tree and store it in the current node as the splitting information; that is, during the splitting process, the results of multiple logical operations will be packed in the same byte or int structure, saving the memory space and transmission volume consumed by storing logical operations.
[0094] Furthermore, the computing node can use the LEAF tree to restore the result, and the traversal method can adopt the above-mentioned method of traversing the NODE tree; specifically, the determining of the target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree includes:
[0095] Traverse all nodes in the target leaf full binary tree based on a preset traversal rule, and extract the decision tree number and the stored encrypted value corresponding to each node;
[0096] Determine the corresponding splitting information from the splitting result based on the decision tree number;
[0097] Determine the target prediction result corresponding to the task to be processed according to the decision tree number, the stored encrypted value and the splitting information.
[0098] In practical applications, in an alternative implementation, when using the LEAF tree recovery result, the traversal method is the same. For each traversed LEAF, all the tree numbers and the corresponding encrypted values in the LEAF are extracted. Then, the splitting information and splitting results of the trees corresponding to these numbers are taken out from the parent NODE, that is, from the large NODE tree. Then, the ciphertext prediction result in the LEAF tree is taken out. Then, the splitting result and the ciphertext prediction result are stored according to each tree. Furthermore, the target prediction result corresponding to the task to be processed is obtained.
[0099] In addition, the computing node reports the calculated target prediction result to the initiator. After the initiator decrypts it, the prediction result of each tree is obtained. Then, the initiator can perform Bagging processing on the prediction results of each tree, and thus obtain the result of Bagging prediction.
[0100] It should be noted that when generating the NODE tree and the LEAF tree, splitting them into two is just one processing method, and they can also be combined into a single tree containing NODE and LEAF; this embodiment does not limit this. In fact, due to the characteristics of MPC, when the traversal of the NODE tree ends, the LEAF tree is also completed at the same time, and vice versa. Only a simple conversion is needed. Since the traversal process requires the participation of multiple parties, and the conversion is just a local calculation that can be ignored, it has almost no impact on performance; in the actual prediction stage, at each depth, the content of each NODE is packed. At the same time, this embodiment points out that all the content at this depth can be packed at each depth, so as to obtain better results when the data volume is small and the maximum depth is large.
[0101] Figure 3 It is the flowchart of the model preprocessing stage of a data processing method provided by an embodiment of the present application; it should be noted that each computing node in the computing node set can refer to this method for model preprocessing, but it is not specifically limited, and other processing methods can also be adopted.
[0102] Step 302: Each computing node respectively parses the NODE nodes and LEAF information of each tree, and then stores them.
[0103] Step 304: Each computing node generates two large trees, one is the NODE tree and the other is the LEAF tree.
[0104] Step 306: Each computing node stores the results in the large trees according to the NODE and LEAF of each obtained tree. The NODEs of all trees are placed in the corresponding positions in the large NODE tree according to the position and depth. Similarly, the same is true for the LEAFs.
[0105] Step 308: Store the encrypted model parameters distributed by the initiator during the data distribution phase in the NODE tree and the LEAF tree according to NODE and LEAF.
[0106] Figure 4 It is the flowchart of the actual prediction phase of a data processing method provided by an embodiment of the present application; it should be noted that each computing node in the computing node set can refer to this method for processing in the actual prediction phase, but it is not specifically limited, and other processing methods can also be used.
[0107] Step 402: All computing nodes split according to the two large trees.
[0108] Step 404: If it is detected that there is content in the NODE tree at the current position, pack the content to be split together for splitting.
[0109] Step 406: If it is detected that there is content in the LEAF tree at the current position, place the values split from the previous NODE at this position in the LEAF, corresponding one by one to the content of the LEAF tree.
[0110] Step 408: After all content is completed, the initiator calls the LEAF tree to restore the result and summarizes and returns the prediction result.
[0111] In summary, the data processing method provided by this embodiment proposes an optimization scheme for the prediction of the Bagging tree model in the privacy computing scenario. It uses the MPC technology, consumes less computing resources than the conventional MPC parallelization scheme, has a shorter operation time, and the performance is improved from 20 seconds for 1000 samples on a 4-core CPU to 3 seconds for 1000 samples on a single-core CPU; in addition, a data structure of the Bagging tree model is used to reduce the interactive complexity (the interactive complexity is the most important aspect restricting the performance in MPC, rather than the actual transmission complexity or the CPU operation complexity), such as the complexity of O(N2^k) to O(2^k) complexity, where N is the number of trees and k is the maximum tree depth; for the splitting method of the PACK ciphertext of the Bagging tree, the non-interfering splitting information of the current layer depth is packed to form one splitting information, and multiple splits are solved in one round of interaction, and finally the effect of reducing the communication interaction rounds is achieved; for the vectorized MPC ciphertext processing method of the Bagging tree, the ciphertext operation uses simple splitting and merging methods, which greatly reduces the communication interaction volume and memory occupancy while slightly increasing the calculation amount.
[0112] Corresponding to the above method embodiment, the present application also provides an embodiment of a data processing device. Figure 5 It shows the structural schematic diagram of a data processing device provided by an embodiment of the present application. As Figure 5As shown, the device is applied to each computing node in a set of computing nodes, and includes:
[0113] A data receiving module 502, configured to receive model processing parameters corresponding to a task to be processed and data to be processed, where both the model processing parameters and the data to be processed are encrypted;
[0114] A first structure tree generation module 504, configured to determine target decision tree structure information corresponding to a set of preset decision tree models, and generate an initial node full binary tree and an initial leaf full binary tree according to the target decision tree structure information;
[0115] A second structure tree generation module 506, configured to add the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively, to generate a target node full binary tree and a target leaf full binary tree;
[0116] A prediction result determination module 508, configured to perform splitting processing on the target node full binary tree, and determine a target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree.
[0117] Optionally, the first structure tree generation module 504 is further configured to:
[0118] Determine at least one decision tree model in the set of preset decision tree models;
[0119] Obtain initial decision tree structure information corresponding to each decision tree model, where the initial decision tree structure information includes node position information and leaf position information;
[0120] Based on the initial decision tree structure information corresponding to each decision tree, determine the target decision tree structure information.
[0121] Optionally, the first structure tree generation module 504 is further configured to:
[0122] Process the at least one decision tree model according to the target decision tree structure information to generate an initial node full binary tree and an initial leaf full binary tree, where the layer depth of the initial leaf full binary tree is 1 greater than the layer depth of the initial node full binary tree.
[0123] Optionally, the device further includes:
[0124] A structure tree processing module, configured to traverse all nodes of the initial node full binary tree, and store a decision tree number corresponding to the node position information consistent with the current node position information in each node;
[0125] Traverse all leaf nodes of the initial leaf full binary tree, and store the decision tree number corresponding to the leaf position information consistent with the current leaf node position information in each leaf node.
[0126] Optionally, the model processing parameters include the splitting information of each tree node and the predicted values of each leaf node;
[0127] Optionally, the second structure tree generation module 506 is further configured to:
[0128] Determine the splitting information of each tree node at the first storage position corresponding to the initial node full binary tree, and add the splitting information of each tree node to the initial node full binary tree;
[0129] Determine the predicted values of each leaf node at the second storage position corresponding to the initial leaf full binary tree, and add the predicted values of each leaf node to the initial leaf full binary tree.
[0130] Optionally, the prediction result determination module 508 is further configured to:
[0131] Traverse all nodes in the target node full binary tree based on a preset traversal rule, and extract the decision tree number and the stored splitting information corresponding to each node;
[0132] Determine the variable value of the data to be processed stored in each node, and perform splitting processing on the target node full binary tree based on the variable value, the decision tree number, and the stored splitting information.
[0133] Optionally, the prediction result determination module 508 is further configured to:
[0134] Traverse all nodes in the target leaf full binary tree based on a preset traversal rule, and extract the decision tree number and the stored encrypted value corresponding to each node;
[0135] Determine the corresponding splitting information from the splitting result based on the decision tree number;
[0136] Determine the target prediction result corresponding to the task to be processed according to the decision tree number, the stored encrypted value, and the splitting information.
[0137] The data processing device provided by the embodiment of the present application constructs an initial node full binary tree and an initial leaf full binary tree in each computing node according to the target decision tree structure information in the preset decision tree model set, adds model processing parameters and data to be processed to these two full binary trees, and uses a splitting method to split the target node full binary tree, and then uses the target leaf full binary tree to restore the split result. This method slightly increases the computational amount while greatly reducing the communication interaction amount and memory occupancy. At the same time, it is applicable to the privacy computing scenario of multi-party secure computing and improves the data processing efficiency.
[0138] The above is a schematic solution of a data processing device in this embodiment. It should be noted that the technical solution of this data processing device and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above data processing method.
[0139] Figure 6 FIG. shows a structural block diagram of a computing device 600 according to an embodiment of the present application. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 through a bus 630, and a database 650 is used to store data.
[0140] The computing device 600 further includes an access device 640, and the access device 640 enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interfaces (for example, a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0141] In one embodiment of the present application, the above components of the computing device 600 and Figure 6 other components not shown therein may also be connected to each other, for example, via a bus. It should be understood that Figure 6 the block diagram of the computing device shown is merely for illustrative purposes and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.
[0142] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smart phones), wearable computing devices (e.g., smart watches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs, Personal Computers). The computing device 600 can also be a mobile or stationary server.
[0143] Wherein, when the processor 620 executes the computer instructions, the steps of the data processing method are implemented.
[0144] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above data processing method.
[0145] An embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions, and when the computer instructions are executed by a processor, the steps of the data processing method as described above are implemented.
[0146] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above data processing method.
[0147] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0148] The computer instructions include computer program code, which may be in the form of source code, object code, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0149] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, some steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily all essential to this application.
[0150] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0151] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The optional embodiments do not elaborate on all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this application. The present application selects and specifically describes these embodiments in order to better explain the principle and practical application of the present application, so that those skilled in the art can understand and utilize the present application well. The present application is only limited by the claims and their full scope and equivalents.
Claims
1. A data processing method, characterized in that, Applied to each computing node in the set of computing nodes, including: Receiving model processing parameters corresponding to the task to be processed and the data to be processed, wherein both the model processing parameters and the data to be processed are encrypted; Determining the target decision tree structure information corresponding to the preset decision tree model set, and generating an initial node full binary tree and an initial leaf full binary tree according to the target decision tree structure information; Adding the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively to generate a target node full binary tree and a target leaf full binary tree, wherein the model processing parameters include the splitting information of each tree node and the prediction values of each leaf node. Adding the model processing parameters to the initial node full binary tree and the initial leaf full binary tree includes: determining the splitting information of each tree node at the first storage position corresponding to the initial node full binary tree, and adding the splitting information of each tree node to the initial node full binary tree; determining the prediction values of each leaf node at the second storage position corresponding to the initial leaf full binary tree, and adding the prediction values of each leaf node to the initial leaf full binary tree; Performing a splitting process on the target node full binary tree, and determining the target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree; Wherein, performing a splitting process on the target node full binary tree includes traversing all nodes in the target node full binary tree based on a preset traversal rule, and extracting the decision tree number and the stored splitting information corresponding to each node; determining the variable value of the data to be processed stored in each node, and performing a splitting process on the target node full binary tree based on the variable value, the decision tree number and the stored splitting information; Determining the target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree includes traversing all nodes in the target leaf full binary tree based on a preset traversal rule, and extracting the decision tree number and the stored encrypted value corresponding to each node; determining the corresponding splitting information from the splitting result based on the decision tree number; and determining the target prediction result corresponding to the task to be processed according to the decision tree number, the stored encrypted value and the splitting information.
2. The method according to claim 1, wherein The determining the target decision tree structure information corresponding to the preset decision tree model set includes: Determining at least one decision tree model in the preset decision tree model set; Obtaining the initial decision tree structure information corresponding to each decision tree model, wherein the initial decision tree structure information includes node position information and leaf position information; Based on the initial decision tree structure information corresponding to each decision tree, determining the target decision tree structure information.
3. The method according to claim 2, characterized in that, The generating an initial node full binary tree and an initial leaf full binary tree according to the target decision tree structure information includes: Processing the at least one decision tree model according to the target decision tree structure information to generate an initial node full binary tree and an initial leaf full binary tree, wherein the depth of the initial leaf full binary tree is 1 greater than the depth of the initial node full binary tree.
4. The method according to claim 3, wherein Before adding the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively, it further includes: Traversing all nodes of the initial node full binary tree, and storing in each node the decision tree number corresponding to the node position information consistent with the current node position information; Traversing all leaf nodes of the initial leaf full binary tree, and storing in each leaf node the decision tree number corresponding to the leaf position information consistent with the current leaf node position information.
5. A data processing system, characterized in that, The system includes a service initiator, at least one data provider, and at least one computing node; The service initiator is configured to determine a task to be processed and send the model processing parameters corresponding to the task to be processed to each computing node, wherein the model processing parameters are encrypted data; The at least one data provider is configured to determine the original data corresponding to the task to be processed, perform encryption processing on the original data to obtain data to be processed, and send the data to be processed to each computing node; The at least one computing node is configured to receive the model processing parameters and the data to be processed, determine the target decision tree structure information corresponding to the preset decision tree model set, and generate an initial node full binary tree and an initial leaf full binary tree according to the target decision tree structure information; add the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively to generate a target node full binary tree and a target leaf full binary tree, wherein the model processing parameters include the splitting information of each tree node and the prediction values of each leaf node. Adding the model processing parameters to the initial node full binary tree and the initial leaf full binary tree includes: determining the splitting information of each tree node at the first storage position corresponding to the initial node full binary tree, and adding the splitting information of each tree node to the initial node full binary tree; determining the prediction values of each leaf node at the second storage position corresponding to the initial leaf full binary tree, and adding the prediction values of each leaf node to the initial leaf full binary tree; performing splitting processing on the target node full binary tree, and determining the data prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree; sending the data prediction result to the service initiator; Wherein, performing splitting processing on the target node full binary tree includes traversing all nodes in the target node full binary tree based on a preset traversal rule, and extracting the decision tree number and the stored splitting information corresponding to each node; determining the variable value of the data to be processed stored in each node, and performing splitting processing on the target node full binary tree based on the variable value, the decision tree number, and the stored splitting information; Determining the target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree includes traversing all nodes in the target leaf full binary tree based on a preset traversal rule, and extracting the decision tree number and the stored encrypted value corresponding to each node; determining the corresponding splitting information from the splitting result based on the decision tree number; and determining the target prediction result corresponding to the task to be processed according to the decision tree number, the stored encrypted value, and the splitting information.
6. The system according to claim 5, characterized in that The service initiator is further configured to receive the data prediction results sent by each computing node; and process the data prediction results sent by each computing node based on the random forest algorithm to obtain the target prediction result corresponding to the task to be processed.
7. A data processing device, characterized in that, Applied to each computing node in the computing node set, it includes: A data receiving module configured to receive the model processing parameters and the data to be processed corresponding to the task to be processed, where both the model processing parameters and the data to be processed are encrypted. A first structure tree generation module configured to determine the target decision tree structure information corresponding to the preset decision tree model set, and generate an initial node full binary tree and an initial leaf full binary tree according to the target decision tree structure information. A second structure tree generation module configured to add the model processing parameters and the data to be processed to the initial node full binary tree and the initial leaf full binary tree respectively to generate a target node full binary tree and a target leaf full binary tree, where the model processing parameters include the splitting information of each tree node and the predicted value of each leaf node. Adding the model processing parameters to the initial node full binary tree and the initial leaf full binary tree includes: determining the splitting information of each tree node at the first storage position corresponding to the initial node full binary tree, and adding the splitting information of each tree node to the initial node full binary tree; determining the predicted value of each leaf node at the second storage position corresponding to the initial leaf full binary tree, and adding the predicted value of each leaf node to the initial leaf full binary tree. A prediction result determination module configured to perform a splitting process on the target node full binary tree, and determine the target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree. Among them, performing a splitting process on the target node full binary tree includes traversing all nodes in the target node full binary tree based on a preset traversal rule, and extracting the decision tree number and the stored splitting information corresponding to each node; determining the variable value of the data to be processed stored in each node, and performing a splitting process on the target node full binary tree based on the variable value, the decision tree number, and the stored splitting information. Determining the target prediction result corresponding to the task to be processed according to the splitting result and the target leaf full binary tree includes traversing all nodes in the target leaf full binary tree based on a preset traversal rule, and extracting the decision tree number and the stored encrypted value corresponding to each node; determining the corresponding splitting information from the splitting result based on the decision tree number; and determining the target prediction result corresponding to the task to be processed according to the decision tree number, the stored encrypted value, and the splitting information.
8. A computing device, comprising a memory, a processor, and computer instructions stored on the memory and executable on the processor, characterized in that, When the processor executes the computer instructions, the steps of the method according to any one of claims 1-4 are implemented.
9. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by the processor, the steps of the method according to any one of claims 1-4 are implemented.
Citation Information
Patent Citations
Huffman material purchasing decision-making algorithm
CN106611284A
Single-lead electroencephalography (EEG) automatic sleep staging method
CN109674468A