Data processing method and related device
By changing the comparison of data feature parameters with child node selection indicator parameters into local operations, the existing anonymous vertical federated algorithm is solved, and more efficient anonymous vertical federated reasoning is achieved.
Patent Information
- Application Number
- CN202510198900.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing anonymous vertical federated algorithm is less efficient in privacy comparison operations, resulting in an overall inference efficiency.
The comparison of data feature parameters with the child node selection indicator parameters is changed into a local operation to avoid data exchange with other participants, thereby saving the overhead of the privacy comparison process.
It improves the efficiency of privacy comparison and improves the overall efficiency of anonymous vertical federal reasoning.
Smart Images

Figure CN119990263A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data processing method and related devices. Background Art
[0002] As various enterprises and institutions collect, store and maintain different data, there are more and more demands to break through the data walls between enterprises and institutions to achieve more comprehensive and accurate data analysis. Vertical federated learning is widely used due to its efficiency and explainability in finance, medical and other fields. For example, in financial scenarios, banks and payment institutions each hold different behavioral characteristics of the same user. The two parties jointly construct a federated gradient boosting decision tree (GBDT) to determine whether there is non-compliant behavior.
[0003] At the beginning of the algorithm, different participants need to align the same group of users. Therefore, the Private Set Intersection (PSI) algorithm is needed to obtain the data of intersection users. After the intersection is completed, no data of non-intersection users will be exposed to the other party. However, this will expose the data of intersection users to each other, and there is a risk of label leakage.
[0004] To deal with this situation, anonymous federation algorithms have been proposed, and most of these solutions are based on Circuit Private Set Intersection (Circuit-PSI). However, the anonymous federation algorithms provided by related technologies require a large number of privacy comparison operations between data feature parameters and the sub-node selection indicator parameters of GBDT nodes, and the overhead of privacy comparison operations is very large, resulting in low efficiency of anonymous federation algorithms provided by related technologies. Summary of the invention
[0005] In order to solve the above technical problems, the present application provides a data processing method and related devices, which transform the comparison between data feature parameters and sub-node selection indication parameters into local operations, thereby eliminating the need to exchange data with other parties, saving a large amount of overhead caused by data exchange in the privacy comparison process, improving the efficiency of privacy comparison, and thereby improving the efficiency of anonymous vertical federated reasoning.
[0006] The embodiments of the present application disclose the following technical solutions:
[0007] On the one hand, an embodiment of the present application provides a data processing method, which is performed by a first participant and includes:
[0008] Obtaining business data and a first decision tree, where the first decision tree runs on the first participant, the first decision tree is obtained through vertical federated learning training, and includes N trees, where the N trees respectively include a plurality of leaf nodes and a plurality of non-leaf nodes, the plurality of non-leaf nodes are respectively pre-configured with a child node selection indication parameter, and the child node selection indication parameter is used to divide the child nodes of the non-leaf nodes used to process the business data;
[0009] For a non-leaf node j of the t-th tree among the N trees, in response to the child node selection indication parameter of the non-leaf node j and the reference child node selection indication parameter stored by the first participant satisfying a preset matching condition, the data feature parameter of the business data is compared with the child node selection indication parameter of the non-leaf node j to obtain a comparison result, and a child node of the non-leaf node j for processing the business data is determined according to the comparison result, where t is a positive integer less than or equal to N, and j is an integer;
[0010] Determine a first partition result of the business data on the tth tree based on the determined child nodes of the non-leaf node j for processing the business data;
[0011] Acquire a first Boolean intersection slice held by the first participant, where the first Boolean intersection slice is used to indicate whether the business data identifier of the business data is in the identifier intersection result;
[0012] A business label corresponding to the business data is determined according to the first Boolean intersection slice and the first division result.
[0013] On the one hand, an embodiment of the present application provides a data processing device, which is deployed at a first participant, and includes an acquisition unit and a determination unit:
[0014] The acquisition unit is used to acquire business data and a first decision tree, the first decision tree runs on the first participant, the first decision tree is obtained through vertical federated learning training, and includes N trees, the N trees respectively include a plurality of leaf nodes and a plurality of non-leaf nodes, the plurality of non-leaf nodes are respectively pre-configured with a child node selection indication parameter, and the child node selection indication parameter is used to divide the child nodes of the non-leaf nodes used to process the business data;
[0015] The determining unit is used for comparing the data characteristic parameters of the business data with the child node selection indication parameters of the non-leaf node j of the t-th tree among the N trees, in response to the child node selection indication parameters of the non-leaf node j and the reference child node selection indication parameters stored by the first participant satisfying a preset matching condition, to obtain a comparison result, and determining a child node of the non-leaf node j for processing the business data according to the comparison result, where t is a positive integer less than or equal to N, and j is an integer;
[0016] The determining unit is further configured to determine a first partitioning result of the business data on the tth tree based on the determined child nodes of the non-leaf node j for processing the business data;
[0017] The acquisition unit is further used to acquire a first Boolean intersection slice held by the first participant, where the first Boolean intersection slice is used to indicate whether the business data identifier of the business data is in the identifier intersection result;
[0018] The determining unit is further configured to determine a service label corresponding to the service data according to the first Boolean intersection slice and the first division result.
[0019] In one aspect, an embodiment of the present application provides a computer device, the computer device comprising a processor and a memory:
[0020] The memory is used to store a computer program and transmit the computer program to the processor;
[0021] The processor is configured to execute the method described in any one of the preceding aspects according to instructions in the computer program.
[0022] On the one hand, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the processor executes the method described in any of the aforementioned aspects.
[0023] In one aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the method described in any of the aforementioned aspects when executed by a processor.
[0024] It can be seen from the above technical solution that when performing anonymous vertical federated reasoning, the present application changes the comparison between data feature parameters and sub-node selection indication parameters into a local operation, that is, after obtaining the business data and the first decision tree, the data feature parameters are first compared with the sub-node selection indication parameters locally. Specifically, since the first participant obtains the first decision tree by participating in the vertical federated learning training, the first participant stores the parameter subnode selection indication parameter of the first decision tree, and the first decision tree includes N trees, the N trees respectively include multiple leaf nodes and multiple non-leaf nodes, and the multiple non-leaf nodes are respectively pre-configured with subnode selection indication parameters, and the subnode selection indication parameters are used to divide the subnodes of the non-leaf nodes used to process business data. Therefore, for the non-leaf node j of the tth tree among the N trees, in response to the subnode selection indication parameter of the non-leaf node j and the reference subnode selection indication parameter satisfying the preset matching condition, that is, the subnode selection indication parameter is locally in the first participant, then the data feature parameters of the business data are compared with the subnode selection indication parameter of the non-leaf node j locally to obtain a comparison result, and the subnode of the non-leaf node j used to process the business data is determined according to the comparison result, t is a positive integer less than or equal to N, j is an integer, and then based on the determined subnode of the non-leaf node j used to process the business data, the first division result of the business data on the tth tree is determined. Only when the sub-node selection indication parameter is local, the comparison between the data feature parameter and the sub-node selection indication parameter is performed locally, thereby saving a lot of overhead caused by data exchange in the privacy comparison process and improving the efficiency of privacy comparison. Afterwards, label reasoning is performed. Specifically, the first Boolean intersection shard held by the first participant can be obtained. The first Boolean intersection shard is used to indicate whether the business data identifier of the business data is in the identifier intersection result, so as to determine the business label corresponding to the business data according to the first Boolean intersection shard and the first division result. Compared with the related art, the present application changes the comparison between the data feature parameter and the sub-node selection indication parameter into a local operation, thereby eliminating the need to exchange data with other participants, saving a lot of overhead caused by data exchange in the privacy comparison process, improving the efficiency of privacy comparison, and thus improving the efficiency of anonymous vertical federated reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technical members in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0026] Figure 1 An application scenario architecture diagram of a data processing method provided in an embodiment of the present application;
[0027] Figure 2 A flowchart of a data processing method provided in an embodiment of the present application;
[0028] Figure 3a An example diagram of obtaining a first partition result by traversing provided in an embodiment of the present application;
[0029] Figure 3b An example diagram of obtaining a first partition result in a non-traversal manner provided in an embodiment of the present application;
[0030] Figure 4 A schematic diagram of a scenario for performing hash mapping provided in an embodiment of the present application;
[0031] Figure 5 An example diagram of pipeline optimization provided in an embodiment of the present application;
[0032] Figure 6 An example diagram of the overall process of a data processing method provided in an embodiment of the present application;
[0033] Figure 7 A structural diagram of a data processing device provided in an embodiment of the present application;
[0034] Figure 8 A structural diagram of a terminal provided in an embodiment of the present application;
[0035] Fig. 9 A structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] The embodiments of the present application are described below in conjunction with the accompanying drawings.
[0037] To facilitate understanding, the terms involved in this application are first explained.
[0038] Federated Learning (FL) is a distributed machine learning technology for privacy protection. It solves the problem of how to jointly train a global model on virtual "aggregated" data while protecting data privacy when sensitive data exists in the hands of multiple independent institutions, groups, and individuals (for example, the first participant and the second participant). Federated Learning can include Vertical Federated Learning and Horizontal Federated Learning.
[0039] Among them, vertical federated learning can include anonymous vertical federated learning. Anonymous vertical federated learning refers to a method of training a vertical federated learning model without exposing the identity (ID) set after the intersection. That is, the sample alignment process is a vertical federated learning without data leakage, and the feature data will be used indiscriminately to train the model parameters. In federated practice, the vertical federated Xgboost (eXtreme Gradient Boosting) algorithm (i.e., the anonymous vertical federated Xgboost training algorithm) is the most widely used federated machine learning algorithm (i.e., the vertical federated learning model training algorithm). Among them, XGBoost is an enhanced version of the gradient boosting decision tree. XGBoost uses classification trees or regression trees as the basic classifier's boosting method, which has higher scalability and faster and higher efficiency.
[0040] Secure multi-party computation (MPC) is a technology that securely solves a global problem when the secret is distributed among multiple parties.
[0041] Secret Sharing (SS) is a method of distributing the secret to various participants. Each participant gets a part of the secret, which becomes a share. Only when more than a certain number of participants aggregate the shares can the secret be restored.
[0042] Secret sharing can include arithmetic secret sharing and Boolean secret sharing. Secret sharing can be used to generate secret shards, arithmetic secret sharing can be used to generate arithmetic shards, and Boolean secret sharing can be used to generate Boolean shards. Arithmetic secret sharing: For the integer ring Z P The number x on the and It is held by two parties, namely Similarly, Arithmetic slicing has additive homomorphism, that is, Multiplication requires the consumption of multiplication triplets. λ ) is the modulus, for example, p is equal to 2 128 (i.e. modulo 2 128 Plus), indicating that the maximum value of an integer is 2 128 -1, at this time, the addition operation refers to Z P Addition and subtraction operations on Z P Boolean secret sharing: Generally, the Boolean value x is split into two slices (i.e., Boolean slices) based on XOR. and Divided into two parties holding, that is, for the Boolean value x, In other words, if the secret sharing method is XOR, the random number obtained by the calculation party is called Boolean sharding; if the secret sharing method is arithmetic addition, the random number obtained by the calculation party is called arithmetic sharding. Among them, Boolean sharding and arithmetic sharding can be collectively referred to as secret sharding (sharding for short).
[0043] It should be understood that for the first participant (ie, P0) and the second participant (ie, P1), the embodiment of the present application can use <·> to represent the shard, the shard superscript A represents the arithmetic shard, the shard superscript B represents the Boolean shard, the shard subscript 0 represents the shard held by the first participant, and the shard subscript 1 represents the shard held by the second participant; when the shard subscript is abbreviated, it means that the calculation needs to be performed by both the first participant and the second participant. In other words, <x> A represents the arithmetic sharding of x, that is, each party has a random value, and the sum of these two random values on Zp is x; <x> B represents a Boolean slice of x, that is, each party has a random bit, and the XOR of these two random bits is x. Assume that each party has a slice of input x, and P0 holds P1 holding Note that x is a floating point number, so it needs to be converted into a slice on the integer ring by multiplying it by a magnification factor.
[0044] Multi-party secure computation operator: (1) Two-party shard addition (ADD), here we use " <x> A + <y> A " means: input <x> A and <y> A , output<x+y> A (2) Two-way piecewise multiplication (MUL), here we use " <x> A · <y> A " means: input <x> A and <y> A , output<x·y> A ; If the input is two matrix slices, it is also expressed as element-wise multiplication, that is, the slice elements corresponding to the same rows and columns of the two matrices are multiplied to obtain a matrix slice of the same size. (3) Two slices and (AND), here we use "AND( <x> B , <y> B )" means: Input <x> B and <y> B , output<x&y> B (4) XOR of two slices, here we use "XOR( <x> B , <y> B )" means: Input <x> B and <y> B , output<x⊕y> B (5) Two-way selection (Multiplexer), here we use "MUX ( <x> B , <y> A )" means: Input <x> B and <y> A , when x=0, the output is <y> A ; Otherwise, output the slice corresponding to 0 <0> A (6) Two-way piecewise Sigmoid function, <y> A =Sigmoid( <x> A ), Sigmoid can be converted into two-way piecewise addition and multiplication through polynomial fitting. (7) Two-way comparison function <z> B =greater( <x> A ,y), when x>=y, z=1, otherwise z=0.
[0045] In anonymous vertical federated reasoning, each participant has its own business data. Taking the first participant P0 and the second participant P1 as examples, P0 has business data X0 and the corresponding user identifier ID0, and P1 has business data X1 and the corresponding user identifier ID1. They want to jointly calculate the reasoning y[i]=f(X0[i]||X1[j]), where ID0[i]=ID1[j]∈ID0∩ID1, and f(·) is the GBDT reasoning process. In this process, both parties will not disclose their business data and ID to each other, and can only know the final label and the public parameters of f(·); the parameters include multiple trees with a depth of D. The leaf node weight w of f(·) obtained by training is the sharding state, which is supported by both parties. It is worth noting that the child node selection indication parameter is only known by the participant who owns the feature and the other party does not know it. For a business data, taking the decision tree as a binary tree as an example, the specific process of f(·) is as follows:
[0046] For each non-leaf node of each tree, compare the size of the corresponding data feature parameter with the child node selection indicator parameter. If it is less than the child node selection indicator parameter, put the sample in the right subtree, otherwise put it in the left subtree, and repeat this process until the sample reaches the leaf node. Then, add the weights of the leaf nodes where the business data of each tree is located to calculate the final predicted label.
[0047] In the above-mentioned anonymous vertical federated reasoning process, since the sub-node selection indication parameter is only known by the participant who has the feature and the other party does not know it, when a participant needs to perform a privacy comparison operation between the data feature parameters and the sub-node selection indication parameter, it needs to exchange data with other participants, which results in a large overhead of the privacy comparison operation, and thus leads to the low efficiency of the above-mentioned anonymous vertical federated reasoning.
[0048] In order to solve the above technical problems, an embodiment of the present application provides a data processing method. When performing anonymous vertical federated reasoning, the method converts the comparison between data feature parameters and sub-node selection indication parameters into local operations, that is, only when the sub-node selection indication parameters are local to the first participant, the comparison between the data feature parameters and the sub-node selection indication parameters is performed locally, thereby eliminating the need to exchange data with other participants, saving a large amount of overhead caused by data exchange in the privacy comparison process, improving the efficiency of privacy comparison, and thereby improving the efficiency of anonymous vertical federated reasoning.
[0049] It should be noted that the data processing method provided in the embodiment of the present application can be applied to cross-company, cross-institutional scenarios to protect user IDs while hiding vertical federated reasoning. It can be used in business scenarios such as financial risk control, medical disease prediction, unmanned driving (for example, improving driving experience), and multimedia data recommendation (for example, video recommendation). For example, in the medical scenario, rare diseases are a common problem in medical research. At present, there are practical difficulties such as samples being scattered in different hospitals, which greatly hinders the diagnosis and treatment work. Through vertical federated learning, samples from different hospitals can be integrated to achieve the diagnosis and treatment of rare diseases. For another example, in the financial scenario, the credit risk control of small and micro enterprises uses federated transfer learning, and can use the existing models of financial institutions in previous applications (for example, the credit model of medium and large enterprises or the marketing model of small and micro enterprises) for transfer learning to improve the effect of the application. For another example, in the unmanned driving scenario, the interactive learning between the vehicle and the system environment can be assisted by other information in the city (for example, city cameras, traffic lights, and future smart roads). Through the vertical federation method, information from different sources can be better integrated under privacy protection to improve the unmanned driving experience. For example, in the multimedia data recommendation scenario, data from different platforms can learn interactively, thereby improving the accuracy of multimedia data recommendations.
[0050] The data processing method provided in the embodiment of the present application can be executed by the first participant, which is a computer device of a party participating in vertical federated learning. The computer device can be, for example, a server or a terminal. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Terminals include but are not limited to vehicle-mounted terminals, smart phones, tablet computers, computers, intelligent voice interaction devices, smart home appliances, aircraft, etc.
[0051] like Figure 1 As shown, Figure 1 The application scenario architecture diagram of a data processing method is shown. In this application scenario, the first participant is a server 100 as an example for introduction.
[0052] The anonymous federation algorithm may include an anonymous longitudinal federation training process and an anonymous longitudinal federation reasoning process. In the anonymous longitudinal federation training process, multiple participants (i.e., the abbreviation of the computer device corresponding to the object (e.g., company, institution, enterprise)) may obtain their own models based on anonymous federation algorithm training, such as a tree-like decision tree model. The tree-like decision tree model is a neural network model with a tree structure, which may also be called a decision tree, so that each participant maintains its own decision tree. Since multiple participants contribute their own business data to train the decision tree during anonymous longitudinal federation training, the key step in constructing the decision tree is continuous splitting. The so-called splitting is to construct different branches at a certain node according to different divisions of a certain data feature parameter until the last layer of branches, thereby obtaining a decision tree. The decision tree obtained in this way has a sub-node selection indication parameter, which determines how to branch. In the anonymous longitudinal federation training process, each participant only has a sub-node selection indication parameter obtained by training based on its own unique business data. Among them, the sub-node selection indication parameter may refer to the splitting boundary determined for the data feature parameter of a certain dimension when constructing the decision tree, such as age 30 years old, height 175 cm, etc. The subnode selection indicator parameter is used to divide the subnodes of non-leaf nodes used to process business data. The subnode selection indicator parameter can also be called a split point. The split point refers to the critical point for dividing a data set into different subsets based on feature selection. In the process of building a decision tree, the goal of split point selection is to find a critical point that can make the subsets as different as possible, so as to clarify the critical point for dividing different molecules. In this way, when using a decision tree for data reasoning, for a business data, it can be determined which branch the business data is divided into based on the data feature parameters and split point of the business data.
[0053] It should be noted that the computer device for obtaining the decision tree based on the anonymous federation algorithm training may be the same as or different from the computer device for executing the data processing method provided in the embodiment of the present application, and the embodiment of the present application does not limit this.
[0054] For example, multiple participants include a first participant and a second participant, and the first participant and the second participant belong to computer devices of different objects respectively. In this example, the first participant and the second participant respectively contribute their business data to train their respective decision trees, and the first participant executes the data processing method provided by the embodiment of the present application. The decision tree maintained by the first participant can be referred to as the first decision tree. For the first participant, some of the sub-node selection indication parameters of all sub-node selection indication parameters of the first decision tree are determined based on the business data of the first participant, and some sub-node selection indication parameters are determined based on the business data of the second participant. The first participant stores sub-node selection indication parameters determined based on the business data of the first participant, and these sub-node selection indication parameters can be referred to as reference sub-node selection indication parameters, and the remaining sub-node selection indication parameters determined based on the business data of the second participant are stored on the second participant.
[0055] In the anonymous vertical federated reasoning process, each participant can use the decision tree maintained by itself to perform label reasoning to obtain the business label of the business data.
[0056] This application mainly improves the anonymous vertical federated reasoning process. When performing anonymous vertical federated reasoning, the comparison between the data feature parameters and the sub-node selection indication parameters is changed to a local operation. Taking the first participant among multiple participants as an example, after obtaining the business data and the first decision tree of the first participant, the server 100 as the first participant first compares the data feature parameters with the sub-node selection indication parameters locally.
[0057] Specifically, since the first decision tree includes N trees, each of the N trees includes leaf nodes and non-leaf nodes, the leaf nodes are the last layer nodes on each tree, and the non-leaf nodes are other nodes on each tree except the leaf nodes. Each non-leaf node has a corresponding child node, and the child node of the non-leaf node is a non-leaf node or a leaf node. When the child node of the non-leaf node is the last layer node of the tree, the child node of the non-leaf node is a leaf node, otherwise, the child node of the non-leaf node is still a non-leaf node. Therefore, the server 100 determines whether the child node selection indication parameter of the non-leaf node j and the reference child node selection indication parameter meet the preset matching condition for the non-leaf node j of the t-th tree among the N trees, that is, determines whether the child node selection indication parameter of the non-leaf node j is stored locally on the first participant. Wherein, t is a positive integer less than or equal to N, and j is an integer.
[0058] In response to the subnode selection indication parameter of non-leaf node j and the reference subnode selection indication parameter satisfying the preset matching condition, i.e., being located locally at the first participant, the server 100 compares the data characteristic parameter of the business data with the subnode selection indication parameter of non-leaf node j locally to obtain a comparison result, and determines the subnode of non-leaf node j for processing the business data according to the comparison result. Figure 1 In, for Figure 1 The non-leaf node j shown in , takes the case where the child node for processing business data is determined to be the left child node of the non-leaf node j according to the comparison result. Then, based on the child node of the non-leaf node j determined to process business data, the first partition result of the business data in the tth tree is determined. That is to say, only when the child node selection indication parameter is local, the server 100 performs the comparison between the data feature parameter and the child node selection indication parameter locally, thereby saving a lot of overhead caused by data exchange in the privacy comparison process and improving the efficiency of privacy comparison.
[0059] Afterwards, the server 100 performs label inference. Specifically, the server 100 can obtain the first Boolean intersection slice held by the first participant, and the first Boolean intersection slice is used to indicate whether the business data identifier of the business data is in the identifier intersection result, so as to determine the business label corresponding to the business data according to the first Boolean intersection slice and the first division result.
[0060] It should be noted that in the specific implementation of this application, user information and other related data may be involved in the entire process. When the above embodiments of this application are applied to specific products or technologies, the user's separate consent or separate permission is required, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.
[0061] Next, the data processing method provided by the embodiment of the present application will be introduced with reference to the accompanying drawings, taking the first participant executing the data processing method as an example. Figure 2 , Figure 2 A flow chart of a data processing method is shown, and the method may include S201-S205, which are specifically as follows:
[0062] S201. Obtain business data and a first decision tree, where the first decision tree runs on the first participant. The first decision tree is obtained through vertical federated learning training and includes N trees, where the N trees respectively include multiple leaf nodes and multiple non-leaf nodes, where the multiple non-leaf nodes are respectively pre-configured with child node selection indication parameters, and the child node selection indication parameters are used to divide the child nodes of the non-leaf nodes used to process business data.
[0063] Multiple participants in anonymous vertical federated learning can obtain their own models, such as decision trees, based on anonymous federated algorithm training, so that each participant maintains its own decision tree. In the anonymous vertical federated reasoning process, each participant can use the decision tree maintained by itself to perform label reasoning, thereby obtaining the business label of the business data. Among them, the multiple participants can be two participants, or more than two participants, and the embodiments of the present application are not limited to this. The first participant can be any participant among the multiple participants. It can be understood that the multiple participants can be divided into a receiver and a sender. The sender is the party holding the privacy set, and the receiver is the party that hopes to obtain the intersection result of the identifier with the sender. Therefore, in one possible implementation method, the first participant can be a receiver, denoted as P0, and the sender is called the second participant, denoted as P1.
[0064] The first participant obtains a first decision tree by participating in the vertical federated learning training, and performs label reasoning by running the first decision tree. The first participant stores a reference subnode selection indication parameter of the first decision tree, and the reference subnode selection indication parameter is a subnode selection indication parameter determined on the first decision tree based on the business data of the first participant.
[0065] The embodiment of the present application does not limit the decision tree maintained. The decision tree may be GBDT, which is a decision tree algorithm that improves the prediction of results by optimizing through differentiable loss in a staged manner. The most commonly used one is extreme gradient boosting (eXtreme Gradient Boosting, XGboost).
[0066] The business data acquired by the embodiment of the present application is the business data that needs to be inferred from the business label, and the business label is used to indicate the category of the business data. In different business scenarios, the business data is different, and the obtained business labels also represent different meanings. For example, the business data can be symptom information in a medical scenario, then the first decision tree can be used to perform business processing on the symptom information, and the result of the business processing (i.e., the business label) can be the auxiliary suggestion information corresponding to the symptom information, and the auxiliary suggestion information can be used to provide to the doctor so that the doctor can diagnose and treat the disease. For another example, the business data can be vehicle driving data in an unmanned driving scenario, the first decision tree can be used to perform business processing on the vehicle driving data (e.g., vehicle driving state analysis), and the result of the business processing (i.e., the business label) can be the driving state corresponding to the vehicle driving data, thereby realizing unmanned driving of the vehicle. For another example, the business data can be operation data in a multimedia data recommendation scenario, the first decision tree can be used to perform business processing on the operation data (e.g., classification label identification), and the result of the business processing (i.e., the business label) can be the classification label corresponding to the operation data, thereby realizing multimedia data recommendation for the classification label.
[0067] S202. For the non-leaf node j of the tth tree among N trees, in response to the child node selection indication parameter of the non-leaf node j and the reference child node selection indication parameter satisfying a preset matching condition, the data feature parameters of the business data are compared with the child node selection indication parameter of the non-leaf node j to obtain a comparison result, and the child node of the non-leaf node j for processing the business data is determined based on the comparison result.
[0068] Among them, the tth tree can be any tree among the N trees, or a tree among the N trees that meets the first preset filtering condition, and the non-leaf node j can be any non-leaf node on the tth tree, or a non-leaf node that meets the second preset filtering condition. The embodiment of the present application does not limit this.
[0069] If the subnode selection indication parameter of the non-leaf node j and the reference subnode selection indication parameter meet the preset matching condition, the data characteristic parameter of the business data is compared with the subnode selection indication parameter of the non-leaf node j to obtain a comparison result, and the subnode of the non-leaf node j used to process the business data is determined according to the comparison result. However, if it is determined that the subnode selection indication parameter of the non-leaf node j and the reference subnode selection indication parameter do not meet the preset matching condition, that is, the subnode selection indication parameter is not local to the first participant, there is no need to compare the subnode selection indication parameter with the data characteristic parameter of the business data, but directly consider that each subnode of the non-leaf node j is used to process the business data, and directly determine that each subnode of the non-leaf node j is used to process the business data. In other words, even if the subnode selection indication parameter is not local to the first participant, there is no need to exchange data, but directly consider that each subnode of the non-leaf node j is used to process the business data, thereby saving a lot of overhead caused by data exchange in the privacy comparison process and improving the efficiency of privacy comparison.
[0070] Among them, the preset matching condition can be a condition for determining that the child node selection indication parameter of the non-leaf node j is a reference child node selection indication parameter stored by the first participant. For example, when the first participant stores the reference child node selection indication parameter, it can store the correspondence between the reference child node selection indication parameter and the identifier of the non-leaf node j to which it belongs. For the non-leaf node j, the identifier of the non-leaf node j can be searched in the corresponding relationship. If the identifier of the non-leaf node j exists in the corresponding relationship, it is determined that the child node selection indication parameter of the non-leaf node j and the reference child node selection indication parameter of the non-leaf node j meet the preset matching condition, thereby determining that the child node selection indication parameter of the non-leaf node j is stored locally on the first participant.
[0071] It can be understood that the sub-node selection indication parameter and the data characteristic parameter can be respectively reflected by the corresponding parameter values. Therefore, in one possible implementation method, comparing the data characteristic parameter of the business data with the sub-node selection indication parameter of the non-leaf node j can refer to comparing the parameter value of the data characteristic parameter with the parameter value of the sub-node selection indication parameter, thereby obtaining a comparison result reflecting the size relationship between the two parameter values.
[0072] In the embodiment of the present application, the tree included in the first decision tree may be a binary tree, a ternary tree, etc. For different trees, a non-leaf node on a tree includes different numbers of child nodes. For example, in a binary tree, the number of child nodes included in a non-leaf node is 2, that is, a non-leaf node has 2 child nodes; for another example, in a ternary tree, the number of child nodes included in a non-leaf node is 3, that is, a non-leaf node has 3 child nodes; and so on to an M-ary tree, the number of child nodes included in a non-leaf node is M, that is, a non-leaf node has M child nodes, and M can be any positive integer.
[0073] For the M child nodes of a non-leaf node (for example, non-leaf node j), it is necessary to determine which child node of the M child nodes the business data is divided into based on the comparison result. When constructing the first decision tree, it is already known how each child node is divided, that is, each child node has a corresponding comparison result. Therefore, in one possible implementation method, the method of determining the child node of non-leaf node j used to process business data based on the comparison result can be that if the comparison result is the comparison result corresponding to the kth child node among the M child nodes, it is determined that the business data is processed by the kth child node of non-leaf node j, where k is a positive integer less than or equal to M.
[0074] Among them, among the comparison results corresponding to the M subnodes, the comparison results corresponding to the M subnodes can be different, or some of the comparison results can be the same (for example, the comparison result corresponding to the first subnode among the M subnodes is the first comparison result, and the comparison results corresponding to the remaining subnodes are all the second comparison results). When the comparison results corresponding to the M subnodes are different, if the comparison result is the comparison result corresponding to which subnode among the M subnodes, it is determined to which subnode of the non-leaf node j the business data belongs. When the comparison results corresponding to the M subnodes have some of the same comparison results, for example, the comparison result corresponding to the first subnode among the M subnodes is the first comparison result, and the comparison results corresponding to the remaining subnodes are all the second comparison results, if the comparison result is the comparison result corresponding to the first subnode, it is determined that the first subnode of the non-leaf node j processes the business data, and if the comparison result is the comparison result corresponding to the remaining subnodes, it is determined that a certain subnode among the remaining subnodes processes the business data, and the certain subnode can be any subnode among the remaining subnodes, or it can be the subnode closest to the first subnode, and so on. The embodiment of the present application matches the comparison result with the division strategy corresponding to different branches determined in the first decision tree construction process (i.e., the comparison results corresponding to different sub-nodes), thereby determining the sub-node used to process the business data, and can more accurately determine the sub-node used to process the business data.
[0075] It is understandable that in order to clarify the sub-nodes used to process business data, it is necessary to record which sub-node of the non-leaf node j the business data is divided to for subsequent use. The embodiment of the present application does not limit the recording form of the sub-nodes used to process business data. In one possible implementation, a division indication mark can be set for each sub-node of the non-leaf node j, and the division indication mark reflects which sub-node of the non-leaf node j is used to process the business data.
[0076] The embodiment of the present application does not limit the representation form of the partition indicator mark, and the representation form of the partition indicator mark can be a symbol, a number, etc. The partition indicator mark of the subnode used to process business data is different from the partition indicator mark of the subnode not used to process business data. For example, the partition indicator mark of the subnode used to process business data is the first indicator mark, and the partition indicator mark of the subnode not used to process business data is the second indicator mark, so that it can be determined whether the subnode is used to process business data through the partition indicator mark. Taking the representation form of the partition indicator mark as a number as an example, the first indicator mark can be 1 and the second indicator mark can be 0.
[0077] At this time, if the comparison result is the comparison result corresponding to the kth child node among the M child nodes, the method for determining that the kth child node of the non-leaf node j processes the business data can be that if the comparison result is the comparison result corresponding to the kth child node among the M child nodes, the partition indication mark of the kth child node is determined to be the first indication mark, and the partition indication marks of the remaining child nodes are determined to be the second indication mark, the first indication mark is used to indicate that the corresponding child node is used to process business data, the second indication mark is used to indicate that the corresponding child node is not used to process business data, and the remaining child nodes are the child nodes other than the kth child node among the M child nodes.
[0078] The embodiment of the present application determines a partition indication mark for each child node of the non-leaf node j, so that the child node used to process the business data can be directly distinguished based on the partition indication mark, which is simpler, more convenient and intuitive.
[0079] In an embodiment of the present application, the tree included in the first decision tree may be a binary tree, a ternary tree, etc. Taking the case where the tree included in the first decision tree is a binary tree, the M child nodes included in the non-leaf node j are the first child node and the second child node. If the comparison result is the comparison result corresponding to the kth child node among the M child nodes, the method for determining that the kth child node of the non-leaf node j processes the business data may be that if the comparison result is that the data characteristic parameter of the business data is less than the child node selection indication parameter of the non-leaf node j, the first child node is determined to be used to process the business data, and the first child node is the kth child node of the non-leaf node j, and k=1; if the comparison result is that the data characteristic parameter of the business data is greater than the child node selection indication parameter of the non-leaf node j, the second child node is determined to be used to process the business data, and the second child node is the kth child node of the non-leaf node j, and k=2.
[0080] For a binary tree, each non-leaf node has two child nodes, located on the left and right of the non-leaf node, respectively. Based on this, in a possible implementation, the first child node can be a child node located on the left, i.e., a left child node, and the second child node can be a child node located on the right, i.e., a right child node. Each tree in the first decision tree includes a root node, an internal node, and a leaf node, wherein the root node represents the entire data set, initially without any division; the internal node represents the splitting boundary of a specific data feature parameter, and the entire data set is divided into multiple child nodes according to the splitting boundary; the leaf node represents the final classification or regression output. Both the root node and the internal node can be called non-leaf nodes. If the number of non-leaf nodes is M, the root node can be non-leaf node 0, and the other internal nodes can be recorded as non-leaf node 1, non-leaf node 2, non-leaf node 3, ..., non-leaf node M-1 in sequence.
[0081] Therefore, for the service data i and the non-leaf node j of the t-th tree, when the above-mentioned partition indicator is represented by a number, the first indicator is 1, and the second indicator is 0, P l Set the root node's partition indicator c l [i, t, 0] are all 1, that is, the business data i is on the root node, l can be 0 or 1. When the child node selection indication parameter of the non-leaf node j is on the first participant's side, the data characteristic parameter of the business data is calculated locally. When the data characteristic parameter is less than the child node selection indication parameter, the partition indication mark of the left child node is determined to be c. l [i, t, j*2+1] = 1, where j*2+1 represents the sequence number of the current left child node as a non-leaf node, and the partition indicator mark of the right child node is determined to be c l [i, t, j*2+2] = 0, where j*2+2 represents the sequence number of the current left child node as a non-leaf node; otherwise, the division indicator mark of the left child node is determined to be c l [i, t, j*2+1] = 0, determine the partition indicator of the right child node as c l [i,t,j*2+2]=1.
[0082] In the embodiment of the present application, when the tth tree is a binary tree, different partition indication marks are determined for the child nodes of the non-leaf node j based on different comparison results, so that the child node used to process the business data can be directly distinguished based on the partition indication mark, making the management of the business data simpler, more convenient and intuitive.
[0083] Correspondingly, when the child node selection indication parameter of the non-leaf node j is not in the first participant, determining which child node of the non-leaf node j is used to process the business data can also be represented by a partition indication mark. At this time, the method of determining that each child node of the non-leaf node j is used to process the business data can be to determine the partition indication mark of each child node of the non-leaf node j as a first indication mark, and the first indication mark is used to indicate that the corresponding child node is used to process the business data.
[0084] That is to say, when the child node selection indication parameter of the non-leaf node j is not local to the first participant, it is directly considered that each child node of the non-leaf node j is used to process business data, and it is directly determined that the partition indication mark of each child node is used to process business data, so the partition indication mark of each child node of the non-leaf node j is set to the first indication mark.
[0085] Taking the example that the tree included in the first decision tree is a binary tree, the first child node is the left child node, and the second child node is the right child node, when the representation form of the above-mentioned partition indicator mark is a number, the first indicator mark is 1, and the second indicator mark is 0, when the child node selection indicator parameter of the non-leaf node i is not on the first participant's side, there is no need to compare the child node selection indicator parameter with the data feature parameter of the business data, but it is directly considered that each child node of the non-leaf node j has business data, and the partition indicator mark of the left child node is directly determined to be c. l [i, t, j*2+1] = 1, determine the partition indicator of the right child node as c l [i,t,j*2+2]=1.
[0086] The embodiment of the present application indicates which child node of the non-leaf node j processes the business data by determining a partition indication mark for each child node of the non-leaf node j, thereby making the management of the business data simpler, more convenient and more intuitive. At the same time, when the child node selection indication parameter of the non-leaf node j is not in the first participant, the same partition indication mark indicating that the business data exists in the child node is directly set for each child node, thereby improving the efficiency of privacy comparison while ensuring the accuracy of the partition as much as possible.
[0087] S203: Determine a first partition result of the business data on the tth tree based on the determined child nodes of the non-leaf node j for processing the business data.
[0088] The first partition result is used to indicate the node for processing business data when the decision stops. When the child node of the non-leaf node j for processing business data is represented by a partition indication mark, the first partition result can also be represented by a partition indication mark, which is not repeated here.
[0089] It should be noted that the implementation of S203 is different depending on the decision mode of the first decision tree. If the decision mode of the first decision tree is a traversal mode, then after obtaining the business data, the N trees in the first decision tree can be traversed based on the business data. Traversal can refer to obtaining the child node selection indication parameters of each non-leaf node of each tree in turn, thereby determining which child node of the non-leaf node processes the business data based on the child node selection indication parameters, until the child node of the non-leaf node is a leaf node.
[0090] At this time, the t-th tree is the tree traversed at the current moment in the process of traversing N trees based on business data, and the non-leaf node j is the non-leaf node traversed at the current moment in the process of traversing all non-leaf nodes of the t-th tree based on business data. For each traversed tree, a similar traversal is performed on the traversed tree. Taking the traversed tree as the t-th tree among N trees as an example, t=1, 2, 3, ..., N, traverse all non-leaf nodes of the t-th tree, and for the traversed non-leaf node j, determine whether the child node selection indication parameter of the non-leaf node j and the reference child node selection indication parameter meet the preset matching condition, that is, whether they are located locally on the first participant, and the traversal is a local calculation. In response to the subnode selection indication parameter of the non-leaf node j and the reference subnode selection indication parameter satisfying the preset matching condition, that is, the subnode selection indication parameter is locally located at the first participant, the data feature parameter of the business data is compared with the subnode selection indication parameter of the non-leaf node j locally to obtain a comparison result, and the subnode of the non-leaf node j used to process the business data is determined according to the comparison result, until the subnode of the non-leaf node j is a leaf node of the t-th tree, thereby completing the traversal of the t-th tree, and then obtaining the first partitioning result of the business data in the t-th tree. At this time, the implementation method of S203 can be that if the subnode of the non-leaf node j determined to be used to process the business data is a leaf node of the t-th tree, then the first partitioning result is obtained, and the first partitioning result is used to indicate the leaf node on the t-th tree used to process the business data.
[0091] See also Figure 3a As shown, in order to simplify the drawings, Figure 3a Only part of the tree in the first decision tree and part of the nodes of the tree are embodied as an example, and do not constitute a limitation on the first decision tree. Starting from the non-leaf node 1 of the first tree, determine whether the sub-node selection indication parameter of the non-leaf node 1 and the reference sub-node selection indication parameter meet the preset matching conditions. If so, compare the data feature parameters of the business data with the sub-node selection indication parameter of the non-leaf node 1 locally to obtain a comparison result, and determine the sub-node of the non-leaf node 1 for processing the business data based on the comparison result. For example, Figure 3a In the example, the child nodes of non-leaf node 1 include non-leaf node 2 and non-leaf node 3. The comparison result determines that the business data is processed by non-leaf node 2. Next, traverse the next non-leaf node, such as non-leaf node 2, and determine whether the child node selection indication parameter of non-leaf node 2 and the reference child node selection indication parameter meet the preset matching condition. If so, compare the data feature parameters of the business data with the child node selection indication parameter of non-leaf node 2 locally to obtain the comparison result, and determine the child node of non-leaf node 2 for processing the business data according to the comparison result. Figure 3a The child nodes of non-leaf node 2 are not shown. Assume that it is determined that the rightmost child node of non-leaf node 2 is used to process the business data. Then continue to traverse. If the non-leaf node j is traversed, it is determined whether the child node selection indication parameter of non-leaf node j and the reference child node selection indication parameter meet the preset matching condition. If so, the data feature parameters of the business data are compared with the child node selection indication parameter of non-leaf node j locally to obtain the comparison result, and the child node of non-leaf node j used to process the business data is determined according to the comparison result. For example, in Figure 3a In the figure, the child nodes of non-leaf node j include leaf node 1 and leaf node 2. By comparing the results, it is determined that leaf node 2 processes the business data. At this time, the last layer of the first tree has been reached, so the traversal can be stopped to obtain the first partition result of the first tree.
[0092] Then traverse the second tree. The traversal method of the second tree is similar to that of the first tree, still starting from non-leaf node 1, until the child node of the non-leaf node j used to process business data is determined to be a leaf node, and the first partition result of the second tree is obtained. Similarly, traverse the Nth tree and obtain the first partition result of the Nth tree.
[0093] By traversing the first decision tree, the first partitioning result can be obtained more accurately, and then the first partitioning results of the N trees can be combined to determine the business label, thereby improving the accuracy of subsequent business label determination. If the decision method of the first decision tree is a non-traversal method, only the first partitioning result of the tth tree can be used to determine the business label, or the first partitioning results of some of the N trees can be used to determine the business label. At this time, the implementation method of S203 can be that if the child node of the non-leaf node j determined to be used to process business data is a non-leaf node, and the child node of the non-leaf node j meets the preset stop condition, the first partitioning result is obtained, and the first partitioning result is used to indicate the non-leaf node on the tth tree used to process business data.
[0094] See also Figure 3b As shown, in order to simplify the drawings, Figure 3b Only part of the trees in the first decision tree and part of the nodes of the trees are reflected as an example, and do not constitute a limitation on the first decision tree. Assuming that only the first division result of the tth tree is used to determine the business label, for the tth tree, some nodes can be used to determine the first division result. For example, for non-leaf node j, it can be determined whether the child node selection indication parameter of non-leaf node j and the reference child node selection indication parameter meet the preset matching conditions. If so, the data feature parameters of the business data are compared with the child node selection indication parameters of non-leaf node j locally to obtain a comparison result, and the child nodes of non-leaf node j for processing business data are determined according to the comparison result. For example, in Figure 3b In the example, the child nodes of non-leaf node j include left child nodes and right child nodes. The comparison result determines that the business data is processed by the left child node. Then, it is determined whether the left child node of non-leaf node j meets the preset stop condition. If so, the left child node of non-leaf node j on the tth tree can be directly used as the first partition result, so that the business label can be determined later using only the output result associated with the left child node of non-leaf node j on the tth tree. The preset stop condition can be that the child node of non-leaf node j is a specific node, or the output result associated with the child node of non-leaf node j reaches a specific threshold, etc.
[0095] The embodiment of the present application determines the first division result by using part or a single tree, or even part of the nodes on the tree, which can simplify the processing process, improve the processing efficiency, and further improve the processing efficiency of service label determination.
[0096] S204. Obtain a first Boolean intersection slice held by the first participant, where the first Boolean intersection slice is used to indicate whether the business data identifier of the business data is in the identifier intersection result.
[0097] It is understandable that Circuit-PSI came into being when there was a risk of label leakage in the PSI algorithm. PSI allows participants to obtain the intersection of two parties' data through cryptography without exposing data elements outside the intersection. Circuit-PSI means that the participants input their respective user IDs, and ultimately both parties can only obtain the shard information about the result of the intersection of the IDs, that is, whether the user ID of the sender of Circuit-PSI is in the Boolean shard value of the result of the intersection of the IDs. Neither the result of the intersection of the IDs nor the non-intersection data can be obtained.
[0098] The embodiment of the present application is an improvement based on the anonymous vertical federated reasoning implemented based on Circuit-PSI. Therefore, the embodiment of the present application also needs to obtain the first Boolean intersection fragment held by the first participant in order to use the first Boolean intersection fragment for subsequent label reasoning to obtain the business label of the business data.
[0099] It should be understood that the specific process of obtaining the first Boolean intersection shard held by the first participating party can be described as follows: The first participating party performs Cuckoo Hash mapping on the first business data identifier of the first participating party to obtain the first hash table corresponding to the first business data identifier. Further, the first participating party can obtain the first Boolean intersection shard associated with the first hash table and the second hash table. Among them, the first hash table and the second hash table are used to perform hash table matching through a pseudo-random function to generate the first Boolean intersection shard and the second Boolean intersection shard held by the second participating party. The first Boolean intersection shard and the second Boolean intersection shard are shards identifying the intersection result, and the second hash table is obtained by the second participating party participating in vertical federated learning performing hash mapping on the second business data identifier of the second participating party.
[0100] Among them, Cuckoo Hash mapping refers to a method of mapping m elements (i.e., business data identifiers) to n positions through k hashes (i.e., k hash functions), and it is required that different elements are mapped to different positions. Simply put, if the i-th element has been mapped to the s-th position under the action of the t-th (t < k) hash, when the j-th (j > i) element is also mapped to the s-th position, a conflict occurs at the s-th position at this time. Then, the j-th element is mapped to the s-th position, and the i-th element is remapped to a new position under the action of the t + 1-th hash. If a conflict occurs at the position, this process is repeated.
[0101] In the embodiments of the present application, the hash mapping can be simple hash mapping. Simple Hash mapping (referred to as hash mapping for short) refers to a method of mapping m elements (i.e., business data identifiers) to n positions through k hashes (i.e., k hash functions). At this time, positions can conflict, that is, different elements can be mapped to the same position. Through simple hash mapping, different data can be mapped to the same position, and the same data can be mapped to at least two positions based on at least two hash functions, so that no additional processing (i.e., no need to resolve hash conflicts) is required when hash conflicts occur, improving the efficiency of hash mapping. It should be understood that the embodiments of the present application do not limit the number of hash functions. For ease of understanding, the embodiments of the present application take the number of hash functions as 2 as an example for illustration. Two hash functions mean that each business data identifier can be mapped to 2 positions.
[0102] Wherein, the hash table dimensions of the first hash table and the second hash table are the same, the first hash table includes a first hash mapping bucket, and the second hash table includes a second hash mapping bucket with the same hash table dimension as the first hash mapping bucket. If the first hash mapping bucket includes a first business data identifier, the first random number corresponding to the first business data identifier in the first hash mapping bucket is obtained by randomly processing the first business data identifier in the first hash mapping bucket based on a pseudo-random function. If the second hash mapping bucket includes a second business data identifier, the second random number corresponding to the second business data identifier in the second hash mapping bucket is obtained by randomly processing the second business data identifier in the second hash mapping bucket based on a pseudo-random function. The first random number and the second random number are used to perform random number matching to generate a first Boolean intersection parameter corresponding to the first hash mapping bucket and a second Boolean intersection parameter corresponding to the second hash mapping bucket; the first Boolean intersection parameter and the second Boolean intersection parameter are merged to obtain a merged Boolean intersection parameter; the merged Boolean intersection parameter is used to indicate the Boolean matching result of the first hash mapping bucket and the second hash mapping bucket (that is, the first Boolean intersection parameter and the second Boolean intersection parameter can be partial parameters indicating whether the first business data identifier of the business data is in the identifier intersection result); the first Boolean intersection parameter corresponding to each hash table dimension in the first hash table is used to form a first Boolean intersection shard corresponding to the first business data identifier, and the second Boolean intersection parameter corresponding to each hash table dimension in the second hash table is used to form a second Boolean intersection shard corresponding to the first business data identifier.
[0103] It should be understood that the pseudo-random function used in the embodiments of the present application may be an oblivious programmable pseudo-random function (OPPRF) or an oblivious pseudo-random function (OPRF). OPPRF may be an OPPRF protocol, which indicates that one party inputs a set X = {(x1, y1), (x2, y2), ..., (xn, yn)}, and the other party inputs x. After executing the protocol, the other party obtains the value y corresponding to x, satisfying that when x = xi, y = yi; otherwise y is a random number.
[0104] Among them, if the first random number and the second random number are the same, the embodiment of the present application can determine the matching success parameter (for example, the matching success parameter can be 1) as the merged Boolean intersection parameter corresponding to the first hash mapping bucket and the second hash mapping bucket. If the first random number and the second random number are different, the embodiment of the present application can determine the matching failure parameter (for example, the matching failure parameter can be 0) as the merged Boolean intersection parameter corresponding to the first hash mapping bucket and the second hash mapping bucket. The first hash mapping bucket and the second hash mapping bucket can both be called buckets (bins).
[0105] For easier understanding, see Figure 4 , Figure 4 Schematic diagram of a scenario for performing hash mapping provided by an embodiment of the present application. Figure 4 The figure shows the implementation of business data identification anonymization based on Circuit-PSI. The protocol includes a first participant and a second participant. Both the first participant and the second participant can be senders. When the business data identification is in the identification intersection result, the sender has a hash mapping bucket containing the business data identification, and the corresponding hash mapping bucket of the receiver also contains the business data identification.
[0106] like Figure 4 As shown, the first participant can perform circuit privacy set intersection as a receiver, at which time, the second participant inserts its own data (i.e., the second business data identifier 403) into the ordinary hash table (i.e., the second hash table 404), and the first participant fills its own data (i.e., the first business data identifier 401) into the cuckoo hash table (i.e., the first hash table 402). Similarly, the second participant can perform circuit privacy set intersection as a receiver, at which time, the first participant inserts its own data (i.e., the first business data identifier 401) into the ordinary hash table, and the second participant fills its own data (i.e., the second business data identifier 403) into the cuckoo hash table.
[0107] like Figure 4 As shown, for ease of understanding, the number of first business data identifiers is 4 as an example for explanation, for example, the first business data identifier may include 2, 3, 7, and 10; for ease of understanding, the number of second business data identifiers is 5 as an example for explanation, for example, the second business data identifier may include 2, 4, 5, 7, and 9.
[0108] like Figure 4 As shown, the hash table dimensions of the first hash table 402 and the second hash table 404 are the same, the first hash table 402 may include a first hash map bucket, and the second hash table 404 may include a second hash map bucket. For ease of understanding, the embodiment of the present application takes the hash table dimensions of the first hash table 402 and the second hash table 404 as 6 dimensions as an example for explanation. Among them, the embodiment of the present application can fill in "*" in the hash map bucket (bin) that is not filled with data (i.e., business data identifier), "*" represents a random number, refers to an empty business data identifier, and the data feature parameters and business tags corresponding to "*" can be 0.
[0109] like Figure 4 As shown, for example, the first hash mapping bucket may be the second hash mapping bucket in the first hash table 402, the second hash mapping bucket may be the second hash mapping bucket in the second hash table 404, the first hash mapping bucket may include 7, and the second hash mapping bucket may include 7 and "*". At this time, the first random number corresponding to 7 in the first hash mapping bucket may be 17, the second random number corresponding to "*" in the second hash mapping bucket may be 14, and the second random number corresponding to 7 may be 17. When the first random number and the second random number are matched, the first random number 17 is equal to the second random number 17. In the embodiment of the present application, the matching success parameter may be determined as the merged Boolean intersection parameter corresponding to the first hash mapping bucket and the second hash mapping bucket, that is, the matching success parameter is determined as the merged Boolean intersection parameter corresponding to the second hash mapping bucket.
[0110] like Figure 4 As shown, the identification intersection result (i.e., q0) can be {0, 1, 0, 1, 0, 0}, and the dimension of the identification intersection result is associated with the number of first business data identifiers 401. Accordingly, the first Boolean intersection slice may include 6 first Boolean intersection parameters, and the second Boolean intersection slice may include 6 second Boolean intersection parameters. Among them, the first first Boolean intersection parameter can be <0> , the first and second Boolean intersection parameters can be <0> , the XOR result of the first first Boolean intersection parameter and the first first Boolean intersection parameter is the merged Boolean intersection parameter (ie, 0). For example, the first first Boolean intersection parameter can be <0> , the first and second Boolean intersection parameters can be <1> , and so on, to obtain the first Boolean intersection slice and the second Boolean intersection slice.
[0111] It should be noted that the embodiment of the present application does not limit the timing of generating the first Boolean intersection fragment (i.e., realizing business data identification anonymization based on Circuit-PSI), and can be generated at any time before executing S205, for example, it can be executed in the first step of the data processing method provided in the embodiment of the present application, or it can be generated at any time before executing S205. Figure 2 The corresponding embodiment is executed at the position shown.
[0112] S205: Determine a service label corresponding to the service data according to the first Boolean intersection fragment and the first division result.
[0113] After obtaining the first Boolean intersection fragment and the first partition result, label reasoning may be performed based on the first Boolean intersection fragment and the first partition result. The process of label reasoning may refer to S205 . Label reasoning is federated computing.
[0114] In the process of label reasoning, some non-leaf nodes may have child node selection indication parameters in the first participant, while some non-leaf nodes may not have child node selection indication parameters in the first participant. In this case, for non-leaf nodes whose child node selection indication parameters are not in the first participant, a more accurate division result may be obtained by other participants, such as the second participant. Therefore, in order for the first participant to obtain an accurate division result, the second participant needs to synchronize its division result with the first participant. Based on this, in a possible implementation, according to the first Boolean intersection shard and the first division result, the method of determining the business label corresponding to the business data can be to receive the second division result synchronized by the second participant, and the second division result is when the child node selection indication parameter of the non-leaf node j is located in the second participant, and the second participant obtains the division result by comparing the data feature parameters of the business data with the child node selection indication parameter of the non-leaf node j. Then, the first division result and the second division result are ANDed to obtain the target division result, so as to correct the first division result based on the second division result to obtain a more accurate target division result. Then, according to the first Boolean intersection shard and the target division result, the business label corresponding to the business data is determined.
[0115] In an embodiment of the present application, each participant locally determines the division result. When the first participant needs to determine the business tag corresponding to the business data, the second participant synchronizes the second division result to the first participant through the second participant, thereby using the second division result to correct the first division result, thereby obtaining a more accurate target division result, so as to improve the accuracy of business tag determination.
[0116] It should be noted that when the second participant synchronizes the second partition result with the first participant, in order to avoid exposing the data of the second participant to the first participant, the two parties perform MPC calculations, for example, an OPPRF synchronization method may be used. At this time, the method of receiving the second partition result synchronized by the second participant participating in the vertical federated learning may be receiving the result partition shard synchronized by the second participant, where the result partition shard is a shard used to indicate the second partition result.
[0117] Specifically, the second participant puts the second partition result into the second hash map bucket corresponding to the second business data identifier in the Circuit-PSI process, and uses the OPPRF synchronization method to synchronize the second partition result to the first participant, and both parties finally obtain the result partition shards. The result partition shards can be expressed as<b1[i,t,j]> B .
[0118] Correspondingly, the first partition result and the second partition result are performed an AND operation to obtain the target partition result. The first partition result and the result partition slice are performed an AND operation to obtain the target partition result.
[0119] The embodiment of the present application synchronizes the second partition result to the first participant through the OPPRF synchronization method, thereby avoiding exposing the data of the second participant to the first participant, which is conducive to realizing anonymous vertical federated reasoning.
[0120] It is understandable that in some possible implementations, the business data obtained may be one or multiple business data in a batch. When the business data obtained is multiple business data in a batch, the first partitioning result of each business data in the t-th tree can be obtained for the business data in the batch. When the first partitioning result is represented by a partitioning indication mark, the partitioning indication mark of all business data is obtained. The partitioning indication marks of all business data are arranged in the Cuckoo hash order corresponding to the first business data identifier in the Circuit-PSI process, denoted as b0. Then the first partitioning result of business data i in the batch in the t-th tree can be expressed as b0[i,t,j], as the partitioning indication mark of the j-th leaf node of business data i in the t-th tree.
[0121] It should be noted that the implementation of S205 may be different depending on the implementation of S203. If in S203, the first partition result of each tree is obtained by traversal, then the implementation of S205 may be to determine the weight of the leaf node for processing business data on the t-th tree according to the first Boolean intersection fragment and the first partition result, and then determine the business label corresponding to the business data according to the weight of the leaf node for processing business data on the t-th tree.
[0122] In the process of determining the weight, if the second participant is required to synchronize the second partition result with the first participant, the calculation formula of the weight of the leaf node for processing business data on the tth tree can be as follows:
[0123] <v[i,t]> A =MUX(AND(AND(b0[i,t,j],<b1[i,t,j]> B ),<Q0[i]> B ),<w0[t,j]> A )
[0124] in,<v[i,t]> A represents the weight of the leaf node for processing business data on the t-th tree, b0[i,t,j] represents the partition indication mark of the j-th leaf node of the t-th tree for the business data i obtained by the first participant,<b1[i,t,j]> B Indicates the result of synchronization by the second participant is divided into shards,<Q0[i]> B represents the first Boolean intersection slice of the first party,<w0[t,j]> A Represents the weight of the j-th node of the t-th tree of the first participant.
[0125] The weight of the leaf node where the business data is located on the tth tree is determined. Since t = 1, 2, 3, ..., N, the weight of the leaf node where the business data is located on each tree can be obtained. Then, the business label corresponding to the business data can be determined according to the weight of the leaf node used to process the business data on each tree. Specifically, the weight of the leaf node where the business data is located on each tree can be added to calculate the final business label.
[0126] Add the weights of the leaf nodes on each tree used to process business data to get the cumulative result<s[i]> A ,<s[i]> A The calculation formula can be shown as follows:
[0127]
[0128] Finally, the business label is calculated. Taking binary classification as an example, the probability is finally calculated.<p[i]> A =Sigmoid(<s[i]> A ), and finally obtain the label fragment<y[i]> A =greater(<p[i]> A ,0.5), both parties restore the fragments and obtain the business label y.
[0129] If in S203, the first partitioning result of a partial tree or the first partitioning result of a single tree is obtained by a non-traversal method, the implementation method of S205 may be partially similar to the implementation method of S205 in the traversal method, except that the weights of the non-leaf nodes on the tth tree for processing business data are determined, and the weights of the non-leaf nodes on the unused trees for processing business data are 0.
[0130] It can be seen from the above technical solution that when performing anonymous vertical federated reasoning, the present application changes the comparison between data feature parameters and sub-node selection indication parameters into a local operation, that is, after obtaining the business data and the first decision tree, the data feature parameters are first compared with the sub-node selection indication parameters locally. Specifically, since the first participant obtains the first decision tree by participating in the vertical federated learning training, the first participant stores the parameter subnode selection indication parameter of the first decision tree, and the first decision tree includes N trees, the N trees respectively include multiple leaf nodes and multiple non-leaf nodes, and the multiple non-leaf nodes are respectively pre-configured with subnode selection indication parameters, and the subnode selection indication parameters are used to divide the subnodes of the non-leaf nodes used to process business data. Therefore, for the non-leaf node j of the tth tree among the N trees, in response to the subnode selection indication parameter of the non-leaf node j and the reference subnode selection indication parameter satisfying the preset matching condition, that is, the subnode selection indication parameter is locally in the first participant, then the data feature parameters of the business data are compared with the subnode selection indication parameter of the non-leaf node j locally to obtain a comparison result, and the subnode of the non-leaf node j used to process the business data is determined according to the comparison result, t is a positive integer less than or equal to N, j is an integer, and then based on the determined subnode of the non-leaf node j used to process the business data, the first division result of the business data on the tth tree is determined. Only when the sub-node selection indication parameter is local, the comparison between the data feature parameter and the sub-node selection indication parameter is performed locally, thereby saving a lot of overhead caused by data exchange in the privacy comparison process and improving the efficiency of privacy comparison. Afterwards, label reasoning is performed. Specifically, the first Boolean intersection shard held by the first participant can be obtained. The first Boolean intersection shard is used to indicate whether the business data identifier of the business data is in the identifier intersection result, so as to determine the business label corresponding to the business data according to the first Boolean intersection shard and the first division result. Compared with the related art, the present application changes the comparison between the data feature parameter and the sub-node selection indication parameter into a local operation, thereby eliminating the need to exchange data with other participants, saving a lot of overhead caused by data exchange in the privacy comparison process, improving the efficiency of privacy comparison, and thus improving the efficiency of anonymous vertical federated reasoning.
[0131] It should be noted that the traversal of the embodiment of the present application is a local calculation, and the label reasoning is a federated calculation, so the pipeline optimization method can be used to fully utilize the network and local computing resources. In one possible implementation method, the method of traversing N trees based on business data can be to complete the traversal of the tth tree based on business data, determine the weight of the leaf node on the tth tree for processing business data, that is, while performing label reasoning, the t+1th tree in the first decision tree can be traversed based on business data. In this way, the network and local computing resources can be fully utilized, and the network bandwidth can be maximized to reduce the delay of the results.
[0132] Similarly, for the processing of multiple batches of business data, if the business data is the business data of the k+1th batch, the way to traverse the N trees based on the business data can be to complete the traversal of the N trees based on the business data of the kth batch and determine the weight of the leaf node on the tth tree for processing the business data of the kth batch, that is, while performing label inference on the business data of the kth batch, the N trees can be traversed based on the business data of the k+1th batch, where k is a positive integer. In this way, network and local computing resources can be fully utilized, and network bandwidth can be maximized to reduce the delay of the results.
[0133] See also Figure 5 As shown in the figure, starting from the first tree, after completing the traversal, the second tree traversal is immediately started while label inference is being performed, and so on, which can maximize the use of network bandwidth to reduce the delay of the result. Similarly, for the processing of multiple batches of business data, this method can also be used. After the traversal of the first batch of business data is completed, the second batch of business data is immediately traversed while label inference is being performed, and so on, until all batches of business data are processed.
[0134] The above embodiment introduces the data processing method provided by the embodiment of the present application in detail. Next, the overall data processing process will be introduced by taking XGBoost, the most commonly used in GBDT, and multiple participants including the first participant and the second participant as an example. Figure 6 As shown, for both the first and second parties, data processing mainly includes three steps:
[0135] S601, finding the intersection of circuit privacy sets.
[0136] The first participant and the second participant both execute Circuit-PSI. For example, the first participant as the receiver (P0) inputs the first business data identifier ID0, and the second participant as the sender (P1) inputs the second business data identifier ID1. Both parties execute Circuit-PSI, and finally perform the first Boolean intersection sharding according to the order of Cuckoohash of the first business data identifier of the receiver (for example, the first participant).
[0137] S602: Local traversal.
[0138] The local traversal is performed to divide the business data into nodes corresponding to the t-th tree. Specifically, when the t-th tree is traversed, for business data i, when the representation form of the above-mentioned division indicator mark is a number, the first indicator mark is 1, and the second indicator mark is 0, P l Set the root node's partition indicator c l [i,t,0] is all 1, l can be 0 or 1.
[0139] Next, all non-leaf nodes j are traversed. When the child node selection indication parameter of the non-leaf node i is on the first participant's side, the data characteristic parameter of the business data is calculated locally. When the data characteristic parameter of the business data is less than the child node selection indication parameter, the division indication mark of the left child node is determined to be c. l [i, t, j*2+1] = 1, determine the partition indicator of the right child node as c l [i, t, j*2+2] = 0; otherwise, the partition indicator of the left child node is determined to be c l [i, t, j*2+1] = 0, determine the partition indicator of the right child node as c l [i, t, j*2+2] = 1. When the child node selection indication parameter of the non-leaf node i is not on the first participant's side, there is no need to compare the child node selection indication parameter with the data characteristic parameter of the business data. Instead, it is directly assumed that each child node of the non-leaf node j has business data, and the division indication mark of the left child node is directly determined to be c. l [i, t, j*2+1] = 1, determine the partition indicator of the right child node as c l [i,t,j*2+2]=1.
[0140] S603: Label inference.
[0141] Arrange the partition indicator marks of all business data in the Cuckoo hash order corresponding to the first business data identifier in the Circuit-PSI process, denoted as b0. Then the first partition result of business data i in a batch in the tth tree can be expressed as b0[i,t,j].
[0142] The second participant uses the OPPRF synchronization method to synchronize the second partition result to the first participant, and both parties finally obtain the result partition shards.<b1[i,t,j]> B .
[0143] Next, b0[i,t,j] and<b1[i,t,j]> B Substitute the above weight calculation formula to calculate the weight of the leaf node on the tth tree used to process business data i, and then add the weights of the leaf nodes on the tth tree used to process business data i based on the above formula to obtain the cumulative result, and then finally calculate the business label based on the cumulative result.
[0144] It should be noted that, based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods.
[0145] Based on the data processing method provided in the above embodiment, the present application embodiment further provides a data processing device 700. The data processing device 700 is deployed at the first participant, see Figure 7 As shown, the data processing device 700 includes an acquisition unit 701 and a determination unit 702:
[0146] The acquisition unit 701 is used to acquire business data and a first decision tree, where the first decision tree runs on the first participant, is obtained through vertical federated learning training, and includes N trees, where the N trees respectively include a plurality of leaf nodes and a plurality of non-leaf nodes, where the plurality of non-leaf nodes are respectively pre-configured with a child node selection indication parameter, and the child node selection indication parameter is used to divide the child nodes of the non-leaf nodes used to process the business data;
[0147] The determining unit 702 is used for comparing the data feature parameters of the business data with the child node selection indication parameters of the non-leaf node j of the t-th tree among the N trees, in response to the child node selection indication parameters of the non-leaf node j and the reference child node selection indication parameters stored by the first participant satisfying a preset matching condition, to obtain a comparison result, and determining a child node of the non-leaf node j for processing the business data according to the comparison result, where t is a positive integer less than or equal to N, and j is an integer;
[0148] The determining unit 702 is further configured to determine a first partitioning result of the service data on the tth tree based on the determined child nodes of the non-leaf node j for processing the service data;
[0149] The acquisition unit 701 is further used to acquire a first Boolean intersection slice held by the first participant, where the first Boolean intersection slice is used to indicate whether the business data identifier of the business data is in the identifier intersection result;
[0150] The determining unit 702 is further configured to determine a service label corresponding to the service data according to the first Boolean intersection slice and the first division result.
[0151] In a possible implementation, the non-leaf node includes M child nodes, and the determining unit 702 is configured to:
[0152] If the comparison result is the comparison result corresponding to the kth child node among the M child nodes, it is determined that the business data is processed by the kth child node of the non-leaf node j, where k is a positive integer less than or equal to M.
[0153] In a possible implementation manner, the determining unit 702 is configured to:
[0154] If the comparison result is the comparison result corresponding to the kth subnode among the M subnodes, the partition indication mark of the kth subnode is determined to be the first indication mark, and the partition indication marks of the remaining subnodes are determined to be the second indication mark, the first indication mark is used to indicate that the corresponding subnode is used to process the business data, the second indication mark is used to indicate that the corresponding subnode is not used to process the business data, and the remaining subnodes are the subnodes among the M subnodes except the kth subnode.
[0155] In a possible implementation manner, the M child nodes include a first child node and a second child node, and the determining unit 702 is configured to:
[0156] If the comparison result is that the data characteristic parameter of the service data is less than the child node selection indication parameter of the non-leaf node j, it is determined that the first child node is used to process the service data, and the first child node is the kth child node of the non-leaf node j, k=1;
[0157] If the comparison result is that the data characteristic parameter of the business data is greater than the child node selection indication parameter of the non-leaf node j, it is determined that the second child node is used to process the business data, and the second child node is the kth child node of the non-leaf node j, k=2.
[0158] In a possible implementation manner, the determining unit 702 is further configured to:
[0159] In response to the child node selection indication parameter and the reference child node selection indication parameter of the non-leaf node j not satisfying a preset matching condition, it is determined that each child node of the non-leaf node j is used to process the business data.
[0160] In a possible implementation, the second division result is represented by a division indication mark, and the determining unit 702 is configured to:
[0161] Determine that the partition indication mark of each child node of the non-leaf node j is a first indication mark, and the first indication mark is used to indicate that the corresponding child node is used to process the business data.
[0162] In a possible implementation manner, the determining unit 702 is configured to:
[0163] receiving a second partition result synchronized by a second participant, wherein the second partition result is a partition result obtained by the second participant by comparing the data feature parameter of the service data with the subnode selection indication parameter of the non-leaf node j when the subnode selection indication parameter of the non-leaf node j is located in the second participant;
[0164] Performing an AND operation on the first division result and the second division result to obtain a target division result;
[0165] A business label corresponding to the business data is determined according to the first Boolean intersection slice and the target partition result.
[0166] In a possible implementation manner, the determining unit 702 is configured to:
[0167] receiving a result partition shard synchronized by the second participant, wherein the result partition shard is a shard used to indicate the second partition result;
[0168] An AND operation is performed on the first partition result and the result partition slice to obtain the target partition result.
[0169] In a possible implementation, the t-th tree is a tree traversed at the current moment in the process of traversing the N trees based on the business data, and the non-leaf node j is a non-leaf node traversed at the current moment in the process of traversing all non-leaf nodes of the t-th tree based on the business data, then the determining unit 702 is used to:
[0170] If the child node of the non-leaf node j determined to be used for processing the business data is a leaf node of the t-th tree, the first partition result is obtained, and the first partition result is used to indicate the leaf node on the t-th tree used for processing the business data;
[0171] Determine, according to the first Boolean intersection fragment and the first partition result, a weight of a leaf node on the t-th tree for processing the business data;
[0172] The service label corresponding to the service data is determined according to the weight of the leaf node on the tth tree used to process the service data.
[0173] In a possible implementation manner, the determining unit 702 is configured to:
[0174] If the child node of the non-leaf node j determined to be used to process the business data is a non-leaf node, and the child node of the non-leaf node j meets the preset stop condition, the first partitioning result is obtained, and the first partitioning result is used to indicate the non-leaf node on the tth tree used to process the business data.
[0175] In a possible implementation manner, the determining unit 702 is configured to:
[0176] After completing the traversal of the tth tree based on the business data and determining the weight of the leaf node on the tth tree for processing the business data, the t+1th tree in the first decision tree is traversed based on the business data.
[0177] In a possible implementation, the service data is the service data of the k+1th batch, and the determining unit 702 is configured to:
[0178] After completing the traversal of the N trees based on the kth batch of business data and determining the weight of the leaf node on the tth tree used to process the kth batch of business data, the N trees are traversed based on the k+1th batch of business data, where k is a positive integer.
[0179] In a possible implementation, the acquiring unit 701 is configured to:
[0180] Performing cuckoo hash mapping on the first business data identifier of the first participant to obtain a first hash table corresponding to the first business data identifier;
[0181] Obtain the first Boolean intersection shard associated with the first hash table and the second hash table, the first hash table and the second hash table are used to perform hash table matching through a pseudo-random function to generate the first Boolean intersection shard and the second Boolean intersection shard held by the second participant; the first Boolean intersection shard and the second Boolean intersection shard are shards of the identifier intersection result; the second hash table is obtained by the second participant participating in vertical federated learning performing hash mapping on the second business data identifier of the second participant.
[0182] In a possible implementation, the first hash table and the second hash table have the same hash table dimension; the first hash table includes a first hash map bucket, and the second hash table includes a second hash map bucket having the same hash table dimension as the first hash map bucket;
[0183] Wherein, if the first hash mapping bucket includes a first business data identifier, the first random number corresponding to the first business data identifier in the first hash mapping bucket is obtained by randomly processing the first business data identifier in the first hash mapping bucket based on the pseudo-random function;
[0184] Among them, if the second hash mapping bucket includes a second business data identifier, the second random number corresponding to the second business data identifier in the second hash mapping bucket is obtained by randomly processing the second business data identifier in the second hash mapping bucket based on the pseudo-random function; the first random number and the second random number are used to perform random number matching to generate a first Boolean intersection parameter corresponding to the first hash mapping bucket and a second Boolean intersection parameter corresponding to the second hash mapping bucket; the first Boolean intersection parameter and the second Boolean intersection parameter are merged to obtain a merged Boolean intersection parameter; the merged Boolean intersection parameter is used to indicate the Boolean matching result of the first hash mapping bucket and the second hash mapping bucket; the first Boolean intersection parameter corresponding to each hash table dimension in the first hash table is used to form a first Boolean intersection shard corresponding to the first business data identifier, and the second Boolean intersection parameter corresponding to each hash table dimension in the second hash table is used to form a second Boolean intersection shard corresponding to the first business data identifier.
[0185] It can be seen from the above technical solution that when performing anonymous vertical federated reasoning, the present application changes the comparison between data feature parameters and sub-node selection indication parameters into a local operation, that is, after obtaining the business data and the first decision tree, the data feature parameters are first compared with the sub-node selection indication parameters locally. Specifically, since the first participant obtains the first decision tree by participating in the vertical federated learning training, the first participant stores the parameter subnode selection indication parameter of the first decision tree, and the first decision tree includes N trees, the N trees respectively include multiple leaf nodes and multiple non-leaf nodes, and the multiple non-leaf nodes are respectively pre-configured with subnode selection indication parameters, and the subnode selection indication parameters are used to divide the subnodes of the non-leaf nodes used to process business data. Therefore, for the non-leaf node j of the tth tree among the N trees, in response to the subnode selection indication parameter of the non-leaf node j and the reference subnode selection indication parameter satisfying the preset matching condition, that is, the subnode selection indication parameter is locally in the first participant, then the data feature parameters of the business data are compared with the subnode selection indication parameter of the non-leaf node j locally to obtain a comparison result, and the subnode of the non-leaf node j used to process the business data is determined according to the comparison result, t is a positive integer less than or equal to N, j is an integer, and then based on the determined subnode of the non-leaf node j used to process the business data, the first division result of the business data on the tth tree is determined. Only when the sub-node selection indication parameter is local, the comparison between the data feature parameter and the sub-node selection indication parameter is performed locally, thereby saving a lot of overhead caused by data exchange in the privacy comparison process and improving the efficiency of privacy comparison. Afterwards, label reasoning is performed. Specifically, the first Boolean intersection shard held by the first participant can be obtained. The first Boolean intersection shard is used to indicate whether the business data identifier of the business data is in the identifier intersection result, so as to determine the business label corresponding to the business data according to the first Boolean intersection shard and the first division result. Compared with the related art, the present application changes the comparison between the data feature parameter and the sub-node selection indication parameter into a local operation, thereby eliminating the need to exchange data with other participants, saving a lot of overhead caused by data exchange in the privacy comparison process, improving the efficiency of privacy comparison, and thus improving the efficiency of anonymous vertical federated reasoning.
[0186] The embodiment of the present application further provides a computer device, which can execute the data processing method. The computer device can be a terminal, Figure 8 The structure diagram of a terminal provided by an embodiment of the present application is shown. Figure 8 In the example, the terminal is a smartphone:
[0187] refer to Figure 8 The smartphone includes: a radio frequency (RF) circuit 810, a memory 820, an input unit 830, a display unit 840, a sensor 850, an audio circuit 860, a wireless fidelity (WiFi) module 870, a processor 880, and a power supply 890. The input unit 830 may include a touch panel 831 and other input devices 832, the display unit 840 may include a display panel 841, and the audio circuit 860 may include a speaker 861 and a microphone 862. It is understood that Figure 8 The structure of the smartphone shown in the figure does not constitute a limitation of the smartphone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0188] The memory 820 can be used to store software programs and modules. The processor 880 executes various functional applications and data processing of the smartphone by running the software programs and modules stored in the memory 820. The memory 820 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, a phone book, etc.), etc. In addition, the memory 820 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0189] The processor 880 is the control center of the smartphone, which uses various interfaces and lines to connect various parts of the entire smartphone, and executes various functions of the smartphone and processes data by running or executing software programs and / or modules stored in the memory 820, and calling data stored in the memory 820. Optionally, the processor 880 may include one or more processing units; preferably, the processor 880 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 880.
[0190] In this embodiment, the processor 880 in the smart phone can execute the methods provided in the embodiments of the present application.
[0191] The computer device provided in the embodiment of the present application may also be a server. Fig. 9 As shown, Fig. 9 The structural diagram of the server 900 provided in the embodiment of the present application, the server 900 may have relatively large differences due to different configurations or performances, and may include one or more processors, such as a central processing unit (CPU) 922, and a memory 932, one or more storage media 930 (such as one or more mass storage devices) storing application programs 942 or data 944. Among them, the memory 932 and the storage medium 930 can be temporary storage or permanent storage. The program stored in the storage medium 930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the server. Furthermore, the central processing unit 922 can be configured to communicate with the storage medium 930 and execute a series of instruction operations in the storage medium 930 on the server 900.
[0192] The server 900 may also include one or more power supplies 926, one or more wired or wireless network interfaces 950, one or more input and output interfaces 958, and / or one or more operating systems 941, such as Windows Server 2000. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM etc.
[0193] In this embodiment, the central processor 922 in the server 900 can execute the methods provided in the embodiments of the present application.
[0194] According to one aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the methods described in the above embodiments.
[0195] According to one aspect of the present application, a computer program product is provided, the computer program product comprising a computer program, the computer program being stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method provided in various optional implementations of the above-mentioned embodiments.
[0196] The descriptions of the processes or structures corresponding to the above-mentioned figures have different emphases. For parts that are not described in detail in a certain process or structure, please refer to the relevant descriptions of other processes or structures.
[0197] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0198] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0199] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0200] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0201] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a terminal, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store computer programs.
[0202] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0203] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, ordinary technical members in the art should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.< / x> < / z> < / x> < / y> < / y> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / x> < / x>
Claims
1. A data processing method, characterized in that: The method is performed by a first participant, and includes: Obtaining business data and a first decision tree, where the first decision tree runs on the first participant, the first decision tree is obtained through vertical federated learning training, and includes N trees, where the N trees respectively include a plurality of leaf nodes and a plurality of non-leaf nodes, the plurality of non-leaf nodes are respectively pre-configured with a child node selection indication parameter, and the child node selection indication parameter is used to divide the child nodes of the non-leaf nodes used to process the business data; For a non-leaf node j of the t-th tree among the N trees, in response to the child node selection indication parameter of the non-leaf node j and the reference child node selection indication parameter stored by the first participant satisfying a preset matching condition, the data feature parameter of the business data is compared with the child node selection indication parameter of the non-leaf node j to obtain a comparison result, and a child node of the non-leaf node j for processing the business data is determined according to the comparison result, where t is a positive integer less than or equal to N, and j is an integer; Determine a first partition result of the business data on the tth tree based on the determined child nodes of the non-leaf node j for processing the business data; Acquire a first Boolean intersection slice held by the first participant, where the first Boolean intersection slice is used to indicate whether the business data identifier of the business data is in the identifier intersection result; A business label corresponding to the business data is determined according to the first Boolean intersection slice and the first division result.
2. The method according to claim 1, characterized in that The non-leaf node j includes M child nodes, and determining the child node of the non-leaf node j for processing the service data according to the comparison result includes: If the comparison result is the comparison result corresponding to the kth child node among the M child nodes, it is determined that the business data is processed by the kth child node of the non-leaf node j, where k is a positive integer less than or equal to M.
3. The method according to claim 2, characterized in that If the comparison result is a comparison result corresponding to the kth child node among the M child nodes, determining that the kth child node of the non-leaf node j processes the service data includes: If the comparison result is the comparison result corresponding to the kth subnode among the M subnodes, the partition indication mark of the kth subnode is determined to be the first indication mark, and the partition indication marks of the remaining subnodes are determined to be the second indication mark, the first indication mark is used to indicate that the corresponding subnode is used to process the business data, the second indication mark is used to indicate that the corresponding subnode is not used to process the business data, and the remaining subnodes are the subnodes among the M subnodes except the kth subnode.
4. The method according to claim 2, characterized in that: The M child nodes include a first child node and a second child node, and if the comparison result is a comparison result corresponding to a k-th child node among the M child nodes, determining that the k-th child node of the non-leaf node j processes the service data includes: If the comparison result is that the data characteristic parameter of the service data is less than the child node selection indication parameter of the non-leaf node j, it is determined that the first child node is used to process the service data, and the first child node is the k-th child node of the non-leaf node j, k=1; If the comparison result is that the data characteristic parameter of the business data is greater than the child node selection indication parameter of the non-leaf node j, it is determined that the second child node is used to process the business data, and the second child node is the kth child node of the non-leaf node j, k=2.
5. The method according to claim 1, characterized in that The method further comprises: In response to the child node selection indication parameter and the reference child node selection indication parameter of the non-leaf node j not satisfying a preset matching condition, it is determined that each child node of the non-leaf node j is used to process the business data.
6. The method according to claim 5, characterized in that The determining that each child node of the non-leaf node j is used to process the service data includes: Determine that the partition indication mark of each child node of the non-leaf node j is a first indication mark, and the first indication mark is used to indicate that the corresponding child node is used to process the business data.
7. The method according to claim 5, characterized in that The determining, according to the first Boolean intersection slice and the first division result, a service label corresponding to the service data includes: receiving a second partition result synchronized by the second participant, the second partition result being a partition result obtained by the second participant by comparing the data feature parameter of the service data with the subnode selection indication parameter of the non-leaf node j when the subnode selection indication parameter of the non-leaf node j is located in the second participant; Performing an AND operation on the first division result and the second division result to obtain a target division result; A business label corresponding to the business data is determined according to the first Boolean intersection fragment and the target partition result.
8. The method according to claim 7, characterized in that The receiving a second division result synchronized by the second participant includes: receiving a result partition shard synchronized by the second participant, wherein the result partition shard is a shard used to indicate the second partition result; The performing an AND operation on the first division result and the second division result to obtain a target division result includes: An AND operation is performed on the first partition result and the result partition slice to obtain the target partition result.
9. The method according to claim 1, characterized in that: The t-th tree is a tree traversed at the current moment in the process of traversing the N trees based on the business data, and the non-leaf node j is a non-leaf node traversed at the current moment in the process of traversing all non-leaf nodes of the t-th tree based on the business data, then the first division result of the business data on the t-th tree is determined based on the child node of the non-leaf node j determined for processing the business data, including: If the child node of the non-leaf node j determined to be used for processing the business data is a leaf node of the t-th tree, the first partition result is obtained, and the first partition result is used to indicate the leaf node on the t-th tree used for processing the business data; The determining, according to the first Boolean intersection slice and the first division result, a service label corresponding to the service data includes: Determine, according to the first Boolean intersection fragment and the first partition result, a weight of a leaf node on the t-th tree for processing the business data; The service label corresponding to the service data is determined according to the weight of the leaf node on the tth tree used to process the service data.
10. The method according to claim 1, characterized in that The determining, based on the determined child nodes of the non-leaf node j for processing the service data, a first partition result of the service data on the tth tree comprises: If the child node of the non-leaf node j determined to be used to process the business data is a non-leaf node, and the child node of the non-leaf node j meets the preset stop condition, the first partitioning result is obtained, and the first partitioning result is used to indicate the non-leaf node on the tth tree used to process the business data.
11. The method according to claim 9, characterized in that The traversing the N trees based on the business data includes: After completing the traversal of the tth tree based on the business data and determining the weight of the leaf node on the tth tree for processing the business data, the t+1th tree in the first decision tree is traversed based on the business data.
12. The method according to claim 9, characterized in that The business data is business data of the k+1th batch, and traversing the N trees based on the business data includes: After completing the traversal of the N trees based on the kth batch of business data and determining the weight of the leaf node on the tth tree used to process the kth batch of business data, the N trees are traversed based on the k+1th batch of business data, where k is a positive integer.
13. The method according to any one of claims 1 to 12, characterized in that: The obtaining of the first Boolean intersection slice held by the first participant includes: Performing cuckoo hash mapping on the first business data identifier of the first participant to obtain a first hash table corresponding to the first business data identifier; Obtain the first Boolean intersection shard associated with the first hash table and the second hash table, the first hash table and the second hash table are used to perform hash table matching through a pseudo-random function to generate the first Boolean intersection shard and the second Boolean intersection shard held by the second participant; the first Boolean intersection shard and the second Boolean intersection shard are shards of the identifier intersection result; the second hash table is obtained by the second participant participating in vertical federated learning performing hash mapping on the second business data identifier of the second participant.
14. The method according to claim 13, characterized in that The first hash table and the second hash table have the same hash table dimension; the first hash table includes a first hash map bucket, and the second hash table includes a second hash map bucket having the same hash table dimension as the first hash map bucket; Wherein, if the first hash mapping bucket includes a first business data identifier, the first random number corresponding to the first business data identifier in the first hash mapping bucket is obtained by randomly processing the first business data identifier in the first hash mapping bucket based on the pseudo-random function; Among them, if the second hash mapping bucket includes a second business data identifier, the second random number corresponding to the second business data identifier in the second hash mapping bucket is obtained by randomly processing the second business data identifier in the second hash mapping bucket based on the pseudo-random function; the first random number and the second random number are used to perform random number matching to generate a first Boolean intersection parameter corresponding to the first hash mapping bucket and a second Boolean intersection parameter corresponding to the second hash mapping bucket; the first Boolean intersection parameter and the second Boolean intersection parameter are merged to obtain a merged Boolean intersection parameter; the merged Boolean intersection parameter is used to indicate the Boolean matching result of the first hash mapping bucket and the second hash mapping bucket; the first Boolean intersection parameter corresponding to each hash table dimension in the first hash table is used to form a first Boolean intersection shard corresponding to the first business data identifier, and the second Boolean intersection parameter corresponding to each hash table dimension in the second hash table is used to form a second Boolean intersection shard corresponding to the first business data identifier.
15. A data processing device, characterized in that: The device is deployed at the first participant, and includes an acquisition unit and a determination unit: The acquisition unit is used to acquire business data and a first decision tree, the first decision tree runs on the first participant, the first decision tree is obtained through vertical federated learning training, and includes N trees, the N trees respectively include a plurality of leaf nodes and a plurality of non-leaf nodes, the plurality of non-leaf nodes are respectively pre-configured with a child node selection indication parameter, and the child node selection indication parameter is used to divide the child nodes of the non-leaf nodes used to process the business data; The determining unit is used for comparing the data characteristic parameters of the business data with the child node selection indication parameters of the non-leaf node j of the t-th tree among the N trees, in response to the child node selection indication parameters of the non-leaf node j and the reference child node selection indication parameters stored by the first participant satisfying a preset matching condition, to obtain a comparison result, and determining a child node of the non-leaf node j for processing the business data according to the comparison result, where t is a positive integer less than or equal to N, and j is an integer; The determining unit is further configured to determine a first partitioning result of the business data on the tth tree based on the determined child nodes of the non-leaf node j for processing the business data; The acquisition unit is further used to acquire a first Boolean intersection slice held by the first participant, where the first Boolean intersection slice is used to indicate whether the business data identifier of the business data is in the identifier intersection result; The determining unit is further configured to determine a service label corresponding to the service data according to the first Boolean intersection slice and the first division result.
16. A computer device, characterized in that: The computer device comprises a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1 to 14 according to the instructions in the computer program.
17. A computer-readable storage medium, the computer-readable storage medium being used to store a computer program, the computer program, when executed by a processor, causing the processor to perform the method according to any one of claims 1 to 14.
18. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 14 is implemented.
Citation Information
Patent Citations
Multi-party secret sharing data privacy comparison method based on efficient ciphertext confusion technology
CN116743376A
Data processing method and device, computer equipment and readable storage medium
CN116975017A
Method and device for dividing circuit, and computer-readable recording medium where circuit dividing program is recorded
JP1997330353A
Class-specific predictor importance
US20230316125A1
Cited By
Rule execution method and device, storage medium and electronic equipment
CN121094171A
Data processing method and apparatus, and device, medium and program product
WO2026175018A1