A data processing method and device, computer equipment and readable storage medium

CN116975018BActive Publication Date: 2026-09-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310587313.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-22
Publication Date
2026-09-22
Estimated Expiration
2043-05-22

AI Technical Summary

Technical Problem

[0003]然而,第一业务样本标识和第二业务样本标识可能为隐私数据(例如,手机号),直接使用隐私数据(例如,手机号)生成决策树的过程不可避免地会暴露这些隐私数据(例如,将第一参与方的第一业务样本标识暴露至第二参与方),从而降低了纵向联邦学习算法的参与方所拥有的数据的安全性

Benefits of technology

[0062]由此可见,本申请实施例可以无需获取第一业务样本标识和第二业务样本标识之间的交集业务样本标识,从而无需直接获取交集梯度向量,而是通过第一参与方和第二参与方分别获取交集梯度向量的分片(即第一交集梯度分片和第二交集梯度分片),无需直接对第一业务样本标识和第二业务样本标识进行比较。进一步地,第一参与方可以直接基于第一业务特征对第一业务样本标识进行业务样本划分,得到第一特征数据划分矩阵,第二参与方直接可以基于第二业务特征对第二业务样本标识进行业务样本划分,得到第二特征数据划分矩阵,从而实现第一参与方和第二参与方分别在本地进行业务样本划分(即在本地计算待选特征划分点),使得第一参与方无需处理第二业务样本标识,第二参与方无需处理第一业务样本标识。可以理解的是,第一特征数据划分矩阵和第二特征数据划分矩阵可以共同用于对交集梯度向量进行梯度划分,得到第一特征梯度分片和第二特征梯度分片,从而基于第一特征梯度分片和第二特征梯度分片确定针对第一参与方和第二参与方的目标决策树,从而可以在实现通过收敛后的目标决策树对业务数据进行业务处理的同时,提高第一参与方和第二参与方所分别拥有的数据的安全性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116975018B_ABST
    Figure CN116975018B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data processing method and device, computer equipment and a readable storage medium. The method comprises: obtaining a first identification quantity corresponding to a first business sample identification of a first participant; if a comparison result indicates that the first identification quantity is less than or equal to a second identification quantity, obtaining a first feature data division matrix of a first business feature for the first business sample identification, and a first feature data division fragment of a second business feature for a second business sample identification; obtaining a first intersection gradient fragment associated with the first business sample identification and the second business sample identification; obtaining a first feature gradient fragment associated with the intersection gradient vector; and the first feature gradient fragment and the second feature gradient fragment are used to determine an optimal split result for the first business feature and the second business feature. By using the present application, the security of data owned by the participant can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technology, and in particular to a data processing method, apparatus, computer equipment, and readable storage medium. Background Technology

[0002] The vertical federated learning algorithm obtains the first business sample identifier of the first participant and the second business sample identifier of the second participant. It directly compares the first and second business sample identifiers (for example, the first participant sends the first business sample identifier to the second participant, and the second participant compares the first and second business sample identifiers) to generate the intersection business sample identifier between the first and second business sample identifiers. Then, based on the business features of the first participant and the business features of the second participant, it directly generates the intersection gradient vector. Based on the intersection gradient vector, it determines the optimal splitting result for the first and second business features, and thus generates a decision tree for business processing of business data based on the optimal splitting result.

[0003] However, the first business sample identifier and the second business sample identifier may be private data (e.g., mobile phone numbers). The process of directly using private data (e.g., mobile phone numbers) to generate a decision tree will inevitably expose this private data (e.g., exposing the first business sample identifier of the first participant to the second participant), thereby reducing the security of the data owned by the participants in the longitudinal federated learning algorithm. Summary of the Invention

[0004] This application provides a data processing method, apparatus, computer equipment, and readable storage medium, which can improve the security of data owned by participating parties.

[0005] This application provides a data processing method, executed by a first participant in vertical federated learning, comprising:

[0006] Obtain the number of first identifiers corresponding to the first business sample identifier of the first participant; the number of first identifiers is used to compare with the number of second identifiers corresponding to the second business sample identifiers of the second participant in the vertical federated learning, and obtain the comparison result;

[0007] If the comparison result indicates that the number of first identifiers is less than or equal to the number of second identifiers, then a first feature data partitioning matrix for the first business feature of the first business sample identifier and a first feature data partitioning slice for the second business feature of the second business sample identifier are obtained; the first feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the first participant in partitioning the first business sample identifier based on the first business feature; the first feature data partitioning slice and the second feature data partitioning slice held by the second participant are slices of the second feature data partitioning matrix; the second feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the second participant in partitioning the second business sample identifier based on the second business feature;

[0008] Obtain the first intersection gradient slice that is associated with both the first business sample identifier and the second business sample identifier; the first intersection gradient slice and the second intersection gradient slice held by the second participant are slices of the intersection gradient vector; the intersection gradient vector is used to characterize the sample gradient of the intersection business sample identifier between the first business sample identifier and the second business sample identifier.

[0009] Obtain the first feature gradient slice associated with the intersection gradient vector; the first feature data partitioning matrix and the second feature data partitioning matrix are used to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient slice and the second feature gradient slice held by the second participant; the first feature gradient slice and the second feature gradient slice are used to determine the optimal splitting result for the first business feature and the second business feature; the optimal splitting result is used to generate the target decision tree for the first participant and the second participant; the converged target decision tree is used to perform business processing on the business data.

[0010] One embodiment of this application provides a data processing apparatus, which operates on a first participant in vertical federated learning, including:

[0011] The quantity acquisition module is used to acquire the number of first identifiers corresponding to the first business sample identifiers of the first participant; the number of first identifiers is used to compare with the number of second identifiers corresponding to the second business sample identifiers of the second participant in the vertical federated learning to obtain the comparison result.

[0012] The first acquisition module is configured to, if the comparison result indicates that the number of first identifiers is less than or equal to the number of second identifiers, acquire a first feature data partitioning matrix for the first business feature of the first business sample identifier, and a first feature data partitioning slice for the second business feature of the second business sample identifier; the first feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the first participant in partitioning the first business sample identifier based on the first business feature; the first feature data partitioning slice and the second feature data partitioning slice held by the second participant are slices of the second feature data partitioning matrix; the second feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the second participant in partitioning the second business sample identifier based on the second business feature;

[0013] The second acquisition module is used to acquire a first intersection gradient slice that is associated with both the first business sample identifier and the second business sample identifier; the first intersection gradient slice and the second intersection gradient slice held by the second participant are slices of the intersection gradient vector; the intersection gradient vector is used to characterize the sample gradient of the intersection business sample identifier between the first business sample identifier and the second business sample identifier.

[0014] The third acquisition module is used to acquire the first feature gradient slice associated with the intersection gradient vector; the first feature data partitioning matrix and the second feature data partitioning matrix are used to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient slice and the second feature gradient slice held by the second participant; the first feature gradient slice and the second feature gradient slice are used to determine the optimal splitting result for the first business feature and the second business feature; the optimal splitting result is used to generate the target decision tree for the first participant and the second participant; the converged target decision tree is used to perform business processing on the business data.

[0015] Among them, the number of first business features is M1, and the number of second business features is M2; M1 is a positive integer, and M2 is a positive integer; the M1 first business features include first business feature V. d d is a non-negative integer less than M1;

[0016] The first acquisition module includes:

[0017] The sample partitioning unit is used to obtain the first business feature V for the first business sample identifier. d Feature data, based on the first business feature V d The feature data is used to divide the first business sample identifier into business samples, generating the first business feature V. d The feature data is a bucketing indicator vector for feature bucketing;

[0018] The matrix determination unit is used to determine the feature data of the M1 first business features according to the bucketing indicator vector of the feature bucketing, and to form the first feature data partitioning matrix of the M1 first business features for the first business sample identifier.

[0019] The sharding acquisition unit is used to acquire the first feature data of M2 second business features for the second business sample identifier and divide them into shards.

[0020] Among them, for the first business feature V d The number of feature buckets is K, where K is a positive integer;

[0021] The sample partitioning unit is specifically used to partition the first business feature V d Feature data is distributed to K feature buckets; the K feature buckets include feature bucket K u u is a non-negative integer less than K;

[0022] The sample partitioning unit is specifically used if the first business sample identifier is in the first business feature V d The feature data below belongs to feature bucket K u Then the data existence parameter is determined as the first business sample identifier in the first business feature V. d Next, for feature binning K u The bucketing indicator parameter, if the first business sample is identified in the first business feature V d The feature data below does not belong to feature bucket K u If the data does not exist, the parameter is determined as the first business sample identifier in the first business feature V. d Next, for feature binning K u The binning indicator parameters;

[0023] The sample partitioning unit is specifically used to divide each first business sample identifier into K buckets based on its features. u The bucketing indicator parameters are determined as the first business feature V. d Feature data for feature binning K u Bucket indicator vector.

[0024] The second acquisition module includes:

[0025] The first acquisition unit is used to acquire a first sample gradient slice that is jointly associated with the first business sample identifier and the second business sample identifier; the first sample gradient slice and the second sample gradient slice held by the second participant are slices of the sample gradient vector; the sample gradient vector is used to characterize the sample gradient of the first business sample identifier;

[0026] The second acquisition unit is used to acquire the first Boolean intersection fragment corresponding to the first business sample identifier; the first Boolean intersection fragment and the second Boolean intersection fragment held by the second participant are fragments of the Boolean intersection vector; the Boolean intersection vector is used to indicate the intersection state of the first business sample identifier with respect to the second business sample identifier;

[0027] The third acquisition unit is used to acquire the first intersection gradient slice that is associated with both the first business sample identifier and the second business sample identifier; the Boolean intersection vector is used to perform gradient filtering on the sample gradient vector to generate the first intersection gradient slice and the second intersection gradient slice held by the second participant.

[0028] Specifically, the first acquisition unit is used to acquire, if the first participant is a tag holder, the first cumulative sample prediction slice of the first business sample identifier and the first business tag vector corresponding to the first business sample tag of the first business sample identifier; and to determine the first sample gradient slice jointly associated with the first business sample identifier and the second business sample identifier based on the first cumulative sample prediction slice and the first business tag vector; the first cumulative sample prediction slice and the second cumulative sample prediction slice held by the second participant are slices of the cumulative sample prediction vector; the cumulative sample prediction vector is used to characterize the historical cumulative sample prediction value when generating the target decision tree through the first business sample identifier; the second participant is used to determine the second sample gradient slice jointly associated with the first business sample identifier and the second business sample identifier based on the second cumulative sample prediction slice;

[0029] The first acquisition unit is specifically configured to, if the second participant is a tag holder, acquire the first cumulative sample prediction slice of the first business sample identifier and the first business tag slice corresponding to the second business sample tag of the second business sample identifier; and determine the first sample gradient slice jointly associated with the first business sample identifier and the second business sample identifier based on the first cumulative sample prediction slice and the first business tag slice; the first business tag slice and the second business tag slice held by the second participant are slices of the second business tag vector corresponding to the second business sample tag; the first cumulative sample prediction slice and the second cumulative sample prediction slice held by the second participant are slices of the cumulative sample prediction vector; the cumulative sample prediction vector is used to characterize the historical cumulative sample prediction value when generating the target decision tree through the first business sample identifier; and the second participant is used to determine the second sample gradient slice jointly associated with the first business sample identifier and the second business sample identifier based on the second cumulative sample prediction slice and the second business tag slice.

[0030] The second acquisition unit is specifically used to perform a cuckoo hash mapping on the first business sample identifier to obtain the first hash table corresponding to the first business sample identifier.

[0031] The second acquisition unit is specifically used to acquire the first Boolean intersection fragment associated with the first hash table and the second hash table; the first hash table and the second hash table are used to perform hash table matching through an unintentionally programmable pseudo-random function to generate the first Boolean intersection fragment and the second Boolean intersection fragment held by the second participant; the second hash table is obtained by the second participant through hash mapping of the second business sample identifier.

[0032] The first hash table and the second hash table have the same hash table dimension; the first hash table includes a first hash mapping bucket, and the second hash table includes a second hash mapping bucket with the same hash table dimension as the first hash mapping bucket;

[0033] Wherein, if the first hash mapping bucket includes the first business sample identifier, the first random number corresponding to the first business sample identifier in the first hash mapping bucket is obtained by randomly processing the first business sample identifier in the first hash mapping bucket based on an unintentionally programmable pseudo-random function;

[0034] Wherein, if the second hash mapping bucket includes the second business sample identifier, the second random number corresponding to the second business sample identifier in the second hash mapping bucket is obtained by randomly processing the second business sample identifier in the second hash mapping bucket based on an unintentionally programmable pseudo-random function; the first random number and the second random number are used to perform random number matching between the first participant and the second participant to generate the first Boolean intersection parameter corresponding to the first hash mapping bucket and the second Boolean intersection parameter corresponding to the second hash mapping bucket; the first Boolean intersection parameter and the second Boolean intersection parameter are the fragments of the merged Boolean intersection parameter; the merged Boolean intersection parameter is used to indicate the Boolean matching result of the first hash mapping bucket and the second hash mapping bucket; the first Boolean intersection parameter corresponding to each hash table dimension in the first hash table is used to form the first Boolean intersection fragment corresponding to the first business sample identifier, and the second Boolean intersection parameter corresponding to each hash table dimension in the second hash table is used to form the second Boolean intersection fragment corresponding to the first business sample identifier.

[0035] The first business sample identifier and the second business sample identifier together include M business features belonging to different business dimensions; M is a positive integer greater than 1; the M business features include business feature V. c c is a non-negative integer less than M;

[0036] Wherein, if business characteristic V c If it belongs to the first business feature, then the first feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector to obtain the data for business feature V. c First feature gradient slicing and targeting business feature V c The second feature gradient partitioning;

[0037] Wherein, if business characteristic Vc If it belongs to the second business feature, then the second feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector to obtain the data for business feature V. c First feature gradient slicing and targeting business feature V c The second feature gradient slice; the first feature gradient slice for M business features is used to form the first feature gradient slice held by the first participant; the second feature gradient slice for M business features is used to form the second feature gradient slice held by the second participant.

[0038] Among them, business feature V c The number of feature buckets is K, where K is a positive integer; if business feature V c If the second feature belongs to the second business sample identifier, then the second feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector, resulting in a first feature gradient sub-partition for K feature buckets and a second feature gradient sub-partition for K feature buckets; the first feature gradient sub-partition for K feature buckets is used to form the business feature V. c The first feature gradient partitioning; the second feature gradient sub-partitioning for K feature buckets is used to form a sub-partition for business feature V. c The second feature gradient slice.

[0039] Among them, the first feature gradient slice held by the first participant includes M first feature gradient slices for M business features; for business feature V c The first feature gradient partition includes K first feature gradient sub-partitions for K feature buckets;

[0040] The device also includes:

[0041] The first generation module is used to generate business features V based on the K first feature gradient sub-slices for K feature buckets. c The corresponding K-1 first-partition feature gradient slices;

[0042] The second generation module is used to generate a first cumulative gradient slice corresponding to the first participant based on the K first feature gradient sub-slices for K feature buckets. The M*KM first partitioning feature gradient slices, the first cumulative gradient slices, the M*KM second partitioning feature gradient slices corresponding to the M business features, and the second cumulative gradient slice corresponding to the second participant are used together to determine the information gain for the M business features and K feature buckets. The second cumulative gradient slice is generated by the second participant based on the K second feature gradient sub-slices for the K feature buckets. The M*KM second partitioning feature gradient slices corresponding to the M business features include business feature V. c The corresponding K-1 second-partition gradient slices; business feature Vc The corresponding K-1 second partition feature gradient slices are generated by the second participant based on the K second feature gradient sub-slices for the K feature buckets; the optimal information gain among the information gains for the M business features and the K feature buckets is used to determine the optimal splitting result for the first business feature and the second business feature; the optimal splitting result includes the optimal splitting feature and the optimal splitting point.

[0043] The device also includes:

[0044] The fourth acquisition module is used to acquire the first split point partitioning matrix for the first business feature and the first split point partitioning fragment for the second business feature of the second business sample identifier; the first split point partitioning matrix is ​​used to characterize the partitioning result of the first feature data partitioning matrix in feature bucketing; the first split point partitioning fragment and the second split point partitioning fragment held by the second participant are fragments of the second feature data partitioning matrix; the second split point partitioning matrix is ​​used to characterize the partitioning result of the second feature data partitioning matrix in feature bucketing;

[0045] The first processing module is used to obtain the first split point partition vector corresponding to the optimal split result from the first split point partition matrix if the optimal split feature belongs to the first business feature, perform XOR processing on the first split point partition vector through the first random blind partition, generate the first synchronous blind partition, and send the first synchronous blind partition to the second participant; the first random blind partition and the first synchronous blind partition are partitions of the first split point partition vector;

[0046] The second processing module is used to perform XOR processing on the first split point partitioning segment through the second random blinding segmentation if the optimal splitting feature belongs to the second business feature, thereby generating a random blinding matrix; the random blinding matrix is ​​used to obtain the random blinding vector corresponding to the optimal splitting result from the random blinding matrix through an unintentional transmission protocol; the second participant is used to obtain the splitting point partitioning sub-segment corresponding to the optimal splitting result from the second splitting point partitioning segmentation, and perform XOR processing on the splitting point partitioning sub-segment through the random blinding vector to generate a second synchronous blinding segmentation; the second random blinding segmentation and the second synchronous blinding segmentation are segments of the second splitting point partitioning vector corresponding to the optimal splitting result in the second splitting point partitioning matrix.

[0047] The device also includes:

[0048] The fifth acquisition module is used to acquire the first node Boolean fragment of the first child node of the node to be partitioned in the target decision tree of the first participant and the second participant, and the second node Boolean fragment of the second child node of the node to be partitioned; the first node Boolean fragment and the third node Boolean fragment held by the second participant are fragments of the first node Boolean vector; the second node Boolean fragment and the fourth node Boolean fragment held by the second participant are fragments of the second node Boolean vector; the first node Boolean vector is used to represent the partitioning state of the intersection business sample identifier being partitioned to the first child node, and the second node Boolean vector is used to represent the partitioning state of the intersection business sample identifier being partitioned to the second child node;

[0049] Wherein, if the optimal splitting feature belongs to the first business feature, then the first node Boolean partition, the second node Boolean partition, the third node Boolean partition, and the fourth node Boolean partition are obtained by the first random blind partition and the first synchronous blind partition jointly performing Boolean partitioning on the Boolean intersection vector;

[0050] If the optimal splitting feature belongs to the second business feature, then the first node Boolean partition, the second node Boolean partition, the third node Boolean partition, and the fourth node Boolean partition are obtained by the second random blind partition and the second synchronous blind partition jointly performing Boolean partitioning on the Boolean intersection vector.

[0051] The device also includes:

[0052] The sixth acquisition module is used to acquire the first node gradient slice of the first child node and the second node gradient slice of the second child node; the first node gradient slice and the third node gradient slice held by the second participant are slices of the first node gradient vector; the second node gradient slice and the fourth node gradient slice held by the second participant are slices of the second node gradient vector; the first node gradient vector is used to represent the sample gradient of the first child node, and the second node gradient vector is used to represent the sample gradient of the second child node;

[0053] The gradient partitioning of the first node and the gradient partitioning of the third node are both obtained by partitioning the first feature gradient partition and the second feature gradient partition using the Boolean vector of the first node; the gradient partitioning of the second node and the gradient partitioning of the fourth node are both obtained by partitioning the first feature gradient partition and the second feature gradient partition using the Boolean vector of the second node.

[0054] The device also includes:

[0055] The weight determination module is used to obtain the node Boolean partition and node gradient partition of the leaf nodes in the converged target decision tree when the target decision tree converges.

[0056] The weight determination module is used to determine the first child node weight partition of a leaf node based on the node Boolean partition and node gradient partition. The first child node weight partition and the second child node weight partition held by the second participant are partitions of the child node weight vector. The child node weight vector is used to characterize the weight parameters of the leaf node.

[0057] The first child node weighted shard is used to update the first cumulative sample prediction shard of the first business sample identifier, and the second child node weighted shard is used to update the second cumulative sample prediction shard of the first business sample identifier; the updated first cumulative sample prediction shard and the updated second cumulative sample prediction shard are used to generate the next decision tree of the target decision tree.

[0058] One embodiment of this application provides a computer device, including: a processor and a memory;

[0059] The processor is connected to a memory, which stores a computer program. When the computer program is executed by the processor, it causes the computer device to perform the method provided in the embodiments of this application.

[0060] One aspect of this application provides a computer-readable storage medium storing a computer program adapted to be loaded and executed by a processor, so that a computer device having the processor performs the method provided in this application.

[0061] One embodiment of this application provides a computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the method provided in this application embodiment.

[0062] Therefore, this embodiment of the application eliminates the need to obtain the intersection business sample identifier between the first business sample identifier and the second business sample identifier, thus avoiding the direct acquisition of the intersection gradient vector. Instead, the first and second participants respectively obtain slices of the intersection gradient vector (i.e., the first intersection gradient slice and the second intersection gradient slice), without directly comparing the first and second business sample identifiers. Furthermore, the first participant can directly perform business sample segmentation on the first business sample identifier based on the first business feature to obtain the first feature data segmentation matrix, and the second participant can directly perform business sample segmentation on the second business sample identifier based on the second business feature to obtain the second feature data segmentation matrix. This allows the first and second participants to perform business sample segmentation locally (i.e., calculate the candidate feature segmentation points locally), so that the first participant does not need to process the second business sample identifier, and the second participant does not need to process the first business sample identifier. It is understandable that the first feature data partitioning matrix and the second feature data partitioning matrix can be used together to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient partition and the second feature gradient partition. Based on the first feature gradient partition and the second feature gradient partition, the target decision tree for the first participant and the second participant can be determined. This can improve the security of the data owned by the first participant and the second participant while realizing business processing of business data through the converged target decision tree. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0064] Figure 1 This is a schematic diagram of a network architecture provided in an embodiment of this application;

[0065] Figure 2 This is a schematic diagram of a data interaction scenario provided in an embodiment of this application;

[0066] Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application;

[0067] Figure 4 This is a schematic flowchart of a module processing method provided in an embodiment of this application;

[0068] Figure 5 This is a flowchart illustrating a data processing method provided in an embodiment of this application;

[0069] Figure 6 This is a schematic diagram of a hash mapping scenario provided in an embodiment of this application;

[0070] Figure 7 This is a flowchart illustrating a data processing method provided in an embodiment of this application;

[0071] Figure 8 This is a schematic diagram of a stealth training process provided in an embodiment of this application;

[0072] Figure 9 This is a schematic diagram of a privacy arrangement scenario provided in an embodiment of this application;

[0073] Figure 10 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;

[0074] Figure 11 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0075] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0076] For details, please see Figure 1 , Figure 1 This is a schematic diagram of a network architecture provided in an embodiment of this application. Figure 1 As shown, the network architecture may include service server 2000, service server 5000, terminal device cluster 3000, and terminal device cluster 4000. Specifically, terminal device cluster 3000 may include one or more terminal devices; the number of terminal devices in terminal device cluster 3000 is not limited here. Figure 1 As shown, the multiple terminal devices may specifically include terminal device 3000a, terminal device 3000b, terminal device 3000c, ..., terminal device 3000n; terminal device 3000a, terminal device 3000b, terminal device 3000c, ..., terminal device 3000n can be directly or indirectly connected to the business server 2000 via wired or wireless communication, so that each terminal device can interact with the business server 2000 through the network connection.

[0077] Specifically, the terminal device cluster 4000 may include one or more terminal devices; the number of terminal devices in the terminal device cluster 4000 will not be limited here. For example... Figure 1 As shown, the multiple terminal devices may specifically include terminal device 4000a, terminal device 4000b, terminal device 4000c, ..., terminal device 4000n; terminal device 4000a, terminal device 4000b, terminal device 4000c, ..., terminal device 4000n can be directly or indirectly connected to the business server 5000 via wired or wireless communication, so that each terminal device can interact with the business server 5000 through the network connection.

[0078] Each terminal device in terminal device clusters 3000 and 4000 can include: smartphones, tablets, laptops, desktop computers, intelligent voice interaction devices, smart home appliances (e.g., smart TVs), wearable devices, vehicle terminals, aircraft, and other intelligent terminals with data processing capabilities. For ease of understanding, embodiments of this application can... Figure 1 From the cluster of 3000 terminal devices shown, one terminal device is selected as the first terminal device. The first terminal device can be a terminal device participating in vertical federated learning (i.e., the first participant). Figure 1 From the terminal device cluster 4000 shown, one terminal device is selected as the second terminal device. The second terminal device can be a terminal device participating in vertical federated learning (i.e., a second participant). For example, in the embodiments of this application, a terminal device can be selected as the second terminal device. Figure 1 The terminal device 3000c shown is used as the first terminal device, and will Figure 1 The terminal device 4000n shown serves as the second terminal device. The first terminal device and the second terminal device can be directly or indirectly connected to the network via wired or wireless communication to facilitate data interaction.

[0079] Specifically, business server 2000 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Business server 5000 can also be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0080] The first terminal device can obtain a first service sample identifier and feature data of a first service feature for the first service sample identifier from the service server 2000. The first service sample identifier can be an identifier (Identity Document, ID) that uniquely identifies a sample in the first terminal device. The second terminal device can obtain a second service sample identifier and feature data of a second service feature for the second service sample identifier from the service server 5000. The second service sample identifier can be an identifier that uniquely identifies a sample in the second terminal device. For ease of understanding, this embodiment uses the example of the first terminal device being the tag holder (i.e., the first terminal device may also include a service tag for the first service sample identifier) ​​and the second terminal device being a non-tag holder. Furthermore, the number of second terminal devices can be one or more, i.e., the number of second participants can be one or more.

[0081] Federated Learning (FL) is a privacy-preserving distributed machine learning technique that addresses the problem of how to collaboratively train a global model on virtually "aggregated" data when sensitive data is held by multiple independent institutions, groups, or individuals (e.g., first and second participants) while protecting data privacy. Federated Learning can be categorized into Vertical Federated Learning and Horizontal Federated Learning.

[0082] Vertical federated learning can include anonymous vertical federated learning. Anonymous vertical federated learning refers to training a vertical federated learning model without exposing the ID set after intersection. In other words, the sample alignment process is information-free, and feature data is used indiscriminately to train model parameters. In federated learning practice, the vertical federated XGBoost (eXtreme Gradient Boosting) algorithm (i.e., the anonymous vertical federated XGBoost training algorithm) is the most widely used federated machine learning algorithm (i.e., the vertical federated learning model training algorithm). XGBoost is an enhanced version of gradient boosting decision trees. XGBoost uses classification trees or regression trees as the base classifier, offering higher scalability and faster, more efficient performance.

[0083] SMPC (Secure Multi-party Computation) is a technique for securely solving a global problem when a secret is distributed among multiple parties (e.g., a first party and a second party). SS (Secret Sharing) works by distributing the secret among the parties, each receiving a portion of the secret, known as a share. Only when a certain number of parties aggregate the shares can the secret be recovered.

[0084] Secret sharing can include arithmetic secret sharing and Boolean secret sharing. Secret sharing can be used to generate secret slices, arithmetic secret sharing can be used to generate arithmetic slices, and Boolean secret sharing can be used to generate Boolean slices. Arithmetic secret sharing: This involves sharing the integer ring Z... p (p(i.e. 2) λ () is the modulus, for example, p equals 2 128 (i.e., modulo 2) 128 Adding ), indicates that the maximum value of the integer is 2. 128 -1, at this time, the addition operation refers to Z. p The addition and subtraction operations on the Z-axis refer to the Z-axis. p The number x in the subtraction problem is split into two parts (i.e., arithmetic parts). and It is held by two parties, namely Similarly, it can be seen that Arithmetic piecewise partitioning has an additive homomorphism, i.e. Boolean Secret Sharing: Splitting a Boolean value x into two parts (i.e., Boolean partitioning) based on XOR. and It is held by two parties, namely In other words, if the random number is secretly shared using XOR, the result is called a Boolean slice; if it is secretly shared using arithmetic addition, the result is called an arithmetic slice. Boolean slices and arithmetic slices can be collectively referred to as secret slices (or simply slices).

[0085] It should be understood that, for the first participant (i.e., P0) and the second participant (i.e., P1), the embodiments of this application can use <·> to represent fragmentation, where the fragment superscript A represents an arithmetic fragment, the fragment superscript B represents a Boolean fragment, the fragment subscript 0 indicates that the fragment is held by the first participant, and the fragment subscript 1 indicates that the fragment is held by the second participant; when the fragment subscript is omitted, it indicates that both the first and second participants need to perform the calculation. In other words, <x> A This represents an arithmetic partition of x, where each participant has a random value, and these two random values ​​are in Z. p The addition on is x; <x> B This represents a Boolean partition of x, where each of the participating parties has a random bit, and the XOR of these two random bits is x.

[0086] Multi-party secure computation operators: (1) Two-party piecewise addition (ADD), here using " <x> A + <y> A "Indicates: Input" <x> A and <y> A Output<x+y> A (2) Partial multiplication (MUL), here we use " <x> A · <y> A "Indicates: Input" <x> A and <y> A Output<x·y> A If the input is two matrix partitions, this is also represented as element-wise multiplication, that is, multiplying the partition elements of the same row and column of the two matrices to obtain a matrix partition of the same size. (3) Two-way partition division (DIV), here using " <x> A / <y> A "Indicates: Input" <x> A and <y> A Output <x y> A (4) Compare the two-part segments, here using "( <x> A , <y> A ")" indicates: Input <x> A and <y> A Output<x≤y> B That is, if x ≤ y is true, output 1; otherwise, output 0. <x> A and v, output<x≤v> B If x≤v is true, output 1; otherwise, output 0. (5) Maximum value of two-sided partition (MAX), here we use "Max( <x> A ")" means: Input a set of numbers <x> A Output the index / subscript of the largest number among these segments. A (6) AND operation on both sides, here we use "AND ( <x> B , <y> B ")" indicates: Input <x> B and <y> B Output<x&y> B Another way is to input <x> B and y, output<x&y> B When x and y are vectors, it means that the components are ANDed. (7) Two-way piecewise XOR, here we use "XOR( <x> B , <y> B ")" indicates: Input <x> B and <y> B Output (8) Multiple choice, here we use "MUX" <x> B , <y> A ")" indicates: Input <x> B and <y> A Output<x·y> A That is, if x = 0, output the y corresponding to the slice; otherwise, output 0 corresponding to the slice. When both x and y are vectors, it means that the component performs multiplexing. When x is a vector and y is a scalar, it means that each element of x and y performs multiplexing. (9) Two-way slice subtraction (SUB), here we use " <x> A - <y> A "Indicates: Input" <x> A and <y> A Output <x-y> A (10) Two-sided Sigmoid: Input <x> A Output<sigmoid(x)> A Here, Sigmoid(x) = 1 / (1+exp(-x)).

[0087] Understandably, the aforementioned network framework can be applied to various business scenarios, including finance (e.g., improving credit risk control), healthcare (e.g., multinational medical research), autonomous driving (e.g., improving the driving experience), and multimedia data recommendation (e.g., video recommendation). For example, in the healthcare context, rare diseases are frequently encountered in medical research. Currently, practical difficulties exist, such as samples being scattered across different hospitals, which significantly hinders diagnosis and treatment. Vertical federated learning can integrate samples from different hospitals, enabling the diagnosis and treatment of rare diseases. In the financial context, for example, credit risk control for micro and small enterprises can utilize federated transfer learning. This leverages existing models from previous applications by financial institutions (e.g., credit models for medium and large enterprises or marketing models for micro and small enterprises) to improve application effectiveness. In the autonomous driving scenario, the interaction between the vehicle and its system environment can be learned, supplemented by other urban information (e.g., city cameras, traffic lights, and future smart roads). Vertical federation can better integrate information from different sources while protecting privacy, improving the autonomous driving experience. Finally, in the multimedia data recommendation scenario, data from different platforms can interact and learn, thereby improving the accuracy of multimedia data recommendations.

[0088] For better understanding, please refer to [link / reference]. Figure 2 , Figure 2 This is a schematic diagram illustrating a data interaction scenario provided in an embodiment of this application. For example... Figure 2 The terminal device 20a shown can be the one described above. Figure 1 The first participant in the corresponding embodiment, such as Figure 2 The terminal device 20b shown can be the one described above. Figure 1 In the corresponding embodiment, the second participant, the first participant, and the second participant can jointly participate in vertical federated learning. Terminal device 20a may include a first service sample identifier and feature data of the first service sample identifier in a first service feature, while terminal device 20b may include a second service sample identifier and feature data of the second service sample identifier in a second service feature.

[0089] like Figure 2 As shown, terminal device 20a can obtain a first feature data partitioning matrix for a first service feature of a first service sample identifier, and terminal device 20b can obtain a second feature data partitioning matrix for a second service feature of a second service sample identifier. The first feature data partitioning matrix represents the partitioning result obtained by terminal device 20a based on the first service feature to partition the first service sample identifier into service samples, and the second feature data partitioning matrix represents the partitioning result obtained by terminal device 20b based on the second service feature to partition the second service sample identifier into service samples.

[0090] like Figure 2 As shown, terminal device 20a can obtain the number of identifiers corresponding to the first service sample identifier (i.e., the first identifier number), and terminal device 20b can obtain the number of identifiers corresponding to the second service sample identifier (i.e., the second identifier number). Terminal device 20a can provide the first identifier number, and terminal device 20b can provide the second identifier number. In this way, the first identifier number and the second identifier number can be used for comparison (i.e., privacy labeling) to generate a comparison result (i.e., privacy comparison result).

[0091] Understandably, if the comparison result indicates that the number of first identifiers is less than or equal to the number of second identifiers, terminal device 20b can perform fragmentation processing on the second feature data partitioning matrix to obtain first feature data partitioning fragments and second feature data partitioning fragments, and then send the first feature data partitioning fragment to terminal device 20a. Here, the first feature data partitioning fragment and the second feature data partitioning fragment are fragments of the second feature data partitioning matrix. Optionally, if the comparison result indicates that the number of first identifiers is greater than the number of second identifiers, terminal device 20a can perform fragmentation processing on the first feature data partitioning matrix to obtain third feature data partitioning fragments and fourth feature data partitioning fragments, and then send the fourth feature data partitioning fragment to terminal device 20b. Here, the third feature data partitioning fragment and the fourth feature data partitioning fragment are fragments of the first feature data partitioning matrix.

[0092] Furthermore, such as Figure 2 As shown, terminal device 20a can obtain a first intersection gradient slice associated with both the first service sample identifier and the second service sample identifier, and terminal device 20b can obtain a second intersection gradient slice associated with both the first service sample identifier and the second service sample identifier. The first and second intersection gradient slices are slices of the intersection gradient vector, which is used to characterize the sample gradient of the intersection service sample identifier between the first and second service sample identifiers.

[0093] It is understandable that the first feature data partitioning matrix and the first feature data partitioning slice held by terminal device 20a, and the second feature data partitioning slice held by terminal device 20b, can be used to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient slice and the second feature gradient slice associated with the intersection gradient vector. In other words, the first feature data partitioning matrix and the second feature data partitioning matrix can be used to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient slice and the second feature gradient slice associated with the intersection gradient vector.

[0094] like Figure 2 As shown, terminal device 20a can acquire a first feature gradient slice, and terminal device 20b can acquire a second feature gradient slice. The first and second feature gradient slices can be used to determine the optimal splitting result for the first and second business features. The optimal splitting result is used to partition the node to be partitioned in the target decision tree, generating child nodes (e.g., the first and second child nodes) of the node to be partitioned. Then, based on the node to be partitioned, the first child node, and the second child node, a target decision tree is generated for the first and second participants. The converged target decision tree is used for business processing of the business data.

[0095] For example, business data could be symptom information in a medical scenario. A converged target decision tree can be used to process this symptom information, and the result can be supplementary suggestions corresponding to the symptom. These suggestions can then be provided to doctors to aid in diagnosis and treatment. Similarly, business data could be vehicle driving data in an autonomous driving scenario. A converged target decision tree can be used to process this data (e.g., vehicle driving status analysis), and the result can be the corresponding driving status, thus enabling autonomous driving. Furthermore, business data could be operational data in a multimedia data recommendation scenario. A converged target decision tree can be used to process this data (e.g., classification label recognition), and the result can be the corresponding classification label, enabling multimedia data recommendations based on that label.

[0096] For ease of understanding, this application uses a multimedia data recommendation scenario as an example. For instance, the first and second business sample identifiers can be object identifiers (i.e., user identifiers). The first business feature can be the video business feature (e.g., hobbies) of the object identifier in the video client (i.e., the first participant), and the second business feature can be the news business feature (e.g., comments) of the object identifier in the news client (i.e., the second participant). By performing longitudinal federated learning on the feature data of the object identifier in the video client and the news client, a converged first decision tree and a converged second decision tree can be generated. Thus, by inputting operation data containing video business features and news business features, classification labels for the business data can be determined, thereby achieving multimedia data recommendation for the business objects corresponding to the business data. In the specific embodiments of this application, the video business features and news business features involved, when applied to specific products or technologies, require user permission or consent, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0097] Therefore, the embodiments of this application can perform business sample partitioning by the first participant and the second participant respectively on their respective business sample identifiers. Then, based on the partitioning results (i.e., the first feature data partitioning matrix and the second feature data partitioning matrix), gradient partitioning is performed on the first intersection gradient partition owned by the first participant and the second intersection gradient partition owned by the second participant, resulting in the first feature gradient partition and the second feature gradient partition. The first feature gradient partition and the second feature gradient partition can be used to generate a target decision tree, thus eliminating the need for direct comparison of the first business sample identifier and the second business sample identifier, thereby improving the security of the data owned by the first participant and the second participant respectively.

[0098] Further, please see Figure 3 , Figure 3 This is a flowchart illustrating a data processing method provided in an embodiment of this application. The method can be executed by a first terminal device (i.e., a first participant), by a second terminal device (i.e., a second participant), or by both the first and second terminal devices. The first terminal device can be one of the aforementioned... Figure 2 The terminal device 20a in the corresponding embodiment can be the second terminal device described above. Figure 2 The corresponding embodiment is terminal device 20b. For ease of understanding, this application embodiment is described using the method executed by a first terminal device as an example. The data processing method may include the following steps S101-S104:

[0099] Step S101: Obtain the number of first identifiers corresponding to the first business sample identifier of the first participant;

[0100] The first identifier quantity (i.e. the number of first business sample identifiers) is used to compare with the second identifier quantity (i.e. the number of second business sample identifiers) corresponding to the second business sample identifiers of the second participant in the vertical federated learning (i.e., the first identifier quantity provided by the first participant and the second identifier quantity provided by the second participant can be compared) to obtain the comparison result.

[0101] It should be understood that, in this embodiment of the application, the number of first identifiers and the number of second identifiers can be compared using Compare(n0,n1) (i.e., privacy comparison). n0 can represent the number of first identifiers, n1 can represent the number of second identifiers, and the number of first identifiers and the number of second identifiers can represent the number of samples. The number of samples is equal to the number of business sample identifiers, and one sample can have one business sample identifier. During the privacy comparison process, the first participant cannot know the number of second identifiers, and the second participant cannot know the number of first identifiers.

[0102] If the number of first identifiers is less than or equal to the number of second identifiers, this embodiment of the application can generate a comparison result indicating that the number of first identifiers is less than or equal to the number of second identifiers; optionally, if the number of first identifiers is greater than the number of second identifiers, this embodiment of the application can generate a comparison result indicating that the number of first identifiers is greater than the number of second identifiers.

[0103] Specifically, the number of first business features identified for the first business sample is M1, and the number of second business features identified for the second business sample is M2. Here, M1 and M2 can both be positive integers, and they can be the same or different. The M1 first business features include first business feature V. d Here, d can be a non-negative integer less than M1.

[0104] Step S102: If the comparison result indicates that the number of first identifiers is less than or equal to the number of second identifiers, then obtain the first feature data partitioning matrix for the first business feature of the first business sample identifier, and the first feature data partitioning segment for the second business feature of the second business sample identifier.

[0105] Specifically, the first participating party can obtain the first business feature V for the identifier of the first business sample. d Feature data, based on the first business feature V d The feature data is used to divide the first business sample identifier into business samples, generating the first business feature V. d The feature data is targeted at the bucketing indicator vector for feature bucketing. Further, the first participant can determine the first feature data partitioning matrix (B) for the M1 first business features targeting the bucketing indicator vectors for each of the M1 first business features, as a first feature data partitioning matrix (B) for the M1 first business features targeting the first business sample identifier. The first feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the first participant based on the first business features to partition the first business sample identifier. Further, the first participant can obtain the first feature data partitioning slices for the M2 second business features targeting the second business sample identifier. The first feature data partitioning slices (i.e. ) and the second feature data held by the second participant are divided into fragments (i.e. The first feature data partitioning segment is a fragment of the second feature data partitioning matrix. The second feature data partitioning matrix is ​​used to represent the partitioning result obtained by the second participant in partitioning the second business sample identifier based on the second business features. The first feature data partitioning segment and the second feature data partitioning segment are obtained by the second participant in partitioning the second feature data partitioning matrix.

[0106] The specific process by which the second participant obtains the second feature data partitioning matrix for the second business feature of the second business sample identifier can be found in the description of the first participant obtaining the first feature data partitioning matrix for the first business feature of the first business sample identifier, and will not be repeated here.

[0107] Among them, for the first business feature V d The number of feature buckets is K, where K can be a positive integer. It should be understood that the first participant, based on the first business feature V... d The feature data is used to divide the first business sample identifier into business samples, generating the first business feature V. d The specific process of bucketing the feature data for feature bucketing can be described as follows: The first participant can assign the first business feature V to the bucketing indicator vector. d The feature data is distributed into K feature buckets. These K feature buckets include feature bucket K... u Here, u can be a non-negative integer less than K. Further, if the first business sample identifier is in the first business feature V... d The feature data below belongs to feature bucket K u Then the first participant can determine the data presence parameter (e.g., the data presence parameter is 1) as the first business sample identifier in the first business feature V. d Next, for feature binning K u The binning indicator parameter; optionally, if the first business sample identifier is in the first business feature V d The feature data below does not belong to feature bucket K u Then the first participant can determine the absence of data parameter (e.g., the absence of data parameter is 0) as the first business sample identifier in the first business feature V. d Next, for feature binning K u The binning indicator parameters. Furthermore, the first participant can assign each first business sample identifier to a feature bin K. u The bucketing indicator parameters are determined as the first business feature V. d Feature data for feature binning K u Bucket indicator vector.

[0108] In other words, for the i-th sample (i.e., the sample identified by the i-th first business sample), the first participant can calculate feature j (e.g., the first business feature V). d The kth bucket [xl] j,k ,xr j,k The partitioning state b) i,j,k =1{xl j,k ≤x <xr j,k }. Among them, xl j,k ,xr j,k The first business feature V can be represented. d The feature data, 1{·}, indicates that the result is 1 when the expression inside the parentheses is true and 0 otherwise, that is, the feature data x of the i-th sample under feature j belongs to [xl]. j,k ,xr j,k When ), the feature data x of the i-th sample under feature j is divided into the k-th bucket.

[0109] Among them, the M1 first service features can each correspond to the same number of feature buckets (e.g., K feature buckets), or they can each correspond to different numbers of feature buckets (e.g., 2 feature buckets for the first service feature V1 and 3 feature buckets for the first service feature V2). For ease of understanding, this application embodiment uses the example of all M1 first service features corresponding to the same number of feature buckets (i.e., K feature buckets) for illustration.

[0110] Therefore, the first business characteristic V d The number of bucketing indicator vectors for feature data bucketing can be K. The dimension of the bucketing indicator vectors (e.g., 1.27*n0) is associated with the first business sample identifier. The first feature data partitioning matrix for the M1 first business features of the first business sample identifier is composed of K*M1 bucketing indicator vectors, that is, the dimension of the first feature data partitioning matrix can be K*M1*(1.27*n0). Similarly, the dimension of the second feature data partitioning matrix can be K*M2*(1.27*n0).

[0111] Step S103: Obtain the first intersection gradient slice that is associated with both the first business sample identifier and the second business sample identifier;

[0112] Among them, the first intersection gradient slice and the second intersection gradient slice held by the second participant are slices of the intersection gradient vector. The intersection gradient vector is used to characterize the sample gradient of the intersection business sample identifier between the first business sample identifier and the second business sample identifier.

[0113] The specific process of the first participant obtaining the first intersection gradient fragment and the second participant obtaining the second intersection gradient fragment can be found below. Figure 5 The description of steps S1031-S1033 in the corresponding embodiment.

[0114] Step S104: Obtain the first feature gradient slice associated with the intersection gradient vector.

[0115] The first feature data partitioning matrix and the second feature data partitioning matrix are used to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient partition and the second feature gradient partition held by the second participant. The first feature gradient partition and the second feature gradient partition are partitions of the feature gradient vector. The first feature gradient partition and the second feature gradient partition are used to determine the optimal splitting result for the first business feature and the second business feature. The optimal splitting result is used to generate the target decision tree for the first participant and the second participant. The converged target decision tree is used to perform business processing on the business data.

[0116] The first business sample identifier and the second business sample identifier together include M business features belonging to different business dimensions (i.e., M equals the sum of M1 and M2), where M can be a positive integer greater than 1; the M business features include business feature V. c Here, c can be a non-negative integer less than M. Wherein, if business feature V... c If it belongs to the first business feature, then the first feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector to obtain the data for business feature V. c First feature gradient slicing and targeting business feature V c The second feature gradient partitioning. Optionally, if the business feature V c If it belongs to the second business feature, then the second feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector to obtain the data for business feature V. c First feature gradient slicing and targeting business feature V c The second feature gradient slice. Specifically, the first feature gradient slice for M business features is used to form the first feature gradient slice held by the first participant; the second feature gradient slice for M business features is used to form the second feature gradient slice held by the second participant.

[0117] Among them, business feature V c The number of feature buckets can be K, where K can be a positive integer. Wherein, if the business feature V... c If the second feature belongs to the second business sample identifier, then the second feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector, resulting in a first feature gradient sub-partition for K feature buckets and a second feature gradient sub-partition for K feature buckets; the first feature gradient sub-partition for K feature buckets is used to form the business feature V. c The first feature gradient partitioning, and the second feature gradient sub-partitioning for K feature buckets, are used to form a subset of the business feature V. c The second feature gradient partitioning. Optionally, if the business feature V c If the first feature belongs to the first business feature identified by the first business sample, then the first feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector, resulting in a first feature gradient sub-partition for K feature buckets and a second feature gradient sub-partition for K feature buckets; the first feature gradient sub-partition for K feature buckets is used to form a subset of the data for business feature V. c The first feature gradient partitioning, and the second feature gradient sub-partitioning for K feature buckets, are used to form a subset of the business feature V. c The second feature gradient slice.

[0118] In other words, for feature j (e.g., business feature V) c Given the candidate split point k, this embodiment of the application can calculate the initial bucket gradient value of the root node of the target decision tree. <g i,j,k > A =MUX(t, <G i > B )and <h i,j,k > A =MUX(t, <H i > B ),in, <G i > B and <H i > B It can represent the intersection gradient vector, where t can be selected as the corresponding sample gradient value. <g i,j,k > A and <h i,j,k > A This can represent the first feature gradient slice. <g i,j,k > A and <h i,j,k > A This can represent the gradient partitioning of the second feature. Specifically, when the j-th feature is local (i.e., the j-th feature belongs to the first business feature), t = B. i,j,k (i.e., the first feature data partitioning matrix); when the j-th feature is not local (i.e., the j-th feature belongs to the second business feature), t = i,j,k > A (i.e., the second feature data partitioning matrix).

[0119] Among them, the first feature gradient slice held by the first participant includes M first feature gradient slices for M business features, and for business feature V c The first feature gradient partition comprises K first feature gradient sub-partitions for the K feature buckets. It should be understood that the first participant can generate the business feature V based on the K first feature gradient sub-partitions for the K feature buckets. c The corresponding K-1 first partitioning feature gradient slices. Further, the first participant can generate a first cumulative gradient slice corresponding to itself based on the K first feature gradient sub-slices for the K feature buckets. Among them, the M*KM first partitioning feature gradient slices, the first cumulative gradient slices, the M*KM second partitioning feature gradient slices corresponding to the M business features, and the second cumulative gradient slice corresponding to the second participant are used together to determine the information gain for the M business features and the K feature buckets; the second cumulative gradient slice is generated by the second participant based on the K second feature gradient sub-slices for the K feature buckets; the M*KM second partitioning feature gradient slices corresponding to the M business features include business feature V. c The corresponding K-1 second-partition gradient slices, business feature V c The corresponding K-1 second partition feature gradient slices are generated by the second participant based on the K second feature gradient sub-slices for the K feature buckets; the optimal information gain among the information gains of the M business features and the K feature buckets is used to determine the optimal splitting result for the first business feature and the second business feature. The optimal splitting result includes the optimal splitting feature and the optimal splitting point.

[0120] In other words, K-1 first-partition feature gradient slices can represent the gradient sum of the left and right subtrees, and the gradient sum of the left subtree (first-order gradient sum). and second gradient and The number of subtrees is K-1, and the gradient sum of the right subtree (first-order gradient sum) and second gradient and ​The number of gradient sums is K-1, thus achieving dense-state gradient histogram computation. Specifically, the number of first-order gradient sums in the left subtree is K-1, and the number of second-order gradient sums in the left subtree is K-1; the number of first-order gradient sums in the right subtree is K-1, and the number of second-order gradient sums in the right subtree is K-1. Therefore, the embodiments of this application enable participants in the training process to perform dense-state computation on sample partitioning in a sliced ​​state.

[0121] It should be understood that the gradient sum of the K-1 left subtrees is determined by the K first feature gradient sub-partitions, and the gradient sum of the K-1 right subtrees is determined by the K first feature gradient sub-partitions. The first feature gradient sub-partition can represent the gradient sum of all samples in one or more feature buckets for a certain feature (i.e., the first feature gradient sub-partition can represent the sum of the first feature gradient sub-partitions). For example, the number of feature buckets can be 3 (i.e., K equals 3), and the 3 feature buckets can include feature bucket K0, feature bucket K1, and feature bucket K2. For example, in this embodiment, the first feature gradient sub-partition for feature bucket K0 can be determined as the gradient sum of the left subtree, and the first feature gradient sub-partitions for feature bucket K1 and feature bucket K2 can be summed to obtain the gradient sum of the right subtree. For example, the gradient sub-slice for feature bucket K0 and the gradient sub-slice for feature bucket K1 are summed to obtain the gradient sum of the left subtree, and the gradient sub-slice for feature bucket K2 is determined as the gradient sum of the right subtree. In other words, the first participant can divide the K feature buckets into two parts according to the order of the K feature buckets. One part of the feature buckets is used to calculate the gradient sum of the left subtree, and the other part of the feature buckets is used to calculate the gradient sum of the right subtree. The K feature buckets can have K-1 partitioning methods, thereby generating K-1 gradient sums of the left subtree and K-1 gradient sums of the right subtree.

[0122] Furthermore, the first cumulative gradient slice can represent the sum of the cumulative gradients of all samples (first-order gradient sum). <g P > A =∑ i ∑ k <g i,0,k > A and second gradient and <h P > A =∑ i ∑ k <h i,0,k > A The cumulative gradient sum is the result of accumulating all samples, calculated based on different features j. <g P > A and <h P > A They are equal. Here, we take j=0 as an example to illustrate that the first cumulative gradient slice can represent the sum of the gradients of all samples in all feature buckets for a certain feature.

[0123] In this embodiment of the application, the payout score for each split point can be calculated: Thus find the j-th opt The k-th feature (i.e., the optimal splitting feature) opt One split value (i.e., the optimal split point): <j opt > A , That is, the optimal information gain can be the maximum value of the information gains for M business features and K feature buckets, the optimal split feature can be the business feature corresponding to the optimal information gain, and the optimal split point can be the split point corresponding to the optimal information gain (the split point can represent the split value of adjacent feature buckets).

[0124] At this point, the participant with the optimal splitting feature (e.g., the first participant) can recover j in plaintext. opt and k opt The other participant (e.g., the second participant) will not receive any j opt and k opt The information provided allows the participants of the optimal splitting feature to represent the participants to which the optimal splitting feature belongs, thereby enabling dense-state search for the optimal splitting point. Optionally, the first participant can generate random splitting vectors for the optimal splitting feature and the optimal splitting point, add the random splitting vectors together to generate a synchronous splitting vector, and then send the synchronous splitting vector to the second participant. Since the random splitting vector and the synchronous splitting vector are fragments of the vector composed of the optimal splitting feature and the optimal splitting point, this embodiment of the application can achieve the joint storage of fragments of the vector composed of the optimal splitting feature and the optimal splitting point by the first and second participants. For example, the vector composed of the optimal splitting feature and the optimal splitting point can be (3,4), the random splitting vector can be (1,3), and the synchronous splitting vector can be (2,1).

[0125] In this embodiment, different feature identifiers can be assigned to the first business feature and the second business feature. The feature identifier of the first business feature can be a first feature identifier, and the feature identifier of the second business feature can be a second feature identifier. In this way, when determining the optimal splitting feature through information gain, the participant to which the optimal splitting feature belongs can be determined based on the feature identifier of the optimal splitting feature. For example, if the feature identifier of the optimal splitting feature belongs to the first feature identifier, it is determined that the optimal splitting feature belongs to the first participant.

[0126] For easier understanding, please refer to Figure 4 , Figure 4 This is a schematic flowchart illustrating a module processing method provided in an embodiment of this application. For example... Figure 4 As shown, the participants in this embodiment may include a first participant 40a and a second participant 40b. The first participant 40a may include three modules: a sample partitioning module, a circuit privacy intersection module, and a stealth training module. The second participant 40b may include three modules: a sample partitioning module, a circuit privacy intersection module, and a stealth training module.

[0127] like Figure 4 As shown, the sample partitioning module can locally calculate the split point situation and partition the samples according to the situation. That is, the first participant 40a can obtain the first feature data partitioning matrix, and the second participant 40b can obtain the second feature data partitioning matrix. Further, the circuit privacy intersection module can input the business sample identifier, business label, and sample partitioning into this module to obtain the stealth partitioning intersection result (i.e., feature gradient vector). That is, the first participant 40a can input the first business sample identifier, business label (i.e., the first business sample label), and sample partitioning (i.e., the first feature data partitioning matrix) to obtain the first feature gradient partition; the second participant 40b can input the second business sample identifier, business label (i.e., the second business sample label), and sample partitioning (i.e., the second feature data partitioning matrix) to obtain the second feature gradient partition. Further, the stealth training module can perform stealth XGBoost training on the partitions (i.e., the first feature data partitioning matrix, the first feature gradient partition, and the second feature gradient partition).

[0128] Optionally, if the comparison result indicates that the number of first identifiers is greater than the number of second identifiers, the first participant can obtain a first feature data partitioning matrix for the first business feature of the first business sample identifier, perform fragmentation on the first feature data partitioning matrix to obtain a third feature data partitioning fragment and a fourth feature data partitioning fragment, and then send the fourth feature data partitioning fragment to the second participant. The specific process by which the first participant generates the converged target decision tree based on the third feature data partitioning fragment, and the second participant generates the converged target decision tree based on the fourth feature data partitioning fragment and the second feature data partitioning matrix, can be found in the description above, which will not be repeated here.

[0129] Therefore, this embodiment of the application eliminates the need to obtain the intersection business sample identifier between the first business sample identifier and the second business sample identifier, thus avoiding the direct acquisition of the intersection gradient vector. Instead, the first and second participants respectively obtain slices of the intersection gradient vector (i.e., the first intersection gradient slice and the second intersection gradient slice), without directly comparing the first and second business sample identifiers. Furthermore, the first participant can directly perform business sample segmentation on the first business sample identifier based on the first business feature to obtain the first feature data segmentation matrix, and the second participant can directly perform business sample segmentation on the second business sample identifier based on the second business feature to obtain the second feature data segmentation matrix. This allows the first and second participants to perform business sample segmentation locally (i.e., calculate the candidate feature segmentation points locally), so that the first participant does not need to process the second business sample identifier, and the second participant does not need to process the first business sample identifier. It is understandable that the first feature data partitioning matrix and the second feature data partitioning matrix can be used together to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient partition and the second feature gradient partition. Based on the first feature gradient partition and the second feature gradient partition, the target decision tree for the first participant and the second participant can be determined. This can improve the security of the data owned by the first participant and the second participant while realizing business processing of business data through the converged target decision tree.

[0130] Further, please see Figure 5 , Figure 5 This is a flowchart illustrating a data processing method provided in an embodiment of this application. The data processing method may include the following steps S1031-S1033, and steps S1031-S1033 are... Figure 3 This is a specific embodiment of step S103 in the corresponding embodiment. In this embodiment, the entity with the smaller sample size can be designated as the receiver of the circuit privacy intersection, and the entity with the larger sample size can be designated as the sender of the circuit privacy intersection, thus achieving anonymity based on the entity with the smaller sample size.

[0131] Step S1031: Obtain the first sample gradient slice that is jointly associated with the first business sample identifier and the second business sample identifier;

[0132] Specifically, if the first participant is the tag holder, then the first participant can obtain the first cumulative sample prediction slice of the first business sample identifier, and the first business tag vector corresponding to the first business sample tag of the first business sample identifier (the first business tag vector represents the vector obtained by sorting the first business sample tags according to the Cuckoo Hash mapping). Based on the first cumulative sample prediction slice and the first business tag vector, the first sample gradient slice jointly associated with the first business sample identifier and the second business sample identifier is determined. The first cumulative sample prediction slice and the second cumulative sample prediction slice held by the second participant are slices of the cumulative sample prediction vector, which is used to represent the historical cumulative sample prediction values ​​when generating the target decision tree through the first business sample identifier. The second participant uses the second cumulative sample prediction slice to determine the second sample gradient slice jointly associated with the first business sample identifier and the second business sample identifier. The first sample gradient slice and the second sample gradient slice held by the second participant are slices of the sample gradient vector, which is used to represent the sample gradient of the first business sample identifier.

[0133] Optionally, if the second participant is a tag holder, the first participant can obtain the first cumulative sample prediction fragment of the first business sample identifier, and the first business tag fragment corresponding to the second business sample tag of the second business sample identifier (i.e., Based on the first cumulative sample prediction shard and the first business label shard, a first sample gradient shard is determined that is jointly associated with the first business sample identifier and the second business sample identifier. Among these, the first business label shard and the second business label shard held by the second participating party (i.e., The first cumulative sample prediction segment and the second cumulative sample prediction segment held by the second participant are segments of the cumulative sample prediction vector, which is used to characterize the historical cumulative sample prediction value when generating the target decision tree through the first business sample identifier. The second participant is used to determine the second sample gradient segment that is associated with both the first business sample identifier and the second business sample identifier based on the second cumulative sample prediction segment and the second business label segment.

[0134] In other words, the first business tag fragment and the second business tag fragment can be Boolean fragments of tags. When the first business sample identifier of the i-th sample of the first participant matches the second business sample identifier of a sample of the second participant, y i = The label of the i-th sample, otherwise y i =0. Where y i It can represent the i-th label in the second business label vector.

[0135] Wherein, if the target decision tree is the first decision tree, the historical cumulative sample prediction value is determined by the initial value (e.g., 0); optionally, if the target decision tree is not the first decision tree, the historical cumulative sample prediction value is obtained by accumulating the outputs of the decision trees before the target decision tree.

[0136] Wherein, the first cumulative sample prediction slice and the second cumulative sample prediction slice can represent the sample prediction value (i.e., the historical cumulative sample prediction value): f here i (X) is the i-th tree (f i The predicted value of ) can be represented by the first cumulative sample prediction slice and the second cumulative sample prediction slice, which can be expressed as the sum of the predicted values ​​of the 1st tree to the (t-1st)th tree. In this case, the target decision tree is the t-th tree. Wherein, if the embodiments of this application are applied to a regression problem, the predicted value can be expressed as: <f i (X)> A = <w l > A Optionally, if the embodiments of this application are applied to a classification problem, the predicted value can be expressed as:

[0137] The first sample gradient slice may include a first-order gradient and a second-order gradient, and the second sample gradient slice may also include a first-order gradient and a second-order gradient, wherein the second-order gradient can be... Optionally, the first participant can obtain a first cumulative sample prediction slice, and the second participant can obtain a second cumulative sample prediction slice. The first and second cumulative sample prediction slices can be used to determine a first sample gradient slice and a second sample gradient slice that are jointly associated with the first business sample identifier and the second business sample identifier. In this case, the first and second sample gradient slices can represent the second derivative gradient.

[0138] Understandably, when the first participant is the label holder, the first-order gradient can be... (Right now (representing the gradient partitioning of the first sample) and (Right now (This represents the gradient slice of the second sample); optionally, when the second participant is the label holder, the first derivative gradient can be... (Right now (representing the gradient partitioning of the first sample) and (Right now (This represents the gradient slicing of the second sample). Here, Y can represent the first business label vector. This can represent the first business tag segment. may represent the second service label fragment.

[0139] It should be understood that the first-order gradient may be the first derivative of the loss function, and the second-order gradient may be the second derivative of the loss function. The embodiments of the present application do not limit the specific form of the loss function. For example, the loss function may be a binary cross-entropy loss function: y may represent a label, may represent a predicted value.

[0140] Step S1032, acquiring a first Boolean intersection fragment corresponding to the first service sample identifier;

[0141] Specifically, the first participant may perform Cuckoo hash mapping on the first service sample identifier to obtain a first hash table corresponding to the first service sample identifier. Further, the first participant may acquire the first Boolean intersection fragment associated with the first hash table and the second hash table. The first hash table and the second hash table are used for hash table matching through an oblivious programmable pseudorandom function to generate the first Boolean intersection fragment and a second Boolean intersection fragment held by the second participant; the second hash table is obtained by the second participant performing hash mapping on the second service sample identifier. In the embodiments of the present application, the hash mapping may be simple hash mapping.

[0142] wherein the first Boolean intersection fragment (i.e., ) and the second Boolean intersection fragment held by the second participant (i.e., ) are fragments of a Boolean intersection vector (i.e., Q), and the Boolean intersection vector is used to indicate an intersection state of the first service sample identifier with respect to the second service sample identifier.

[0143] wherein Cuckoo Hash refers to a method of mapping m elements (i.e., service sample identifiers) to n positions through k hashes (i.e., k hash functions), which requires different elements to be mapped to different positions. In simple terms, if the i-th element has been mapped to the s-th position under the action of the t-th (t < k) hash, when the j-th (j > i) element is also mapped to the s-th position, a conflict occurs at the s-th position. Then the j-th element is mapped to the s-th position, and the i-th element is re-mapped to a new position under the action of the (t+1)-th hash. If a position conflict occurs, this process is repeated.

[0144] Simple hash mapping (or simply hash mapping) refers to a method of mapping m elements (i.e., business sample identifiers) to n locations using k hashes (i.e., k hash functions). Location collisions are possible, meaning different elements can be mapped to the same location. Simple hash mapping can map different data to the same location, and by using at least two hash functions, it maps the same data to at least two locations. This eliminates the need for additional processing (i.e., resolving hash collisions) when collisions occur, thus improving the efficiency of hash mapping. It should be understood that this application does not limit the number of hash functions. For ease of understanding, this application uses two hash functions as an example, meaning each business sample identifier can be mapped to two locations.

[0145] The first hash table and the second hash table have the same hash table dimension. The first hash table includes a first hash mapping bucket, and the second hash table includes a second hash mapping bucket with the same hash table dimension as the first hash mapping bucket. If the first hash mapping bucket includes a first business sample identifier, the first random number corresponding to the first business sample identifier in the first hash mapping bucket is obtained by randomly processing the first business sample identifier in the first hash mapping bucket using an unintentionally programmable pseudo-random function. Similarly, if the second hash mapping bucket includes a second business sample identifier, the second random number corresponding to the second business sample identifier in the second hash mapping bucket is obtained by randomly processing the second business sample identifier in the second hash mapping bucket using an unintentionally programmable pseudo-random function. Specifically, the first random number and the second random number are used to perform random number matching between the first participant and the second participant to generate the first Boolean intersection parameter corresponding to the first hash mapping bucket and the second Boolean intersection parameter corresponding to the second hash mapping bucket. The first Boolean intersection parameter and the second Boolean intersection parameter are fragments of the merged Boolean intersection parameter. The merged Boolean intersection parameter is used to indicate the Boolean matching result of the first hash mapping bucket and the second hash mapping bucket (i.e., the first Boolean intersection parameter and the second Boolean intersection parameter can be fragments indicating whether the first business sample identifier is in the intersection). The first Boolean intersection parameter corresponding to each hash table dimension in the first hash table is used to form the first Boolean intersection fragment corresponding to the first business sample identifier, and the second Boolean intersection parameter corresponding to each hash table dimension in the second hash table is used to form the second Boolean intersection fragment corresponding to the first business sample identifier.

[0146] It should be understood that an oblivious programmable pseudo-random function can be an OPPRF protocol. An OPPRF protocol represents a set of inputs X = {(x1,y1),(x2,y2),…,(x...}}. n ,y n )}, the other party inputs x, and after executing the protocol, the other party obtains the value y corresponding to x, satisfying the condition that when x = x i When, y = y i Otherwise, y is a random number. Optionally, embodiments of this application may also use an Oblivious Pseudo-Random Function (OPRF) to randomly process the business sample identifier.

[0147] If the first random number and the second random number are the same, then in this embodiment of the application, the successfully matched parameter (e.g., 1) can be determined as the merged Boolean intersection parameter corresponding to the first hash mapping bucket and the second hash mapping bucket; optionally, if the first random number and the second random number are different, then in this embodiment of the application, the unsuccessfully matched parameter (e.g., 0) can be determined as the merged Boolean intersection parameter corresponding to the first hash mapping bucket and the second hash mapping bucket.

[0148] In other words, the first Boolean intersection partition and the second Boolean intersection partition can be Boolean partitions indicating whether a sample is in the intersection. When the first business sample identifier of the i-th sample of the first participant matches the second business sample identifier of a sample of the second participant, q i =1, otherwise q i =0. Where, q i It can represent the i-th merged Boolean intersection parameter in the Boolean intersection vector.

[0149] For easier understanding, please refer to Figure 6 , Figure 6 This is a schematic diagram illustrating a hash mapping scenario provided in an embodiment of this application. For example... Figure 6 The diagram illustrates how PSI-Circuit (Private Set Intersection Circuit) anonymizes service sample identifiers. The protocol includes a sender and a receiver. The sender (second participant 60b) inserts its own data (service sample identifier) ​​into a regular hash table (second hash table), while the receiver (first participant 60a) fills its own data (service sample identifier) ​​into a cuckoo hash table (first hash table). Here, circuit privacy set intersection means that both participants input sets, and ultimately, they only receive fragment information about the intersection (intersection service sample identifier). This is a Boolean fragment value indicating whether the receiver's data is in the intersection; neither intersection data nor non-intersection data is obtained.

[0150] like Figure 6 As shown, the first participant 60a may include a first business sample identifier. For ease of understanding, we will take an example where the number of first business sample identifiers is 4, such as 2, 3, 7, and 10. The second participant 60b may include a second business sample identifier. For ease of understanding, we will take an example where the number of second business sample identifiers is 5, such as 2, 4, 5, 7, and 9.

[0151] like Figure 6 As shown, the first participant 60a can insert the first business sample identifier into the first hash table using a cuckoo hash, and the second participant 60b can insert the second business sample identifier into the second hash table using a simple hash. The first and second hash tables have the same hash table dimension. The first hash table may include a first hash mapping bucket, and the second hash table may include a second hash mapping bucket. For ease of understanding, this embodiment uses a 6-dimensional hash table for both the first and second hash tables as an example. In this embodiment, "*" can be entered into hash mapping buckets (bins) that are not yet filled with data (i.e., business sample identifiers), where "*" represents a random number.

[0152] like Figure 6 As shown, for example, the first hash mapping bucket can be the fourth hash mapping bucket in the first hash table, and the second hash mapping bucket can be the fourth hash mapping bucket in the second hash table. The first hash mapping bucket can include 2, and the second hash mapping bucket can include 2 and 9. In this case, the first random number corresponding to 2 can be 12, the second random number corresponding to 2 can be 12, and the second random number corresponding to 9 can be 19. When the first random number and the second random number are matched, the first random number 12 is equal to the second random number 12. In this embodiment, the successfully matched parameter can be determined as the merged Boolean intersection parameter corresponding to the first hash mapping bucket and the second hash mapping bucket, that is, the successfully matched parameter can be determined as the merged Boolean intersection parameter corresponding to the fourth hash mapping bucket; similarly, in this embodiment, the successfully matched parameter can be determined as the merged Boolean intersection parameter corresponding to the second hash mapping bucket, and the unsuccessfully matched parameter can be determined as the merged Boolean intersection parameters corresponding to the first hash mapping bucket, the third hash mapping bucket, the fifth hash mapping bucket, and the sixth hash mapping bucket, respectively.

[0153] like Figure 6 As shown, the first Boolean intersection partition can include 6 first Boolean intersection parameters, and the second Boolean intersection partition can include 6 second Boolean intersection parameters. The first Boolean intersection parameter can... <0> The first and second Boolean intersection parameters can be <0> The XOR result of the first Boolean intersection parameter and the first Boolean intersection parameter is the merged Boolean intersection parameter (i.e., 0). For example, the first Boolean intersection parameter can be 0, and the first Boolean intersection parameter can be 1.

[0154] Step S1033: Obtain the first intersection gradient slice that is associated with both the first business sample identifier and the second business sample identifier.

[0155] The Boolean intersection vector is used to perform gradient filtering on the sample gradient vector, generating a first intersection gradient slice and a second intersection gradient slice held by the second participant. The first and second intersection gradient slices are slices of the intersection gradient vector; the intersection gradient vector is used to represent the sample gradient of the intersection business sample identifier between the first and second business sample identifiers.

[0156] Gradient filtering can be expressed as: <g> A =MUX( <s> B , <g> A )and <h> A =MUX( <s> B , <h> A ), <s> B It can represent Boolean intersection vectors ( <s> B It can include and This can represent the first Boolean intersection partition. It can represent the second Boolean intersection (partition), the right side of the equals sign <g> A and <h> A It can represent the sample gradient vector, the left side of the equals sign. <g> A and <h> A The intersection gradient vector can be represented; here, we do not distinguish between the representation of the sample gradient vector and the intersection gradient vector.

[0157] Therefore, the embodiments of this application can realize local calculation of candidate feature splitting points, obtain local results (i.e., different splitting results of samples under different splitting points), and then send the local results (i.e., sample splitting state) to the circuit privacy intersection module for fragmentation, avoiding the privacy comparison operation with huge overhead in the training phase. In addition, the embodiments of this application can dynamically adjust the roles of the protocol participants based on the number of samples (i.e., the initiator and receiver of circuit privacy intersection are determined according to the sample size of both parties), and reduce the number of occult samples participating in subsequent training by unilateral circuit privacy intersection (i.e., one circuit privacy intersection). (i.e., based on the data of the participant with less sample size, the number of occult samples during training is close to the number of the original data of the party with less sample size. The number of occult samples represents the number of samples after the circuit privacy set intersection. For example, the number of occult samples can be...) Figure 4 (The six corresponding to the embodiments) thereby reducing the computational overhead during training.

[0158] Further, please see Figure 7 , Figure 7 This is a flowchart illustrating a data processing method provided in an embodiment of this application. The method can be executed by a first terminal device (i.e., a first participant), by a second terminal device (i.e., a second participant), or by both the first and second terminal devices. The first terminal device can be one of the aforementioned... Figure 2 The terminal device 20a in the corresponding embodiment can be the second terminal device described above. Figure 2 The corresponding embodiment is terminal device 20b. For ease of understanding, this application embodiment is described using the method executed by a first terminal device as an example. The data processing method may include the following steps S201-S207:

[0159] Step S201: Obtain the first split point partitioning matrix for the first business feature, and the first split point partitioning slice for the second business feature of the second business sample identifier.

[0160] The first split point partitioning matrix is ​​used to characterize the partitioning result (i.e., bucketing state) of the first feature data partitioning matrix in feature bucketing; the first split point partitioning shard and the second split point partitioning shard held by the second participant are shards of the second feature data partitioning matrix, and the second split point partitioning matrix is ​​used to characterize the partitioning result (i.e., bucketing state) of the second feature data partitioning matrix in feature bucketing.

[0161] Wherein, when the j-th feature is local (i.e., the j-th feature belongs to the first business feature), the first split point partition matrix can be represented as BB i,j,k When the j-th feature is not local (i.e., the j-th feature belongs to the second business feature), the second split point partition matrix can be represented as follows: <BB i,j,k > A .

[0162] For example, the number of feature buckets can be three (i.e., K equals 3). These three feature buckets can include feature bucket K0, feature bucket K1, and feature bucket K2. Here, we take a business sample identifier as an example. If the business sample identifier's j-th feature belongs to feature bucket K0, then B... 0,j,0 =1, B 0,j,1 =0, B 0,j,2 =0; If feature bucket K0 and feature bucket K1 are left subtrees, then the first split point partition matrix can include: BB = (1,1,0); If feature bucket K0 is a left subtree, then the first split point partition matrix can include: BB = (1,0,0).

[0163] For example, the number of feature buckets can be three (i.e., K equals 3). These three feature buckets can include feature bucket K0, feature bucket K1, and feature bucket K2. Here, we take four business sample identifiers as an example. If the second business sample identifier belongs to feature bucket K0 for the j-th feature, the third and fourth business sample identifiers belong to feature bucket K1 for the j-th feature, and the first business sample identifier belongs to feature bucket K2 for the j-th feature, then B... i,j,0 = (0,1,0,0), B i,j,1 = (0,0,1,1), B i,j,2 = (1,0,0,0); If feature bucket K0 and feature bucket K1 are left subtrees, then the first split point partition matrix can include: BB = (0,1,1,1); If feature bucket K0 is a left subtree, then the first split point partition matrix can include: BB = (0,1,0,0).

[0164] It should be understood that the optimal splitting result includes the optimal splitting feature (i.e., j) opt ) and the optimal split point (i.e., k) opt The optimal splitting feature (i.e., the feature to be split) can belong to either the first service feature or the second service feature; this application does not limit this. Specifically, when the optimal splitting feature belongs to the first service feature, the first participant (i.e., the receiver of the circuit privacy intersection) can directly obtain the optimal splitting feature and the optimal splitting point, while the second participant (i.e., the initiator of the circuit privacy intersection) cannot directly obtain the optimal splitting feature and the optimal splitting point. Optionally, when the optimal splitting feature belongs to the second service feature, the second participant can directly obtain the optimal splitting feature and the optimal splitting point, while the first participant cannot directly obtain the optimal splitting feature and the optimal splitting point.

[0165] Step S202: If the optimal splitting feature belongs to the first business feature, then obtain the first splitting point partitioning vector corresponding to the optimal splitting result from the first splitting point partitioning matrix, perform XOR processing on the first splitting point partitioning vector through the first random blind partitioning, generate the first synchronous blind partitioning, and send the first synchronous blind partitioning to the second participant.

[0166] Among them, the first randomized blinding partition and the first synchronous blinding partition are the first split point partition vectors (i.e.<BB′> B The first randomized blinded partition (i.e., the blinded vector) has a dimension (e.g., 1.27*n0) that is associated with the number of first identifiers.

[0167] Among them, the first participant can calculate It can represent the partition vector of the first split point corresponding to the optimal split result. This can represent the first randomized blinded partition. This can represent the first synchronous blinded fragment.

[0168] Step S203: If the optimal splitting feature belongs to the second business feature, then the first splitting point is divided into segments by the second random blinding segmentation and XORed to generate a random blinding matrix.

[0169] The random blinding matrix is ​​used to obtain the random blinding vector corresponding to the optimal splitting result from the random blinding matrix through an unintentional transmission protocol; the second participant is used to obtain the splitting point partitioning sub-partition corresponding to the optimal splitting result from the second splitting point partitioning partition, and to perform XOR processing on the splitting point partitioning sub-partition through the random blinding vector to generate the second synchronous blinding partition.

[0170] Wherein, the second randomized blind partition and the second synchronous blind partition are the second split point partition vectors corresponding to the optimal split result in the second split point partition matrix (i.e.,<BB′> B The second randomized blinded partition (i.e., the blinded vector) has a dimension (e.g., 1.27*n0) that is associated with the number of the first identifiers.

[0171] It is understandable that the first and second participants can execute... OT (Oblivious Transfer) is a type of protocol in cryptography widely used in secure multi-party computation. It allows the sender to transmit one of many potential pieces of information to the receiver without knowing the specific information received. Specifically... This means that player A inputs m*K numbers, player B inputs the index i of the desired number, and player B ultimately obtains the i-th number from the m*K numbers, but does not know the other numbers, and player A does not know the number chosen by player B. The first participant outputs the m*K matrix R corresponding to each sample. i (i.e., randomized blinding matrix) Input 1-out-of-(m*K), the second participant will j opt ,k opt Input 1-out-of-(m*K), and the second participant ultimately receives... (i.e., randomized blinding vector). Where, j opt It can represent the optimal splitting feature, k opt It can represent the optimal split point.

[0172] Among them, the first participant can calculate This can represent the partitioning of the data at the first split point. This can represent the second randomized blinded partition, R i,j,k This can represent a randomized blinding matrix. The second participant can compute... and therefore, (Right now This can represent split points dividing into sub-parts. This can represent the second synchronous blinding fragment.

[0173] Step S204: Obtain the first node Boolean partition of the first child node of the node to be partitioned in the target decision tree for the first participant and the second participant, and the second node Boolean partition of the second child node of the node to be partitioned.

[0174] In this structure, the first node Boolean partition and the third node Boolean partition held by the second participant are partitions of the first node Boolean vector; the second node Boolean partition and the fourth node Boolean partition held by the second participant are partitions of the second node Boolean vector. The first node Boolean vector represents the partitioning state where the intersection business sample identifier is assigned to the first child node, and the second node Boolean vector represents the partitioning state where the intersection business sample identifier is assigned to the second child node. The first child node can be the left subtree of the node to be partitioned, and the second child node can be the right subtree of the node to be partitioned. The first node Boolean vector can also be called the indicator mask of the left subtree (i.e.,...). l > B The Boolean vector of the second node can also be called the indicator mask of the right subtree (i.e., r > B By using the Boolean vectors of the first and second nodes, the intersection business sample identifiers can be divided into the left and right subtrees, thereby achieving dense split nodes. ​​

[0175] Specifically, if the optimal splitting feature belongs to the first business feature, then the first node Boolean partition, the second node Boolean partition, the third node Boolean partition, and the fourth node Boolean partition are obtained by jointly performing Boolean partitioning on the Boolean intersection vector by the first randomized blind partition and the first synchronous blind partition. That is, the first randomized blind partition and the first synchronous blind partition can be used to partition the node to be partitioned, obtaining the first child node and the second child node. Optionally, if the optimal splitting feature belongs to the second business feature, then the first node Boolean partition, the second node Boolean partition, the third node Boolean partition, and the fourth node Boolean partition are obtained by jointly performing Boolean partitioning on the Boolean intersection vector by the second randomized blind partition and the second synchronous blind partition. That is, the second randomized blind partition and the second synchronous blind partition can be used to partition the node to be partitioned, obtaining the first child node and the second child node.

[0176] The first node Boolean partition and the third node Boolean partition are both obtained by performing a Boolean partition on the Boolean intersection vector jointly by the first randomized blind partition and the first synchronous blind partition (i.e., the indicator mask of the left subtree is...). l > B =AND(<BB′> B , <s> B The second node Boolean partition and the fourth node Boolean partition are both obtained by XORing the first node Boolean vector and the Boolean intersection vector (i.e., ...). r > B =XOR( l > B , <s> B )).in, <s> B It can represent Boolean intersection vectors.<BB′> B This can represent the first randomized blinding partition and the first synchronous blinding partition. l > B It can represent the Boolean vector of the first node. r > B It can represent the Boolean vector of the second node.

[0177] Please see again. Figure 6 For example, the intersection Boolean vector can be (0,1,0,1,0,0), the first node Boolean vector can be (0,1,0,0,0,0), and the second node Boolean vector can be (0,0,0,1,0,0). This means that the node to be divided can include 2 and 7, the first child node can include 7, and the second child node can include 2. Similarly, the intersection Boolean vector can be (0,1,0,1,0,0), the first node Boolean vector can be (0,0,0,1,0,0), and the second node Boolean vector can be (0,1,0,0,0,0). This means that the node to be divided can include 2 and 7, the first child node can include 2, and the second child node can include 7.

[0178] Step S205: Obtain the first node gradient slice of the first child node and the second node gradient slice of the second child node;

[0179] In this design, the gradient partition of the first node and the gradient partition of the third node held by the second participant are partitions of the gradient vector of the first node; the gradient partition of the second node and the gradient partition of the fourth node held by the second participant are partitions of the gradient vector of the second node. The gradient vector of the first node is used to represent the sample gradient of the first child node (thus realizing the gradient update of the left subtree), and the gradient vector of the second node is used to represent the sample gradient of the second child node (thus realizing the gradient update of the right subtree). Both the gradient partition of the first node and the gradient partition of the third node are obtained by partitioning the first feature gradient partition and the second feature gradient partition (i.e., the feature gradient vector) using the Boolean vector of the first node. and The second node gradient partitioning and the fourth node gradient partitioning are both obtained by partitioning the first feature gradient partitioning and the second feature gradient partitioning (i.e., feature gradient vectors) using the second node Boolean vector. and ).

[0180] Among them, when but otherwise and The gradient partitioning is the same as that of the parent node. Similarly, when but ​​otherwise and The gradient is the same as that of the parent node.

[0181] Step S206: When the target decision tree converges, obtain the node Boolean partition and node gradient partition of the leaf nodes in the converged target decision tree.

[0182] Wherein, if the first child node and the second child node are leaf nodes in the converged target decision tree, then the node Boolean partition of the leaf node may include the first node Boolean partition (i.e., the node Boolean partition of the first child node) and the second node Boolean partition (i.e., the node Boolean partition of the second child node), and the node gradient partition of the leaf node may include the first node gradient partition (i.e., the node gradient partition of the first child node) and the second node gradient partition (i.e., the node gradient partition of the second child node). Optionally, if the first child node and the second child node are not leaf nodes in the converged target decision tree, then in this embodiment, the node partitioning of the first child node and the second child node may continue until the target decision tree converges.

[0183] Understandably, when the target decision tree converges, both the first and second participants can obtain all nodes in the converged target decision tree (for a tree of depth d, the converged target decision tree requires traversing 2...). d -1 node) of the indicator mask fragment (i.e. <s> B The indicator mask fragment can be used to indicate the samples contained in a node, where the sample order follows the receiver order of the circuit privacy intersection. For example, the indicator mask fragment of the root node can be... <q> B This indicates that the root node contains all samples in the intersection (i.e., the samples corresponding to the intersection business sample identifiers).

[0184] It is understood that, in this embodiment of the application, the convergence of the target decision tree can be determined when the level (i.e., depth) of the target decision tree reaches a level threshold, thus obtaining the converged target decision tree. However, this embodiment of the application does not limit the specific value of the level threshold; for example, the level threshold can be equal to 4 (i.e., the convergence of the target decision tree is determined when the target decision tree has 4 levels).

[0185] Step S207: Determine the weight partition of the first child node of the leaf node based on the node Boolean partition and node gradient partition of the leaf node.

[0186] In this context, the first child node weight slice and the second child node weight slice held by the second participant are slices of the child node weight vector, where the child node weight vector represents the weight parameters of the leaf nodes. The first and second child node weight slices can be represented as follows: You can perform weighted sharding for the first child node. The second child node can be weighted and partitioned. Here, η is the learning rate, used to reduce the influence of the target decision tree and make the iteration more stable.

[0187] The first child node weight shard is used to update the first cumulative sample prediction shard of the first business sample identifier, and the second child node weight shard is used to update the second cumulative sample prediction shard of the first business sample identifier. If the target decision tree is not the last decision tree, the updated first cumulative sample prediction shard and the updated second cumulative sample prediction shard are used to generate the next decision tree of the target decision tree.

[0188] It should be understood that the first and second participants can jointly store the converged target decision tree. That is, the converged target decision tree can include the first decision subtree held by the first participant and the second decision subtree held by the second participant. Nodes at the same position in the first and second decision subtrees have a one-to-one correspondence (i.e., a node in the first decision subtree can correspond to a node in the second decision subtree). For example, the first and second decision subtrees can include nodes at the same position. Taking node 1 in the first decision subtree and node 2 in the second decision subtree as an example, if node 1 (i.e., node 2) is not a leaf node, then node 1 (i.e., node 2) can have the optimal splitting feature and the optimal splitting point. Specifically, when the optimal splitting feature of node 1 (i.e., node 2) belongs to the first participant, node 1 records the optimal splitting feature and the optimal splitting point, and node 2 records "-1"; optionally, when the optimal splitting feature of node 1 (i.e., node 2) belongs to the second participant, node 2 records the optimal splitting feature and the optimal splitting point, and node 1 records "-1".

[0189] It should be understood that the number of decision trees (e.g., regression trees, classification trees) can be set in this embodiment, starting from the generation of the first decision tree and continuing until the generation of the last decision tree. The converged target decision tree can be any decision tree from the first to the last decision tree. Each decision tree can be jointly stored by the first and second participants. When performing business processing on business data, this embodiment can input the feature data of the business data (which can include feature data of the business data in the first business feature or feature data of the business data in the second business feature) into each decision tree to obtain the output of each decision tree (wherein, the output of each decision tree needs to be jointly determined by the first and second participants), thereby realizing business processing of the business data (e.g., classifying the business data) based on the output of each decision tree.

[0190] For easier understanding, please refer to Figure 8 , Figure 8 This is a schematic diagram of a stealth training process provided in an embodiment of this application. Figure 8 As shown, embodiments of this application can perform dense-state training using a multi-party (e.g., two-party) secure computation method. The two parties may include a first participant 80a and a second participant 80b, which are symmetrical. The first participant 80a can execute steps S11-S19, and the second participant 80b can execute steps S21-S29.

[0191] like Figure 8 As shown, in steps S11-S19, the process begins, and the first participant 80a can execute step S11 to determine whether all trees have been established. If not, step S12 is executed to initialize a new tree, i.e., to obtain the first intersection gradient slice associated with the first business sample identifier and the second business sample identifier. Optionally, if all trees have been established, the process ends. Further, the first participant 80a can execute step S13 to perform dense gradient update, i.e., to obtain the first feature gradient slice associated with the intersection gradient vector. Further, the first participant 80a can execute step S14 to determine whether all nodes have been traversed, i.e., whether the level of the target decision tree has reached the level threshold. If the level of the target decision tree has reached the level threshold, it is determined that all nodes have been traversed, and step S11 is then executed. Optionally, if the level of the target decision tree has not reached the level threshold, it is determined that not all nodes have been traversed, and step S15 is then executed. Further, the first participant 80a can determine whether the current node is a leaf node through step S15; wherein, if the current node is a leaf node, then step S19 is executed, and the weights are calculated in a closed state through step S19, that is, the node gradient slices of the leaf node are obtained, until all nodes have been traversed; optionally, if the current node is not a leaf node, then step S16 is executed, and the gradient histogram is calculated in a closed state through step S16, that is, the first partitioning feature gradient slices and the first cumulative gradient slices are obtained. Further, the first participant 80a can execute step S17, and find the optimal split point in a closed state through step S17, that is, determine the optimal splitting feature and the optimal splitting point for the first business feature and the second business feature, thereby realizing the node partitioning of the target decision tree. Further, the first participant 80a can execute step S18, and split the node in a closed state through step S18, that is, obtain the first node Boolean slices of the first child node of the node to be partitioned, and the second node Boolean slices of the second child node of the node to be partitioned, until all nodes have been traversed.

[0192] like Figure 8 As shown, in steps S21-S29, the process begins, and the second participant 80b can execute step S21 to determine whether all trees have been established. If not, step S22 is executed to initialize a new tree, i.e., to obtain the second intersection gradient slice associated with the first business sample identifier and the second business sample identifier. Optionally, if all trees have been established, the process ends. Further, the second participant 80b can execute step S23 to perform dense gradient update, i.e., to obtain the second feature gradient slice associated with the intersection gradient vector. Further, the second participant 80b can execute step S24 to determine whether all nodes have been traversed, i.e., whether the level of the target decision tree has reached the level threshold. If the level of the target decision tree has reached the level threshold, it is determined that all nodes have been traversed, and step S21 is executed. Optionally, if the level of the target decision tree has not reached the level threshold, it is determined that not all nodes have been traversed, and step S25 is executed. Further, the second participant 80b can determine whether the current node is a leaf node through step S25; wherein, if the current node is a leaf node, then step S29 is executed, and the weights are calculated in a closed state through step S29, that is, the node gradient slices of the leaf node are obtained, until all nodes have been traversed; optionally, if the current node is not a leaf node, then step S26 is executed, and the gradient histogram is calculated in a closed state through step S26, that is, the second partitioning feature gradient slices and the second cumulative gradient slices are obtained. Further, the second participant 80b can execute step S27, and find the optimal split point in a closed state through step S27, that is, determine the optimal splitting feature and the optimal splitting point for the first business feature and the second business feature, thereby realizing the node partitioning of the target decision tree. Further, the second participant 80b can execute step S28, and split the node in a closed state through step S28, that is, obtain the third node Boolean slices of the first child node of the node to be partitioned, and the fourth node Boolean slices of the second child node of the node to be partitioned, until all nodes have been traversed.

[0193] like Figure 8 As shown, in steps S16 and S26, the first partitioning feature gradient slice held by the first participant 80a and the second partitioning feature gradient slice held by the second participant 80b can be used together to execute steps S17 and S28 (i.e., the first participant 80a can provide the first feature gradient sub-slice to obtain the first partitioning feature gradient slice and the first cumulative gradient slice; the second participant 80b can provide the second feature gradient sub-slice to obtain the second partitioning feature gradient slice and the second cumulative gradient slice); in steps S17 and S27, one of the first participant 80a or the second participant 80b can directly obtain the optimal splitting feature and the optimal splitting point, while the other participant cannot directly obtain the optimal splitting feature and the optimal splitting point (i.e., the first participant 80a can provide the first partitioning feature gradient slice and the first cumulative gradient slice, while the second participant 80b...). (A second partitioning feature gradient partition and a second cumulative gradient partition can be provided to obtain the optimal splitting result); In steps S18 and S28, the first node Boolean partition held by the first participant 80a and the third node Boolean partition held by the second participant 80b can represent the node Boolean partition of the first child node, and the second node Boolean partition held by the first participant 80a and the fourth node Boolean partition held by the second participant 80b can represent the node Boolean partition of the second child node; In steps S19 and S29, the first node gradient partition held by the first participant 80a and the third node gradient partition held by the second participant 80b can represent the node gradient partition of the first child node, and the second node gradient partition held by the first participant 80a and the fourth node gradient partition held by the second participant 80b can represent the node gradient partition of the second child node.

[0194] For example, the optimal splitting feature could be a video service feature (e.g., video viewing duration) in a multimedia data recommendation scenario, with the optimal split point being 60 minutes. In this case, the left subtree could identify objects with a video viewing duration of less than or equal to 60 minutes, and the right subtree could identify objects with a video viewing duration greater than 60 minutes. As another example, the optimal splitting feature could be a news service feature (e.g., number of comments) in a multimedia data recommendation scenario, with the optimal split point being 5 comments. In this case, the left subtree could identify objects with a number of comments less than or equal to 5, and the right subtree could identify objects with a number of comments greater than 5.

[0195] Therefore, the embodiments of this application provide a stealthy two-party longitudinal federated XGBoost training method based on the intersection of circuit privacy sets for imbalanced scenarios (i.e., one party has a large amount of data while the other party has only a small amount of data, for example, the first party has a small number of samples and the second party has a large number of samples). This method can avoid the expensive privacy comparison operation, and the number of samples is only similar to that of the party with the smaller number of samples. When the data of the two parties is very imbalanced, the training efficiency is greatly improved (i.e., the XGBoost training overhead in imbalanced scenarios is greatly reduced).

[0196] For easier understanding, please refer to Figure 9 , Figure 9 This is a schematic diagram illustrating a privacy arrangement scenario provided in an embodiment of this application. For example... Figure 9 As shown, embodiments of this application can perform encrypted training using a multi-party (e.g., two-party) secure computation method. The two parties may include a first participant 90a and a second participant 90b. The circuit privacy set intersection method in embodiments of this application can associate the number of samples with the number of the smaller party (e.g., making the number of samples 1.27 times the number of the smaller party), while the additional number (e.g., 0.27 times the number) is a random number. Therefore, embodiments of this application can perform encrypted training as described above. Figure 4 In the corresponding embodiment, an Oblivious Switching Network (OSN) module is added between the circuit privacy intersection module and the stealth training module to remove the extra data.

[0197] like Figure 9 As shown, the first participant 90a can input input 91a and sequence 90c (i.e., sequence π) into the unintentional exchange network, and the second participant 90b can input input 91b into the unintentional exchange network. Inputs 91a and 91b can be fragments of the input. Furthermore, the unintentional exchange network can implement privacy permutation, obtaining the output 92a of the first participant 90a and the output 92b of the second participant 90b. Outputs 92a and 92b can be fragments of the output (i.e., the output represents the rearrangement of the input according to sequence π).

[0198] Understandably, the receiver of the circuit privacy intersection (e.g., the first participant 90a) can allocate the first service sample identifier to different hash mapping buckets according to the cuckoo hash. Then, the first participant 90a can know the corresponding order π of this allocation. At this time, the first participant 90a and the second participant 90b can perform a privacy permutation through an unintentional exchange network. That is, all the non-random numbers n0 (i.e., the number of first identifiers) fragments corresponding to the first service sample identifiers are arranged in the first n0, and 0.27*n0 random numbers are arranged at the end of the result. Thus, the first n0 are retained, the last 0.27*n0 are removed, and the n0 sample data are sent to the stealth training module to reduce training overhead.

[0199] like Figure 9 As shown, the first participant 90a and the second participant 90b will not operate on plaintext. Therefore, the first participant 90a and the second participant 90b must send their respective feature data partitions B (i.e., the first feature data partition, the second feature data partition, the third feature data partition, and the fourth feature data partition) and labels Y (e.g., the first service label partition and the second service label partition corresponding to the second service sample label) to the circuit privacy intersection calculation. That is, inputs 91a, 91b, outputs 92a, and 92b are all in the form of partitions. Among them, input 91a may include the first service label partition, the first feature data partition, the third feature data partition, and the first Boolean intersection partition, and input 91b may include the second service label partition, the second feature data partition, the fourth feature data partition, and the second Boolean intersection partition. Therefore, output 92a can include a first service label segment, a first feature data segment, a third feature data segment, and a first Boolean intersection segment, while output 92b can include a second service label segment, a second feature data segment, a fourth feature data segment, and a second Boolean intersection segment. The difference between inputs 91a, 91b, outputs 92a, and output 92b lies in their quantity; for example, the number of inputs 91a is 1.27 times the number of outputs 92a, and the number of inputs 91b is 1.27 times the number of outputs 92b.

[0200] Further, please see Figure 10 , Figure 10 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. The data processing device 1 runs on a first participant in vertical federated learning. The data processing device 1 may include: a quantity acquisition module 11, a first acquisition module 12, a second acquisition module 13, and a third acquisition module 14; further, the data processing device 1 may also include: a first generation module 15, a second generation module 16, a fourth acquisition module 17, a first processing module 18, a second processing module 19, a fifth acquisition module 20, a sixth acquisition module 21, and a weight determination module 22.

[0201] The quantity acquisition module 11 is used to acquire the number of first identifiers corresponding to the first business sample identifiers of the first participant; the number of first identifiers is used to compare with the number of second identifiers corresponding to the second business sample identifiers of the second participant in the vertical federated learning to obtain a comparison result.

[0202] The first acquisition module 12 is configured to, if the comparison result indicates that the number of first identifiers is less than or equal to the number of second identifiers, acquire a first feature data partitioning matrix for the first business feature of the first business sample identifier, and a first feature data partitioning fragment for the second business feature of the second business sample identifier; the first feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the first participant in partitioning the first business sample identifier based on the first business feature; the first feature data partitioning fragment and the second feature data partitioning fragment held by the second participant are fragments of the second feature data partitioning matrix; the second feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the second participant in partitioning the second business sample identifier based on the second business feature;

[0203] Among them, the number of first business features is M1, and the number of second business features is M2; M1 is a positive integer, and M2 is a positive integer; the M1 first business features include first business feature V. d d is a non-negative integer less than M1;

[0204] The first acquisition module 12 includes: a sample partitioning unit 121, a matrix determination unit 122, and a slice acquisition unit 123;

[0205] Sample segmentation unit 121 is used to obtain the first business feature V for the first business sample identifier. d Feature data, based on the first business feature V d The feature data is used to divide the first business sample identifier into business samples, generating the first business feature V. d The feature data is a bucketing indicator vector for feature bucketing;

[0206] Among them, for the first business feature V d The number of feature buckets is K, where K is a positive integer;

[0207] Sample partitioning unit 121 is specifically used to partition the first business feature V d Feature data is distributed to K feature buckets; the K feature buckets include feature bucket K u u is a non-negative integer less than K;

[0208] Sample partitioning unit 121 is specifically used to partition the first business sample identifier in the first business feature V. d The feature data below belongs to feature bucket K u Then the data existence parameter is determined as the first business sample identifier in the first business feature V. d Next, for feature binning K u The bucketing indicator parameter, if the first business sample is identified in the first business feature V d The feature data below does not belong to feature bucket K u If the data does not exist, the parameter is determined as the first business sample identifier in the first business feature V. d Next, for feature binning K u The binning indicator parameters;

[0209] Sample partitioning unit 121 is specifically used to divide each first business sample identifier into feature buckets K. u The bucketing indicator parameters are determined as the first business feature V. d Feature data for feature binning K u Bucket indicator vector.

[0210] The matrix determination unit 122 is used to determine the feature data of the M1 first business features according to the bucketing indicator vector of the feature bucketing, and to form the first feature data partitioning matrix of the M1 first business features for the first business sample identifier.

[0211] The sharding acquisition unit 123 is used to acquire the first feature data of M2 second business features for the second business sample identifier and divide them into shards.

[0212] The specific implementation methods of the sample partitioning unit 121, the matrix determination unit 122, and the slice acquisition unit 123 can be found in the above description. Figure 3 The description of step S102 in the corresponding embodiments will not be repeated here.

[0213] The second acquisition module 13 is used to acquire a first intersection gradient slice that is associated with both the first business sample identifier and the second business sample identifier; the first intersection gradient slice and the second intersection gradient slice held by the second participant are slices of the intersection gradient vector; the intersection gradient vector is used to characterize the sample gradient of the intersection business sample identifier between the first business sample identifier and the second business sample identifier.

[0214] The second acquisition module 13 includes: a first acquisition unit 131, a second acquisition unit 132, and a third acquisition unit 133;

[0215] The first acquisition unit 131 is used to acquire a first sample gradient slice that is jointly associated with the first business sample identifier and the second business sample identifier; the first sample gradient slice and the second sample gradient slice held by the second participant are slices of the sample gradient vector; the sample gradient vector is used to characterize the sample gradient of the first business sample identifier;

[0216] Specifically, the first acquisition unit 131 is used to acquire, if the first participant is a tag holder, the first cumulative sample prediction slice of the first business sample identifier and the first business tag vector corresponding to the first business sample tag of the first business sample identifier, and to determine the first sample gradient slice jointly associated with the first business sample identifier and the second business sample identifier based on the first cumulative sample prediction slice and the first business tag vector; the first cumulative sample prediction slice and the second cumulative sample prediction slice held by the second participant are slices of the cumulative sample prediction vector; the cumulative sample prediction vector is used to characterize the historical cumulative sample prediction value when generating the target decision tree through the first business sample identifier; the second participant is used to determine the second sample gradient slice jointly associated with the first business sample identifier and the second business sample identifier based on the second cumulative sample prediction slice;

[0217] The first acquisition unit 131 is specifically used to, if the second participant is a tag holder, acquire the first cumulative sample prediction slice of the first business sample identifier and the first business tag slice corresponding to the second business sample tag of the second business sample identifier; and determine the first sample gradient slice jointly associated with the first business sample identifier and the second business sample identifier based on the first cumulative sample prediction slice and the first business tag slice; the first business tag slice and the second business tag slice held by the second participant are slices of the second business tag vector corresponding to the second business sample tag; the first cumulative sample prediction slice and the second cumulative sample prediction slice held by the second participant are slices of the cumulative sample prediction vector; the cumulative sample prediction vector is used to characterize the historical cumulative sample prediction value when generating the target decision tree through the first business sample identifier; and the second participant is used to determine the second sample gradient slice jointly associated with the first business sample identifier and the second business sample identifier based on the second cumulative sample prediction slice and the second business tag slice.

[0218] The second acquisition unit 132 is used to acquire the first Boolean intersection fragment corresponding to the first business sample identifier; the first Boolean intersection fragment and the second Boolean intersection fragment held by the second participant are fragments of the Boolean intersection vector; the Boolean intersection vector is used to indicate the intersection state of the first business sample identifier with respect to the second business sample identifier;

[0219] The second acquisition unit 132 is specifically used to perform a cuckoo hash mapping on the first business sample identifier to obtain the first hash table corresponding to the first business sample identifier.

[0220] The second acquisition unit 132 is specifically used to acquire the first Boolean intersection fragment associated with the first hash table and the second hash table; the first hash table and the second hash table are used to perform hash table matching through an unintentionally programmable pseudo-random function to generate the first Boolean intersection fragment and the second Boolean intersection fragment held by the second participant; the second hash table is obtained by the second participant through hash mapping of the second business sample identifier.

[0221] The first hash table and the second hash table have the same hash table dimension; the first hash table includes a first hash mapping bucket, and the second hash table includes a second hash mapping bucket with the same hash table dimension as the first hash mapping bucket;

[0222] Wherein, if the first hash mapping bucket includes the first business sample identifier, the first random number corresponding to the first business sample identifier in the first hash mapping bucket is obtained by randomly processing the first business sample identifier in the first hash mapping bucket based on an unintentionally programmable pseudo-random function;

[0223] Wherein, if the second hash mapping bucket includes the second business sample identifier, the second random number corresponding to the second business sample identifier in the second hash mapping bucket is obtained by randomly processing the second business sample identifier in the second hash mapping bucket based on an unintentionally programmable pseudo-random function; the first random number and the second random number are used to perform random number matching between the first participant and the second participant to generate the first Boolean intersection parameter corresponding to the first hash mapping bucket and the second Boolean intersection parameter corresponding to the second hash mapping bucket; the first Boolean intersection parameter and the second Boolean intersection parameter are the fragments of the merged Boolean intersection parameter; the merged Boolean intersection parameter is used to indicate the Boolean matching result of the first hash mapping bucket and the second hash mapping bucket; the first Boolean intersection parameter corresponding to each hash table dimension in the first hash table is used to form the first Boolean intersection fragment corresponding to the first business sample identifier, and the second Boolean intersection parameter corresponding to each hash table dimension in the second hash table is used to form the second Boolean intersection fragment corresponding to the first business sample identifier.

[0224] The third acquisition unit 133 is used to acquire the first intersection gradient slice that is associated with the first business sample identifier and the second business sample identifier; the Boolean intersection vector is used to perform gradient filtering on the sample gradient vector to generate the first intersection gradient slice and the second intersection gradient slice held by the second participant.

[0225] The specific implementation methods of the first acquisition unit 131, the second acquisition unit 132, and the third acquisition unit 133 can be found in the above description. Figure 3 In the corresponding embodiment, steps S103 and Figure 5 The descriptions of steps S1031-S1033 in the corresponding embodiments will not be repeated here.

[0226] The third acquisition module 14 is used to acquire the first feature gradient slice associated with the intersection gradient vector; the first feature data partitioning matrix and the second feature data partitioning matrix are used to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient slice and the second feature gradient slice held by the second participant; the first feature gradient slice and the second feature gradient slice are used to determine the optimal splitting result for the first business feature and the second business feature; the optimal splitting result is used to generate the target decision tree for the first participant and the second participant; the converged target decision tree is used to perform business processing on the business data.

[0227] The first business sample identifier and the second business sample identifier together include M business features belonging to different business dimensions; M is a positive integer greater than 1; the M business features include business feature V. c c is a non-negative integer less than M;

[0228] Wherein, if business characteristic V c If it belongs to the first business feature, then the first feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector to obtain the data for business feature V. c First feature gradient slicing and targeting business feature V c The second feature gradient partitioning;

[0229] Wherein, if business characteristic V c If it belongs to the second business feature, then the second feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector to obtain the data for business feature V. c First feature gradient slicing and targeting business feature V c The second feature gradient slice; the first feature gradient slice for M business features is used to form the first feature gradient slice held by the first participant; the second feature gradient slice for M business features is used to form the second feature gradient slice held by the second participant.

[0230] Among them, business feature V c The number of feature buckets is K, where K is a positive integer; if business feature V c If the second feature belongs to the second business sample identifier, then the second feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector, resulting in a first feature gradient sub-partition for K feature buckets and a second feature gradient sub-partition for K feature buckets; the first feature gradient sub-partition for K feature buckets is used to form the business feature V. c The first feature gradient partitioning; the second feature gradient sub-partitioning for K feature buckets is used to form a sub-partition for business feature V. c The second feature gradient slice.

[0231] Optionally, the first feature gradient slice held by the first participant includes M first feature gradient slices for M business features; for business feature V c The first feature gradient partition includes K first feature gradient sub-partitions for K feature buckets;

[0232] The first generation module 15 is used to generate business features V based on the K first feature gradient sub-slices for K feature buckets. c The corresponding K-1 first-partition feature gradient slices;

[0233] The second generation module 16 is used to generate a first cumulative gradient slice corresponding to the first participant based on the K first feature gradient sub-slices for the K feature buckets; the M*KM first partitioning feature gradient slices, the first cumulative gradient slices, the M*KM second partitioning feature gradient slices, and the second cumulative gradient slice corresponding to the second participant for the M business features are used together to determine the information gain for the M business features and the K feature buckets; the second cumulative gradient slice is generated by the second participant based on the K second feature gradient sub-slices for the K feature buckets; the M*KM second partitioning feature gradient slices for the M business features include business feature V c The corresponding K-1 second-partition gradient slices; business feature V c The corresponding K-1 second partition feature gradient slices are generated by the second participant based on the K second feature gradient sub-slices for the K feature buckets; the optimal information gain among the information gains for the M business features and the K feature buckets is used to determine the optimal splitting result for the first business feature and the second business feature; the optimal splitting result includes the optimal splitting feature and the optimal splitting point.

[0234] Optionally, the fourth acquisition module 17 is used to acquire a first split point partitioning matrix for the first business feature and a first split point partitioning slice for the second business feature of the second business sample identifier; the first split point partitioning matrix is ​​used to characterize the partitioning result of the first feature data partitioning matrix in feature bucketing; the first split point partitioning slice and the second split point partitioning slice held by the second participant are slices of the second feature data partitioning matrix; the second split point partitioning matrix is ​​used to characterize the partitioning result of the second feature data partitioning matrix in feature bucketing;

[0235] The first processing module 18 is used to obtain the first split point partition vector corresponding to the optimal split result from the first split point partition matrix if the optimal split feature belongs to the first business feature, perform XOR processing on the first split point partition vector through the first random blind partition, generate the first synchronous blind partition, and send the first synchronous blind partition to the second participant; the first random blind partition and the first synchronous blind partition are partitions of the first split point partition vector;

[0236] The second processing module 19 is used to perform XOR processing on the first split point partitioning segment through the second random blinding segmentation if the optimal splitting feature belongs to the second business feature, thereby generating a random blinding matrix; the random blinding matrix is ​​used to obtain the random blinding vector corresponding to the optimal splitting result from the random blinding matrix through an unintentional transmission protocol; the second participant is used to obtain the splitting point partitioning sub-segment corresponding to the optimal splitting result from the second splitting point partitioning segmentation, and perform XOR processing on the splitting point partitioning sub-segment through the random blinding vector to generate a second synchronous blinding segmentation; the second random blinding segmentation and the second synchronous blinding segmentation are segments of the second splitting point partitioning vector corresponding to the optimal splitting result in the second splitting point partitioning matrix.

[0237] Optionally, the fifth acquisition module 20 is used to acquire the first node Boolean fragment of the first child node of the node to be partitioned in the target decision tree of the first participant and the second participant, and the second node Boolean fragment of the second child node of the node to be partitioned; the first node Boolean fragment and the third node Boolean fragment held by the second participant are fragments of the first node Boolean vector; the second node Boolean fragment and the fourth node Boolean fragment held by the second participant are fragments of the second node Boolean vector; the first node Boolean vector is used to represent the partitioning state of the intersection business sample identifier being partitioned to the first child node, and the second node Boolean vector is used to represent the partitioning state of the intersection business sample identifier being partitioned to the second child node;

[0238] Wherein, if the optimal splitting feature belongs to the first business feature, then the first node Boolean partition, the second node Boolean partition, the third node Boolean partition, and the fourth node Boolean partition are obtained by the first random blind partition and the first synchronous blind partition jointly performing Boolean partitioning on the Boolean intersection vector;

[0239] If the optimal splitting feature belongs to the second business feature, then the first node Boolean partition, the second node Boolean partition, the third node Boolean partition, and the fourth node Boolean partition are obtained by the second random blind partition and the second synchronous blind partition jointly performing Boolean partitioning on the Boolean intersection vector.

[0240] Optionally, the sixth acquisition module 21 is used to acquire the first node gradient slice of the first child node and the second node gradient slice of the second child node; the first node gradient slice and the third node gradient slice held by the second participant are slices of the first node gradient vector; the second node gradient slice and the fourth node gradient slice held by the second participant are slices of the second node gradient vector; the first node gradient vector is used to characterize the sample gradient of the first child node, and the second node gradient vector is used to characterize the sample gradient of the second child node;

[0241] The gradient partitioning of the first node and the gradient partitioning of the third node are both obtained by partitioning the first feature gradient partition and the second feature gradient partition using the Boolean vector of the first node; the gradient partitioning of the second node and the gradient partitioning of the fourth node are both obtained by partitioning the first feature gradient partition and the second feature gradient partition using the Boolean vector of the second node.

[0242] Optionally, the weight determination module 22 is used to obtain the node Boolean partition and node gradient partition of the leaf nodes in the converged target decision tree when the target decision tree converges.

[0243] The weight determination module 22 is used to determine the first child node weight partition of the leaf node based on the node Boolean partition and node gradient partition of the leaf node; the first child node weight partition and the second child node weight partition held by the second participant are partitions of the child node weight vector; the child node weight vector is used to characterize the weight parameters of the leaf node.

[0244] The first child node weighted shard is used to update the first cumulative sample prediction shard of the first business sample identifier, and the second child node weighted shard is used to update the second cumulative sample prediction shard of the first business sample identifier; the updated first cumulative sample prediction shard and the updated second cumulative sample prediction shard are used to generate the next decision tree of the target decision tree.

[0245] The specific implementation methods of the quantity acquisition module 11, the first acquisition module 12, the second acquisition module 13, the third acquisition module 14, the first generation module 15, and the second generation module 16 can be found in the above description. Figure 3 In the corresponding embodiment, steps S101-S104 and Figure 5 The descriptions of steps S1031-S1033 in the corresponding embodiments will not be repeated here. The specific implementation methods of the fourth acquisition module 17, the first processing module 18, the second processing module 19, the fifth acquisition module 20, the sixth acquisition module 21, and the weight determination module 22 can be found above. Figure 7 The descriptions of steps S201-S207 in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.

[0246] Further, please see Figure 11 , Figure 11 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device can be a first terminal device or a second terminal device. Figure 11 As shown, the computer device 1000 may include a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer device 1000 may also include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to enable communication between these components. In some embodiments, the user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. Optionally, the network interface 1004 may include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the processor 1001. Figure 11 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.

[0247] In such Figure 11 In the computer device 1000 shown, the network interface 1004 provides network communication functionality; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0248] Obtain the number of first identifiers corresponding to the first business sample identifier of the first participant; the number of first identifiers is used to compare with the number of second identifiers corresponding to the second business sample identifiers of the second participant in the vertical federated learning, and obtain the comparison result;

[0249] If the comparison result indicates that the number of first identifiers is less than or equal to the number of second identifiers, then a first feature data partitioning matrix for the first business feature of the first business sample identifier and a first feature data partitioning slice for the second business feature of the second business sample identifier are obtained; the first feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the first participant in partitioning the first business sample identifier based on the first business feature; the first feature data partitioning slice and the second feature data partitioning slice held by the second participant are slices of the second feature data partitioning matrix; the second feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the second participant in partitioning the second business sample identifier based on the second business feature;

[0250] Obtain the first intersection gradient slice that is associated with both the first business sample identifier and the second business sample identifier; the first intersection gradient slice and the second intersection gradient slice held by the second participant are slices of the intersection gradient vector; the intersection gradient vector is used to characterize the sample gradient of the intersection business sample identifier between the first business sample identifier and the second business sample identifier.

[0251] Obtain the first feature gradient slice associated with the intersection gradient vector; the first feature data partitioning matrix and the second feature data partitioning matrix are used to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient slice and the second feature gradient slice held by the second participant; the first feature gradient slice and the second feature gradient slice are used to determine the optimal splitting result for the first business feature and the second business feature; the optimal splitting result is used to generate the target decision tree for the first participant and the second participant; the converged target decision tree is used to perform business processing on the business data.

[0252] It should be understood that the computer device 1000 described in the embodiments of this application can execute the foregoing text. Figure 3 The description of the data processing method in the corresponding embodiments can also be performed as described above. Figure 9 The description of the data processing device 1 in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated here.

[0253] Furthermore, it should be noted that this application embodiment also provides a computer-readable storage medium, which stores a computer program executed by the aforementioned data processing device 1. When the processor executes the computer program, it can execute the aforementioned... Figure 3 , Figure 5 or Figure 7 The description of the data processing method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments related to this application, please refer to the description of the method embodiments of this application.

[0254] Furthermore, it should be noted that this application also provides a computer program product, which may include a computer program that can be stored in a computer-readable storage medium. The processor of a computer device reads the computer program from the computer-readable storage medium, and the processor can execute the computer program, causing the computer device to perform the aforementioned... Figure 3 , Figure 5 or Figure 7 The description of the data processing method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer program product embodiments related to this application, please refer to the description of the method embodiments of this application.

[0255] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0256] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.< / q> < / s> < / s> < / s> ​​< / s> ​< / h> < / g> < / h> < / g> < / s> < / s> < / h> < / s> < / h> < / g> < / s> < / g> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / x> < / y> < / x> < / y> < / x> < / x> < / x> < / x> < / y> < / x> < / y> < / x> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / y> < / x> < / x> < / x>

Claims

1. A data processing method, characterized in that, The method is performed by the first participant in the longitudinal federated learning and includes: Obtain the number of first identifiers corresponding to the first business sample identifier of the first participant; the number of first identifiers is used to compare with the number of second identifiers corresponding to the second business sample identifiers of the second participant in the vertical federated learning, and obtain the comparison result; If the comparison result indicates that the number of the first identifiers is less than or equal to the number of the second identifiers, then a first feature data partitioning matrix for the first business feature of the first business sample identifier and a first feature data partitioning fragment for the second business feature of the second business sample identifier are obtained; the first feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the first participant in partitioning the first business sample identifier based on the first business feature; the first feature data partitioning fragment and the second feature data partitioning fragment held by the second participant are fragments of the second feature data partitioning matrix; the second feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the second participant in partitioning the second business sample identifier based on the second business feature; Obtain a first intersection gradient slice that is jointly associated with the first business sample identifier and the second business sample identifier; the first intersection gradient slice and the second intersection gradient slice held by the second participant are slices of the intersection gradient vector; the intersection gradient vector is used to characterize the sample gradient of the intersection business sample identifier between the first business sample identifier and the second business sample identifier; Obtain a first feature gradient slice associated with the intersection gradient vector; the first feature data partitioning matrix and the second feature data partitioning matrix are used to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient slice and the second feature gradient slice held by the second participant; the first feature gradient slice and the second feature gradient slice are used to determine the optimal splitting result for the first business feature and the second business feature; the optimal splitting result is used to generate a target decision tree for the first participant and the second participant; the converged target decision tree is used to perform business processing on the business data.

2. The method according to claim 1, characterized in that, The number of the first service features is M1, and the number of the second service features is M2; where M1 is a positive integer and M2 is a positive integer; the M1 first service features include the first service feature V. d The d is a non-negative integer less than M1; The step of obtaining the first feature data partitioning matrix for the first business feature of the first business sample identifier, and the first feature data partitioning slice for the second business feature of the second business sample identifier, includes: Obtain the first service feature V for the first service sample identifier. d Feature data, based on the first business feature V d The feature data is used to segment the first service sample identifier into service samples, generating the first service feature V. d The feature data is a bucketing indicator vector for feature bucketing; The feature data of the M1 first business features are respectively determined according to the bucketing indicator vector of feature bucketing, and the first feature data partitioning matrix of the M1 first business features for the first business sample identifier is determined. Obtain the first feature data segmentation of M2 second business features for the second business sample identifier.

3. The method according to claim 2, characterized in that, Regarding the first business feature V d The number of feature buckets is K, where K is a positive integer; Based on the first business feature V d The feature data is used to segment the first service sample identifier into service samples, generating the first service feature V. d The feature data includes a bucketing indicator vector for feature bucketing, which includes: The first business feature V d Feature data is allocated to K feature buckets; the K feature buckets include feature bucket K u u is a non-negative integer less than K; If the first service sample is identified in the first service feature V d The feature data below belongs to the feature bucket K. u Then the data existence parameter is determined as the first business sample identifier in the first business feature V. d The following describes the feature binning K. u The bucketing indicator parameter, if the first service sample is identified in the first service feature V d The feature data below does not belong to the feature bucket K. u If the data does not exist, the parameter is determined to be the first business sample identifier in the first business feature V. d The following describes the feature binning K. u The binning indicator parameters; Each of the first business sample identifiers is divided into buckets K based on the features. u The bucketing indicator parameter is determined as the first service feature V. d Feature data for the feature bins K u Bucket indicator vector.

4. The method according to claim 1, characterized in that, The step of obtaining the first intersection gradient slice associated with both the first service sample identifier and the second service sample identifier includes: Obtain a first sample gradient slice that is jointly associated with the first business sample identifier and the second business sample identifier; the first sample gradient slice and the second sample gradient slice held by the second participant are slices of sample gradient vectors; the sample gradient vector is used to characterize the sample gradient of the first business sample identifier; Obtain the first Boolean intersection fragment corresponding to the first service sample identifier; the first Boolean intersection fragment and the second Boolean intersection fragment held by the second participant are fragments of the Boolean intersection vector; the Boolean intersection vector is used to indicate the intersection state of the first service sample identifier with respect to the second service sample identifier; Obtain a first intersection gradient slice that is associated with both the first business sample identifier and the second business sample identifier; the Boolean intersection vector is used to perform gradient filtering on the sample gradient vector to generate the first intersection gradient slice and the second intersection gradient slice held by the second participant.

5. The method according to claim 4, characterized in that, The step of obtaining the first sample gradient slice associated with both the first service sample identifier and the second service sample identifier includes: If the first participant is a tag holder, then the first cumulative sample prediction slice of the first business sample identifier and the first business tag vector corresponding to the first business sample tag of the first business sample identifier are obtained. Based on the first cumulative sample prediction slice and the first business tag vector, the first sample gradient slice commonly associated with the first business sample identifier and the second business sample identifier is determined. The first cumulative sample prediction slice and the second cumulative sample prediction slice held by the second participant are slices of the cumulative sample prediction vector. The cumulative sample prediction vector is used to characterize the historical cumulative sample prediction value when generating the target decision tree through the first business sample identifier. The second participant is used to determine the second sample gradient slice commonly associated with the first business sample identifier and the second business sample identifier based on the second cumulative sample prediction slice. If the second participant is a tag holder, then the first cumulative sample prediction slice of the first business sample identifier and the first business tag slice corresponding to the second business sample tag of the second business sample identifier are obtained. Based on the first cumulative sample prediction slice and the first business tag slice, a first sample gradient slice commonly associated with the first business sample identifier and the second business sample identifier is determined. The first business tag slice and the second business tag slice held by the second participant are slices of the second business tag vector corresponding to the second business sample tag. The first cumulative sample prediction slice and the second cumulative sample prediction slice held by the second participant are slices of the cumulative sample prediction vector. The cumulative sample prediction vector is used to characterize the historical cumulative sample prediction value when generating the target decision tree through the first business sample identifier. The second participant is used to determine the second sample gradient slice commonly associated with the first business sample identifier and the second business sample identifier based on the second cumulative sample prediction slice and the second business tag slice.

6. The method according to claim 4, characterized in that, The step of obtaining the first Boolean intersection fragment corresponding to the first service sample identifier includes: Perform a cuckoo hash mapping on the first service sample identifier to obtain the first hash table corresponding to the first service sample identifier; Obtain a first Boolean intersection fragment associated with the first hash table and the second hash table; the first hash table and the second hash table are used to perform hash table matching through an unintentionally programmable pseudo-random function to generate the first Boolean intersection fragment and the second Boolean intersection fragment held by the second participant; the second hash table is obtained by the second participant through hash mapping of the second business sample identifier.

7. The method according to claim 6, characterized in that, The first hash table and the second hash table have the same hash table dimension; the first hash table includes a first hash mapping bucket, and the second hash table includes a second hash mapping bucket with the same hash table dimension as the first hash mapping bucket; Wherein, if the first hash mapping bucket includes the first business sample identifier, the first random number corresponding to the first business sample identifier in the first hash mapping bucket is obtained by randomly processing the first business sample identifier in the first hash mapping bucket based on an unintentionally programmable pseudo-random function; Wherein, if the second hash mapping bucket includes a second business sample identifier, the second random number corresponding to the second business sample identifier in the second hash mapping bucket is obtained by randomly processing the second business sample identifier in the second hash mapping bucket based on the unintentionally programmable pseudo-random function; the first random number and the second random number are used to perform random number matching between the first participant and the second participant to generate a first Boolean intersection parameter corresponding to the first hash mapping bucket and a second Boolean intersection parameter corresponding to the second hash mapping bucket; the first Boolean intersection parameter and the second Boolean intersection parameter are fragments of merged Boolean intersection parameters; the merged Boolean intersection parameter is used to indicate the Boolean matching result of the first hash mapping bucket and the second hash mapping bucket; the first Boolean intersection parameter corresponding to each hash table dimension in the first hash table is used to form a first Boolean intersection fragment corresponding to the first business sample identifier, and the second Boolean intersection parameter corresponding to each hash table dimension in the second hash table is used to form a second Boolean intersection fragment corresponding to the first business sample identifier.

8. The method according to claim 1, characterized in that, The first business sample identifier and the second business sample identifier together include M business features belonging to different business dimensions; M is a positive integer greater than 1; the M business features include business feature V. c Where c is a non-negative integer less than M; Wherein, if the business feature V c If the data belongs to the first business feature, then the first feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector to obtain the data for the business feature V. c The first feature gradient sharding and the business feature V c The second feature gradient partitioning; Wherein, if the business feature V c If it belongs to the second business feature, then the second feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector to obtain the data for the business feature V. c The first feature gradient sharding and the business feature V c The second feature gradient slice; the first feature gradient slice for the M business features is used to form the first feature gradient slice held by the first participant; the second feature gradient slice for the M business features is used to form the second feature gradient slice held by the second participant.

9. The method according to claim 8, characterized in that, The business feature V c The number of feature buckets is K, where K is a positive integer; if the business feature V c If the second feature belongs to the second business feature identified by the second business sample, then the second feature data partitioning matrix is ​​used to perform gradient partitioning on the intersection gradient vector to obtain a first feature gradient sub-partition for the K feature buckets and a second feature gradient sub-partition for the K feature buckets; the first feature gradient sub-partition for the K feature buckets is used to form a feature for the business feature V. c The first feature gradient partition; the second feature gradient sub-partition for the K feature buckets is used to form the feature V. c The second feature gradient slice.

10. The method according to claim 9, characterized in that, The first feature gradient slice held by the first participant includes M first feature gradient slices for M said business features; for said business feature V c The first feature gradient partition includes K first feature gradient sub-partitions for the K feature buckets; The method further includes: Based on the K first feature gradient sub-slices for the K feature buckets, the business feature V is generated. c The corresponding K-1 first-partition feature gradient slices; Based on the K first feature gradient sub-slices for the K feature buckets, a first cumulative gradient slice corresponding to the first participant is generated; the M*KM first partitioning feature gradient slices, the first cumulative gradient slice, the M*KM second partitioning feature gradient slices, and the second cumulative gradient slice corresponding to the M business features are used together to determine the information gain for the M business features and the K feature buckets; the second cumulative gradient slice is generated by the second participant based on the K second feature gradient sub-slices for the K feature buckets; the M*KM second partitioning feature gradient slices corresponding to the M business features include the business feature V c The corresponding K-1 second partitioning feature gradient slices; the business feature V c The corresponding K-1 second partitioning feature gradient slices are generated by the second participant based on the K second feature gradient sub-slices for the K feature buckets; the optimal information gain among the information gains of the M service features and the K feature buckets is used to determine the optimal splitting result for the first service feature and the second service feature; the optimal splitting result includes the optimal splitting feature and the optimal splitting point.

11. The method according to claim 10, characterized in that, The method further includes: Obtain a first split point partitioning matrix for the first business feature, and a first split point partitioning slice for the second business feature of the second business sample identifier; the first split point partitioning matrix is ​​used to characterize the partitioning result of the first feature data partitioning matrix in feature bucketing; the first split point partitioning slice and the second split point partitioning slice held by the second participant are slices of the second feature data partitioning matrix; the second split point partitioning matrix is ​​used to characterize the partitioning result of the second feature data partitioning matrix in feature bucketing; If the optimal splitting feature belongs to the first business feature, then the first splitting point partitioning vector corresponding to the optimal splitting result is obtained from the first splitting point partitioning matrix, and the first splitting point partitioning vector is XORed through the first random blinding partitioning to generate the first synchronous blinding partitioning, and the first synchronous blinding partitioning is sent to the second participant; the first random blinding partitioning and the first synchronous blinding partitioning are partitions of the first splitting point partitioning vector; If the optimal splitting feature belongs to the second service feature, then the first splitting point partitioning fragment is XORed with the second randomized blinding fragment to generate a randomized blinding matrix; the randomized blinding matrix is ​​used to obtain the randomized blinding vector corresponding to the optimal splitting result from the randomized blinding matrix through an unintentional transmission protocol; the second participant is used to obtain the splitting point partitioning sub-fragment corresponding to the optimal splitting result from the second splitting point partitioning fragment, and XOR the splitting point partitioning sub-fragment with the randomized blinding vector to generate a second synchronous blinding fragment; the second randomized blinding fragment and the second synchronous blinding fragment are fragments of the second splitting point partitioning vector corresponding to the optimal splitting result in the second splitting point partitioning matrix.

12. The method according to claim 11, characterized in that, The method further includes: Obtain the first node Boolean fragment of the first child node of the node to be partitioned in the target decision tree for the first participant and the second participant, and the second node Boolean fragment of the second child node of the node to be partitioned; the first node Boolean fragment and the third node Boolean fragment held by the second participant are fragments of the first node Boolean vector; the second node Boolean fragment and the fourth node Boolean fragment held by the second participant are fragments of the second node Boolean vector; the first node Boolean vector is used to represent the partitioning state in which the intersection service sample identifier is partitioned to the first child node, and the second node Boolean vector is used to represent the partitioning state in which the intersection service sample identifier is partitioned to the second child node; Wherein, if the optimal splitting feature belongs to the first business feature, then the first node Boolean partition, the second node Boolean partition, the third node Boolean partition, and the fourth node Boolean partition are obtained by the first random blind partition and the first synchronous blind partition jointly performing Boolean partitioning on the Boolean intersection vector; Wherein, if the optimal splitting feature belongs to the second business feature, then the first node Boolean partition, the second node Boolean partition, the third node Boolean partition, and the fourth node Boolean partition are obtained by the second random blind partition and the second synchronous blind partition jointly performing Boolean partitioning on the Boolean intersection vector.

13. The method according to claim 12, characterized in that, The method further includes: Obtain the first node gradient slice of the first child node and the second node gradient slice of the second child node; the first node gradient slice and the third node gradient slice held by the second participant are slices of the first node gradient vector; the second node gradient slice and the fourth node gradient slice held by the second participant are slices of the second node gradient vector; the first node gradient vector is used to characterize the sample gradient of the first child node, and the second node gradient vector is used to characterize the sample gradient of the second child node; The first node gradient partition and the third node gradient partition are both obtained by partitioning the first feature gradient partition and the second feature gradient partition using the first node Boolean vector; the second node gradient partition and the fourth node gradient partition are both obtained by partitioning the first feature gradient partition and the second feature gradient partition using the second node Boolean vector.

14. The method according to claim 13, characterized in that, The method further includes: When the target decision tree converges, obtain the node Boolean partition and node gradient partition of the leaf nodes in the converged target decision tree; Based on the node Boolean partition and node gradient partition of the leaf node, the first child node weight partition of the leaf node is determined; the first child node weight partition and the second child node weight partition held by the second participant are partitions of the child node weight vector; the child node weight vector is used to characterize the weight parameters of the leaf node; Wherein, the first child node weighted shard is used to update the first cumulative sample prediction shard of the first business sample identifier, and the second child node weighted shard is used to update the second cumulative sample prediction shard of the first business sample identifier; the updated first cumulative sample prediction shard and the updated second cumulative sample prediction shard are used to generate the next decision tree of the target decision tree.

15. A data processing apparatus, characterized in that, The device operates on the first participant in the vertical federated learning process and includes: The quantity acquisition module is used to acquire the number of first identifiers corresponding to the first business sample identifier of the first participant; the number of first identifiers is used to compare with the number of second identifiers corresponding to the second business sample identifiers of the second participant in the vertical federated learning to obtain a comparison result. The first acquisition module is configured to, if the comparison result indicates that the number of the first identifiers is less than or equal to the number of the second identifiers, acquire a first feature data partitioning matrix for the first business feature of the first business sample identifier, and a first feature data partitioning fragment for the second business feature of the second business sample identifier; the first feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the first participant in partitioning the first business sample identifier based on the first business feature; the first feature data partitioning fragment and the second feature data partitioning fragment held by the second participant are fragments of the second feature data partitioning matrix; the second feature data partitioning matrix is ​​used to characterize the partitioning result obtained by the second participant in partitioning the second business sample identifier based on the second business feature; The second acquisition module is used to acquire a first intersection gradient slice that is associated with both the first service sample identifier and the second service sample identifier; the first intersection gradient slice and the second intersection gradient slice held by the second participant are slices of the intersection gradient vector; the intersection gradient vector is used to characterize the sample gradient of the intersection service sample identifier between the first service sample identifier and the second service sample identifier. The third acquisition module is used to acquire a first feature gradient slice associated with the intersection gradient vector; the first feature data partitioning matrix and the second feature data partitioning matrix are used to perform gradient partitioning on the intersection gradient vector to obtain the first feature gradient slice and the second feature gradient slice held by the second participant; the first feature gradient slice and the second feature gradient slice are used to determine the optimal splitting result for the first business feature and the second business feature; the optimal splitting result is used to generate a target decision tree for the first participant and the second participant; the converged target decision tree is used to perform business processing on the business data.

16. A computer device, characterized in that, include: Processor and memory; The processor is connected to the memory, wherein the memory is used to store a computer program, and the processor is used to invoke the computer program to cause the computer device to perform the method according to any one of claims 1-14.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-14.

18. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium and adapted to be read and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-14.

Citation Information

Patent Citations

  • Multi-party XGBoost security prediction model training method based on secret sharing and federated learning

    CN112464287A

  • Outsourcing data alignment method based on function secret sharing

    CN116015653A