Feature-based binning processing method, device, equipment, and medium

By determining the boxing type in the federated learning system and using multi-party security computing technology to perform feature boxing processing, the feature boxing problem in multi-participant node scenarios is solved, the stability and accuracy of the model are improved, and the training efficiency is improved.

CN114841371BActive Publication Date: 2025-08-29BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210470050.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-28
Publication Date
2025-08-29
Estimated Expiration
2042-04-28

AI Technical Summary

Technical Problem

In the existing federated learning system, effective feature binning processing is lacking, especially in the multi-participating node scenario, which is difficult to achieve accurate feature binning processing, which affects the stability and accuracy of the model.

Method used

A feature-based binning processing method is provided. By obtaining the distribution of features and binning requirements, combining federated learning scenarios, the binning type is determined, and multi-party security computing technology is used to perform binning processing locally on each participant node to ensure data security and accuracy.

Benefits of technology

It improves the stability and accuracy of the model in the federated learning system, prevents overfitting, improves the model training speed, and effectively deals with null and missing values.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114841371B_ABST
    Figure CN114841371B_ABST
Patent Text Reader

Abstract

The present disclosure provides a feature-based binning processing method, device, equipment and medium, which relates to technical fields such as artificial intelligence and can be applied in distributed data processing scenarios such as federated learning. The specific implementation scheme is: obtaining the features to be referenced for binning processing; determining the scenario of federated learning based on the fields of the features in the nodes of each participant in the federated learning system and the distribution of the sample data corresponding to the features; determining the binning type based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each participant's node or the preset binning requirements, and referring to the scenario of federated learning; using the binning type, binning the features on the nodes of each participant in the federated learning system. The present disclosure can provide a feature-based binning scheme that can be applied to a federated learning system with multiple participating nodes, and can accurately and effectively bin the features in the nodes of each participant in the federated learning system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, specifically to artificial intelligence and other technical fields, and can be applied in distributed data processing scenarios such as federated learning. In particular, it relates to a feature-based binning processing method, apparatus, device, and medium. Background Art

[0002] Federated Learning is an emerging basic artificial intelligence technology designed to carry out efficient machine learning among multiple parties or computing nodes while ensuring information security during big data exchange, protecting terminal data and personal data privacy, and ensuring legality and compliance.

[0003] Federated learning has become a research hotspot in data science. To improve the effectiveness and accuracy of federated learning training models, data preprocessing is often required before model training. Binning is a key data preprocessing method, improving model stability and robustness, preventing overfitting, accelerating model training, and handling null and missing values. Summary of the Invention

[0004] The present disclosure provides a feature-based binning processing method, apparatus, device, and medium.

[0005] According to one aspect of the present disclosure, a feature-based binning method is provided, comprising:

[0006] Get the features to be referenced for binning;

[0007] Determine the federated learning scenario based on the characteristics of the nodes of each participant in the federated learning system and the distribution of sample data corresponding to the characteristics;

[0008] Determine the binning type based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each of the participant nodes or the preset binning requirements, and with reference to the federated learning scenario;

[0009] The binning type is used to perform binning processing on the features on each of the participant nodes in the federated learning system.

[0010] According to another aspect of the present disclosure, a feature-based binning processing apparatus is provided, comprising:

[0011] Feature acquisition module, used to obtain the features to be referenced for binning processing;

[0012] A first determination module is configured to determine a federated learning scenario based on the fields of the features in the nodes of each participant in the federated learning system and the distribution of sample data corresponding to the features;

[0013] A second determination module is configured to determine a binning type based on a distribution property of a feature value corresponding to the feature in the sample data on each of the participant nodes or a preset binning requirement, and with reference to the federated learning scenario;

[0014] A binning processing module is used to use the binning type to perform binning processing on the features on each of the participant nodes in the federated learning system.

[0015] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0016] at least one processor; and

[0017] a memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any possible implementation manner and the aspects described above.

[0019] According to yet another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method of the above-mentioned aspect and any possible implementation manner.

[0020] According to yet another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the above-mentioned aspects and any possible implementation method when executed by a processor.

[0021] According to the technology disclosed in the present invention, a feature-based binning scheme applicable to a federated learning system with multiple participating nodes can be provided, which can accurately and effectively bin the features in the sample data of each participating node in the federated learning system.

[0022] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0024] Figure 1is a schematic diagram according to a first embodiment of the present disclosure;

[0025] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;

[0026] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;

[0027] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure;

[0028] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure;

[0029] Figure 6 is a schematic diagram according to a sixth embodiment of the present disclosure;

[0030] Figure 7 is a schematic diagram according to a seventh embodiment of the present disclosure;

[0031] Figure 8 is a schematic diagram according to an eighth embodiment of the present disclosure;

[0032] Figure 9 is a block diagram of an electronic device for implementing the method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0033] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0034] Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0035] It should be noted that the terminal devices involved in the embodiments of the present disclosure may include but are not limited to mobile phones, personal digital assistants (PDAs), wireless handheld devices, tablet computers and other smart devices; display devices may include but are not limited to personal computers, televisions and other devices with display functions.

[0036] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0037] Existing research on federated learning has almost exclusively focused on modeling algorithms. However, relatively little research has been conducted on data preprocessing for federated learning. Furthermore, current research on feature binning primarily focuses on a single-node model and cannot be applied to multi-node scenarios to effectively bin sample data based on features. Therefore, the present disclosure provides a technical solution for effectively binning samples based on features, which can be applied in a multi-node federated learning system.

[0038] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure; Figure 1 As shown, this embodiment provides a feature-based binning method that can be applied in a federated learning system. Specifically, the method may include the following steps:

[0039] S101, obtaining features to be referenced for binning processing;

[0040] S102. Determine a federated learning scenario based on the distribution of feature fields and sample data corresponding to the features in the nodes of each participant in the federated learning system.

[0041] S103. Determine the binning type based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each participant's node or the preset binning requirements, and with reference to the federated learning scenario;

[0042] S104: Use the binning type to bin the features on the nodes of each participant in the federated learning system.

[0043] The feature-based binning method of this embodiment can be executed by a feature-based binning device. This device can be an electronic entity or a software-integrated application. During use, the device can access each participating node in the federated learning system and perform binning on all feature values ​​corresponding to the referenced feature in the sample data on each participating node.

[0044] In this embodiment, the features to be referenced for binning can be any one, two, or more features included in the sample data. Specifically, the features to be referenced for binning can be received as external input. However, during binning, binning needs to be performed separately based on each feature.

[0045] In this disclosure, the distribution of feature fields in each participant node in a federated learning system may include, for example, whether different participant nodes have the same feature identifier, i.e., the same feature field. The distribution of sample data corresponding to a feature may include, for example, whether different participant nodes have sample data corresponding to the feature and with the same sample identifier. Based on the distribution of feature fields and sample data corresponding to the features in each participant node, the federated learning scenario in the architecture can be inferred.

[0046] Furthermore, in the embodiment of the present disclosure, the distribution properties of the feature values ​​corresponding to the feature to be referenced in the sample data on each participant's node or the preset binning requirements can also be determined together with the scenario of federated learning. Finally, the determined binning type is used to perform binning on the features in the sample data on each participant's node in the federated learning system. After adopting the feature-based binning process of this embodiment, each bin includes the feature values ​​corresponding to the feature in multiple sample data. For ease of description, the feature values ​​in the bins can also be referred to as feature data. Each bin based on the feature does not include other feature data in the sample data. Specifically, the feature data in each bin may include one, two or more feature values ​​of the feature. In subsequent model training, the feature data of each bin after binning can be used to train the model. Since the feature data within a bin has certain commonalities in this feature, using the binned feature data to train the model can effectively improve the stability and robustness of the model and prevent overfitting; it can also accelerate the model training speed and the ability to handle null values ​​and missing values, resulting in a more stable and accurate model.

[0047] The feature-based binning method of this embodiment, by employing the aforementioned technical solution, can automatically bin features in the sample data of each participating node in a federated learning system. This effectively addresses the shortcomings of existing technologies and provides a feature-based binning solution applicable to a multi-node federated learning system, enabling accurate and effective binning of features in the sample data of each participating node in the federated learning system. Furthermore, the binned feature values ​​can be used to train a model, effectively improving the model's stability and accuracy.

[0048] It should be noted that, in one embodiment of the present disclosure, based on the distribution of feature fields and sample data corresponding to the features in the nodes of each participant in the federated learning system, the federated learning scenario may include the following situations:

[0049] In scenario 1, if the overlap ratio of feature fields in sample data with different identifiers from different participating nodes exceeds a preset threshold, the federated learning scenario is determined to be horizontal federated learning. In this embodiment, the preset threshold can be set to 80%, 85%, 90%, or even above 95%, and can be set to a percentage close to 1 based on actual needs.

[0050] In this scenario, the sample data from each participating node contains a significant overlap of feature fields. Overlapping feature fields here refers to the same feature identifiers, such as feature names, while the corresponding values ​​can be the same or different, without limitation. The distribution of sample data in federated learning in this case can be called horizontal distribution, and the corresponding federated learning scenario can be called horizontal federated learning.

[0051] It should be noted that in actual applications, after determining that the federated learning scenario is horizontal federated learning in scenario 1, step 104 uses the binning type. Before binning the features on each participant node in the federated learning system, it is necessary to perform feature alignment on all sample data included in the different participant nodes in the federated learning system to clean the sample data, remove sample data from each participant node that does not include the features to be referenced by the binning process, and retain sample data from each participant node that includes the features to be referenced by the binning process, so as to facilitate subsequent accurate binning. After the feature alignment process, the sample data with different sample data identifiers included on different participant nodes all include features with the same feature identifier, i.e., feature name.

[0052] Scenario 2: If the overlap ratio of the identifiers of the sample data included in the nodes of different participants in the federated learning system is greater than the preset ratio threshold, the federated learning scenario is determined to be vertical federated learning.

[0053] Unlike scenario 1 above, in this scenario, the sample data of each participant's nodes overlaps significantly, while the feature fields included in the sample data overlap less. The label data for the same sample data resides in one participant's node, while the other features of the sample data reside in another participant's node. The corresponding distribution of sample data in federated learning can be called vertical distribution, and the corresponding federated learning scenario can be called vertical federated learning.

[0054] It should be noted that in actual applications, after determining that the scenario of federated learning is vertical federated learning in situation two, step 104 adopts the binning type. Before binning the features on the nodes of each participant in the federated learning system, it is also necessary to perform sample alignment on all sample data included in the different participant nodes in the federated learning system to clean the sample data and remove the sample data of the sample data identifier that is not included in any node of the participant nodes. At the same time, the sample data corresponding to the features to be referenced for binning processing should also be removed. After the sample alignment processing, different features of the sample data with the same sample data identifier are included on different participant nodes, and some participant nodes include features to be referenced for binning processing. After the sample alignment processing, the sample data with different sample data identifiers included on different participant nodes are convenient for subsequent accurate binning processing.

[0055] This approach allows for accurate identification of the federated learning scenario, enabling the subsequent selection of accurate and appropriate binning types to improve binning efficiency. Furthermore, tailored alignment preprocessing is performed for different federated learning scenarios to effectively filter out irrelevant data and improve subsequent binning efficiency.

[0056] In one embodiment of the present disclosure, step S103 determines the binning type based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each participant's node or the preset binning requirements, and with reference to the federated learning scenario. Specifically, the binning type may be as follows:

[0057] (1) If the federated learning scenario is horizontal federated learning, the eigenvalues ​​corresponding to the features in the sample data on the nodes of each participant in the federated learning system are evenly distributed, and the binning type is determined to be horizontal equal-width binning;

[0058] (2) If the federated learning scenario is horizontal federated learning, the feature values ​​corresponding to the features in the sample data on the nodes of each participant in the federated learning system are concentrated in multiple preset intervals, and the binning type is determined to be horizontal equal-frequency binning;

[0059] (3) If the federated learning scenario is horizontal federated learning, the pre-set binning requirements include goodness of fit and / or independence test requirements, and the binning type is determined to be horizontal chi-square binning; or

[0060] (4) If the federated learning scenario is longitudinal federated learning, the preset binning requirements include goodness of fit and / or independence test requirements, and the binning type is determined to be longitudinal chi-square binning.

[0061] Different binning methods can have different advantages and disadvantages and can be configured based on different requirements. For example, in horizontal federated learning, if the distribution of the eigenvalues ​​corresponding to a feature tends to be uniform, equal-width binning can be used, which achieves excellent binning results. However, equal-width binning can easily lead to the eigenvalues ​​in the sample data being concentrated within a certain interval, resulting in essentially identical bin numbers and significant information loss. If the eigenvalues ​​corresponding to a feature in the sample data are concentrated within multiple preset intervals, equal-frequency binning can be used to overcome the drawbacks of equal-width binning. Alternatively, if goodness-of-fit and independence tests are determined to be necessary before modeling, chi-square binning can be used. In this case, the distribution properties of the eigenvalues ​​corresponding to the features of the sample data can be disregarded. Alternatively, weights can be assigned to the distribution properties of the eigenvalues ​​corresponding to the features of the sample data and the preset binning requirements, such that when the preset binning requirements are included, the weight of the preset binning requirements is greater than the weight of the distribution properties of the eigenvalues ​​corresponding to the features, e.g., a weight of 0.8 for the former and only 0.2 for the latter. Therefore, when the preset binning requirements are included, the binning type is primarily based on the requirements of the preset binning requirements.

[0062] Optionally, in one embodiment of the present disclosure, a pre-trained bin type determination model may be used to determine the bin type based on the distribution properties of the feature values ​​corresponding to the feature in the sample data on each participant's node and / or preset bin requirements, with reference to the federated learning scenario. The pre-trained bin type determination model is used to implement the determination of the bin type.

[0063] The binning type determination model is a pre-trained multi-classification neural network model. When used, it outputs the distribution properties of the feature values ​​corresponding to the features in the sample data on each participant's node, the preset binning requirements, and the federated learning scenario. The model can predict and output a relatively reasonable binning type.

[0064] It should be noted that if the binning type determines the model training without using the preset binning requirements or the distribution properties of the feature values ​​corresponding to the features, no relevant information is required for the corresponding prediction.

[0065] In practical applications, in addition to the above four situations, the following two situations can also be included:

[0066] If the federated learning scenario is vertical federated learning, the eigenvalues ​​corresponding to the features in the sample data on each participant's node in the federated learning system are evenly distributed, and the binning type is determined to be vertical equal-width binning;

[0067] If the federated learning scenario is vertical federated learning, the eigenvalues ​​corresponding to the features in the sample data on the nodes of each participant in the federated learning system are concentrated in multiple preset intervals, and the binning type is determined to be vertical equal-frequency binning.

[0068] It should be noted that the above-mentioned vertical equal-width binning and vertical equal-frequency binning, including the sample data of the features to be referenced for binning processing, are all on the same participant node, and do not involve binning different features in the sample data of multiple participant nodes at the same time. Please refer to the feature binning processing scheme on the relevant single node and will not go into details here.

[0069] Moreover, the above-mentioned longitudinal equal-width binning and longitudinal equal-frequency binning can also adopt a pre-trained binning type determination model and use a similar method to the above-mentioned embodiment to determine the binning type.

[0070] By adopting the above methods, the binning type can be determined effectively and accurately, which facilitates the subsequent accurate binning of the features of each participating node based on the binning type.

[0071] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure; Figure 2 As shown, this embodiment provides a feature-based binning method. Based on the above embodiment, taking horizontal equal-width binning as an example, this binning type is used to bin features on each participant node in the federated learning system. Specifically, the following steps may be included:

[0072] S201. Calculate the maximum and minimum values ​​of the eigenvalues ​​locally at each participant's node;

[0073] At each participant's local node, the eigenvalues ​​corresponding to each feature to be referenced for binning processing in all sample data are analyzed to obtain the maximum and minimum values ​​of the eigenvalues.

[0074] S202. Using secure multi-party computation (MPC) technology, based on the maximum and minimum local eigenvalues ​​of each participant, obtain the global maximum and minimum eigenvalues ​​of each participant's local nodes.

[0075] MPC technology is a theoretical framework for collaborative computing between a group of mutually distrusting nodes, while protecting privacy and without a trusted third party. Based on this MPC technology and the local maximum and minimum eigenvalues ​​obtained by each participating node, the global maximum and minimum eigenvalues ​​can be obtained locally on each participating node.

[0076] S203. At each participant's local node, determine the eigenvalue corresponding to the binning point of each bin based on the preset number of bins and the global maximum and global minimum values ​​corresponding to the eigenvalues, and obtain a binning point set;

[0077] The preset number of bins in this embodiment can be set by the staff according to actual needs and input from the outside. Correspondingly, each participant node in the federated learning system can obtain the preset number of bins. For example, the global maximum value of a certain feature value is max, the global minimum value is min, and the preset number of bins is b. The eigenvalues ​​corresponding to the bin points of each bin can be expressed as: s = min + (max-min) / b * t, t = 1, 2, ..., b. The eigenvalues ​​corresponding to the bin points of all bins based on the feature constitute the set of bin points corresponding to the feature.

[0078] S204. At each participant's local node, bin the local eigenvalues ​​according to the bin point set.

[0079] Each participating node can obtain the set of binning points corresponding to the feature locally in the manner described above. During the specific binning process, each participating node can perform binning on all feature values ​​corresponding to the feature field in the local sample data, based on the binning points of each bin in the obtained binning point set. According to the binning processing method of this embodiment, the binning standards for all participating nodes are the same, and the binning results are very reasonable and accurate.

[0080] The executor of the feature-based binning method of this embodiment is a feature-based binning device. The binning module for binning in the device can be embedded in each participant node in the federated learning system, so as to implement binning of all feature values ​​corresponding to the feature field in the sample data based on the feature locally in each participant node.

[0081] In practical applications, when building a model, it is usually necessary to refer to multiple features of the sample data. In this case, it is necessary to use the method of this embodiment to bin the feature values ​​corresponding to each feature in the sample data. The binning results corresponding to each feature are also different.

[0082] The following describes the horizontal equal-width binning process of this embodiment, taking m features as an example:

[0083] (a1) At each participant’s node P i Locally, calculate the maximum and minimum values ​​of each feature max ij 、min ij ; where i = 1, 2, 3…p, p is the number of participating nodes, j = 1, 2, 3…m, m is the number of features;

[0084] (b1) At each participant’s node P i Locally, use MPC technology to calculate and obtain the global maximum value of each feature max ij and the global minimum minij , j = 1, 2, 3…m, m is the number of features;

[0085] (c1) At each participant's node P i Locally, according to the preset number of bins b i ,i=1,2,3…m and the following formula (1) calculates the bin point set S={s ij |i=1,2,3…m,j=1,2,3…b i};

[0086]

[0087] (d1) At each participant’s node P i Locally, the samples are divided according to the feature bin point set S and the following formula (2), and the binning results are:

[0088] B={B ij |i=1,2,3…m,j=1,2,3…b i};

[0089] B ij ={f ij |s ij <f ik ≤s ij+1 ,i=1,2,3…m,k=1,2,3…n} (2)

[0090] Among them, f ij Represents the characteristic value of the jth bin of the i-th feature, which can also be called the characteristic data of the j-th bin of the i-th feature; s ij Indicates the lower limit of the binning point of the jth bin of the i-th feature, s ij+1 Indicates the upper limit of the binning point of the jth bin of the i-th feature, f ik Represents the kth feature data of the i-th feature, k = 1, 2, 3...n, and n is the number of feature data.

[0091] The feature-based binning method of this embodiment, by adopting the above-mentioned scheme, can adopt the horizontal equal-width binning method to effectively bin all the eigenvalues ​​of the feature on each participant's node based on the feature. During the binning process, the MPC technology can be used to obtain the global maximum and global minimum values ​​of each eigenvalue in each participant's node, which can effectively ensure the accuracy and effectiveness of obtaining the global maximum and global minimum values ​​of each eigenvalue on the basis of ensuring the data security of each participant's node. Furthermore, based on the global maximum and global minimum values ​​of each eigenvalue, the binning point set of the corresponding feature can be obtained, and based on the feature binning point set, all the eigenvalues ​​of the feature in the sample data can be effectively binned locally at each participant's node, which can effectively ensure the rationality and accuracy of the binning process and provide an effective basis for subsequent model training.

[0092] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure; Figure 3 As shown, this embodiment provides a feature-based binning method. Based on the above embodiment, taking horizontal equal-frequency binning as an example, this binning type is used to bin features on each participant node in the federated learning system. Specifically, the following steps may be included:

[0093] S301. At each participant's local node, sort the local eigenvalues ​​according to their magnitudes, and count the number of samples.

[0094] For example, the characteristic values ​​in the local sample data of each participant node can be sorted in ascending order of characteristic values. Of course, in actual applications, it is also possible to sort in descending order of characteristic values, with a similar principle. In this embodiment, after sorting, it is also necessary to count the number of characteristic data, i.e., characteristic values, of each participant node.

[0095] S302. At each participant's local node, based on the preset number of bins and number of samples, determine the characteristic values ​​corresponding to the bin points of each bin in an equal-frequency binning manner to obtain a bin point set;

[0096] Similarly, the preset number of bins can be set by the staff according to actual needs and input from the outside. The equal frequency binning method means that the number of samples included in each bin is the same, which is equal to the reciprocal of the number of bins. For example, if the preset number of bins for a certain feature is b, the corresponding eigenvalues ​​of the bin points of each bin can be obtained by traversing each eigenvalue locally on each participating node, so that the following formula is satisfied: Where t=1,2,3…,b,|B t | represents the feature data of the t-th bin of the feature, that is, the number of eigenvalues.

[0097] S303. At each participant's node, using MPC technology, the local bin point set is updated based on the bin point set of each participant's node.

[0098] S304. At each participant's local node, the feature values ​​in the local sample data are binned according to the updated binning point set.

[0099] That is, the initially determined binning point set at each participant's local location does not take into account the information of other participant nodes and may therefore be inaccurate. Therefore, in this embodiment, each participant's local node also needs to use MPC technology to update the local binning point set based on the binning point set of each participant node, while ensuring the data security of each participant's node. This allows the resulting binning point combination to reference the global information of each participant's node and be more accurate. Finally, each participant's local node can perform binning processing on the feature values ​​in the local sample data based on the feature values ​​of each binning point in the updated binning point set.

[0100] The feature-based binning method of this embodiment, by adopting the above-mentioned scheme, can effectively bin the sample data on each participant's node based on the features using a horizontal equal-frequency binning method. During this binning process, the local binning point set can be updated based on the binning point set of each participant's node through MPC technology, which can effectively ensure the accuracy of the updated binning point set while ensuring the data security of each participant's node. Furthermore, each participant's node can effectively bin the feature values ​​in the sample data locally based on the updated binning point set, effectively ensuring the rationality and accuracy of the binning process and providing an effective foundation for subsequent model training.

[0101] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure; Figure 4 As shown, this embodiment provides a feature-based binning method. Figure 3 Based on the embodiments, the technical solution of the present disclosure is described in more detail. Figure 4 As shown, in this embodiment, horizontal equal-frequency binning is still used as an example. This binning type is used to bin features on nodes of each participant in the federated learning system. Specifically, the following steps may be included:

[0102] S401. At each participant's local node, sort the local eigenvalues ​​according to their magnitudes, and count the number of samples.

[0103] S402. At each participant's local node, based on the preset number of bins and number of samples, determine the characteristic values ​​corresponding to the bin points of each bin in an equal-frequency binning manner to obtain a bin point set;

[0104] S403. Using the MPC technology, at each participant's local node, obtain a reference eigenvalue of the corresponding bin point based on the maximum eigenvalue and the minimum eigenvalue corresponding to each bin point of the same frequency in the bin point set of each participant's node;

[0105] Because equal-frequency binning is used, the number of bins in each participant's local binning point set is the same, and the proportion of feature data (i.e., eigenvalues) included in each bin is the same. However, the eigenvalues ​​in the sample data vary across different participant nodes, resulting in different eigenvalues ​​for bins of the same frequency across different participant nodes. For example, if the bins of each participant node are labeled 1, 2, 3, ..., b in sequential order, then the same binning point identifier corresponds to bins of the same frequency.

[0106] For example, in this embodiment, the reference eigenvalue of each binning point may be equal to the average of the maximum eigenvalue and the minimum eigenvalue corresponding to the binning point.

[0107] S404. At each participant's local node, based on the reference characteristic value of each bin point, re-count the number of characteristic values ​​of each bin;

[0108] The characteristic data of each bin refers to the characteristic value in the corresponding bin. The number of characteristic data in each bin refers to the number of characteristic values ​​in the corresponding bin. In other words, at each participating node, the reference characteristic value of each bin point is used as the characteristic value of the bin point, and the characteristic values ​​in the local sample data are re-binned to re-calculate the number of characteristic data in each bin.

[0109] S405. At each participant's local node, based on the re-counted number of eigenvalues ​​for each bin, the reference eigenvalue, the maximum eigenvalue, the minimum eigenvalue, the preset number of bins, and the number of samples, the eigenvalue of each bin point in the local bin point set is updated.

[0110] For example, at each participant node locally, based on the re-counted eigenvalues ​​of each bin, i.e., the number of eigenvalues ​​and the number of samples of each participant node, it can be determined whether the re-counted bins meet the equal frequency binning requirement, i.e., the reciprocal of the preset number of bins. If not, the eigenvalues ​​of the corresponding bin points can be updated based on the reference eigenvalues, maximum eigenvalues, and minimum eigenvalues ​​of each bin point. For example, for any participant node and any bin, if the ratio of the number of eigenvalues ​​of the bin based on the re-count to the number of samples of the current participant node is less than the reciprocal of the preset number of bins, the eigenvalue of the bin point corresponding to the bin can be taken as equal to the average of the maximum eigenvalue and the reference eigenvalue. If the ratio of the number of eigenvalues ​​of the bin based on the re-count to the number of samples of the current participant node is greater than the reciprocal of the preset number of bins, the eigenvalue of the bin point corresponding to the bin can be taken as equal to the average of the minimum eigenvalue and the reference eigenvalue. If the ratio of the number of eigenvalues ​​of the bin based on the re-count to the number of samples of the current participant node is equal to the inverse of the preset number of bins, the eigenvalue of the bin point corresponding to the bin can be taken as equal to the reference eigenvalue.

[0111] S406. At each participant's local node, check whether the absolute value error of the updated feature value of each bin point relative to the feature value of the corresponding bin point before the update is less than a preset threshold; if so, determine that the update is complete and execute step S407; otherwise, return to step S403 to continue the update;

[0112] S407. At each participant's local node, binning is performed on the local feature values ​​based on the updated binning point set and the features.

[0113] The following describes the horizontal equal-frequency binning process of this embodiment, taking m features as an example:

[0114] (a2) At each participant’s node P i Locally, sort the eigenvalues ​​in each sample data in ascending order, and count the number of local samples N i ; Where i = 1, 2, 3…p, p is the number of participating nodes;

[0115] (b2) At each participant’s node P i Locally, traverse the characteristic values ​​of each latitude, according to the number of bins b i (i=1,2,3…m, m is the number of features) and the following formula (2) and formula (3) are used to calculate the feature bin point set S={s ij |i=1,2,3…m,j=1,2,3…b}.

[0116]

[0117] (c2) At each participant P i Locally, the node uses MPC technology to calculate the maximum eigenvalue of the same feature at the corresponding bin point at the same frequency. Minimum eigenvalue segmentation value Calculate the reference eigenvalue corresponding to each bin point according to the following formula (4):

[0118]

[0119] (d2) At each participant P i Locally, the reference eigenvalues ​​corresponding to each bin point As the eigenvalue of the corresponding bin point of the bin, re-count the number of eigenvalues ​​of each bin | B ij |, and update the eigenvalue s of the corresponding bin point according to the following formula (5) ij ;

[0120]

[0121] (e2) Repeat the above steps (c2) and (d2) until the absolute value of the error between the two binning points at the same frequency is less than ε, and the update is completed.

[0122] (f2) At each participant's node P i Locally, the eigenvalues ​​are divided according to the binning point set S and the above formula (2), and the binning result B = {B ij |i=1,2,3…m,j=1,2,3…b i The feature-based binning processing method of this embodiment, by adopting the above-mentioned scheme and the horizontal equal-frequency binning method, can effectively update the feature value of each binning point in the local binning point set of each participating node, thereby ensuring the accuracy of the feature value of each binning point in the final updated binning point combination; furthermore, each participating node can effectively bin the feature values ​​in the sample data based on the updated binning point set, thereby effectively ensuring the rationality and accuracy of the binning processing and providing an effective basis for subsequent model training.

[0123] Figure 5 is a schematic diagram according to the fifth embodiment of the present disclosure; Figure 5 As shown, this embodiment provides a feature-based binning method. Based on the above embodiment, taking horizontal chi-square binning as an example, this binning type is used to bin features on each participant node in the federated learning system. Specifically, the following steps may be included:

[0124] S501. Using MPC technology, based on the size of the local eigenvalues ​​of each participant's node, globally sort the eigenvalues ​​in the sample data of all participant nodes; and count the number of samples;

[0125] First, all eigenvalues ​​in the sample data can be sorted locally on each participant node by their eigenvalues, for example in ascending order, and the number of samples in that participant node can be counted. In practical applications, sorting can also be done in descending order.

[0126] Then, through MPC technology, based on the local sorting of each participant's node, all feature values ​​in the sample data of all participant nodes are globally sorted, and the number of samples is counted.

[0127] S502. At each participant's local node, divide each eigenvalue into a bin based on the global sorting result;

[0128] During the initial binning, each eigenvalue is divided into one bin.

[0129] S503. Based on the MPC technology, the chi-square value between each pair of adjacent bins is calculated locally at each participating node.

[0130] S504: merge a pair of adjacent bins with the smallest chi-square value;

[0131] For example, step S503 may include the following steps during specific implementation:

[0132] (a3) Using MPC technology, the number of samples of each type of label corresponding to the feature value in each bin and the proportion of samples of each type of label in the global sample are counted locally on each participant node;

[0133] For example, each participant node can include many pieces of sample data including the feature. These sample data include labels. And the labels in different sample data can be the same or different. The labels in all sample data can be classified into K categories. The feature data in the bin refers to the feature values ​​included in the bin. And each feature value corresponds to a piece of sample data locally at the participant node. And the sample data includes labels. Therefore, the multiple feature values ​​included in the feature data can correspond to multiple types of label samples. At each participant node, the labels of the samples corresponding to each feature value in each bin can be obtained, and then the number of samples of each type of label corresponding to the feature values ​​in each bin and the proportion of samples of each type of label in the global sample can be counted.

[0134] (b3) Using MPC technology, count the global sample number of each participant node in each bin;

[0135] Furthermore, after knowing the global sample count of each participant's node in each local bin, the global sample count of each bin can be calculated based on MPC technology. In other words, the global sample count refers to the total number of feature samples of all participant nodes included in the same bin.

[0136] (c3) Locally at each participant node, the chi-square value between each pair of adjacent bins is calculated based on the proportion of samples of each type of label in each bin in the global sample, the number of samples of each type of label, the number of global samples of each participant node, and the total number of samples.

[0137] Further optionally, in one embodiment of the present disclosure, after step S504, the merged bins may not yet meet the requirements of chi-square binning. In this case, the following steps may be further included:

[0138] S505. Continuing to calculate the chi-square value between each pair of adjacent bins locally on each participant's node based on the MPC technology;

[0139] S506. At each participant's local node, check whether the minimum chi-square value exceeds a preset specified threshold, or whether the number of bins is less than a preset specified number; if so, determine each bin to obtain the binning result; otherwise, if not, return to step S504 and continue processing.

[0140] In this embodiment, the determination of each bin is to determine the characteristic samples in each bin, and no merging changes will occur.

[0141] The following describes the horizontal chi-square binning process of this embodiment, taking m features as an example:

[0142] (a4) At each participant's node P i Locally, sort the eigenvalues ​​in ascending order and count the number of samples N i , i = 1, 2, 3…p, p is the number of participating nodes; then the MPC technology is used to globally sort the eigenvalues ​​of all participating nodes with the same characteristics;

[0143] (b4) At each participant's node P i , traverse the eigenvalues ​​of each sample data and divide each eigenvalue into a separate bin b jt where j = 1, 2, 3…m is the number of features, t = 1, 2, 3…N i ;

[0144] (c4) At each participant's node P i Locally, using MPC technology, calculate the feature data in each bin, that is, the proportion C of the sample data of the k-th label corresponding to the feature value in the total sample data k(k=1,2,3…K); K is the total number of classifications of the labels of the sample data;

[0145] (d4) At each participant's node P i Local, statistical features f j (j=1,2,3…m is the number of features) in bin b t The number of samples n ijt (1≤t<<|B ij |), using MPC technology to count samples in bin b jt The number of global samples n jt ;

[0146] (e4) At each participant's node P i Locally, for example, the chi-square value of each pair of adjacent bin intervals can be calculated according to the following formulas (6) and (7): (j=1, 2, 3…m is the number of features), and merge the pair of bins with the smallest chi-square value;

[0147]

[0148]

[0149] Among them, A tk is the number of samples of the kth class label in the tth bin, E tk The expected number of samples of the k-th category label in the t-th bin; N jt is the number of eigenvalues ​​in the tth bin of the jth feature. Optionally, the calculation of the chi-square value of each pair of adjacent bin intervals in this embodiment can also refer to the relevant chi-square value calculation method, which is not limited here.

[0150] (f4) Repeat steps (d4) and (e4) until feature f j Minimum chi-square value (j=1,2,3…m is the number of features) Exceeds the specified threshold e or the number of bins is less than the specified number b j So far, we can get the binning result B={B ij |i=1,2,3…m,j=1,2,3…b i}.

[0151] The feature-based binning method of this embodiment, by adopting the above-mentioned solution, can effectively bin the feature values ​​on each participant's node using horizontal chi-square binning. During this binning process, MPC technology can be used to calculate the chi-square value between each pair of adjacent bins locally on each participant's node; and the pair of adjacent bins with the smallest chi-square value is merged. This can effectively bin and merge, while ensuring the data security of each participant's node, to obtain more accurate binning, providing an effective foundation for subsequent model training.

[0152] Figure 6 is a schematic diagram according to the sixth embodiment of the present disclosure; Figure 6 As shown, this embodiment provides a feature-based binning method. Based on the above embodiment, vertical chi-square binning is used as an example. This binning method is used to bin features on each participant node in the federated learning system. Specifically, the following steps may be included:

[0153] S601. At each first participant node, sort the eigenvalues ​​in each sample data according to the size of the eigenvalues; and count the total number of samples;

[0154] After vertical chi-square binning, the sample data of all participating nodes has completed the sample alignment process. In the federated learning system of this embodiment, each participating node includes multiple first participating nodes and second participating nodes, where the sample data of each first participating node includes features, and the second participating nodes include label data.

[0155] S602: At each first participant's local node, divide the characteristic values ​​in each sample data into a bin based on the sorting result;

[0156] In practical applications, each piece of sample data corresponds to a feature field, that is, a feature value corresponding to the feature. During the initial binning, each feature value is divided into a bin.

[0157] S603. Calculate the chi-square value between each pair of adjacent bins locally on each first participant node, with reference to the sample identifier of the label in the second participant node;

[0158] S604: Merge a pair of adjacent bins with the smallest chi-square value;

[0159] Since adjacent bins have the closest eigenvalues, if the chi-square value of the adjacent bins is the smallest, it means that the adjacent bins have strong commonalities, so they can be merged into one bin, which can effectively merge feature samples and improve binning efficiency.

[0160] For example, step S603 calculates the chi-square value between each pair of adjacent bins locally at each first participant node, with reference to the sample identifier of the label in the second participant node. Specifically, the steps may include:

[0161] (a5) Locally, on each first participant node, take the intersection of the identifier of the sample data corresponding to the feature value in each bin and the sample identifier of the label in the second participant node to obtain the intersection sample identifier;

[0162] For example, in this embodiment, the intersection of the identifiers of the sample data corresponding to the eigenvalues ​​in each bin and the sample identifiers of the labels in the second participant node can be calculated based on the GBF-Privacy Preserving Set Intersection (PSI) algorithm. The relevant GBF-PSI algorithm can refer to the relevant existing technology and will not be repeated here.

[0163] (b5) At the local node of the second participant, based on the intersection sample identifier, count the number of samples of each type of label in each bin and the proportion of samples of each type of label in the global sample;

[0164] (c5) sending the number of samples of each type of label in each bin and the proportion of samples of each type of label in the global sample to each first participant node locally at the second participant node;

[0165] (d5) Locally at each first participant node, the chi-square value between each pair of adjacent bins is calculated based on the proportion of samples of each type of label in each bin in the global sample, the number of samples of each type of label, the number of samples local to each first participant node, and the total number of samples.

[0166] Further optionally, in one embodiment of the present disclosure, after step S604, the merged bins may not yet meet the requirements of chi-square binning. In this case, the following steps may be further included:

[0167] S605: Continuing to calculate the chi-square value between each pair of adjacent bins after merging, locally on each first participant node, with reference to the sample identifier of the label in the second participant node;

[0168] For example, the specific calculation process can refer to the above steps (a5)-(d5).

[0169] S606: At each first participant's local node, check whether the minimum chi-square value exceeds a preset threshold, or whether the number of bins is less than a preset number. If so, determine each bin, thus obtaining the binning results. End. If not, return to step S604 and continue processing.

[0170] The following is a binning process with reference to m features, with the first participant node Pi (i=1,2…s), with eigenvalue X i s is the number of nodes of the first participant. Taking the second participant node P0 having the label data Y as an example, the horizontal chi-square binning process of this embodiment is described:

[0171] (a6) At each first participant node P i (i=1,2,3…s) locally, sort the eigenvalues ​​in ascending order, and count the number of eigenvalues ​​N, which is also the corresponding number of samples N;

[0172] (b6) At each first participant node P i Locally, traverse the eigenvalues ​​in each sample data and divide the eigenvalues ​​in each sample data into a separate bin b jt where j = 1, 2, 3…m is the number of features and t = 1, 2, 3…N;

[0173] (c6) At each first participant node P i Locally, according to bin b jt The sample identification id corresponding to the eigenvalue inside x The sample ID corresponding to the label data Y y The intersection is obtained by the GBF-PSI algorithm, and the intersection sample ID is obtained. join ;

[0174] (d6) At the second participant node P0, according to the intersection sample identification id join Statistical bin b jt The distribution of samples corresponding to the eigenvalues ​​in each label category k, and the proportion C of samples with the kth label in the total number of samples N are calculated. k (k=1,2,3…K); K is the total number of classifications of the labels of the sample data;

[0175] (e6) At the second participant node P0, the sub-box b jt The sample distribution information is sent to other participants P i (i=1,2,3…s);

[0176] For example, bin b jt The sample distribution information can include bin b jt The number of samples of each type of label in the dataset and the proportion of samples of each type of label in the global samples.

[0177] (f6) At each first participant node P i Locally, according to the above formula (6) and formula (7), calculate the chi-square value of each pair of adjacent bins And merge the pair of bins with the smallest chi-square value;

[0178] (g6) Repeat steps (c6) to (f6) until feature f j Minimum chi-square value (j=1,2,3…m is the number of features) Exceeds the specified threshold e or the number of bins is less than the specified number b j So far, we can get the binning result B={B ij |i=1,2,3…m,j=1,2,3…b i}.

[0179] The feature-based binning processing method of this embodiment, by adopting the above-mentioned scheme, can adopt the vertical chi-square binning method, based on the GBF-PSI algorithm, to obtain the parameter information required to calculate the chi-square value between adjacent bins, and then can merge adjacent bins based on the minimum chi-square value. It can effectively bin and merge on the basis of ensuring the data security of each participating node, obtain more accurate binning, and provide an effective basis for subsequent model training.

[0180] Figure 7 is a schematic diagram according to the seventh embodiment of the present disclosure; Figure 7 As shown, this embodiment provides a feature-based binning processing device 700, including:

[0181] Feature acquisition module 701, used to obtain features to be referenced for binning processing;

[0182] A first determination module 702 is configured to determine a federated learning scenario based on the distribution of feature fields and sample data corresponding to the features in the nodes of each participant in the federated learning system;

[0183] The second determination module 703 is configured to determine the binning type based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each participant's node or the preset binning requirements, and with reference to the federated learning scenario;

[0184] The binning processing module 704 is configured to perform binning processing on the features on each participant node in the federated learning system using a binning type.

[0185] The feature-based binning processing device 700 of this embodiment implements the feature-based binning processing by adopting the above-mentioned modules, and the implementation principle is the same as that of the above-mentioned related method embodiments. For details, please refer to the records of the above-mentioned related embodiments, which will not be repeated here.

[0186] Figure 8 is a schematic diagram according to the eighth embodiment of the present disclosure; Figure 8 As shown, this embodiment provides a feature-based binning processing device 800, including: Figure 7The modules with the same name and function are the feature acquisition module 801 , the first determination module 802 , the second determination module 803 and the binning processing module 804 .

[0187] In this embodiment, the first determining module 802 is configured to:

[0188] If the overlap ratio of feature fields included in sample data with different identifiers in different participating nodes is greater than a preset ratio threshold, the federated learning scenario is determined to be horizontal federated learning; or

[0189] If the identification overlap ratio of sample data included in the nodes of different participants in the federated learning system is greater than a preset ratio threshold, the federated learning scenario is determined to be vertical federated learning.

[0190] like Figure 8 As shown, in one embodiment of the present disclosure, the feature-based binning processing device 800 further includes:

[0191] An alignment processing module 805 is used to perform feature alignment processing on all sample data included in different participant nodes in the federated learning system; or

[0192] Perform sample alignment on all sample data included in the nodes of different participants in the federated learning system.

[0193] In one embodiment of the present disclosure, the second determining module 803 is configured to:

[0194] If the federated learning scenario is horizontal federated learning, the eigenvalues ​​corresponding to the features in the sample data on each participant's node in the federated learning system are evenly distributed, and the binning type is determined to be horizontal equal-width binning;

[0195] If the federated learning scenario is horizontal federated learning, and the eigenvalues ​​corresponding to the features in the sample data on the nodes of each participant in the federated learning system are concentrated in multiple preset intervals, the binning type is determined to be horizontal equal-frequency binning;

[0196] If the federated learning scenario is horizontal federated learning, the pre-set binning requirements include goodness of fit and / or independence test requirements, and the binning type is determined to be horizontal chi-square binning; or

[0197] If the federated learning scenario is vertical federated learning, the preset binning requirements include goodness of fit and / or independence test requirements, and the binning type is determined to be vertical chi-square binning.

[0198] In one embodiment of the present disclosure, the second determining module 803 is configured to:

[0199] Based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each participant's node and / or the preset binning requirements, and with reference to the federated learning scenario, a pre-trained binning type determination model is used to determine the binning type.

[0200] In one embodiment of the present disclosure, the binning processing module 804 can implement different functions in the following four different scenarios.

[0201] Scenario 1: If the binning type is horizontally equal-width binning, the binning processing module 804 is used to:

[0202] Calculate the maximum and minimum values ​​of the eigenvalues ​​locally on each participant's node;

[0203] Through multi-party secure computing technology, based on the maximum and minimum values ​​of the local feature values ​​of each participant, the global maximum and minimum values ​​of the feature are obtained locally on each participant's node;

[0204] At each participant's local node, based on the preset number of bins and the global maximum and minimum values ​​corresponding to the features, the feature values ​​corresponding to the bin points of each bin are determined to obtain a bin point set;

[0205] At each participant's local node, the local feature values ​​are binned based on the features according to the binning point set.

[0206] Scenario 2: If the binning type is horizontal equal-frequency binning, the binning processing module 804 is further configured to:

[0207] Sort the local eigenvalues ​​of each participating node according to the size of the eigenvalues ​​corresponding to the features; and count the number of samples;

[0208] At each participant's local node, based on the preset number of bins and number of samples, the eigenvalues ​​corresponding to the bin points of each bin are determined in an equal-frequency binning manner to obtain a local bin point set;

[0209] At each participant's local node, the local bin point set is updated based on the bin point set of each participant's node through multi-party secure computing technology;

[0210] At each participant's local node, the local feature values ​​are binned based on the updated bin point set and the features.

[0211] Furthermore, the binning processing module 804 is configured to:

[0212] At each participant's local node, through multi-party secure computing technology, based on the maximum eigenvalue and minimum eigenvalue corresponding to each bin point of the same frequency in the bin point set of each participant's node, the reference eigenvalue of the corresponding bin point is obtained;

[0213] At each participant's local node, the number of eigenvalues ​​of each bin is re-counted based on the reference eigenvalues ​​of each bin point;

[0214] At each participating node, the eigenvalue of each bin point in the local bin point set is updated based on the re-counted number of eigenvalues ​​of each bin, the global eigenvalue of each bin point, the maximum eigenvalue, the minimum eigenvalue, the preset number of bins, and the number of samples.

[0215] Furthermore, the binning processing module is also used to:

[0216] Detect and determine the eigenvalue of each bin point in the updated bin point set, and the absolute value error of the eigenvalue of the bin point before the update is less than a preset threshold.

[0217] Furthermore, the binning processing module is also used to:

[0218] If the absolute value error of the eigenvalue of each binning point in the updated binning point set relative to the eigenvalue of the corresponding binning point before the update is not less than the preset threshold, the local binning point set of each participating node will continue to be updated based on the binning point set of each participating node through multi-party secure computing technology.

[0219] Scenario 3: If the binning type is horizontal chi-square binning, the binning processing module 804 is used to:

[0220] Through multi-party secure computing technology, the eigenvalues ​​of all participants are globally sorted based on the size of the eigenvalues ​​corresponding to the local features of each participant's node; and the number of samples is counted;

[0221] At each participant's local node, each eigenvalue is divided into a bin based on the global sorting result;

[0222] Based on multi-party secure computing technology, the chi-square value of each bin is calculated locally on each participant's node;

[0223] Merge the pair of adjacent bins with the smallest difference in chi-square values.

[0224] Furthermore, the binning processing module 804 is further configured to:

[0225] After merging the adjacent bins with the smallest chi-square value, the chi-square value between each pair of adjacent bins is calculated locally on each participating node based on multi-party secure computing technology.

[0226] Check whether the minimum chi-square value exceeds the preset threshold, or whether the number of bins is less than the preset number;

[0227] If so, determine each bin;

[0228] If not, merge the pair of adjacent bins with the smallest chi-square value.

[0229] Furthermore, the binning processing module 804 is configured to:

[0230] Through multi-party secure computing technology, the number of samples of each type of label corresponding to the characteristic value in each bin and the proportion of samples of each type of label in the global sample are counted locally on each participant's node;

[0231] Through multi-party secure computing technology, the number of global samples corresponding to the characteristic values ​​in each bin in each participating node is counted;

[0232] Locally at each participating node, the chi-square value between each pair of adjacent bins is calculated based on the proportion of samples of each type of label corresponding to the characteristic value in each bin in the global sample, the number of samples of each type of label, the number of global samples at each participating node, and the total number of samples.

[0233] In scenario 4, if the binning type is vertical chi-square binning, each participant node includes multiple first participant nodes and second participant nodes, the sample data in each first participant node includes features, and the second participant node includes label data. The binning processing module 804 is used to:

[0234] At each first-party node, sort the eigenvalues ​​according to the size of the eigenvalues ​​corresponding to the features; and count the total number of samples; each eigenvalue corresponds to one sample;

[0235] At each first-party node, the feature values ​​in each sample data corresponding to the feature are divided into a bin based on the sorting result;

[0236] Calculate the chi-square value of each bin locally on each first participant node, referring to the sample identifier of the label in the second participant node;

[0237] Merge the pair of adjacent bins with the smallest chi-square value.

[0238] Furthermore, the binning processing module 804 is further configured to:

[0239] After merging a pair of adjacent bins with the smallest chi-square value, continue to calculate the chi-square value between each pair of adjacent bins after merging, locally on each first participant node, with reference to the sample identifier of the label in the second participant node;

[0240] Check whether the minimum chi-square value exceeds the preset threshold, or whether the number of bins is less than the preset number;

[0241] If so, determine each bin;

[0242] If not, merge the pair of adjacent bins with the smallest chi-square value.

[0243] Furthermore, the binning processing module 804 is configured to:

[0244] Locally, on each first participant node, take the intersection of the identifiers of the sample data in each bin and the sample identifiers of the labels in the second participant node to obtain an intersection sample identifier;

[0245] At the local node of the second participant, based on the intersection sample identifier, the number of samples of each type of label corresponding to the characteristic value in each bin and the proportion of samples of each type of label in the global sample are counted;

[0246] The second participant node locally sends to each first participant node the number of samples of each type of label corresponding to the characteristic value in each bin and the proportion of samples of each type of label in the global sample;

[0247] Locally at each first participant node, the chi-square value between each pair of adjacent bins is calculated based on the proportion of each type of label samples corresponding to the characteristic values ​​in each bin in the global sample, the number of samples of each type of label, the number of samples local to each first participant node, and the total number of samples.

[0248] The binning processing module 804 of this embodiment can be transplanted into each participant node in the federated learning system to implement effective binning processing of features in the sample data based on features on each participant node.

[0249] The feature-based binning processing device 800 of this embodiment implements the feature-based binning processing by adopting the above-mentioned modules, and the implementation principle is the same as that of the above-mentioned related method embodiments. For details, please refer to the records of the above-mentioned related embodiments, which will not be repeated here.

[0250] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0251] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0252] Figure 9A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0253] like Figure 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0254] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0255] The computing unit 901 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the above-mentioned methods of the present disclosure. For example, in some embodiments, the above-mentioned methods of the present disclosure can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the above-mentioned methods of the present disclosure described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the above-mentioned methods of the present disclosure by any other appropriate means (e.g., by means of firmware).

[0256] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0257] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0258] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0259] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0260] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0261] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0262] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0263] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A feature-based binning method, comprising: Get the features to be referenced for binning; Determine the federated learning scenario based on the fields of the features in the nodes of each participant in the federated learning system and the distribution of sample data corresponding to the features; Determine the binning type based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each of the participant nodes or the preset binning requirements, and with reference to the federated learning scenario; Using the binning type to perform binning processing on the features on each of the participant nodes in the federated learning system; wherein using the binning type to perform binning processing on the features on each of the participant nodes in the federated learning system includes: If the binning type is horizontal equal-frequency binning, the local eigenvalues ​​of each participant node are sorted according to the size of the eigenvalues ​​corresponding to the features; and the number of samples is counted; Locally, at each participant node, based on a preset number of bins and the number of samples, determining the characteristic value corresponding to the binning point of each bin in an equal-frequency binning manner to obtain a local binning point set; At each participant node, using a multi-party secure computing technology, based on the binning point set of each participant node, the local binning point set is updated; At each participant node, binning the local feature values ​​according to the updated binning point set; Wherein, at each participant node, using a multi-party secure computing technology, based on the binning point set of each participant node, updating the local binning point set includes: Locally at each participating node, the eigenvalue corresponding to each binning point in the local binning point set is updated based on the re-counted number of eigenvalues ​​of each binning point, the global eigenvalue, the maximum eigenvalue, the minimum eigenvalue of each binning point, the preset number of bins and the number of samples.

2. The method according to claim 1, wherein Based on the fields of the features in the nodes of each participant in the federated learning system and the distribution of sample data corresponding to the features, the federated learning scenario is determined, including: If the overlap ratio of the feature fields included in the sample data with different identifiers in different participant nodes is greater than a preset ratio threshold, determining that the federated learning scenario is horizontal federated learning; or If the identification overlap ratio of the sample data included in different participant nodes in the federated learning system is greater than the preset ratio threshold, it is determined that the federated learning scenario is vertical federated learning.

3. The method according to claim 2, wherein: After determining that the federated learning scenario is horizontal federated learning, and before using the binning type to bin the features on each of the participant nodes in the federated learning system, the method includes: Performing feature alignment processing on all sample data included in different participant nodes in the federated learning system; or After determining that the federated learning scenario is vertical federated learning, and before using the binning type to bin the feature values ​​on each of the participant nodes in the federated learning system based on the feature, the method includes: Perform sample alignment processing on all sample data included in different participant nodes in the federated learning system.

4. The method according to claim 3, wherein: Based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each of the participant nodes or the preset binning requirements, and with reference to the federated learning scenario, the binning type is determined, including: If the federated learning scenario is horizontal federated learning, and the characteristic values ​​corresponding to the characteristics in the sample data on each of the participant nodes in the federated learning system are evenly distributed, the binning type is determined to be horizontal equal-width binning; If the federated learning scenario is horizontal federated learning, and the feature values ​​corresponding to the features in the sample data on the participating nodes in the federated learning system are concentrated in multiple preset intervals, the binning type is determined to be horizontal equal-frequency binning; If the federated learning scenario is horizontal federated learning, the pre-set binning requirements include goodness of fit and / or independence test requirements, and the binning type is determined to be horizontal chi-square binning; or If the federated learning scenario is vertical federated learning, the preset binning requirements include goodness of fit and / or independence test requirements, and the binning type is determined to be vertical chi-square binning.

5. The method according to claim 3, wherein: Based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each of the participant nodes or the preset binning requirements, and with reference to the federated learning scenario, the binning type is determined, including: Based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each of the participating nodes and / or the preset binning requirements, and with reference to the federated learning scenario, a pre-trained binning type determination model is used to determine the binning type.

6. The method according to claim 4 or 5, wherein: Using the binning type to perform binning processing on the features on each of the participant nodes in the federated learning system, further comprising: If the binning type is horizontal equal-width binning, the maximum and minimum values ​​of the feature values ​​corresponding to the feature are calculated locally on each of the participant nodes; Using a multi-party secure computing technique, based on the maximum and minimum values ​​of the local eigenvalues ​​of each participant, a global maximum and a global minimum value of the eigenvalue are obtained locally at each participant's node; Locally, at each participant node, determine the eigenvalue corresponding to the binning point of each bin according to the preset number of bins and the global maximum and global minimum values ​​of the eigenvalues, and obtain a binning point set; At each participant node, the local feature values ​​are binned according to the binning point set.

7. The method according to claim 1, wherein Before updating the eigenvalue corresponding to each binning point in the local binning point set at each participant node based on the re-counted number of eigenvalues ​​of each binning point, the global eigenvalue, the maximum eigenvalue, the minimum eigenvalue of each binning point, the preset number of bins, and the number of samples, the method further includes: At each participant node, using a multi-party secure computing technique, based on the maximum eigenvalue and the minimum eigenvalue corresponding to each binning point of the same frequency in the binning point set of each participant node, a reference eigenvalue of the corresponding binning point is obtained; At each participant node locally, the number of characteristic values ​​of each bin is recounted based on the reference characteristic value of each bin point.

8. The method according to claim 7, wherein: After each participant node locally updates the local binning point set based on the binning point set of each participant node using a multi-party secure computing technique, and before each participant node locally bins the local feature value based on the feature according to the updated binning point set, the method further includes: Detect and determine the characteristic value corresponding to each binning point in the updated binning point set, and determine that the absolute value error of the characteristic value of the binning point before the update is less than a preset threshold.

9. The method according to claim 8, wherein The method further comprises: If the absolute value error of the characteristic value of each binning point in the updated binning point set relative to the characteristic value of the corresponding binning point before the update is not less than the preset threshold, continue to update the local binning point set based on the binning point set of each participant node through multi-party secure computing technology.

10. The method according to claim 4 or 5, wherein: Using the binning type, binning the features on each of the participant nodes in the federated learning system includes: If the binning type is horizontal chi-square binning, the eigenvalues ​​of all participants are globally sorted based on the eigenvalues ​​corresponding to the local features of each participant's node using multi-party secure computing technology; and the number of samples of each participant's node is counted; Locally, at each participant node, each characteristic value is divided into a bin according to the global sorting result; Based on the multi-party secure computing technology, the chi-square value between each pair of adjacent bins is calculated locally at each participant node; A pair of adjacent bins with the smallest chi-square value is merged.

11. The method according to claim 10, wherein: After merging the pair of adjacent bins with the smallest chi-square value, the method further includes: Continuing to calculate the chi-square value between each pair of adjacent bins locally at each participant node based on the multi-party secure computing technology; Detecting whether the minimum chi-square value exceeds a preset specified threshold, or whether the number of the bins is less than a preset specified number; If so, determining each of the bins; If not, merge the pair of adjacent bins with the smallest chi-square value.

12. The method according to claim 10, wherein: Based on the multi-party secure computing technology, the chi-square value between each pair of adjacent bins is calculated locally at each participant node, including: By using multi-party secure computing technology, the number of samples of each type of label corresponding to the characteristic value in each bin and the proportion of samples of each type of label in the global sample are counted locally on each participant node; By using multi-party secure computing technology, the number of global samples corresponding to the characteristic values ​​in each of the bins in each of the participant nodes is counted; Locally at each participating node, the chi-square value between each pair of adjacent bins is calculated based on the proportion of samples of each type of label corresponding to the characteristic values ​​in each bin in the global sample, the number of samples of each type of label, the number of global samples at each participating node, and the number of samples at each participating node.

13. The method according to claim 4 or 5, wherein: Each of the participant nodes includes a plurality of first participant nodes and second participant nodes, sample data in each of the first participant nodes includes the feature; and the second participant nodes include label data; and using the binning type, binning the features on each of the participant nodes in the federated learning system includes: If the binning type is vertical chi-square binning, at each of the first participant nodes, the eigenvalues ​​are sorted according to the size of the eigenvalues ​​corresponding to the features; and the total number of samples is counted; Locally, at each of the first participant nodes, the feature values ​​in each of the sample data corresponding to the feature are divided into a bin according to the sorting result; Locally, at each of the first participant nodes, referring to the sample identifiers of the labels in the second participant nodes, calculating the chi-square value between each pair of adjacent bins; A pair of adjacent bins with the smallest chi-square value is merged.

14. The method according to claim 13, wherein After merging the pair of adjacent bins with the smallest chi-square value, the method further includes: Continuing to calculate the chi-square value between each pair of adjacent bins after merging, locally on each of the first participant nodes, with reference to the sample identifiers of the labels in the second participant nodes; Detecting whether the minimum chi-square value exceeds a preset specified threshold, or whether the number of the bins is less than a preset specified number; If so, determining each of the bins; If not, merge the pair of adjacent bins with the smallest chi-square value.

15. The method according to claim 13, wherein Calculating, locally at each of the first participant nodes, a chi-square value between each pair of adjacent bins with reference to the sample identifier of the label in the second participant node, including: Locally, on each of the first participant nodes, taking the intersection of the identifier of the sample data corresponding to the feature value in each of the bins and the sample identifier of the label in the second participant node to obtain an intersection sample identifier; Locally, at the second participant node, based on the intersection sample identifier, counting the number of samples of each type of label corresponding to the feature value in each of the bins and the proportion of samples of each type of label in the global sample; The second participant node locally sends to each of the first participant nodes the number of samples of each type of label corresponding to the feature value in each of the bins and the proportion of samples of each type of label in the global sample; Locally at each of the first participant nodes, the chi-square value between each pair of adjacent bins is calculated based on the proportion of samples of each type of label corresponding to the characteristic values ​​in each bin in the global sample, the number of samples of each type of label, the number of samples local to each of the first participant nodes, and the total number of samples.

16. A feature-based binning processing device, comprising: Feature acquisition module, used to obtain the features to be referenced for binning processing; A first determination module is configured to determine a federated learning scenario based on the fields of the features in the nodes of each participant in the federated learning system and the distribution of sample data corresponding to the features; A second determination module is configured to determine a binning type based on a distribution property of a feature value corresponding to the feature in the sample data on each of the participant nodes or a preset binning requirement, and with reference to the federated learning scenario; A binning processing module, configured to perform binning processing on the features on each of the participant nodes in the federated learning system using the binning type; The binning processing module is used to: If the binning type is horizontal equal-frequency binning, the local eigenvalues ​​of each participant node are sorted according to the size of the eigenvalues ​​corresponding to the features; and the number of samples is counted; Locally, at each participant node, based on a preset number of bins and the number of samples, determining the characteristic value corresponding to the binning point of each bin in an equal-frequency binning manner to obtain a local binning point set; At each participant node, using a multi-party secure computing technology, based on the binning point set of each participant node, the local binning point set is updated; At each participant node, binning the local feature values ​​based on the features according to the updated binning point set; Wherein, the binning processing module is used to: Locally at each participating node, the eigenvalue corresponding to each binning point in the local binning point set is updated based on the re-counted number of eigenvalues ​​of each binning point, the global eigenvalue, the maximum eigenvalue, the minimum eigenvalue of each binning point, the preset number of bins and the number of samples.

17. The device according to claim 16, wherein The first determining module is configured to: If the overlap ratio of the feature fields included in the sample data with different identifiers in different participant nodes is greater than a preset ratio threshold, determining that the federated learning scenario is horizontal federated learning; or If the identification overlap ratio of the sample data included in different participant nodes in the federated learning system is greater than the preset ratio threshold, it is determined that the federated learning scenario is vertical federated learning.

18. The device according to claim 17, wherein The device further comprises: an alignment processing module, configured to perform feature alignment processing on all sample data included in different participant nodes in the federated learning system; or Perform sample alignment processing on all sample data included in different participant nodes in the federated learning system.

19. The device according to claim 18, wherein The second determining module is further configured to: If the federated learning scenario is horizontal federated learning, and the characteristic values ​​corresponding to the characteristics in the sample data on each of the participant nodes in the federated learning system are evenly distributed, the binning type is determined to be horizontal equal-width binning; If the federated learning scenario is horizontal federated learning, and the feature values ​​corresponding to the features in the sample data on the participating nodes in the federated learning system are concentrated in multiple preset intervals, the binning type is determined to be horizontal equal-frequency binning; If the federated learning scenario is horizontal federated learning, the pre-set binning requirements include goodness of fit and / or independence test requirements, and the binning type is determined to be horizontal chi-square binning; or If the federated learning scenario is vertical federated learning, the preset binning requirements include goodness of fit and / or independence test requirements, and the binning type is determined to be vertical chi-square binning.

20. The apparatus according to claim 18, wherein The second determining module is configured to: Based on the distribution properties of the feature values ​​corresponding to the features in the sample data on each of the participating nodes and / or the preset binning requirements, and with reference to the federated learning scenario, a pre-trained binning type determination model is used to determine the binning type.

21. The device according to claim 19 or 20, wherein The binning processing module is used to: If the binning type is horizontal equal-width binning, the maximum and minimum values ​​of the feature values ​​corresponding to the feature are calculated locally on each of the participant nodes; Using a multi-party secure computing technique, based on the maximum and minimum values ​​of the local eigenvalues ​​of each participant, a global maximum and a global minimum value of the eigenvalue are obtained locally at each participant's node; Locally, at each participant node, determine the feature value corresponding to the binning point of each bin according to the preset number of bins and the global maximum and global minimum values ​​corresponding to the feature, and obtain a binning point set; At each participant node, the local feature values ​​are binned based on the features according to the binning point set.

22. The apparatus according to claim 16, wherein The binning processing module is used to: At each participant node, using a multi-party secure computing technique, based on the maximum eigenvalue and the minimum eigenvalue corresponding to each binning point of the same frequency in the binning point set of each participant node, a reference eigenvalue of the corresponding binning point is obtained; At each participant node locally, the number of characteristic values ​​of each bin is recounted based on the reference characteristic value of each bin point.

23. The device according to claim 22, wherein The binning processing module is further used to: Detect and determine the characteristic value corresponding to each binning point in the updated binning point set, and determine that the absolute value error of the characteristic value of the binning point before the update is less than a preset threshold.

24. The device according to claim 23, wherein The binning processing module is further used to: If the absolute value error of the characteristic value of each binning point in the updated binning point set relative to the characteristic value of the corresponding binning point before the update is not less than the preset threshold, continue to update the local binning point set based on the binning point set of each participant node through multi-party secure computing technology.

25. The device according to claim 19 or 20, wherein The binning processing module is used to: If the binning type is horizontal chi-square binning, the eigenvalues ​​of all participants are globally sorted based on the size of the eigenvalues ​​corresponding to the local features of each participant's node using multi-party secure computing technology; and counting the number of samples of each participant's nodes; Locally, at each participant node, each characteristic value is divided into a bin according to the global sorting result; Based on the multi-party secure computing technology, the chi-square value between each pair of adjacent bins is calculated locally at each participant node; A pair of adjacent bins with the smallest chi-square value is merged.

26. The device according to claim 25, wherein The binning processing module is further used to: After merging the pair of adjacent bins with the smallest chi-square value, continue to calculate the chi-square value between each pair of adjacent bins locally on each participant node based on the multi-party secure computing technology; Detecting whether the minimum chi-square value exceeds a preset specified threshold, or whether the number of the bins is less than a preset specified number; If so, determining each of the bins; If not, merge the pair of adjacent bins with the smallest chi-square value.

27. The apparatus according to claim 25, wherein The binning processing module is used to: By using multi-party secure computing technology, the number of samples of each type of label corresponding to the characteristic value in each bin and the proportion of samples of each type of label in the global sample are counted locally on each participant node; By using multi-party secure computing technology, the number of global samples corresponding to the characteristic values ​​in each of the bins in each of the participant nodes is counted; Locally at each participating node, the chi-square value between each pair of adjacent bins is calculated based on the proportion of samples of each type of label corresponding to the characteristic value in each bin in the global sample, the number of samples of each type of label, the number of global samples of each participating node, and the number of samples of each participating node.

28. The apparatus according to claim 19 or 20, wherein Each of the participant nodes includes a plurality of first participant nodes and second participant nodes, the sample data in each of the first participant nodes includes the features; the second participant nodes include label data; the binning processing module is used to: If the binning type is vertical chi-square binning, at each of the first participant nodes, the eigenvalues ​​are sorted according to the size of the eigenvalues ​​corresponding to the features; and the total number of samples is counted; Locally, at each of the first participant nodes, the feature values ​​in each of the sample data corresponding to the feature are divided into a bin according to the sorting result; Locally, at each of the first participant nodes, referring to the sample identifiers of the labels in the second participant nodes, calculating the chi-square value between each pair of adjacent bins; A pair of adjacent bins with the smallest chi-square value is merged.

29. The apparatus according to claim 28, wherein The binning processing module is further used to: After merging a pair of adjacent bins with the smallest chi-square value, continue to calculate the chi-square value between each pair of adjacent bins locally on each first participant node, referring to the sample identifier of the label in the second participant node; Detecting whether the minimum chi-square value exceeds a preset specified threshold, or whether the number of the bins is less than a preset specified number; If so, determining each of the bins; If not, merge the pair of adjacent bins with the smallest chi-square value.

30. The apparatus according to claim 28, wherein The binning processing module is used to: Locally, on each of the first participant nodes, taking the intersection of the identifier of the sample data corresponding to the feature value in each of the bins and the sample identifier of the label in the second participant node to obtain an intersection sample identifier; Locally, at the second participant node, based on the intersection sample identifier, counting the number of samples of each type of label corresponding to the feature value in each of the bins and the proportion of samples of each type of label in the global sample; The second participant node locally sends to each of the first participant nodes the number of samples of each type of label corresponding to the feature value in each of the bins and the proportion of samples of each type of label in the global sample; Locally at each of the first participant nodes, the chi-square value between each pair of adjacent bins is calculated based on the proportion of samples of each type of label corresponding to the characteristic values ​​in each bin in the global sample, the number of samples of each type of label, the number of samples local to each of the first participant nodes, and the total number of samples.

31. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 15.

32. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-15.

33. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Automatic data binning method and device

    CN110084376A

  • Rapid chi-square binning method and device

    CN112990487A

  • Decision tree construction method and device based on federated learning system and electronic equipment

    CN113408668A

  • Feature binning method and device and storage medium

    CN114329127A