Feature binning method and apparatus, electronic device, and storage medium
By generating hybrid simulated distribution information and performing feature binning, the problems of low efficiency and information leakage risk in existing technologies are solved, achieving efficient and secure feature binning, which is suitable for collaborative sharing scenarios of multi-source non-independent and identically distributed data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-06
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies suffer from low efficiency, large transmission volume, and information leakage risks during feature binning, especially in collaborative sharing scenarios of multi-source, non-independent, and co-distributed data, making it difficult to guarantee data security and privacy protection.
By generating hybrid simulated distribution information, and using a preset binning strategy to perform feature binning on the hybrid simulated distribution information, multiple iterative adjustments and large amounts of data transmission are avoided. Statistical distribution and data simulation methods are used for feature binning.
It improves the efficiency and accuracy of feature binning, reduces communication volume and cost, and reduces the risk of information leakage, enabling secure and convenient feature binning without exposing individual data.
Smart Images

Figure CN116933308B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a feature binning method, apparatus, electronic device and storage medium. Background Technology
[0002] Feature binning is an important method widely used in data analysis, including statistics, feature engineering, algorithms, and machine learning modeling. With the nation and corporations vigorously promoting cloud-network convergence, the digital economy, big data industrialization, and the integration of data from the east and computing from the west, and with increasing emphasis on data security and privacy protection, the volume of collaboratively shared data continues to grow. Therefore, improving the security of feature binning has become particularly crucial. Summary of the Invention
[0003] This application provides a feature binning method, apparatus, electronic device, and storage medium to improve the security of feature binning.
[0004] In a first aspect, embodiments of this application provide a feature binning method, comprising: a first terminal receiving second distribution information from a second terminal. The second distribution information is obtained by the second terminal extracting second service data; when the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal generates hybrid simulated distribution information based on the second distribution information and the first distribution information using a preset hybrid model. The first distribution information is obtained by the first terminal extracting first service data, and the data types of the first service data and the second service data are the same. The first terminal performs feature binning on the hybrid simulated distribution information using a preset binning strategy to obtain binning results.
[0005] The above method, by employing statistical distribution and data simulation approaches, solves the joint binning problem of multi-source, non-independent, identically distributed data by generating hybrid simulated distribution information and performing feature binning on it, compared to existing technologies that require multiple iterations to adjust distribution information. Furthermore, it eliminates the need for iterative processing, obtaining binning results with only a single transmission of feature distribution information. This eliminates the need for iterative processing and the transmission of large amounts of statistical data, reducing communication volume and cost while mitigating the risk of leakage of sample details, thus improving performance and accuracy. Moreover, compared to existing technologies that require the transmission of numerous statistical results for sample feature intervals, this application performs feature binning by transmitting distribution information directly from the terminal, reducing the risk of information leakage. It allows various terminals to perform secure and convenient feature binning without exposing their own data, protecting privacy.
[0006] Optionally, the above methods also include:
[0007] When the values of the second distribution information and the first distribution information belong to the same distribution, the first terminal uses a preset binning strategy to perform feature binning on the first distribution information to obtain the binning result.
[0008] In the above method, when it is determined that the values of the second distribution information and the first distribution information belong to the same distribution, the first terminal uses a preset binning strategy to perform feature binning on the first distribution information to obtain the binning result. This method can determine the binning result in a timely manner and improve the efficiency of feature binning.
[0009] Optionally, after determining that the values of the second distribution information and the first distribution information do not belong to the same distribution, the method further includes:
[0010] If the first terminal determines that the second distribution information does not include the necessary feature information, the first terminal sends the first information to the second terminal. The necessary feature information is used to generate the hybrid simulation distribution information. The first information is used to instruct the second terminal to send the necessary feature information to the first terminal.
[0011] The second terminal sends the necessary feature information to the first terminal.
[0012] In the above method, the first terminal promptly determines whether the second distribution information includes the necessary feature information that can be used to generate the hybrid simulation distribution information. This allows the second terminal to promptly send the necessary feature information to the first terminal even when the second distribution information does not include it. This facilitates the timely generation of the hybrid simulation distribution information, the determination of feature binning results, and improves the efficiency of feature binning.
[0013] Optionally, the hybrid simulation model satisfies the following formula:
[0014]
[0015] in, This represents the k-th distribution information. The probability density function representing the k-th distribution information. This represents the mixing weights for each distribution k. , .
[0016] Optionally, before the first terminal receives the second distribution information from the second terminal, the method further includes:
[0017] The second terminal adjusts the second business data according to the preset data processing rules;
[0018] The second terminal extracts the adjusted second business data to determine the second distribution information.
[0019] In the above method, the second business data is adjusted according to preset data processing rules; the adjusted second business data is then extracted to determine the second distribution information. This facilitates the first terminal in determining whether the values of the second distribution information and the first distribution information belong to the same distribution, thus improving the efficiency of feature binning in subsequent steps.
[0020] Optionally, before the first terminal receives the second distribution information from the second terminal, the above method further includes:
[0021] The first terminal adjusts the first business data according to the preset data processing rules;
[0022] The first terminal extracts the adjusted first business data and determines the first distribution information.
[0023] In the above method, the first business data is adjusted according to preset data processing rules; the adjusted first business data is then extracted to determine the first distribution information. This facilitates the first terminal in determining whether the values of the second distribution information and the first distribution information belong to the same distribution, thus improving the efficiency of feature binning in subsequent steps.
[0024] Secondly, embodiments of this application provide a feature binning method. This method can be applied to a first terminal. In this method, the first terminal receives second distribution information from a second terminal, which is obtained by the second terminal extracting second service data. When the values of the second distribution information and the first distribution information do not belong to the same distribution, a hybrid simulated distribution information is generated based on the second distribution information and the first distribution information using a preset hybrid model. The first distribution information is obtained by the first terminal extracting first service data, and the data types of the first service data and the second service data are the same. The hybrid simulated distribution information is feature-binned using a preset binning strategy to obtain binning results.
[0025] Optionally, the above method further includes: when the values of the second distribution information and the first distribution information belong to the same distribution, the first terminal uses a preset binning strategy to perform feature binning on the first distribution information to obtain binning results.
[0026] Optionally, after the first terminal determines that the values of the second distribution information and the first distribution information do not belong to the same distribution, the above method further includes:
[0027] If the first terminal determines that the second distribution information does not include the necessary feature information, it sends the first information to the second terminal. The necessary feature information is used to generate the hybrid simulation distribution information, and the first information is used to instruct the second terminal to send the necessary feature information to the first terminal.
[0028] Optionally, the above hybrid simulation model satisfies the following formula:
[0029]
[0030] in, This represents the k-th distribution information. The probability density function representing the k-th distribution information. This represents the mixing weights for each distribution k. , .
[0031] Optionally, before the first terminal receives the second distribution information from the second terminal, the above method further includes:
[0032] The first terminal adjusts the first business data according to the preset data processing rules;
[0033] The first terminal extracts the adjusted first business data and determines the first distribution information.
[0034] Thirdly, embodiments of this application provide a feature binning method. This method can be applied to a second terminal. In this method, the second terminal sends second distribution information to a first terminal, so that when the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal generates hybrid simulated distribution information based on the second distribution information and the first distribution information using a preset hybrid model, and performs feature binning on the hybrid simulated distribution information using a preset binning strategy to obtain a binning result. The second distribution information is obtained by the second terminal from second business data, and the first distribution information is obtained by the first terminal from first business data; the data types of the first business data and the second business data are the same.
[0035] Optionally, the above method further includes: receiving first information from a first terminal, the first information being used to instruct a second terminal to send necessary feature information to the first terminal, the necessary feature information being used to generate hybrid simulation distribution information; and the second terminal sending the necessary feature information to the first terminal.
[0036] Optionally, before the second terminal sends the second distribution information to the first terminal, the above method further includes:
[0037] The second terminal adjusts the second business data according to the preset data processing rules;
[0038] The second terminal extracts the adjusted second business data to determine the second distribution information.
[0039] Fourthly, embodiments of this application provide a feature sorting system, the system comprising: a first terminal and a second terminal.
[0040] The second terminal is used to send second distribution information to the first terminal. The second distribution information is obtained by the second terminal from the second service data.
[0041] The first terminal is used to generate hybrid simulation distribution information based on the second distribution information and the first distribution information when the values of the second distribution information and the first distribution information do not belong to the same distribution. The first distribution information is obtained by the first terminal from the first business data, and the data types of the first business data and the second business data are the same.
[0042] The first terminal is also used to perform feature binning on the mixed simulated distribution information using a preset binning strategy to obtain binning results.
[0043] Fifthly, embodiments of this application provide a feature sorting device, comprising:
[0044] The transceiver module is used to receive second distribution information from the second terminal, which is obtained by the second terminal from the second service data.
[0045] The processing module is used to generate hybrid simulated distribution information based on the second distribution information and the first distribution information when the values of the second distribution information and the first distribution information do not belong to the same distribution. The first distribution information is obtained by the first terminal from the first business data, and the data types of the first business data and the second business data are the same.
[0046] The binning module is used to perform feature binning on the mixed simulated distribution information using a preset binning strategy to obtain binning results.
[0047] Sixthly, embodiments of this application provide a feature sorting device, comprising:
[0048] The processing module is used to extract the second business data and obtain the second distribution information;
[0049] The transceiver module is used to send second distribution information to the first terminal so that when the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal generates hybrid simulated distribution information based on the second distribution information and the first distribution information using a preset hybrid model; and performs feature binning on the hybrid simulated distribution information using a preset binning strategy to obtain binning results.
[0050] The second distribution information is obtained by the second terminal from the second service data, and the first distribution information is obtained by the first terminal from the first service data. The first service data and the second service data have the same data type.
[0051] In a seventh aspect, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the processor implements any of the feature binning methods in the first to third aspects.
[0052] Eighthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any one of the feature binning methods of the first to third aspects.
[0053] In a ninth aspect, embodiments of this application also provide a computer program product, including a computer program executed by a processor to implement the feature binning method as described in any of the first to third aspects above.
[0054] The technical effects of any of the implementation methods in aspects four through eight can be found in the technical effects of the corresponding implementation methods in aspects one through three, and will not be repeated here. Attached Figure Description
[0055] Figure 1 This is a schematic diagram illustrating an application scenario of a feature binning method provided in an embodiment of this application.
[0056] Figure 2 A flowchart of a feature binning method provided in an embodiment of this application;
[0057] Figure 3 A schematic diagram of a probability distribution model provided in an embodiment of this application;
[0058] Figure 4 A flowchart illustrating an exemplary feature binning method provided in this application embodiment;
[0059] Figure 5 A schematic diagram of a feature binning detection device provided in an embodiment of this application;
[0060] Figure 6 A schematic diagram of another feature binning detection device provided in an embodiment of this application;
[0061] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0063] The application scenarios described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. Those skilled in the art will understand that with the emergence of new application scenarios, the technical solutions provided in this application are also applicable to similar technical problems. In the description of this application, unless otherwise stated, "multiple" means two or more.
[0064] Feature binning is an important method widely used in data analysis, including statistics, feature engineering, algorithms, and machine learning modeling. For example, feature binning is used to improve the performance of some tree models, to calculate information value (IV) to evaluate feature importance and use this metric for feature selection, and to perform group profiling and data analysis.
[0065] Currently, existing or commonly used distributed feature binning methods mainly employ an engineering-based algorithm that involves interval information transmission and iterative approximation. These methods require multiple transmissions of quantile or interval statistical data to obtain global quantile or interval statistical data. This data is then used to determine the relationship between the initial binning results of each data participant and the global target binning result. Based on this relationship, the binning results are gradually adjusted through iterative iteration to ultimately approximate the true global target binning result. These methods suffer from low efficiency, large transmission volumes, and a certain risk of data leakage.
[0066] With the nation and corporations vigorously promoting cloud-network convergence, the digital economy, big data industrialization, and the integration of data from the east and computing from the west, and with increasing emphasis on data security and privacy protection, the volume of data that needs to be shared collaboratively is also continuously increasing. Therefore, improving the security of feature binning has become particularly important.
[0067] To address the aforementioned problems, embodiments of this application provide a feature binning method, apparatus, electronic device, and storage medium. For example, a first terminal receives second distribution information from a second terminal. The second distribution information is obtained by the second terminal extracting second service data. When the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal generates hybrid simulated distribution information based on the second distribution information and the first distribution information using a preset hybrid model. The first distribution information is obtained by the first terminal extracting first service data, and the data types of the first service data and the second service data are the same. The first terminal performs feature binning on the hybrid simulated distribution information using a preset binning strategy to obtain binning results.
[0068] like Figure 1 The diagram illustrates an application scenario of an optional feature binning method according to an embodiment of this application, including a first terminal 101, a first server 102, a second terminal 103, and a second server 104. The first terminal 101, the first server 102, the second terminal 103, and the second server 104 can be connected via a network to implement the feature binning method of this application.
[0069] The first terminal 102 and the first server 102 corresponding to the first terminal can be connected through a network to enable data interaction between the first terminal 103 and the first server 102 corresponding to the first terminal.
[0070] The second terminal 103 and the second server 104 corresponding to the second terminal can be connected through a network to enable data interaction between the second terminal 103 and the second server 104 corresponding to the second terminal.
[0071] A communication connection is established between the first server 102 corresponding to the first terminal and the second server 104 corresponding to the second terminal. In this way, data exchange can occur between the first server corresponding to the first terminal and the second server corresponding to the second terminal.
[0072] The first terminal 101 may include, but is not limited to, mobile phones, tablets, personal computers (PCs), or handheld computers (PDAs). The first server 102 may include, but is not limited to, servers or cloud computing. The second terminal 103 may include, but is not limited to, mobile phones, tablets, personal computers, or handheld computers. The second server 104 may include, but is not limited to, servers or cloud computing.
[0073] It is understood that the application scenarios of the optional feature binning method of this application may include multiple terminals. That is, one or more second terminals may send second distribution information to the first terminal, enabling the first terminal to perform subsequent feature binning steps. The first terminal can be pre-configured by those skilled in the art, and the first terminal can be reasonably configured according to the specific application scenario.
[0074] like Figure 2 As shown in the flowchart of a feature binning method provided in this application embodiment, it may specifically include the following steps.
[0075] S201, The first terminal receives the second distribution information from the second terminal.
[0076] The second distribution information is obtained by the second terminal extracting the second business data.
[0077] Distribution information may include, but is not limited to: distribution type, uniform distribution, exponential distribution, logarithmic distribution, distribution parameters, distribution characteristic information (e.g., the intervals of the distribution), local calculation results of distribution comparison tests, and other information.
[0078] It is understandable that each terminal stores business data of different data types, i.e., feature data with different characteristics, in its respective database. In this embodiment, the distribution information of each terminal can be obtained by extracting business data of a single data type individually, or by bundling business data of multiple data types and extracting them in multiple dimensions.
[0079] In one optional embodiment, to facilitate the first terminal in determining whether the second distribution information and the first distribution information belong to the same distribution, the second terminal may adjust the second service data according to preset data processing rules before sending the second distribution information to the first terminal. The second terminal extracts the adjusted second service data to determine the second distribution information. Alternatively, the first terminal may adjust the first service data according to preset data processing rules. The first terminal extracts the adjusted first service data to determine the first distribution information.
[0080] The preset data processing rules can include, but are not limited to, methods for adjusting feature data such as: filling missing values, cleaning data, outlier detection and processing, and adjusting distribution skewness. This application does not impose specific limitations on these methods.
[0081] For outlier detection and processing, the second terminal can employ methods such as the Rheinda criterion (Z-SCORE) and the Grubbs algorithm to detect outliers. The second terminal can also process outliers by deleting them, not adding new observations, deleting them, and adding new observations or appropriate interpolated values. This application does not impose specific limitations on these methods.
[0082] Regarding distribution skewness adjustment, the second terminal can adjust the distribution skewness based on whether data distribution is performed. The data distribution depends on subsequent distribution comparison tests, distribution hypotheses, and distribution information extraction methods. This application does not specify the particular method for adjusting distribution skewness.
[0083] Optionally, since this statistical probability distribution-based method has a certain degree of robustness, not adjusting the data distribution will only have a minor impact on accuracy, and adjustments to the distribution of the second business data can be made through fine-tuning. Therefore, the second terminal can also adjust the second business data in the following ways: 1. When extracting the adjusted second business data, the second terminal can delete some extreme values. 2. The second terminal can also use nonparametric hypothesis testing and kernel density estimation methods that do not require distribution assumptions to perform distribution comparison tests and extract and simulate distribution information, respectively.
[0084] S202. When the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal generates hybrid simulation distribution information based on the second distribution information and the first distribution information using a preset hybrid model.
[0085] The first distribution information is obtained by the first terminal from the first service data. The first service data and the second service data have the same data type.
[0086] In one optional embodiment, after receiving the second distribution information sent by the second terminal, the first terminal can compare the statistical significance of the differences between the distribution information using methods including but not limited to hypothesis testing. That is, whether there are truly systematic differences in the distribution information, rather than differences caused by sampling noise.
[0087] The following describes how the first terminal determines whether the values of the second distribution information and the first distribution information belong to the same distribution:
[0088] Since the distribution of feature data is generally normal, and compared to other methods of detecting distribution information, the normality test can save computational resources and improve detection efficiency during the normality test calculation, the first terminal can use the normality test to determine whether each distribution information follows a normal distribution.
[0089] In one possible scenario, the second distribution information sent by the second terminal to the first terminal may include indication information. This indication information indicates whether the distribution type of the second distribution information is a normal distribution. After receiving the second distribution information from the second terminal, the second terminal can determine, based on the indication information, whether the values of the second distribution information and the first distribution information belong to the same distribution.
[0090] For example, when the second terminal extracts second business data, it can determine whether the feature data belongs to a normal distribution through a normality test. If so, the indication information includes that the distribution type is a normal distribution. The second terminal can carry the above indication information in the second distribution information and send the second distribution information to the first terminal. After receiving the second distribution information from the second terminal, the first terminal can determine whether the values of the second distribution information and the first distribution information belong to the same distribution based on the indication information. If any indication information from the second terminal indicates that the distribution type is not a normal distribution, then the first terminal can determine that the values of the second distribution information and the first distribution information do not belong to the same distribution.
[0091] In another possible scenario, since the distribution of feature data is generally normally distributed, the first terminal can use the distribution assumption to compare normally distributed data. Methods such as one-way ANOVA, which require the normal distribution assumption, can be used to perform more accurate and efficient comparisons.
[0092] The methods for testing the normality hypothesis may include, but are not limited to, one-way ANOVA, Bayesian statistical methods, and other methods for testing the normality hypothesis. This application does not impose specific limitations on these methods.
[0093] In another possible scenario, the first terminal can use methods such as the Kruskal-Wallis test or the chi-square test, which do not require comparing normal distributions, to compare the value distributions of the second distribution information with those of the first distribution information.
[0094] It is understood that the methods for determining whether the values of the second distribution information and the first distribution information belong to the same distribution in this application include, but are not limited to, the methods mentioned above. Furthermore, those skilled in the art can flexibly set the method for determining whether the values of the second distribution information and the first distribution information belong to the same distribution according to specific application scenarios.
[0095] In the above method, hypothesis testing is used to compare the distribution information of multiple terminals to determine whether the business data of the same data type corresponding to different terminals have the same distribution. This makes it easier to determine that the values of the second distribution information and the first distribution information belong to the same distribution in the future. This allows the first terminal to directly perform feature binning based on the first distribution information, thereby accelerating the feature binning process and improving the efficiency of feature binning.
[0096] Optionally, when the first terminal determines that the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal may determine whether the second distribution information includes necessary feature information. The necessary feature information is used to generate the hybrid simulation distribution information. In one possible scenario, if the first terminal determines that the second distribution information does not include necessary feature information, it sends first information to the second terminal. The necessary feature information is used to generate the hybrid simulation distribution information, and the first information instructs the second terminal to send the necessary feature information to the first terminal. After receiving the necessary feature information from the first terminal, the second terminal sends it back to the first terminal.
[0097] For example, when determining whether the values of the second distribution information and the first distribution information belong to the same distribution, only the mean, variance, and distribution are needed to determine whether the values of the second distribution information and the first distribution information belong to the same distribution. However, when the values of the second distribution information and the first distribution information do not belong to the same distribution, the necessary feature information also needs the weights of the second distribution information and the first distribution information, or the proportion of their respective sample sizes to the total sample size.
[0098] It is understood that the necessary feature information may include, but is not limited to, distribution type, distribution parameters, distribution characteristic information, and other necessary feature information used to generate hybrid simulation distribution information. Those skilled in the art can flexibly set these features according to specific application scenarios, and this application does not impose specific limitations on them.
[0099] In the above method, the first terminal promptly determines whether the second distribution information includes the necessary feature information that can be used to generate the hybrid simulation distribution information. This allows the second terminal to promptly send the necessary feature information to the first terminal even when the second distribution information does not include it. This facilitates the timely generation of the hybrid simulation distribution information, the determination of feature binning results, and improves the efficiency of feature binning.
[0100] In another possible scenario, if the first terminal determines that the second distribution information includes the necessary feature information, the first terminal may generate hybrid simulation distribution information based on the second distribution information and the first distribution information using a preset hybrid model.
[0101] The first distribution information is obtained by the first terminal extracting the first service data.
[0102] In an optional embodiment, the first terminal can use a hybrid model to aggregate the second distribution information and the first distribution information to generate a new hybrid distribution information, namely, hybrid simulated distribution information.
[0103] The hybrid simulation model satisfies the following formula:
[0104]
[0105] in, This represents the k-th distribution information. The probability density function represents the k-th distribution information. This represents the mixing weight for each distribution k (the mixing weight can be set as the proportion of the sample size of the participants with that distribution to the total sample size). , .
[0106] In the above method, when the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal generates hybrid simulated distribution information based on the second distribution information and the first distribution information using a preset hybrid model for feature binning, thus solving the problem of non-independent identically distributed data in multi-source data. Furthermore, compared to existing technologies that require iterative approximation and statistical transmission of large amounts of data for feature binning, this application uses the idea of statistical distribution and data simulation, employing a concept similar to reverse-engineering or restoring the original distribution after compressed information transmission to generate hybrid simulated distribution information. This reduces communication volume and cost, and lowers the risk of data transmission leakage.
[0107] In some embodiments, since the values of the second distribution information and the first distribution information belong to the same distribution, that is, the difference between the distributions of the first terminal and the second terminal is not statistically significant, any terminal can be selected to perform feature binning on the local distribution information using a preset binning strategy to obtain the binning result.
[0108] For example, when the first terminal determines that the values of the second distribution information and the first distribution information belong to the same distribution, the first terminal can use a preset binning strategy to perform feature binning on the first distribution information to obtain binning results. Alternatively, the second terminal can use a preset binning strategy to perform feature binning on the second distribution information corresponding to the second terminal to obtain binning results.
[0109] S203. The first terminal uses a preset binning strategy to perform feature binning on the hybrid simulated distribution information to obtain binning results.
[0110] The following section introduces several methods for feature binning based on hybrid simulated distribution information:
[0111] 1. The first binning result is determined directly using the probability distribution of the Gaussian mixture model and the preset binning parameters.
[0112] 2. First, generate the mixed simulation distribution information using a Gaussian mixture model, and then determine the first binning result by performing feature binning on the mixed simulation distribution information.
[0113] 3. Generate only the simulation data for the second terminal, and combine the simulation data with the real data to obtain hybrid simulated distribution information. Then, perform feature binning on the hybrid simulated distribution information to determine the first binning result.
[0114] It is understood that the feature binning methods based on hybrid simulated distribution information in this application include, but are not limited to, the methods described above. Furthermore, those skilled in the art can flexibly configure feature binning based on hybrid simulated distribution information according to specific application scenarios. The preset binning strategy can be any feasible feature binning method. For example, equal-frequency binning, equal-distance binning, etc. This application does not specifically limit this approach.
[0115] Optionally, to ensure all terminals can obtain the binning results for subsequent data analysis, after obtaining the binning results, the first terminal can send the binning results to the second terminal, enabling the sharing results to be synchronized to all terminals participating in feature binning.
[0116] The following is about Figure 2 The following examples illustrate the implementation:
[0117] Assuming the feature data are the most common continuous variable features, when the amount of data is large, according to the central limit theorem and the law of large numbers, we can assume that the random variable approximately follows a distribution / Gaussian distribution.
[0118] Suppose that the feature data involves two participants in joint feature binning: a first terminal and a second terminal. Furthermore, distribution comparison reveals that the values of the second distribution information and the first distribution information of the first and second terminals do not belong to the same distribution. That is, the second distribution information and the first distribution information are two different Gaussian distributions. and .
[0119] Assuming that generating these two Gaussian distributions only requires the expected value (mean) of the first moment and the second moment (variance) of the feature data from the first terminal and the second terminal, respectively, the first terminal can combine the two Gaussian distributions with the first distribution information to generate a Gaussian mixture model after obtaining the mean and variance sent by the second terminal.
[0120] like Figure 3 As shown, the Gaussian mixture model has a probability distribution model of the following form:
[0121]
[0122]
[0123] Since there are two distinct distributions, K=2. Let's define the percentage of samples each terminal has. For example, if the first terminal has 500 samples and the second terminal has 1500 samples, then... =0.25, =0.75. Substituting the mean and variance from the distribution information of the first and second terminals into the above formula, we obtain the probability density formula for the target Gaussian mixture model. Then, we generate the mixture simulation distribution information based on random numbers. Finally, we perform equal-frequency binning on the mixture simulation distribution information according to the probability density to obtain the first binning result.
[0124] The above method employs an innovative approach to feature binning using statistical distribution probability theory, distribution hypothesis testing, and data simulation. This differs from existing distributed engineering systems, which typically involve transmitting interval statistical information to a terminal for data aggregation, adjusting the binning results, and iterating repeatedly to approximate the true or optimal global binning. Furthermore, by utilizing statistical distribution and data simulation, this application addresses the joint binning problem of multi-source, non-independent, identically distributed data by generating hybrid simulated distribution information and performing feature binning on it. This eliminates the need for iterative iterations; binning results are obtained with a single transmission of feature distribution information. Reducing communication volume and cost by eliminating iterative transmission and large amounts of statistical data lowers the risk of sample detail information leakage, while improving performance and accuracy. Moreover, compared to existing technologies requiring the transmission of numerous sample feature interval statistical results, this application performs feature binning solely through terminal transmission of distribution information, further reducing the risk of information leakage. It enables multiple parties or nodes to perform secure and convenient feature binning without exposing their own data and protecting privacy.
[0125] It is understood that this application is applicable to all distributed application scenarios that require feature binning, including but not limited to distributed statistics, federated learning, and multi-party secure computation applications. This application does not impose any specific limitations on these scenarios.
[0126] like Figure 4 As shown in the figure, an exemplary flowchart of a feature binning method provided in this application embodiment may specifically include the following operations:
[0127] S401. The second terminal adjusts the second service data according to the preset data processing rules;
[0128] S402, The second terminal extracts the adjusted second service data and determines the second distribution information;
[0129] S403, The second terminal sends the second distribution information to the first terminal;
[0130] S404. The first terminal determines whether the values of the second distribution information and the first distribution information belong to the same distribution. If yes, proceed to step S405; otherwise, proceed to step S407.
[0131] S405. The first terminal uses a preset binning strategy to perform feature binning on the first distribution information to obtain binning results.
[0132] S406. The first terminal sends the bin sorting result to the second terminal;
[0133] S407. The first terminal determines whether the second distribution information includes necessary feature information. If not, proceed to step S408; if yes, proceed to step S410.
[0134] S408, The first terminal sends the first information to the second terminal;
[0135] S409 The second terminal sends necessary feature information to the first terminal;
[0136] S410. The first terminal generates hybrid simulation distribution information based on the second distribution information and the first distribution information using a preset hybrid model;
[0137] S411. The first terminal uses a preset binning strategy to perform feature binning on the hybrid simulation distribution information, obtains the binning results, and returns to the execution step S406.
[0138] Figure 5 This is a schematic diagram of the feature-dividing device provided in the embodiments of this application, as shown below. Figure 5 As shown, it includes: a transceiver module 501, a processing module 502, and a binning module 503.
[0139] The transceiver module 501 is used to receive second distribution information from the second terminal, which is obtained by the second terminal from the second service data.
[0140] Processing module 502 is used to generate hybrid simulation distribution information based on the second distribution information and the first distribution information when the values of the second distribution information and the first distribution information do not belong to the same distribution. The first distribution information is obtained by the first terminal from the first business data, and the data types of the first business data and the second business data are the same.
[0141] The binning module 503 is used to perform feature binning on the mixed simulated distribution information using a preset binning strategy to obtain binning results.
[0142] Optionally, the binning module 503 is also used for:
[0143] When the values of the second distribution information and the first distribution information belong to the same distribution, the first distribution information is binned using a preset binning strategy to obtain the binning result.
[0144] Optionally, after the first terminal determines that the values of the second distribution information and the first distribution information do not belong to the same distribution, the transceiver module 501 is further configured to:
[0145] If it is determined that the second distribution information does not include the necessary feature information, the first information is sent to the second terminal. The necessary feature information is used to generate the hybrid simulation distribution information, and the first information is used to instruct the second terminal to send the necessary feature information to the first terminal.
[0146] Optionally, the above hybrid simulation model satisfies the following formula:
[0147]
[0148] in, This represents the k-th distribution information. The probability density function representing the k-th distribution information. This represents the mixing weights for each distribution k. , .
[0149] Optionally, before the first terminal receives the second distribution information from the second terminal, the processing module 502 is further configured to:
[0150] The first terminal adjusts the first business data according to the preset data processing rules;
[0151] The first terminal extracts the adjusted first business data and determines the first distribution information.
[0152] Figure 6 This is a schematic diagram of another feature-separating device provided in an embodiment of this application, as shown below. Figure 6 As shown, it includes: a processing module 601 and a transceiver module 602.
[0153] Processing module 601 is used to extract the second business data and obtain the second distribution information;
[0154] The transceiver module 602 is used to send second distribution information to the first terminal so that when the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal generates hybrid simulated distribution information based on the second distribution information and the first distribution information using a preset hybrid model; and performs feature binning on the hybrid simulated distribution information using a preset binning strategy to obtain binning results.
[0155] The second distribution information is obtained by the second terminal from the second service data, and the first distribution information is obtained by the first terminal from the first service data. The first service data and the second service data have the same data type.
[0156] Optionally, the transceiver module 602 is also used for:
[0157] The system receives first information from a first terminal, which instructs a second terminal to send necessary feature information to the first terminal. The necessary feature information is used to generate hybrid simulation distribution information. The second terminal then sends the necessary feature information to the first terminal.
[0158] Before the second terminal sends the second distribution information to the first terminal, the processing module 601 is also used to:
[0159] Adjust the second business data according to the preset data processing rules;
[0160] Extract the adjusted second business data to determine the second distribution information.
[0161] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0162] At least one processor 701 and a memory 702 connected to at least one processor 701. In this embodiment, the specific connection medium between the processor 701 and the memory 702 is not limited. Figure 7 The example shown is the connection between processor 701 and memory 702 via bus 700. Bus 700 is... Figure 7 The connections between other components are indicated by thick lines and are for illustrative purposes only, not as limiting information. The 700 bus can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 7 The term is represented by a single thick line, but this does not imply that there is only one bus or one type of bus. Alternatively, the processor 701 can also be called a controller; there is no restriction on the name.
[0163] In this embodiment, memory 702 stores instructions executable by at least one processor 701. By executing the instructions stored in memory 702, at least one processor 701 can perform the feature binning method described above. Processor 701 can implement... Figure 5 or Figure 6 The functions of each module in the device shown.
[0164] The processor 701 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 702 and calling data stored in memory 702, the processor can perform various functions and process data, thereby monitoring the device as a whole.
[0165] In one possible design, processor 701 may include one or more processing units. Processor 701 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, driver interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 701. In some embodiments, processor 701 and memory 702 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.
[0166] The processor 701 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the feature binning method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0167] Memory 702, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 702 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 702 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. Memory 702 in the embodiments of this application may also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.
[0168] By designing and programming the processor 701, the code corresponding to the feature binning method described in the foregoing embodiments can be embedded into the chip, enabling the chip to execute it during runtime. Figure 2 The feature binning method of the illustrated embodiment. How to design and program the processor 701 is a technique well known to those skilled in the art, and will not be described in detail here.
[0169] It should be noted that the electronic device provided in this application embodiment can implement all the method steps implemented in the above method embodiment and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0170] This application also provides a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the database maintenance method described in the above embodiments.
[0171] This application embodiment also provides a feature binning system, which may include the aforementioned first terminal and second terminal. The operations performed by the first terminal and second terminal can be referred to... Figure 2 The relevant descriptions in the illustrated embodiments.
[0172] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0173] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0174] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0175] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0176] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A feature binning method, characterized in that, The method includes: The first terminal receives second distribution information from the second terminal, the second distribution information being obtained by the second terminal from the second service data; When the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal generates hybrid simulated distribution information based on the second distribution information and the first distribution information using a preset hybrid model. The first distribution information is extracted by the first terminal from the first business data. The first business data and the second business data have the same data type. The hybrid model generates the hybrid simulated distribution information by weighted summation of the probability densities of samples under each distribution based on the hybrid weights of each distribution. If it is determined that the second distribution information does not include necessary feature information, first information is sent to the second terminal to cause the second terminal to send the necessary feature information to the first terminal; the necessary feature information is used to generate the hybrid simulation distribution information, and the first information is used to instruct the second terminal to send the necessary feature information to the first terminal; the necessary feature information includes at least distribution type, distribution parameters and distribution feature information. The first terminal uses a preset binning strategy to perform feature binning on the hybrid simulated distribution information to obtain binning results.
2. The method according to claim 1, characterized in that, The method further includes: When the values of the second distribution information and the first distribution information belong to the same distribution, the first terminal uses a preset binning strategy to perform feature binning on the first distribution information to obtain the binning result.
3. The method according to claim 1, characterized in that, The hybrid model satisfies the following formula: in, This represents the k-th distribution information. The probability density function representing the k-th distribution information. This represents the mixing weights for each distribution k. , .
4. The method according to claim 1, characterized in that, Before the first terminal receives the second distribution information from the second terminal, the method further includes: The second terminal adjusts the second service data according to preset data processing rules; The second terminal extracts the adjusted second service data to determine the second distribution information.
5. The method according to claim 1, characterized in that, Before the first terminal receives the second distribution information from the second terminal, the method further includes: The first terminal adjusts the first service data according to preset data processing rules; The first terminal extracts the adjusted first service data to determine the first distribution information.
6. A feature binning method, characterized in that, Applied to a first terminal, the method includes: Receive second distribution information from the second terminal, wherein the second distribution information is obtained by the second terminal from the second service data; When the values of the second distribution information and the first distribution information do not belong to the same distribution, a hybrid simulation distribution information is generated based on the second distribution information and the first distribution information using a preset hybrid model. The first distribution information is obtained by the first terminal from the first business data. The first business data and the second business data have the same data type. The hybrid model generates the hybrid simulation distribution information by weighted summing of the probability densities of samples under each distribution based on the hybrid weights of each distribution. If it is determined that the second distribution information does not include necessary feature information, first information is sent to the second terminal to cause the second terminal to send the necessary feature information to the first terminal; the necessary feature information is used to generate the hybrid simulation distribution information, and the first information is used to instruct the second terminal to send the necessary feature information to the first terminal; the necessary feature information includes at least distribution type, distribution parameters and distribution feature information. The hybrid simulated distribution information is binned by a preset binning strategy to obtain binning results.
7. A feature binning method, characterized in that, Applied to a second terminal, the method includes The first terminal sends second distribution information to the first terminal so that when the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal generates hybrid simulated distribution information based on the second distribution information and the first distribution information using a preset hybrid model. The second distribution information is obtained by the second terminal from the second business data, and the first distribution information is obtained by the first terminal from the first business data. The first business data and the second business data have the same data type. The hybrid model generates the hybrid simulated distribution information by weighted summation of the probability densities of samples under each distribution based on the hybrid weights of each distribution. If the second distribution information does not include necessary feature information, first information is received from the first terminal, and the necessary feature information is sent to the first terminal so that the first terminal can perform feature binning on the hybrid simulation distribution information using a preset binning strategy to obtain binning results. The necessary feature information is used to generate the hybrid simulation distribution information, and the first information is used to instruct the second terminal to send the necessary feature information to the first terminal. The necessary feature information includes at least distribution type, distribution parameters, and distribution feature information.
8. A feature-based binning system, characterized in that, The system includes: a first terminal and a second terminal; The second terminal is used to send second distribution information to the first terminal, the second distribution information being obtained by the second terminal from the second service data; The first terminal is configured to generate hybrid simulated distribution information based on the second distribution information and the first distribution information when the values of the second distribution information and the first distribution information do not belong to the same distribution. The first distribution information is obtained by the first terminal from the first business data. The first business data and the second business data have the same data type. The hybrid model generates the hybrid simulated distribution information by weighted summation of the probability densities of samples under each distribution based on the mixing weights of each distribution. The first terminal is further configured to send first information to the second terminal when it is determined that the second distribution information does not include necessary feature information, wherein the necessary feature information is used to generate the hybrid simulation distribution information, and the first information is used to instruct the second terminal to send the necessary feature information to the first terminal, wherein the necessary feature information includes at least distribution type, distribution parameters and distribution feature information; The second terminal is also used to send the necessary feature information to the first terminal; The first terminal is also used to perform feature binning on the hybrid simulated distribution information using a preset binning strategy to obtain binning results.
9. A feature-separating device, characterized in that, include: The transceiver module is used to receive second distribution information from the second terminal, wherein the second distribution information is obtained by the second terminal from the second service data; The processing module is used to generate hybrid simulated distribution information based on the second distribution information and the first distribution information when the values of the second distribution information and the first distribution information do not belong to the same distribution. The first distribution information is extracted by the first terminal from the first business data. The first business data and the second business data have the same data type. The hybrid model generates the hybrid simulated distribution information by weighted summation of the probability densities of samples under each distribution based on the mixing weights of each distribution. The transceiver module is further configured to send first information to the second terminal when it is determined that the second distribution information does not include necessary feature information, so that the second terminal sends the necessary feature information to the first terminal; the necessary feature information is used to generate the hybrid simulation distribution information, and the first information is used to instruct the second terminal to send the necessary feature information to the first terminal, and the necessary feature information includes at least distribution type, distribution parameters and distribution feature information; The binning module is used to perform feature binning on the hybrid simulated distribution information using a preset binning strategy to obtain binning results.
10. A feature-separating device, characterized in that, include: The processing module is used to extract the second business data and obtain the second distribution information; The transceiver module is used to send the second distribution information to the first terminal so that when the values of the second distribution information and the first distribution information do not belong to the same distribution, the first terminal generates hybrid simulated distribution information based on the second distribution information and the first distribution information using a preset hybrid model. The first distribution information is extracted by the first terminal from the first business data, and the first business data and the second business data have the same data type. The hybrid model generates the hybrid simulated distribution information by weighted summation of the probability densities of samples under each distribution based on the hybrid weights of each distribution. The transceiver module is further configured to receive first information from the first terminal and send the necessary feature information to the first terminal when the second distribution information does not include necessary feature information, so that the first terminal performs feature binning on the hybrid simulation distribution information using a preset binning strategy to obtain binning results. The necessary feature information is used to generate the hybrid simulation distribution information, and the first information is used to instruct the second terminal to send the necessary feature information to the first terminal. The necessary feature information includes at least distribution type, distribution parameters, and distribution feature information.
11. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-5.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that causes the computer to perform the steps of any of the methods described in claims 1-5.
13. A computer program product, characterized in that, When the computer program product is invoked by a computer, it causes the computer to perform the steps of the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Characteristic box separation method, device and equipment and computer readable storage medium
CN111506485A
Characteristic box separation method, device and equipment and readable storage medium
CN111898765A