Data classification method and device
By using encrypted vectors and multiple iterative operations in vertical federated learning scenarios, the problem that the meaning of the data-party feature cannot be disclosed is solved, the efficiency and accuracy of data classification are improved, and the interpretability of the model is improved.
Patent Information
- Application Number
- CN202210121822.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-09
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-02-09
AI Technical Summary
In the vertical federated learning scenario, the modeling characteristics of the data party cannot be disclosed to the business party, resulting in a reduced interpretability of the model.
By obtaining the candidate sample set and characterizing whether each sample in the candidate sample set is a cryptographic vector that represents the real sample, performing multiple iteration operations, determining the parties and numbers of the target classification characteristics, and dividing the sample set based on the encrypted vector, gradually optimizing the sample set and encrypted vectors to improve the efficiency and accuracy of data classification.
It realizes data classification when the business party does not know the real sample set, improves the efficiency and accuracy of data classification, and improves the interpretability of the model by disclosing the characteristic meaning of the data party.
Smart Images

Figure CN114462535B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, specifically to the field of data processing technology, and in particular to a data classification method and device. Background Art
[0002] Federated learning technology aims to achieve joint modeling among multiple participants, but does not share data, only shares intermediate results, and cannot infer data. It has been widely used and developed in risk control scenarios. At present, the characteristics of vertical federated learning in federated learning are: the label is only held by one party (hereinafter referred to as the business party), while the other parties only have some features of the data (hereinafter referred to as the data party). The business party hopes to improve the effect of the model through cooperation with the data party. Extreme gradient boosting (eXtreme Gradient Boosting, referred to as XGBoost) as a method for building a gradient boosting tree model has been widely used in risk control scenarios because of its good interpretability and strong learning ability.
[0003] In order to extend XGBoost to vertical federation scenarios and protect the privacy and security of data, a lossless privacy-preserving tree enhancement system SecureBoost based on federated learning is proposed. Although the SecureBoost method provides an idea for safely building a gradient boosting tree model, in the process of classifying and labeling data and using accurately labeled data for model training, the meaning of the modeling features of the data party cannot be disclosed to the business party, which greatly reduces the interpretability of the model. Summary of the invention
[0004] The present application provides a data classification method, apparatus, device and storage medium.
[0005] According to a first aspect of the present application, a data classification method is provided, which is applied to a business party, including: obtaining a candidate sample set and an encryption vector characterizing whether each sample in the candidate sample set is a real sample, and performing multiple rounds of iterative operations: based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, determining the party to which the target classification feature belongs and the number of the target classification feature; in response to the target classification feature belonging to the business party, using the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current candidate sample set, dividing the current candidate sample set into at least one sub-candidate sample set, and determining the sub-encryption vector corresponding to the sub-candidate sample set based on the current encryption vector; using the sub-candidate sample set determined in the current iterative operation as the current candidate sample set in the next round of iterative operation, and using the sub-encryption vector determined in the current iterative operation as the current encryption vector in the next round of iterative operation; in response to determining that the iterative operation meets a preset condition, stopping the iterative operation, and determining the target classification feature determined in the multiple rounds of iterative operations as the final classification feature of the candidate sample set.
[0006] In some embodiments, the iteration operation further includes: sending iteration data generated in the current iteration operation to a data source, wherein the iteration data at least includes a subset of candidate samples.
[0007] In some embodiments, the iterative operation also includes: in response to the target classification feature belonging to the data party, sending the number of the target classification feature to the data party, so that the data party divides the current sample set representing the real sample into at least one sub-sample set based on the target classification feature corresponding to the number of the target classification feature; receiving the first encryption vector determined by the data party based on the sub-sample set and the sub-candidate sample set, and using the first encryption vector as the sub-encryption vector determined in the current iterative operation.
[0008] In some embodiments, the encryption vector is generated by using homomorphic encryption technology and encrypting the data party's key sent by the data party; based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, determine the owner of the target classification feature and the number of the target classification feature, including: receiving the following information of each classification feature of the business party determined by the data party based on the first encrypted aggregation information of each classification feature of the business party: information gain and number, the first encrypted aggregation information is generated based on the data party key, and the aggregation information includes the number of each classification feature; receiving the second encrypted aggregation information of each classification feature of the data party, determining the following information of each classification feature of the data party: information gain and number, the second encrypted aggregation information is generated by the data party based on the received business party key; based on the comparison results of all information gains, determine the owner of the target classification feature and the number of the target classification feature.
[0009] In some embodiments, based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, determining the party to which the target classification feature belongs and the number of the target classification feature, includes: receiving the following information of the business party determined by the data party based on the first encrypted aggregate information of each classification feature of the business party: the first information gain maximum value and the number of the classification feature corresponding to the first information gain maximum value; receiving the second encrypted aggregate information of each classification feature of the data party, determining the following information of the data party: the number of the classification feature corresponding to the second information gain maximum value; based on the comparison result of the first information gain maximum value and the second information gain maximum value, determining the party to which the target classification feature belongs and the number of the target classification feature.
[0010] In some embodiments, the first encrypted aggregate information is generated based on the current encryption vector and a data party key.
[0011] In some embodiments, the data classification process is used as a single tree establishment process in tree model construction, and the current encryption vector determined in the iterative operation is used as the encryption vector of each node; the method also includes: based on the encryption vector of each leaf node and the data party key, determining the third encrypted aggregate information of each leaf node, and sending the third encrypted aggregate information to the data party; based on the decryption result of the third encrypted aggregate information returned by the data party, determining the weight of each leaf node.
[0012] According to a second aspect of the present application, a data classification method is provided, which is applied to a data party, including: receiving first encrypted aggregate information of each classification feature of a business party sent by a business party, the aggregate information including a number of each classification feature; based on a decryption result of the first encrypted aggregate information, determining the following information of each classification feature of the business party: information gain and number; and sending the following information of each classification feature of the business party: information gain and number to the business party.
[0013] In some embodiments, the method also includes: receiving iterative data sent by the business party and updating the current sample set, the iterative data at least including a sub-candidate sample set; in response to receiving the number of the target classification feature, using the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current sample set, dividing the current sample set into at least one sub-sample set, and based on the sub-sample set and the sub-candidate sample set, determining a first encryption vector that characterizes whether each sample in the sub-candidate sample set is a real sample; and sending the first encryption vector to the business party.
[0014] According to a third aspect of the present application, a data classification device is provided, which is applied to a business party, including: an iteration unit, configured to obtain a candidate sample set and an encrypted vector representing whether each sample in the candidate sample set is a real sample, and perform multiple rounds of iteration operations; a first determination unit, configured to determine the party to which the target classification feature belongs and the number of the target classification feature based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number; a classification unit, configured to respond to the target classification feature belonging to the business party, using the target classification feature corresponding to the number of the target classification feature to and data features of sample data corresponding to the current candidate sample set, dividing the current candidate sample set into at least one sub-candidate sample set, and determining the sub-encryption vector corresponding to the sub-candidate sample set based on the current encryption vector; an updating unit, configured to use the sub-candidate sample set determined in the current iteration operation as the current candidate sample set in the next round of iteration operation, and use the sub-encryption vector determined in the current iteration operation as the current encryption vector in the next round of iteration operation; a selecting unit, configured to stop the iteration operation in response to determining that the iteration operation satisfies a preset condition, and determine the target classification features determined in multiple rounds of iteration operations as the final classification features of the candidate sample set.
[0015] In some embodiments, the iteration unit further includes: a first sending module configured to send iteration data generated in the current iteration operation to a data source, wherein the iteration data at least includes a subset of candidate samples.
[0016] In some embodiments, the iteration unit also includes: a second sending module, configured to send the number of the target classification feature to the data party in response to the target classification feature belonging to the data party, so that the data party divides the current sample set representing the real sample into at least one sub-sample set based on the target classification feature corresponding to the number of the target classification feature; an update module, configured to receive a first encryption vector determined by the data party based on the sub-sample set and the sub-candidate sample set, and use the first encryption vector as the sub-encryption vector determined in the current iteration operation.
[0017] In some embodiments, the encryption vector in the device is generated by using homomorphic encryption technology and encrypting the data party's key sent by the data party; the first determination unit includes: a first receiving module, configured to receive the following information of each classification feature of the business party determined by the data party based on the first encrypted aggregation information of each classification feature of the business party: information gain and number, the first encrypted aggregation information is generated based on the data party's key, and the aggregation information includes the number of each classification feature; the second receiving module is configured to receive the second encrypted aggregation information of each classification feature of the data party, and determine the following information of each classification feature of the data party: information gain and number, the second encrypted aggregation information is generated by the data party based on the received business party key; the first determination module is configured to determine the party to which the target classification feature belongs and the number of the target classification feature based on the comparison results of all information gains.
[0018] In some embodiments, the first determination unit includes: a third receiving module, configured to receive the following information of the business party determined by the data party based on the first encrypted aggregate information of each classification feature of the business party: a first information gain maximum value and the number of the classification feature corresponding to the first information gain maximum value; a fourth receiving module, configured to receive the second encrypted aggregate information of each classification feature of the data party, and determine the following information of the data party: the second information gain maximum value and the number of the classification feature corresponding to the second information gain maximum value; and a second determination module, configured to determine the party to which the target classification feature belongs and the number of the target classification feature based on the comparison result of the first information gain maximum value and the second information gain maximum value.
[0019] In some embodiments, the first encrypted aggregate information in the first receiving module and / or the third receiving module is generated based on the current encryption vector and the data party key.
[0020] In some embodiments, the data classification device is used as a device for establishing a single tree in the construction of a tree model, and the current encryption vector determined in the iterative operation in the device is used as the encryption vector of each node; the device also includes: a second determination unit, configured to determine the third encrypted aggregate information of each leaf node based on the encryption vector of each leaf node and the data party key, and send the third encrypted aggregate information to the data party; a third determination unit, configured to determine the weight of each leaf node based on the decryption result of the third encrypted aggregate information returned by the data party.
[0021] According to a fourth aspect of the present application, a data classification device is provided, which is applied to a data party, and includes: a first receiving unit, configured to receive first encrypted aggregate information of each classification feature of a business party sent by a business party, the aggregate information including a number of each classification feature; a determination unit, configured to determine the following information of each classification feature of the business party based on a decryption result of the first encrypted aggregate information: information gain and number; a first sending unit, configured to send the following information of each classification feature of the business party: information gain and number to the business party.
[0022] In some embodiments, the device also includes: a second receiving unit, configured to receive iterative data sent by the business party and update the current sample set, the iterative data at least including a sub-candidate sample set; a classification unit, configured to respond to receiving the number of the target classification feature, using the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current sample set, to divide the current sample set into at least one sub-sample set, and based on the sub-sample set and the sub-candidate sample set, determine a first encryption vector that characterizes whether each sample in the sub-candidate sample set is a real sample; a second sending unit, configured to send the first encryption vector to the business party.
[0023] According to the fifth aspect of the present application, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any implementation manner of the first aspect or the second aspect.
[0024] According to the sixth aspect of the present application, the present application provides a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to enable a computer to execute a method described in any implementation of the first aspect or the second aspect.
[0025] According to the technology of the present application, a candidate sample set and an encryption vector characterizing whether each sample in the candidate sample set is a real sample are obtained, and multiple rounds of iterative operations are performed: based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, the party to which the target classification feature belongs and the number of the target classification feature are determined, in response to the target classification feature belonging to the business party, the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current candidate sample set are used to divide the current candidate sample set into at least one sub-candidate sample set, and the sub-encryption vector corresponding to the sub-candidate sample set is determined based on the current encryption vector, the sub-candidate sample set determined in the current iterative operation is used as the current candidate sample set in the next round of iterative operation, the sub-encryption vector determined in the current iterative operation is used as the current encryption vector in the next round of iterative operation, in response to determining that the iterative operation meets the preset conditions, the iterative operation is stopped, and the target classification feature determined in the multiple rounds of iterative operations is determined as the final classification feature of the candidate sample set, thereby realizing data classification in the case where the business party does not know the real sample set, and improving the efficiency and accuracy of data classification. Applying this method to the model training of federated learning can ensure that the feature meaning of the data party is open to the business party without leaking data privacy, thus achieving the interpretability of the model.
[0026] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present application.
[0028] Figure 1 is a schematic diagram of a first embodiment of a data classification method applied to a business party according to the present application;
[0029] Figure 2 It is a scene diagram that can implement the data classification method of the embodiment of the present application;
[0030] Figure 3 is a schematic diagram of a second embodiment of a data classification method applied to a business party according to the present application;
[0031] Figure 4 is a structural schematic diagram of an embodiment of a data classification method applied to a data party according to the present application;
[0032] Figure 5 An example diagram of a sample set encoding calculation according to the data classification method applied to the business party in this application;
[0033] Figure 6 is a structural schematic diagram of an embodiment of a data classification device applied to a business party according to the present application;
[0034] Figure 7 is a structural schematic diagram of an embodiment of a data classification device applied to a data party according to the present application;
[0035] Figure 8 It is a block diagram of an electronic device used to implement the data classification method of an embodiment of the present application. DETAILED DESCRIPTION
[0036] The following is a description of exemplary embodiments of the present application in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0037] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0038] Figure 1 A schematic diagram 100 showing a first embodiment of a data classification method according to the present application, which is applied to a business party, comprises the following steps:
[0039] Step 101 , obtaining a candidate sample set and an encrypted vector representing whether each sample in the candidate sample set is a real sample.
[0040] In this embodiment, the execution subject (e.g., the server of the business party) can obtain the candidate sample set representing each user and the encrypted vector representing whether each sample in the candidate sample set is a real sample from the local. The encrypted vector can be a vector that is pre-encrypted by other parties and received by the execution subject. The dimension of the encrypted vector is consistent with the number of samples in the candidate sample set. The encrypted vector represents the encoding of whether each sample in the candidate sample set is a real sample. For example, when the vector encoding value is 1, it means that the sample in the candidate sample set is the current real sample. When the vector encoding value is 0, it means that the sample in the candidate sample set is not the current real sample. When the tree model is trained, the current real sample can refer to the real sample of the current node.
[0041] Step 102, perform multiple rounds of iterative operations:
[0042] Step 1021, based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, determine the party to which the target classification feature belongs and the number of the target classification feature.
[0043] In this embodiment, the execution subject can analyze all information gains based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, determine the party with the largest information gain as the party to which the target classification feature belongs, and determine the number corresponding to the classification feature with the largest information gain as the number of the target classification feature. The target classification feature refers to the feature based on which feature the sample data corresponding to the sample set is classified. For example, if the user sample data in the user sample data set is classified based on gender, then gender is the target classification feature; if the user sample data in the user sample data set is classified based on age, then age is the target classification feature; if the item sample data set in the item sample data set is classified based on size, then size is the target classification feature. Since the business party determines the party to which the target classification feature belongs and the number based on the information gain of each classification feature and the number of each classification feature, the business party does not obtain the sample data, and can avoid the risk of leakage of sample data while determining the target classification feature.
[0044] Step 1022, in response to the target classification feature belonging to the business party, the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current candidate sample set are used to divide the current candidate sample set into at least one sub-candidate sample set, and the sub-encryption vector corresponding to the sub-candidate sample set is determined based on the current encryption vector.
[0045] In this embodiment, when the execution subject determines that the target classification feature belongs to the business party (i.e., the party), the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current candidate sample set can be used to classify the sample data, and the current candidate sample set is divided into at least one sub-candidate sample set according to the division result of the sample data, and then the sub-encryption vector corresponding to the sub-candidate sample set is determined based on the current encryption vector. Specifically, each label corresponding to the target classification feature is determined based on the target classification feature, and for each sample data corresponding to the candidate sample set, the sample data is divided into the corresponding sub-sample data set based on the degree of conformity between the data feature of the sample data itself and the label, and the corresponding sub-candidate sample sets are determined according to the divided sub-sample data sets. The sub-encryption vector is obtained by screening the current encryption vector according to the determined sub-candidate sample set, and the dimension of the sub-encryption vector is consistent with the number of samples in the sub-sample data set.
[0046] For example, if the target classification feature is age, and the labels corresponding to the target classification feature are 20-30 years old, 31-40 years old, and 41-50 years old, then the sample data corresponding to the candidate sample set can be divided into the 20-30 years old sub-sample data set, the sample data with the data feature of 31-40 years old can be divided into the 31-40 years old sub-sample data set, and the sample data with the data feature of 41-50 years old can be divided into the 41-50 years old sub-sample data set, and then the samples corresponding to each sub-sample data set are determined as the corresponding sub-candidate sample set. If the candidate sample set is assumed to be user {1,2,3,4,5}, the encryption vector is <[1,1,0,1,1]>, and the sub-candidate sample sets after division are {1,2,3} and {4,5}, then the sub-encryption vector corresponding to the sub-candidate sample set {1,2,3} is <[1,1,0]>, and the sub-encryption vector corresponding to the sub-candidate sample set {4,5} is <[1,1]>.
[0047] In some optional implementations of this embodiment, the iteration operation further includes: sending the iteration data generated in the current iteration operation to the data party, the iteration data at least includes the sub-candidate sample set, and the iteration data may be the number of iterations, the number of layers of leaf nodes generated after the current round of iteration operation (i.e., the depth of the decision tree model containing the leaf nodes), etc. By synchronizing the iteration data to the data party, the data party stores the real sample set, so that the data party can perform subsequent operations based on the real sample set.
[0048] Step 1023: Use the sub-candidate sample set determined in the current iteration operation as the current candidate sample set in the next iteration operation, and use the sub-encryption vector determined in the current iteration operation as the current encryption vector in the next iteration operation.
[0049] In this embodiment, the execution entity can use the sub-candidate sample set determined in the current iterative operation as the current candidate sample set in the next round of iterative operation, and use the sub-encryption vector determined in the current iterative operation as the current encryption vector in the next round of iterative operation, so as to reduce the number of samples in the candidate sample set and reduce the dimension of the encryption vector corresponding to the candidate sample set.
[0050] In multiple rounds of iterative operations, the candidate sample set is divided into sub-candidate sample sets, and the sub-candidate sample sets are further classified into the next level of sub-candidate sample sets. Each classification of the candidate sample set / sub-candidate sample set is based on the target classification feature (such as gender). After each classification, each sub-candidate sample set has a label corresponding to the target classification feature (such as female). After the candidate sample sets are classified layer by layer, the label of the sub-candidate sample set at the last level is a more accurate label that describes the sample data in the sub-candidate sample set.
[0051] Step 103, in response to determining that the iterative operation meets the preset condition, the iterative operation is stopped, and the target classification features determined in multiple rounds of iterative operations are determined as the final classification features of the candidate sample set.
[0052] In this embodiment, if the execution subject determines that the iterative operation meets the preset conditions, the iterative operation is stopped, and the target classification features determined in multiple rounds of iterative operations are determined as the final classification features of the sample data in the sample data set.
[0053] Specifically, the target classification feature determined in the last round of iterative operations can be used as the final classification feature, for example, the target classification feature "gender" determined in the last round of iterative operations can be used as the final classification feature. Multiple target classification features determined in multiple rounds of iterative operations can also be used as the final classification feature, for example, the target classification features "gender" - "age" - "native place" determined one by one in multiple rounds of iterative operations can be used as the final classification feature.
[0054] Continue to see Figure 2 , the data classification method 200 of this embodiment runs in an electronic device 201. First, the electronic device 201 obtains a candidate sample set and an encrypted vector 202 that characterizes whether each sample in the candidate sample set is a real sample. Then, the electronic device 201 performs multiple rounds of iteration operations 203. The iteration process is as follows: first, based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, determine the party to which the target classification feature belongs and the number of the target classification feature 2031. Then, in response to the target classification feature belonging to the business party, use the target classification feature corresponding to the number of the target classification feature and the sample data corresponding to the current candidate sample set. The data feature of the data is used to divide the current candidate sample set into at least one sub-candidate sample set, and determine the sub-encryption vector 2032 corresponding to the sub-candidate sample set based on the current encryption vector. Then, the sub-candidate sample set determined in the current iteration operation is used as the current candidate sample set in the next round of iteration operation, and the sub-encryption vector determined in the current iteration operation is used as the current encryption vector 2033 in the next round of iteration operation. Finally, the electronic device 201 stops the iteration operation in response to determining that the iteration operation meets the preset conditions, and determines the target classification feature determined in multiple rounds of iteration operations as the final classification feature 204 of the candidate sample set.
[0055] The data classification method provided by the above-mentioned embodiment of the present application adopts the method of obtaining a candidate sample set and an encryption vector that characterizes whether each sample in the candidate sample set is a real sample, and performing multiple rounds of iterative operations: based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, determining the party to which the target classification feature belongs and the number of the target classification feature, in response to the target classification feature belonging to the business party, using the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current candidate sample set, dividing the current candidate sample set into at least one sub-candidate sample set, and determining the sub-encryption vector corresponding to the sub-candidate sample set based on the current encryption vector, using the sub-candidate sample set determined in the current iterative operation as the current candidate sample set in the next round of iterative operation, using the sub-encryption vector determined in the current iterative operation as the current encryption vector in the next round of iterative operation, in response to determining that the iterative operation meets the preset conditions, stopping the iterative operation, and determining the target classification feature determined in the multiple rounds of iterative operations as the final classification feature of the candidate sample set, thereby realizing a data classification method in which the business party does not know the real sample set (only holds the encrypted value of the sample set). By gradually optimizing the number of sample sets and reducing the dimension of the encrypted vector, the efficiency and accuracy of classifying candidate sample sets are improved. Applying this method to the model training of federated learning can ensure that the feature meaning of the data party is open to the business party without leaking data privacy, thus achieving the interpretability of the model.
[0056] Further references Figure 3 , which shows a schematic diagram 300 of a second embodiment of a data classification method. The process of the method includes the following steps:
[0057] Step 301 : Obtain a candidate sample set and an encrypted vector representing whether each sample in the candidate sample set is a real sample.
[0058] Step 302, perform multiple rounds of iterative operations:
[0059] Step 3021, based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, determine the party to which the target classification feature belongs and the number of the target classification feature.
[0060] In some optional implementations of the present embodiment, the encryption vector is generated by encrypting using homomorphic encryption technology and a data party key sent by the data party; based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, the party to which the target classification feature belongs and the number of the target classification feature are determined, including: receiving the following information of each classification feature of the business party determined by the data party based on the first encrypted aggregation information of each classification feature of the business party: information gain and number, the first encrypted aggregation information is generated using homomorphic encryption technology based on the data party key, and the aggregation information includes the number of each classification feature; receiving the second encrypted aggregation information of each classification feature of the data party, determining the following information of each classification feature of the data party: information gain and number, the second encrypted aggregation information can be generated by the data party using homomorphic encryption technology based on the received business party key; based on the comparison results of all information gains, determining the party to which the target classification feature belongs and the number of the target classification feature. Two sets of homomorphic encryption public and private key pairs are used. For the feature classification of the data party, the business party decrypts the aggregate derivative for the data party, while for the feature classification of the business party, the data party decrypts the aggregate derivative for the business party. This allows the corresponding information gain to be securely calculated, so that the feature meaning of the data party can be disclosed to the business party.
[0061] In addition, each classification feature of the business party and each classification feature of the data party can be determined based on manual experience by performing statistical analysis on the data features of their own sample data, or can be predetermined using a trained sample feature prediction model.
[0062] In some optional implementations of this embodiment, based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, determine the party to which the target classification feature belongs and the number of the target classification feature, including: receiving the following information of the business party determined by the data party based on the first encrypted aggregation information of each classification feature of the business party: the first information gain maximum value and the number of the classification feature corresponding to the first information gain maximum value, the first encrypted aggregation information is generated based on the current encryption vector and the data party key; receiving the second encrypted aggregation information of each classification feature of the data party, determining the following information of the data party: the number of the classification feature corresponding to the second information gain maximum value and the second information gain maximum value; based on the comparison result of the first information gain maximum value and the second information gain maximum value, determine the party to which the target classification feature belongs and the number of the target classification feature. By only transmitting the information gain maximum value and the number corresponding to the information gain maximum value, the communication overhead is reduced and the system processing efficiency is improved. The current encryption vector is added to the calculation method of the aggregation information of the business party to make the data calculation more accurate.
[0063] Specifically, the calculation method of the aggregation value in the first encrypted aggregation information may be: Among them, i represents the identity of the target classification feature, gi represents the first-order derivative of the loss value, and I G represents the candidate sample set, I represents the current real sample set, To perform encryption operations using the public key of the homomorphic encryption system generated by the data party.
[0064] Step 3022, in response to the target classification feature belonging to the data party, sending the serial number of the target classification feature to the data party.
[0065] In this embodiment, when the execution entity determines that the target classification feature belongs to the data party, the number of the target classification feature is sent to the data party so that the data party divides the current sample set representing the real sample into at least one sub-sample set based on the target classification feature corresponding to the number of the target classification feature.
[0066] Step 3023: The data receiving party determines the first encryption vector based on the sub-sample set and the sub-candidate sample set, and uses the first encryption vector as the sub-encryption vector determined in the current iterative operation.
[0067] In this embodiment, the execution subject may receive the first encryption vector determined by the data party based on the sub-sample set and the sub-candidate sample set, and use the first encryption vector as the sub-encryption vector determined in the current iteration operation, while the sub-candidate sample set remains unchanged.
[0068] Step 3024: Use the sub-candidate sample set determined in the current iteration operation as the current candidate sample set in the next iteration operation, and use the sub-encryption vector determined in the current iteration operation as the current encryption vector in the next iteration operation.
[0069] Step 303, in response to determining that the iterative operation meets the preset condition, the iterative operation is stopped, and the target classification features determined in multiple rounds of iterative operations are determined as the final classification features of the candidate sample set.
[0070] In some optional implementations of this embodiment, the data classification process is used as the process of establishing a single tree in the tree model construction, and the current encryption vector determined in the iterative operation is used as the encryption vector of each node; the method also includes: based on the encryption vector of each leaf node and the data party key, determining the third encrypted aggregate information of each leaf node, and sending the third encrypted aggregate information to the data party; based on the decryption result of the third encrypted aggregate information returned by the data party, determining the weight of each leaf node. A more secure calculation of the leaf node weight is achieved, and the characteristic meaning of the data party can be further disclosed to the business party.
[0071] In this embodiment, the specific operations of steps 301, 3021, 3024 and 303 are the same as Figure 1The operations of steps 101, 1021, 1023 and 103 in the illustrated embodiment are substantially the same and will not be described in detail herein.
[0072] from Figure 3 It can be seen that Figure 1 Compared with the corresponding embodiment, the schematic diagram 300 of the data classification method in this embodiment adopts the method of sending the target classification feature number to the data party in response to the target classification feature belonging to the data party, so that the data party divides the current sample set representing the real sample into at least one sub-sample set based on the target classification feature corresponding to the target classification feature number, receives the first encryption vector determined by the data party based on the sub-sample set and the sub-candidate sample set, and uses the first encryption vector as the sub-encryption vector determined in the current iteration operation, uses the sub-candidate sample set determined in the current iteration operation as the current candidate sample set in the next round of iteration operation, and uses the sub-encryption vector determined in the current iteration operation as the current encryption vector in the next round of iteration operation, thereby ensuring the accuracy of the current encryption vector, making the target classification feature and / or data classification more accurate, and realizing the complete process of data classification by multiple parties. Data transmission is performed only when the target classification feature belongs to the data party, and no data transmission is required when the target classification feature belongs to the business party, which greatly reduces the communication overhead in the data classification process.
[0073] Further references Figure 4 , a schematic diagram 400 of an embodiment of a data classification method according to the present application is shown, applied to a data party, and includes the following steps:
[0074] Step 401: Receive first encrypted aggregate information of various classification features of a business party sent by a business party.
[0075] In this embodiment, the execution entity (eg, the server of the data party) may receive first encrypted aggregate information of each classification feature of the business party sent by the business party, where the aggregate information includes the serial number of each classification feature.
[0076] Step 402: Based on the decryption result of the first encrypted aggregate information, determine the following information of each classification feature of the business party: information gain and number.
[0077] In this embodiment, the execution entity may determine the information gain and the number of each classification feature of the business party by using the information gain calculation method based on the decryption result of the first encrypted aggregate information, that is, the decrypted aggregate information.
[0078] Step 403, sending the following information of each classification feature of the business party: information gain and number to the business party.
[0079] In this embodiment, the execution entity may send the information gain of each classification feature of the business party and the number of each classification feature of the business party to the business party.
[0080] The data classification method provided by the above-mentioned embodiment of the present application adopts the first encrypted aggregate information of each classification feature of the business party sent by the receiving business party, and determines the following information of each classification feature of the business party based on the decryption result of the first encrypted aggregate information: information gain and number, and sends the following information of each classification feature of the business party: information gain and number to the business party. Because the aggregate information of the business party is encrypted by the data party's key, the data party decrypts the aggregate information, and generates the information gain of each classification feature of the business party and the number of each classification feature of the business party, and returns it to the business party, thereby improving the efficiency of data classification of the business party and realizing the data classification process of multiple parties.
[0081] Optionally, the method also includes: receiving iterative data sent by the business party and updating the current sample set, the iterative data at least including the sub-candidate sample set; in response to receiving the number of the target classification feature, using the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current sample set, dividing the current sample set into at least one sub-sample set, and based on the sub-sample set and the sub-candidate sample set, determining the first encrypted vector representing whether each sample in the sub-candidate sample set is a real sample; sending the first encrypted vector to the business party. The data classification of the data party is realized, so that when the data party holds the real sample set at the classification node, the encrypted vector representing the real sample set is sent to the business party, so that the business party has the encrypted value of the sample set. This method is applied to the model training of federated learning. While ensuring the model training, it can also ensure that the feature meaning of the data party is open to the business party without leaking data privacy, thereby realizing the interpretability of the model.
[0082] To further specify, Figure 5 An example diagram of the sample set encoding calculation for data classification. Assume that business party A has candidate sample set I at the current node f1. G ={1,2,3,4,5}, the encrypted vector corresponding to the candidate sample set <π> = <[1,1,1,1,1]>, π = 1 means that the sample belongs to the sample of the current node, that is, it is a real sample, π = 0 means that the sample does not belong to the sample of the current node, that is, it is a non-real sample, and the sample set of data party B at the current node f1 is I = {1,2,3,4,5}. When the target classification feature belongs to the business party, the business party classifies the data for the candidate sample set. In the classified current node f2, the candidate sample set I of business party A G ={3,4,5}, encrypted vector <π> = <[1,1,1]>, data party B receives candidate sample set I G={3,4,5}, the sample set of the current node f2 is updated to obtain the updated sample set I = {3,4,5}. Then for the current node f2, when the target classification feature belongs to the data party, the data party classifies the sample set. In the classified current node f3, the sample set classified by data party B is I = {4,5}, and the candidate sample set of business party A remains unchanged, I G ={3,4,5}, the data side is based on the I and I of the current node G , generate a first encryption vector <π'>=<[0,1,1]> and send it to the business party. The business party determines that the encryption vector of the current node is <π>=<[0,1,1]> according to the received first encryption vector.
[0083] Further references Figure 6 , as a response to the above Figures 1 to 3 In order to realize the method shown in the figure, the present application provides an embodiment of a data classification device, and the device embodiment is Figure 1 Corresponding to the method embodiment shown in FIG. 1 , in addition to the features described below, the device embodiment may also include Figure 1 The same or corresponding features of the method embodiment shown, and the generation of Figure 1 The method embodiment shown has the same or corresponding effects, and the device can be specifically applied to various electronic devices.
[0084] like Figure 6 As shown, the data classification device 600 of this embodiment is applied to the business party, and the device includes: an iteration unit 601, a first determination unit 602, a classification unit 603, an updating unit 604 and a selection unit 605, wherein the iteration unit is configured to obtain a candidate sample set and an encrypted vector representing whether each sample in the candidate sample set is a real sample, and perform multiple rounds of iteration operations: the first determination unit is configured to determine the party to which the target classification feature belongs and the number of the target classification feature based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number; the classification unit is configured to respond to the target classification feature belonging to the business party, The target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current candidate sample set are used to divide the current candidate sample set into at least one sub-candidate sample set, and the sub-encryption vector corresponding to the sub-candidate sample set is determined based on the current encryption vector; the updating unit is configured to use the sub-candidate sample set determined in the current iteration operation as the current candidate sample set in the next round of iteration operation, and the sub-encryption vector determined in the current iteration operation as the current encryption vector in the next round of iteration operation; the selecting unit is configured to stop the iteration operation in response to determining that the iteration operation meets the preset condition, and determine the target classification feature determined in multiple rounds of iteration operations as the final classification feature of the candidate sample set.
[0085] In this embodiment, the specific processing of the iteration unit 601, the first determination unit 602, the classification unit 603, the updating unit 604 and the selection unit 605 of the data classification device 600 and the technical effects thereof can be referred to respectively. Figure 1 The relevant descriptions of steps 101 to 103 in the corresponding embodiment are not repeated here.
[0086] In some optional implementations of this embodiment, the iteration unit further includes: a first sending module configured to send iteration data generated in the current iteration operation to the data party, where the iteration data at least includes a sub-candidate sample set.
[0087] In some optional implementations of this embodiment, the iteration unit also includes: a second sending module, configured to send the number of the target classification feature to the data party in response to the target classification feature belonging to the data party, so that the data party divides the current sample set representing the real sample into at least one sub-sample set based on the target classification feature corresponding to the number of the target classification feature; an update module, configured to receive a first encryption vector determined by the data party based on the sub-sample set and the sub-candidate sample set, and use the first encryption vector as the sub-encryption vector determined in the current iteration operation.
[0088] In some optional implementations of the present embodiment, the encryption vector in the device is generated by using homomorphic encryption technology and encrypting the data party key sent by the data party; the first determination unit includes: a first receiving module, configured to receive the following information of each classification feature of the business party determined by the data party based on the first encrypted aggregation information of each classification feature of the business party: information gain and number, the first encrypted aggregation information is generated based on the data party key, and the aggregation information includes the number of each classification feature; the second receiving module is configured to receive the second encrypted aggregation information of each classification feature of the data party, and determine the following information of each classification feature of the data party: information gain and number, the second encrypted aggregation information is generated by the data party based on the received business party key; the first determination module is configured to determine the party to which the target classification feature belongs and the number of the target classification feature based on the comparison results of all information gains.
[0089] In some optional implementations of the present embodiment, the first determination unit includes: a third receiving module, configured to receive the following information of the business party determined by the data party based on the first encrypted aggregate information of each classification feature of the business party: a first information gain maximum value and the number of the classification feature corresponding to the first information gain maximum value; a fourth receiving module, configured to receive the second encrypted aggregate information of each classification feature of the data party, and determine the following information of the data party: the second information gain maximum value and the number of the classification feature corresponding to the second information gain maximum value; and a second determination module, configured to determine the party to which the target classification feature belongs and the number of the target classification feature based on the comparison result of the first information gain maximum value and the second information gain maximum value.
[0090] In some optional implementations of this embodiment, the first encrypted aggregate information in the first receiving module and / or the third receiving module is generated based on the current encryption vector and the data party key.
[0091] In some optional implementations of the present embodiment, the data classification device is used as a device for establishing a single tree in the tree model construction, and the current encryption vector determined in the iterative operation in the device is used as the encryption vector of each node; the device also includes: a second determination unit, configured to determine the third encrypted aggregate information of each leaf node based on the encryption vector of each leaf node and the data party key, and send the third encrypted aggregate information to the data party; a third determination unit, configured to determine the weight of each leaf node based on the decryption result of the third encrypted aggregate information returned by the data party.
[0092] Further references Figure 7 , as a response to the above Figure 4 In order to realize the method shown in the figure, the present application provides an embodiment of a data classification device, and the device embodiment is Figure 4 Corresponding to the method embodiment shown in FIG. 1 , in addition to the features described below, the device embodiment may also include Figure 4 The same or corresponding features of the method embodiment shown, and the generation of Figure 4 The method embodiment shown has the same or corresponding effects, and the device can be specifically applied to various electronic devices.
[0093] like Figure 7As shown, the data classification device 700 of this embodiment is applied to the data party, and the device includes: a first receiving unit 701, a determining unit 702 and a first sending unit 703, wherein the first receiving unit is configured to receive first encrypted aggregate information of each classification feature of the business party sent by the business party, and the aggregate information includes the number of each classification feature; the determining unit is configured to determine the following information of each classification feature of the business party based on the decryption result of the first encrypted aggregate information: information gain and number; the first sending unit is configured to send the following information of each classification feature of the business party: information gain and number to the business party.
[0094] In this embodiment, the specific processing of the first receiving unit 701, the determining unit 702 and the first sending unit 703 of the data classification device 700 and the technical effects thereof can be referred to in Figure 4 The relevant descriptions of steps 401 to 403 in the corresponding embodiment are not repeated here.
[0095] In some optional implementations of this embodiment, the device also includes: a second receiving unit, configured to receive iterative data sent by the business party and update the current sample set, the iterative data at least including a sub-candidate sample set; a classification unit, configured to respond to receiving the number of the target classification feature, using the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current sample set, to divide the current sample set into at least one sub-sample set, and based on the sub-sample set and the sub-candidate sample set, determine a first encryption vector that characterizes whether each sample in the sub-candidate sample set is a real sample; a second sending unit, configured to send the first encryption vector to the business party.
[0096] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.
[0097] like Figure 8 , is a block diagram of an electronic device according to a data classification method according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0098] like Figure 8As shown, the electronic device includes: one or more processors 801, memory 802, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are interconnected using different buses and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 8 A processor 801 is taken as an example.
[0099] The memory 802 is a non-transient computer-readable storage medium provided in the present application. The memory stores instructions executable by at least one processor to enable at least one processor to perform the data classification method provided in the present application. The non-transient computer-readable storage medium of the present application stores computer instructions, which are used to enable a computer to perform the data classification method provided in the present application.
[0100] The memory 802 is a non-transient computer-readable storage medium that can be used to store non-transient software programs, non-transient computer executable programs and modules, such as program instructions / modules corresponding to the data classification method in the embodiment of the present application (for example, the attached Figure 6 The processor 801 executes various functional applications and data processing of the server by running the non-transient software programs, instructions and modules stored in the memory 802, that is, the data classification method in the above method embodiment is implemented.
[0101] The memory 802 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created based on the use of a data update electronic device based on federated learning, etc. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 802 may optionally include a memory remotely disposed relative to the processor 801, and these remote memories may be connected to the data update electronic device based on federated learning via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0102] The electronic device of the data classification method may further include: an input device 803 and an output device 804. The processor 801, the memory 802, the input device 803 and the output device 804 may be connected via a bus or other means. Figure 8 The example of connecting through bus is taken in the following.
[0103] The input device 803 can receive input digital or character information, and generate key signal input related to updating user settings and function control of the electronic device based on federated learning data, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, an indicator rod, one or more mouse buttons, a trackball, a joystick, and other input devices. The output device 804 may include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0104] Various implementations of the systems and techniques described herein can be realized in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0105] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for programmable processors and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or means (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0107] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0108] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.
[0109] According to the technical solution of the embodiment of the present application, a candidate sample set and an encryption vector characterizing whether each sample in the candidate sample set is a real sample are obtained, and multiple rounds of iterative operations are performed: based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, the party to which the target classification feature belongs and the number of the target classification feature are determined, in response to the target classification feature belonging to the business party, the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current candidate sample set are used to divide the current candidate sample set into at least one sub-candidate sample set, and the sub-encryption vector corresponding to the sub-candidate sample set is determined based on the current encryption vector, the sub-candidate sample set determined in the current iterative operation is used as the current candidate sample set in the next round of iterative operation, the sub-encryption vector determined in the current iterative operation is used as the current encryption vector in the next round of iterative operation, in response to determining that the iterative operation meets the preset conditions, the iterative operation is stopped, and the target classification feature determined in the multiple rounds of iterative operations is determined as the final classification feature of the candidate sample set, thereby realizing data classification in the case where the business party does not know the real sample set, and improving the efficiency and accuracy of data classification. Applying this method to the model training of federated learning can ensure that the feature meaning of the data party is open to the business party without leaking data privacy, thus achieving the interpretability of the model.
[0110] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution disclosed in this application can be achieved, and this document is not limited here.
[0111] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this application should be included in the protection scope of this application.
Claims
1. A data classification method, applied to a business party, including: Obtaining a candidate sample set and an encrypted vector, and performing multiple rounds of iterative operations. The encrypted vector is generated by encrypting an encoding vector representing whether each sample in the candidate sample set is a real sample: Based on the following information of each classification feature of the business party sent by the data party: information gain and number, and the following information of each classification feature of the data party: information gain and number, determining the owner of the target classification feature and the number of the target classification feature; In response to the target classification feature belonging to the business party, using the target classification feature corresponding to the number of the target classification feature and the data features of the sample data corresponding to the current candidate sample set, dividing the current candidate sample set into at least one sub-candidate sample set, and determining a sub-encrypted vector corresponding to the sub-candidate sample set based on the current encrypted vector; Taking the sub-candidate sample set determined in the current iterative operation as the current candidate sample set in the next iterative operation, and taking the sub-encrypted vector determined in the current iterative operation as the current encrypted vector in the next iterative operation; In response to determining that the iterative operation meets a preset condition, stopping the iterative operation, and determining the target classification feature determined in the multiple rounds of iterative operations as the final selected classification feature of the candidate sample set.
2. The method according to claim 1, wherein, the iterative operation further includes: Sending the iterative data generated in the current iterative operation to the data party, and the iterative data at least includes the sub-candidate sample set.
3. The method according to claim 2, wherein, the iterative operation further includes: In response to the target classification feature belonging to the data party, sending the number of the target classification feature to the data party, so that the data party divides the current sample set representing real samples into at least one sub-sample set based on the target classification feature corresponding to the number of the target classification feature; Receiving a first encrypted vector determined by the data party based on the sub-sample set and the sub-candidate sample set, and taking the first encrypted vector as the sub-encrypted vector determined in the current iterative operation.
4. The method according to claim 1, wherein, the encrypted vector is generated by using homomorphic encryption technology and encrypting with the data party key sent by the data party; the determining the owner of the target classification feature and the number of the target classification feature based on the following information of each classification feature of the business party sent by the data party: information gain and number, and the following information of each classification feature of the data party: information gain and number, includes: Receiving the following information of each classification feature of the business party determined by the data party based on the first encrypted aggregation information of each classification feature of the business party: information gain and number, the first encrypted aggregation information is generated based on the data party key, and the aggregation information includes the numbers of each classification feature; Receiving the second encrypted aggregation information of each classification feature of the data party, and determining the following information of each classification feature of the data party: information gain and number, the second encrypted aggregation information is generated by the data party based on the received business party key; Based on the comparison results of all the information gains, the party to which the target classification feature belongs and the serial number of the target classification feature are determined.
5. The method according to claim 4, in, The determining the party to which the target classification feature belongs and the number of the target classification feature based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number, includes: The data receiving party determines the following information of the business party based on the first encrypted aggregate information of each classification feature of the business party: the first information gain maximum value and the number of the classification feature corresponding to the first information gain maximum value; receiving second encrypted aggregation information of each classification feature of the data party, and determining the following information of the data party: a second maximum information gain value and a serial number of the classification feature corresponding to the second maximum information gain value; Based on the comparison result of the first information gain maximum value and the second information gain maximum value, the party to which the target classification feature belongs and the serial number of the target classification feature are determined.
6. The method according to any one of claims 4 and 5, in, The first encrypted aggregate information is generated based on a current encryption vector and the data party key.
7. The method according to claim 4, in, The data classification process is used as a single tree establishment process in tree model construction, and the current encryption vector determined in the iterative operation is used as the encryption vector of each node; the method further includes: Determine third encrypted aggregate information of each leaf node based on the encryption vector of each leaf node and the data party key, and send the third encrypted aggregate information to the data party; Based on the decryption result of the third encrypted aggregate information returned by the data party, the weight of each leaf node is determined.
8. A data classification method, applied to data, include: receiving first encrypted aggregate information of each classification feature of the business party sent by the business party, wherein the aggregate information includes a serial number of each classification feature; Based on the decryption result of the first encrypted aggregate information, determine the following information of each classification feature of the business party: information gain and number; Sending the following information of each classification feature of the business party: information gain and number to the business party; the method also includes: Receiving iterative data sent by the business party and updating the current sample set, wherein the iterative data at least includes a sub-candidate sample set; In response to receiving the number of the target classification feature, using the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current sample set, dividing the current sample set into at least one sub-sample set, and determining, based on the sub-sample set and the sub-candidate sample set, a first encryption vector characterizing whether each sample in the sub-candidate sample set is a true sample; Sending the first encryption vector to the service party.
9. A data classification device, used in business parties, include: The iteration unit is configured to obtain a candidate sample set and an encryption vector, and perform multiple rounds of iteration operations, wherein the encryption vector is generated by encrypting a coding vector representing whether each sample in the candidate sample set is a real sample: A first determining unit is configured to determine the party to which the target classification feature belongs and the number of the target classification feature based on the following information of each classification feature of the business party sent by the data party: information gain and number and the following information of each classification feature of the data party: information gain and number; a classification unit configured to, in response to the target classification feature belonging to the business party, use the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current candidate sample set to divide the current candidate sample set into at least one sub-candidate sample set, and determine a sub-encryption vector corresponding to the sub-candidate sample set based on the current encryption vector; An updating unit is configured to use the sub-candidate sample set determined in the current iteration operation as the current candidate sample set in the next iteration operation, and use the sub-encryption vector determined in the current iteration operation as the current encryption vector in the next iteration operation; The selection unit is configured to stop the iterative operation in response to determining that the iterative operation meets a preset condition, and determine the target classification feature determined in the multiple rounds of iterative operations as the final classification feature of the candidate sample set.
10. The device according to claim 9, wherein the iteration unit further comprises: include: The first sending module is configured to send iterative data generated in the current iterative operation to the data party, wherein the iterative data at least includes the sub-candidate sample set.
11. The device according to claim 10, in, The iteration unit also includes: A second sending module is configured to send the serial number of the target classification feature to the data entity in response to the target classification feature belonging to the data entity, so that the data entity divides the current sample set representing the real sample into at least one sub-sample set based on the target classification feature corresponding to the serial number of the target classification feature; The updating module is configured to receive a first encryption vector determined by the data source based on the sub-sample set and the sub-candidate sample set, and use the first encryption vector as the sub-encryption vector determined in the current iterative operation.
12. A data classification device, applied to data, include: A first receiving unit is configured to receive first encrypted aggregate information of each classification feature of the business party sent by the business party, wherein the aggregate information includes a serial number of each classification feature; a determining unit configured to determine the following information of each classification feature of the business party based on the decryption result of the first encrypted aggregate information: information gain and number; The first sending unit is configured to send the following information of each classification feature of the business party: information gain and number to the business party; the device also includes: a second receiving unit, configured to receive iterative data sent by the business party and update the current sample set, wherein the iterative data at least includes a sub-candidate sample set; a classification unit configured to, in response to receiving the number of the target classification feature, use the target classification feature corresponding to the number of the target classification feature and the data feature of the sample data corresponding to the current sample set to divide the current sample set into at least one sub-sample set, and determine, based on the sub-sample set and the sub-candidate sample set, a first encrypted vector characterizing whether each sample in the sub-candidate sample set is a true sample; The second sending unit is configured to send the first encryption vector to the service party.
13. An electronic device, It is characterized in that include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7 or 8.
14. A non-transitory computer-readable storage medium storing computer instructions, It is characterized in that The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7 or 8.
Citation Information
Patent Citations
Federal learning model training method and device, electronic equipment and storage medium
CN113947211A