Data missing value completion method for longitudinal federated learning and related equipment
By using homomorphic encryption and aggregation techniques in vertical federated learning, taking into account the data distribution situation, the problem of low accuracy of missing value completion in the existing technology is solved, and more efficient data missing value prediction is achieved.
Patent Information
- Application Number
- CN202411228874.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art does not consider the original distribution of data in vertical federated learning, resulting in low accuracy of missing value completion results.
By generating homomorphic encryption public and private keys by collaborative parties, the participants select the initial sample point index from the sample index set, calculate the intercept and intermediate distance matrix, and use public key encryption, the collaborative parties aggregate and decrypt it to obtain the target distance matrix, which is used to divide the sample set and build an index tree, traverse the missing locations to determine the nearest neighbor index set, and predict the missing value.
Taking into account the original distribution of the data, the accuracy of missing value prediction is improved and the effect of data missing value completion is enhanced.
Smart Images

Figure CN120146216A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data prediction, and in particular to a method and related device for filling missing values in vertical federated learning. Background Art
[0002] With the development of the times, people's demand for the privacy protection of personal data has become increasingly strong, and the EU, the United States, and China have successively introduced various relevant policies. In the past few years, artificial intelligence has developed rapidly. Machine learning, computer vision, natural language processing, and deep learning are all based on big data. However, with the increasingly strict legal environment, in many cases, the data scale obtained before modeling cannot meet the training requirements, such as a small amount of data, lack of labels, or some feature values.
[0003] Since data cannot be directly traded now, data islands have formed, and federated learning has emerged as the times require. Federated learning is a machine learning model based on distributed data sets. Existing missing value filling methods based on federated learning include maximum value filling, minimum value filling, and mean value filling, etc. However, methods such as maximum value filling, minimum value filling, and mean value filling are all constant value filling methods, which do not consider the original distribution of the data, resulting in low accuracy of the filling results. Summary of the Invention
[0004] In view of this, the present invention provides a method and related device for filling missing values in vertical federated learning, which are used to solve the problem that the original distribution of the data is not considered in the prior art, resulting in low accuracy of the filling results.
[0005] In a first aspect, an embodiment of the present invention provides a method for filling missing values in vertical federated learning, and the method includes:
[0006] Controlling a collaborating party to generate a public key and a private key for homomorphic encryption, and sending the public key to a participating party corresponding to the collaborating party;
[0007] Controlling the participating party to select an initial sample point index in a sample index set, and determining a target sample index according to the initial sample point index;
[0008] Controlling the participating party to calculate the intercept and the intermediate distance matrix of the participating party through the target sample point, and encrypting the intermediate distance matrix through the public key to obtain an encrypted distance matrix;
[0009] Aggregating the intercept and the encrypted distance matrix respectively by the collaborating party to obtain a global intercept corresponding to the intercept and an aggregated distance matrix corresponding to the encrypted distance matrix, and decrypting the aggregated distance matrix through the private key to obtain a target distance matrix;
[0010] Partition the sample set where the target sample points are located according to the global intercept and the target distance matrix to obtain a left leaf node index set and a right leaf node index set;
[0011] Control the participating parties to construct an index tree according to the left leaf node index set and the right leaf node index set, and traverse the missing positions at each leaf node of the index tree to obtain the target index set of the missing positions at the leaf nodes of the index tree;
[0012] Determine the near neighbor index set according to the target index set, and predict the data missing value of the missing position according to the near neighbor index set.
[0013] Optionally, the step of controlling the participating parties to select an initial sample point index from the sample index set and determine the target sample point according to the initial sample point index includes:
[0014] Control the participating parties to select at least two initial sample point indexes from the sample index set according to a preset rule, and determine the initial sample points corresponding to the initial sample point indexes in the sample set corresponding to the sample index set;
[0015] Select a random sample in the sample set, and calculate the distance between the initial sample and the random sample;
[0016] Update the initial sample point according to the distance and a preset update rule until the initial sample point is updated a preset number of times, and use the initial sample point after being updated a preset number of times as the target sample point.
[0017] Optionally, before the step of controlling the participating parties to construct an index tree according to the left leaf node index set and the right leaf node index set, it further includes:
[0018] Obtain the first quantity information of the indexes in the left leaf node index set and the second quantity information of the indexes in the right leaf node index set;
[0019] If the first quantity information and / or the second quantity information is greater than a threshold, continue to execute the step of controlling the participating parties to select an initial sample point index from the sample index set and determine the target sample index according to the initial sample point index.
[0020] Optionally, before the step of traversing the missing positions at each leaf node of the index tree to obtain the target index set of the missing positions at the leaf nodes of the index tree, it further includes:
[0021] Obtain the first quantity information of the constructed index tree and the second quantity information of the initial index tree corresponding to the sample set;
[0022] When the first number information is less than the second number information, continue to execute the step of controlling the participant to select an initial sample point index from the sample index set and determining a target sample index according to the initial sample point index.
[0023] Optionally, the step of determining the near neighbor index set according to the target index set includes:
[0024] Obtain the missing sample index and the current index where the target feature value is missing according to the target index set;
[0025] Integrate the indexes other than the missing sample index and the current index to obtain the near neighbor index set.
[0026] Optionally, the step of predicting the data missing value at the missing position according to the near neighbor index set includes:
[0027] Determine near neighbor samples according to the near neighbor index set, and obtain the target feature values of a target number of near neighbor samples to obtain a target feature value set;
[0028] Calculate the mean value of the target feature value set, and use the mean value as the data missing value at the missing position.
[0029] Optionally, the step of traversing the missing position at each leaf node of the index tree to obtain the target index set of the leaf node of the index tree where the missing position is located includes:
[0030] Take the positions where the mask matrix takes the value of 1 as the missing positions, and traverse the missing positions at the leaf nodes of each index tree to obtain at least two initial index sets;
[0031] Obtain the union of at least two initial index sets to obtain the target index set.
[0032] On the other hand, an embodiment of the present application provides a data missing value completion device for vertical federated learning, and the device includes:
[0033] A key control module, configured to control the collaborative party to generate a public key and a private key for homomorphic encryption, and send the public key to the participant corresponding to the collaborative party;
[0034] A first control module, configured to control the participant to select an initial sample point index from the sample index set and determine a target sample index according to the initial sample point index;
[0035] The first control module is further configured to control the participant to calculate the intercept and the intermediate distance matrix of the participant through the target sample point, and encrypt the intermediate distance matrix through the public key to obtain an encrypted distance matrix;
[0036] A second control module, configured to aggregate the intercept and the encrypted distance matrix respectively through the collaborating party to obtain a global intercept corresponding to the intercept and an aggregated distance matrix corresponding to the encrypted distance matrix, and decrypt the aggregated distance matrix through the private key to obtain a target distance matrix;
[0037] A partitioning module, configured to partition the sample set where the target sample points are located according to the global intercept and the target distance matrix to obtain a left leaf node index set and a right leaf node index set;
[0038] A construction module, configured to control the participating parties to construct an index tree according to the left leaf node index set and the right leaf node index set, and traverse the missing positions at each leaf node of the index tree to obtain a target index set of the missing positions at the leaf nodes of the index tree;
[0039] An estimation module, configured to determine a nearest neighbor index set according to the target index set, and predict the data missing value at the missing position according to the nearest neighbor index set.
[0040] In a third aspect, an embodiment of the present invention further provides an electronic device, where the electronic device includes:
[0041] One or more processors;
[0042] A storage device, configured to store one or more programs;
[0043] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for filling missing data values for vertical federated learning in any embodiment of the present invention.
[0044] In a fourth aspect, an embodiment of the present invention further provides a storage medium containing computer-executable instructions, where the computer-executable instructions are used to execute the method for filling missing data values for vertical federated learning in any embodiment of the present invention when executed by a computer processor.
[0045] The technical solution of the embodiment of the present invention controls the collaboration party to generate the public key and private key of the homomorphic encryption, and sends the public key to the participating party corresponding to the collaboration party to ensure that the data is used locally by the participating party and ensure the data security; controls the participating party to select the initial sample point index in the sample index set, and determines the target sample index according to the initial sample point index; controls the participating party to calculate the intercept and the intermediate distance matrix of the participating party through the target sample point, and encrypts the intermediate distance matrix through the public key to obtain the encrypted distance matrix; aggregates the intercept and the encrypted distance matrix respectively through the collaboration party to obtain the global intercept corresponding to the intercept and the aggregated distance matrix corresponding to the encrypted distance matrix, and decrypts the aggregated distance matrix through the private key to obtain the target distance matrix; divides the sample set where the target sample point is located according to the global intercept and the target distance matrix to obtain the left leaf node index set and the right leaf node index set; controls the participating party to construct an index tree according to the left leaf node index set and the right leaf node index set, and traverses the missing positions at each leaf node of the index tree to obtain the target index set of the missing positions at the leaf nodes of the index tree; determines the nearest neighbor index set according to the target index set, and predicts the data missing value of the missing position according to the nearest neighbor index set. Considering the original distribution of the data, that is, considering the correlation existing between the data, the prediction of the data missing value is carried out, which improves the accuracy of the prediction result. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0047] Among them:
[0048] Figure 1 It is a schematic flowchart of a method for filling missing data values for vertical federated learning in an embodiment;
[0049] Figure 2 It is a schematic architecture diagram of a method for filling missing data values for vertical federated learning in an embodiment;
[0050] Figure 3 It is another schematic flowchart of a method for filling missing data values for vertical federated learning in an embodiment;
[0051] Figure 4RMSE comparison graph of a data missing value completion method and a statistical completion method for vertical federated learning in an embodiment;
[0052] Figure 5 Schematic structural diagram of a data missing value completion device for vertical federated learning in an embodiment;
[0053] Figure 6 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0055] Before elaborating on the technical solutions of the embodiments of the present invention, the application scenarios of the embodiments of the present invention will be exemplarily described first:
[0056] Definition of missing value completion: On the one hand, in the rapid development of artificial intelligence, machine learning, computer vision, natural language processing, and deep learning in the past few years, these artificial intelligence methods are all based on big data. However, with the increasingly stringent legal environment, the scale of data obtained before modeling often fails to meet the training requirements in many cases, and the problem of feature missing is particularly obvious. Missing data (or missing values) refers to the variable data values that are not stored in the relevant observed values. The problem of missing data is relatively common in almost all studies and will have a significant impact on the conclusions drawn from the data. For example, in the industrial field, industrial time series data is the basis for the research of big data technology in modern industrial systems. Therefore, data integrity is of great significance to the research work in the industrial field. Due to the complexity of the industrial production environment and potential problems in the data transmission process, the problem of data loss often occurs. Another example is in the field of healthcare. Global healthcare institutions are increasingly deploying electronic health record (EHR) systems, which generate electronic medical records for each patient who visits. In addition to supporting administrative processing and routine clinical tasks, the large amount, rich, and diverse patient data in the electronic medical records provides valuable insights for clinical analysis and research, including patient-level analysis and prediction. However, missing values are very common, which greatly limits the use and clinical value of electronic medical records in patient management and cost control. Doctors are often troubled by incomplete patient data, which hinders their effective treatment of patients.
[0057] In view of this, an embodiment of the present invention provides a method for filling missing values in vertical federated learning. First, the coordinator sends the public key required for homomorphic encryption, and the private key is retained locally. When starting to build a tree structure, the tree nodes need to save all the sample indexes included in the current node, that is, the ID set. Each participating party needs the same ID set for calculation to determine how to split the tree nodes. Each participating party uniformly sends the intermediate results calculated according to the ID to the coordinator C. For local variables with lower privacy, such as normal vectors and intercepts, they are directly aggregated into global variables; for locally encrypted variables, such as the calculation results of the distance between sample points and the plane, they are decrypted with the private key after aggregation to ensure that the coordinator C cannot reverse-infer the local data of each party. Compared with the centralized filling method, the biggest difference is that in the centralized method, all-dimensional samples can be used to enter the tree structure for searching, while in the federated structure, each party has global sample fragments and cannot directly enter the tree structure for searching. Therefore, other tree search methods are needed; the method of selecting and positioning the ID is used to find the neighboring samples. When each participating party is looking for neighboring samples to fill in the missing values, it traverses each tree in the set, locates the leaf node where the missing sample ID is located under the tree according to the missing sample ID, and obtains all the sample ID sets included in the node. The union of the sample ID sets of all the trees is taken to obtain all the neighboring sample ID sets. Next, the IDs with missing feature values and the query ID are removed, and the mean value of the feature values of the remaining samples is calculated for filling.
[0058] In one embodiment, an embodiment of the present invention provides a method for filling missing values in vertical federated learning. Considering the original distribution of the data, that is, considering the correlation between the data, the missing values of the data are predicted, and the accuracy of the prediction result is improved. The method for filling missing values in vertical federated learning according to the embodiment of the present invention can be executed by a device for filling missing values in vertical federated learning. The device for filling missing values in vertical federated learning can be implemented by software and / or hardware.
[0059] As Figure 1 shown, the method for filling missing values in vertical federated learning according to the embodiment of the present invention specifically includes the following steps:
[0060] S110. Control the coordinator to generate a public key and a private key for homomorphic encryption, and send the public key to the participating party corresponding to the coordinator;
[0061] S120. Control the participating party to select an initial sample point index from the sample index set, and determine a target sample index according to the initial sample point index;
[0062] S130. Control the participating party to calculate the intercept and the intermediate distance matrix of the participating party through the target sample point, and encrypt the intermediate distance matrix with the public key to obtain an encrypted distance matrix;
[0063] S140. Aggregate the intercept and the encrypted distance matrix respectively by the collaborating party to obtain a global intercept corresponding to the intercept and an aggregated distance matrix corresponding to the encrypted distance matrix, and decrypt the aggregated distance matrix with the private key to obtain a target distance matrix;
[0064] S150. Divide the sample set where the target sample point is located according to the global intercept and the target distance matrix to obtain a left leaf node index set and a right leaf node index set;
[0065] Exemplarily, aggregate the intercepts to obtain a global intercept b, aggregate the encrypted intermediate distance matrices of all parties, decrypt the aggregated distance matrix with the private key, according to the formula:
[0066]
[0067] ID right = ID - ID left
[0068] Obtain the calculation results of each sample to divide the current sample set (the sample set where the target sample point is located), and obtain a left leaf node index set ID left and a right leaf node index set ID right ;
[0069] S160. Control the participating party to construct an index tree according to the left leaf node index set and the right leaf node index set, and traverse the missing positions at each leaf node of the index tree to obtain a target index set of the missing positions at the leaf nodes of the index tree;
[0070] In a possible implementation manner, the step of traversing the missing positions at each leaf node of the index tree to obtain a target index set of the missing positions at the leaf nodes of the index tree includes:
[0071] Take the positions where the mask matrix takes the value of 1 as the missing positions, traverse the missing positions at the leaf nodes of each index tree to obtain at least two initial index sets;
[0072] Obtain the union of at least two initial index sets to obtain the target index set.
[0073] Exemplarily, traverse the positions where the traversal mask matrix takes the value of 1, i.e., M(r, o) = 1, and obtain the sample index r corresponding to the missing position; for each r, traverse each tree and obtain the index set ID of the leaf node where the index r is located t ; the index sets ID of the leaf nodes of all trees t Take the union to obtain
[0074] S170. Determine the near-neighbor index set according to the target index set, and predict the data missing value at the missing position according to the near-neighbor index set.
[0075] In a possible implementation manner, the step of determining the near-neighbor index set according to the target index set includes:
[0076] Obtain the missing sample index and the current index where the target feature value is missing according to the target index set;
[0077] Integrate the indexes other than the missing sample index and the current index to obtain the near-neighbor index set.
[0078] Exemplarily, according to ID r , remove the sample index and the current index r where the o-th feature value (target feature value) in the near-neighbor index set is missing, and use the samples corresponding to the remaining indexes in the near-neighbor index set as near-neighbor samples.
[0079] In a possible implementation manner, the step of predicting the data missing value at the missing position according to the near-neighbor index set includes:
[0080] Determine near-neighbor samples according to the near-neighbor index set, and obtain the target feature values of a target number of near-neighbor samples to obtain a target feature value set;
[0081] Calculate the mean value of the target feature value set, and use the mean value as the data missing value at the missing position.
[0082] Exemplarily, obtain the mean value of the o-th feature value X obs (target feature value) of the remaining K (target number) near-neighbor samples
[0083]
[0084] Use the obtained mean value as the estimator of the missing value (data missing value).
[0085] Considering the original distribution of the data, that is, considering the correlation existing between the data, predicts the data missing value, improving the accuracy of the prediction result.
[0086] In a possible implementation manner, the step of controlling the participating party to select an initial sample point index from the sample index set and determine a target sample point according to the initial sample point index includes:
[0087] Controlling the participating party to select at least two initial sample point indexes from the sample index set according to a preset rule, and determining initial sample points corresponding to the initial sample point indexes in the sample set corresponding to the sample index set;
[0088] Selecting a random sample from the sample set, and calculating the distance between the initial sample and the random sample;
[0089] Updating the initial sample point according to the distance and a preset update rule until the initial sample point is updated a preset number of times, and taking the initial sample point after being updated the preset number of times as the target sample point.
[0090] Exemplarily, select two initial sample indexes, denoted as i and j, satisfying i ∼ U(1, n), j ∼ U(1, n - 1);
[0091] Define sample points p = x i and q = x j , and initialize the counters ic = 1 and jc = 1.
[0092] Perform the sample point position optimization process. Specifically, repeat the following steps L times (preset number of times), where L is the number of iteration steps:
[0093] Randomly select a sample x k , satisfying
[0094] k ∼ U(1, n)
[0095] Calculate the distances from the sample points (the initial sample points) p and q to the sample (random sample) x k :
[0096]
[0097] Update the sample points (the initial sample points) p and q according to the distance magnitudes:
[0098]
[0099] Finally, output the updated sample points p and q as the selected two sample indexes (the indexes of the target sample point, and determine the target sample point according to the indexes of the target sample point).
[0100] In a possible implementation manner, before the step of controlling the participating party to construct an index tree according to the left leaf node index set and the right leaf node index set, it further includes:
[0101] Obtain the first quantity information of the indexes in the left leaf node index set and the second quantity information of the indexes in the right leaf node index set;
[0102] If the first quantity information and / or the second quantity information is greater than the threshold, continue to execute the step of controlling the participating party to select an initial sample point index from the sample index set and determining the target sample index according to the initial sample point index.
[0103] In a possible implementation manner, before the step of traversing the missing positions at each leaf node of the index tree to obtain the target index set of the leaf nodes of the index tree where the missing positions are located, it further includes:
[0104] Obtain the first quantity information of the constructed index tree and the second quantity information (T) of the initial index tree corresponding to the sample set;
[0105] When the first quantity information is less than the second quantity information, continue to execute the step of controlling the participating party to select an initial sample point index from the sample index set and determining the target sample index according to the initial sample point index.
[0106] Exemplarily, the index tree is an Annoy tree, the sample set X includes n samples X = {x 1 ; x 2 ;... x i ;... x n}(|X| = n, x i ∈ R m ), the sample index set is ID = {1, 2,... n}, and the sample set corresponds to T initial index trees.
[0107] In a possible implementation manner, such as Figures 2 - 3As shown in the figure, the architecture of a method for filling missing values in vertical federated learning includes a coordinator C, a participant A, and a participant B. The coordinator C generates a public key and a private key, and sends the public key to the participant A and the participant B respectively. The participant A and the participant B select sample points (target sample points) according to the sample index set, calculate the normal vector and intercept according to the selected sample points, obtain the intermediate distance according to the normal vector and intercept, encrypt the intermediate distance using the public key, aggregate the encrypted intermediate distance and send it to the coordinator C. The coordinator C decrypts the aggregated encrypted intermediate distance, and divides it into a left leaf node index set and a right leaf node index set. The coordinator C returns the left leaf node index set and the right leaf node index set to the participant A and the participant B. The participant A and the participant B construct an index tree according to the left leaf node index set and the right leaf node index set, and traverse the missing positions at each leaf node of the index tree to obtain the target index set of the missing positions at the leaf nodes of the index tree; determine the neighbor index set according to the target index set, and predict the missing value of the missing position according to the neighbor index set.
[0108] In terms of generating data missing, starting from 1000 samples, the number of samples is increased to 15000 with a step of 2000. For each number of samples, an incomplete data set with missing rates of 1%, 5%, and 10% is randomly generated using the MCAR mechanism. In terms of data splitting, the feature x 1 to x 6 is divided into the data set of Party A, and the rest is divided into Party B. The number of neighbors is set to 20. By comparing the RMSE values, the effectiveness of the filling results obtained by the method for filling missing values in vertical federated learning described in this application is tested. As Figure 4 shown, the RMSE value of the filling results obtained by the method for filling missing values in vertical federated learning described in this application is significantly lower than the RMSE values of the other three methods (the maximum value method, the minimum value method, and the mean filling method), and the curve is smoother. At different missing rates, the maximum value method and the minimum value method have the highest RMSE values. Compared with the mean filling method, our method has a lower RMSE value at different missing rates.
[0109] In another embodiment of the present invention, a device for filling missing values in vertical federated learning is provided. Figure 5 As shown in the structural schematic diagram of a device for filling missing values in vertical federated learning provided by an embodiment of the present invention, the device for filling missing values in vertical federated learning provided by an embodiment of the present invention can execute the method for filling missing values in vertical federated learning provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. The device includes:
[0110] The key control module 201 is used to control the collaborating party to generate the public key and private key for homomorphic encryption, and send the public key to the participating party corresponding to the collaborating party;
[0111] The first control module 202 is used to control the participating party to select an initial sample point index from the sample index set, and determine the target sample index according to the initial sample point index;
[0112] The first control module 202 is further used to control the participating party to calculate the intercept and the intermediate distance matrix of the participating party through the target sample point, and encrypt the intermediate distance matrix through the public key to obtain the encrypted distance matrix;
[0113] The second control module 203 is used to aggregate the intercept and the encrypted distance matrix respectively through the collaborating party to obtain the global intercept corresponding to the intercept and the aggregated distance matrix corresponding to the encrypted distance matrix, and decrypt the aggregated distance matrix through the private key to obtain the target distance matrix;
[0114] The partitioning module 204 is used to partition the sample set where the target sample point is located according to the global intercept and the target distance matrix to obtain the left leaf node index set and the right leaf node index set;
[0115] The construction module 205 is used to control the participating party to construct an index tree according to the left leaf node index set and the right leaf node index set, and traverse the missing positions at each leaf node of the index tree to obtain the target index set of the missing positions at the leaf nodes of the index tree;
[0116] The estimation module 206 is used to determine the nearest neighbor index set according to the target index set, and predict the data missing value of the missing position according to the nearest neighbor index set.
[0117] It should be noted that the various modules included in the above device are only divided according to the functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional modules are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present invention.
[0118] In another embodiment of the present invention, an electronic device is further provided. Figure 6 The block diagram of an exemplary electronic device 50 suitable for implementing the embodiments of the present invention is shown. Figure 6 The shown electronic device 50 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.
[0119] As Figure 6As shown, the electronic device 50 is presented in the form of a general-purpose computing device. The components of the electronic device 50 may include, but are not limited to: one or more processors or processing units 501, a system memory 502, and a bus 503 that connects different system components (including the system memory 502 and the processing unit 501).
[0120] The bus 503 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the multiple bus structures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0121] The electronic device 50 typically includes a variety of computer system-readable media. These media can be any available media that can be accessed by the electronic device 50, including volatile and non-volatile media, removable and non-removable media.
[0122] The system memory 502 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 504 and / or cache memory 505. The electronic device 50 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 506 can be used for reading and writing on non-removable, non-volatile magnetic media ( Figure 6 not shown, commonly referred to as a "hard disk drive"). Although Figure 6 not shown in the figure, a disk drive for reading and writing on a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing on a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to the bus 503 through one or more data media interfaces. The memory 502 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments of the present invention.
[0123] A program / utility 508 having a set (at least one) of program modules 507 can be stored, for example, in the memory 502. Such program modules 507 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 507 generally execute the functions and / or methods in the embodiments described in the present invention.
[0124] The electronic device 50 can also communicate with one or more external devices 509 (such as a keyboard, a pointing device, a display 510, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 50, and / or communicate with any device that enables the electronic device 50 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 511. Moreover, the electronic device 50 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 512. As shown in the figure, the network adapter 512 communicates with other modules of the electronic device 50 through a bus 503. It should be understood that although Figure 6 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 50, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0125] The processing unit 501 executes various functional applications and data processing by running programs stored in the system memory 502, such as implementing the method for filling missing data values for vertical federated learning provided in the embodiments of the present invention.
[0126] In another embodiment of the present invention, there is also provided a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute a method for filling missing data values for vertical federated learning when executed by a computer processor. The method includes:
[0127] Controlling the collaborating party to generate a public key and a private key for homomorphic encryption, and sending the public key to the participating party corresponding to the collaborating party;
[0128] Controlling the participating party to select an initial sample point index from a sample index set, and determining a target sample index according to the initial sample point index;
[0129] Controlling the participating party to calculate the intercept and the intermediate distance matrix of the participating party through the target sample points, and encrypting the intermediate distance matrix through the public key to obtain an encrypted distance matrix;
[0130] Respectively aggregating the intercept and the encrypted distance matrix by the collaborating party to obtain a global intercept corresponding to the intercept and an aggregated distance matrix corresponding to the encrypted distance matrix, and decrypting the aggregated distance matrix through the private key to obtain a target distance matrix;
[0131] Partition the sample set where the target sample points are located according to the global intercept and the target distance matrix to obtain a left leaf node index set and a right leaf node index set;
[0132] Control the participating parties to construct an index tree according to the left leaf node index set and the right leaf node index set, and traverse the missing positions at each leaf node of the index tree to obtain a target index set of the missing positions at the leaf nodes of the index tree;
[0133] Determine a neighbor index set according to the target index set, and predict the data missing value of the missing position according to the neighbor index set.
[0134] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0135] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0136] The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0137] Computer program code for performing the operations of the embodiments of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0138] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made in accordance with the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A method for completing missing values in vertical federated learning, characterized in that: include: Controlling the collaborative party to generate a homomorphically encrypted public key and a private key, and sending the public key to a participant corresponding to the collaborative party; Controlling the participant to select an initial sample point index from a sample index set, and determining a target sample index according to the initial sample point index; Controlling the participant to calculate the intercept and the intermediate distance matrix of the participant through the target sample point, and encrypting the intermediate distance matrix through the public key to obtain an encrypted distance matrix; The intercept and the encrypted distance matrix are aggregated by the collaborative party to obtain a global intercept corresponding to the intercept and an aggregated distance matrix corresponding to the encrypted distance matrix, and the aggregated distance matrix is decrypted by the private key to obtain a target distance matrix; Divide the sample set where the target sample point is located according to the global intercept and the target distance matrix to obtain a left leaf node index set and a right leaf node index set; Control the participant to construct an index tree according to the left leaf node index set and the right leaf node index set, and traverse the missing position at each leaf node of the index tree to obtain a target index set of leaf nodes at the missing position on the index tree; A neighbor index set is determined according to the target index set, and the data missing value of the missing position is predicted according to the neighbor index set.
2. The method according to claim 1, characterized in that The step of controlling the participant to select an initial sample point index from a sample index set and determining a target sample point according to the initial sample point index includes: Controlling the participant to select at least two initial sample point indexes from a sample index set according to a preset rule, and determining an initial sample point corresponding to the initial sample point index in a sample set corresponding to the sample index set; Selecting a random sample from the sample set, and calculating the distance between the initial sample and the random sample; The initial sample point is updated according to the distance and a preset update rule until the initial sample point is updated a preset number of times, and the initial sample point after being updated the preset number of times is used as the target sample point.
3. The method according to claim 1, characterized in that Before the step of controlling the participant to construct an index tree according to the left leaf node index set and the right leaf node index set, the step further includes: Obtain first quantity information of indexes in the left leaf node index set and second quantity information of indexes in the right leaf node index set; If the first quantity information and / or the second quantity information is greater than a threshold, the step of controlling the participant to select an initial sample point index from a sample index set and determining a target sample index according to the initial sample point index is continued.
4. The method according to claim 2, characterized in that: Before the step of traversing the missing position of each leaf node of the index tree to obtain the target index set of the leaf nodes of the missing position on the index tree, the method further includes: Obtaining first numerical information of the constructed index tree and second numerical information of the initial index tree corresponding to the sample set; When the first number information is less than the second number information, the step of controlling the participant to select an initial sample point index from a sample index set and determining a target sample index according to the initial sample point index is continued.
5. The method according to claim 1, characterized in that The step of determining a neighbor index set according to the target index set comprises: Obtaining missing sample indexes and current indexes where target feature values are missing according to the target index set; The indexes except the missing sample index and the current index are integrated to obtain the neighbor index set.
6. The method according to claim 5, characterized in that The step of predicting the missing value of the data at the missing position according to the neighbor index set comprises: Determine neighbor samples according to the neighbor index set, and obtain target feature values of a target number of neighbor samples to obtain a target feature value set; The mean of the target feature value set is calculated, and the mean is used as the data missing value of the missing position.
7. The method according to claim 1, characterized in that The step of traversing the missing position at each leaf node of the index tree to obtain a target index set of leaf nodes at the missing position on the index tree comprises: The position where the mask matrix takes a value of 1 is used as the missing position, and the missing position is traversed at the leaf nodes of each index tree to obtain at least two initial index sets; A union of at least two initial index sets is obtained to obtain the target index set.
8. A device for completing missing values of data for vertical federated learning, characterized in that: The device comprises: A key control module, used to control the collaborative party to generate a public key and a private key for homomorphic encryption, and send the public key to the participating party corresponding to the collaborative party; A first control module, used to control the participant to select an initial sample point index from a sample index set, and determine a target sample index according to the initial sample point index; The first control module is further used to control the participant to calculate the intercept and the intermediate distance matrix of the participant through the target sample point, and encrypt the intermediate distance matrix through the public key to obtain an encrypted distance matrix; A second control module is used to aggregate the intercept and the encrypted distance matrix respectively through the collaborative party to obtain a global intercept corresponding to the intercept and an aggregated distance matrix corresponding to the encrypted distance matrix, and decrypt the aggregated distance matrix through the private key to obtain a target distance matrix; A partitioning module, used for partitioning the sample set where the target sample point is located according to the global intercept and the target distance matrix to obtain a left leaf node index set and a right leaf node index set; A construction module, used to control the participant to construct an index tree according to the left leaf node index set and the right leaf node index set, and traverse the missing position in each leaf node of the index tree to obtain a target index set of leaf nodes of the missing position on the index tree; An estimation module is used to determine a neighbor index set based on the target index set, and predict the data missing value of the missing position based on the neighbor index set.
9. An electronic device, characterized in that: The electronic device comprises: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data missing value completion method for vertical federated learning as described in any one of claims 1-7.
10. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions, when executed by a computer processor, are used to execute the data missing value completion method for longitudinal federated learning as described in any one of claims 1-7.