Privacy protection vertical federal learning method and system based on additive integration
Through the additively integrated vertical federated learning (AE-VFL) method, the gradient lifting mechanism is used for model training, which solves the problems of low computational efficiency and insufficient security in vertical federated learning, and achieves higher security and model performance, which is better than centralized training.
Patent Information
- Application Number
- CN202510440924.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
The existing vertical federated learning (VFL) methods have shortcomings in privacy protection and computing efficiency, high computing and communication overhead, insufficient security, vulnerable to side channel attacks, and low model training efficiency.
The additive integrated vertical federated learning (AE-VFL) method is adopted, and the server-side top-level model and participant underlying model are initialized, and the gradient enhancement mechanism is used for training in sequence. The participant trains the underlying model and then transmits it to the top-level model. The server calculates the loss and propagates the gradient. The next participant trains the underlying and top-level models through residuals to achieve model complementarity.
It significantly improves the security and performance of the model, reduces the success rate of tag inference attacks, improves the overall performance of the model, and is better than centralized training, especially on the CIFAR-10 dataset, F-score is increased by 19.26%, and its security is enhanced.
Smart Images

Figure CN120373422A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of information and data security, and particularly to a privacy-preserving vertical federated learning method and system based on additive integration. Background Art
[0002] The statements in this section merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] With the accelerated advancement of global digital transformation, the security of user data and privacy protection have become a common consensus in the international community. Although existing protection regulations safeguard user privacy, they have also given rise to the so-called "data silo" phenomenon, making it difficult for different institutions to share data, which in turn affects the effects of joint modeling and machine learning. To address this dilemma, federated learning, as an innovative distributed machine learning technology, has emerged. Federated learning was first proposed by the Google research team in 2016. Its core idea is to allow each participating party to conduct collaborative training through encrypted model parameters without sharing the original data, thereby protecting data privacy while achieving joint modeling of data.
[0004] Federated learning can be divided into two main paradigms: horizontal federated learning (HFL) and vertical federated learning (VFL) according to the characteristics of the participating parties' data and the different overlapping situations of the sample spaces. Horizontal federated learning is applicable to scenarios where each participating party has the same feature space but different sample spaces, such as multiple medical institutions in different regions sharing data of different patients for joint modeling. Vertical federated learning, on the other hand, is applicable to the situation where the participating parties share sample identifiers but each has a different feature space, such as a financial institution and an e-commerce platform sharing user IDs to improve the modeling ability through different feature data.
[0005] A typical VFL application scenario is the auto insurance field. For example, an insurance company jointly models with cross-industry institutions such as banks and tax departments, and improves the accuracy of customer risk assessment by sharing different feature data (such as credit records, tax information, etc.).
[0006] Currently, horizontal federated learning (HFL) has developed into a relatively mature research system. However, existing vertical federated learning (VFL) methods still have the following problems:
[0007] 1) In order to ensure privacy and security, existing solutions generally introduce complex encryption and security protocols, which significantly increase the computational and communication overhead, resulting in low model training efficiency and insufficient real-time performance; the encryption operations and iterative secure computing steps often restrict the response speed and scalability of the overall system, and have not achieved the ideal application effect.
[0008] 2) Although various encryption technologies are adopted in the existing solutions, there may still be security risks in key management, protocol design, and ciphertext operation processes, increasing the risk of sensitive information leakage. The protection measures against abnormal behaviors and adversarial attacks are not yet perfect, and some solutions may face side-channel attacks or other new security threats in practical applications. Summary of the Invention
[0009] To solve the above problems, the present disclosure proposes a privacy-preserving vertical federated learning method and system based on additive integration. By using additive integration vertical federated learning (AE-VFL), the local models of each participant are regarded as sub-models, and sequential training is carried out guided by the loss function gradient generated by the previous model, driving each sub-model to compensate for the limitations of the predecessor model, and optimizing the efficiency and security of VFL by integrating the gradient boosting mechanism.
[0010] According to some embodiments, the present disclosure adopts the following technical solutions:
[0011] A privacy-preserving vertical federated learning method based on additive integration, comprising:
[0012] Initializing the top-level model of the server side and the bottom-level models of each participant, and constructing an additive integration vertical federated learning framework;
[0013] In the additive integration vertical federated learning framework, the top-level model of the server side is assigned to each participant, and each participant performs sequential training of the bottom-level model through the gradient boosting mechanism;
[0014] Among them, the sequential training of the bottom-level model through the gradient boosting mechanism includes: after the participant trains the bottom-level model using local private data, uploading the output of the bottom model to its corresponding top model, the server calculates the loss of each top model, and propagates the gradient to the bottom model of the next participant. The next participant serializes to train the bottom model and the top model of the next participant by considering the residuals of the previous participant. Through ensemble learning and gradient boosting, a complementary effect between participant models is formed to achieve additive integration vertical federated learning.
[0015] According to some embodiments, the present disclosure adopts the following technical solutions:
[0016] A privacy-preserving vertical federated learning system based on additive integration, comprising:
[0017] An initialization module, configured to initialize the top-level model of the server side and the bottom-level models of each participant, and construct an additive integration vertical federated learning framework;
[0018] A federated learning module is used in an additive integrated vertical federated learning framework, where the top-level model on the server side is assigned to each participant, and each participant serially trains the bottom-level model through a gradient boosting mechanism.
[0019] Among them, the serial training of the bottom-level model through the gradient boosting mechanism includes: after the participant trains the bottom-level model using local private data, the output of the bottom model is uploaded to its corresponding top model, the server calculates the loss of each top model, and propagates the gradient to the bottom model of the next participant. The next participant serializes the training of the bottom model and the top model of the next participant by considering the residuals of the previous participant. Through ensemble learning and gradient boosting, a complementary effect between participant models is formed to achieve additive integrated vertical federated learning.
[0020] According to some embodiments, the present disclosure adopts the following technical solutions:
[0021] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the privacy-preserving vertical federated learning method based on additive integration.
[0022] According to some embodiments, the present disclosure adopts the following technical solutions:
[0023] A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the privacy-preserving vertical federated learning method based on additive integration is implemented.
[0024] According to some embodiments, the present disclosure adopts the following technical solutions:
[0025] An electronic device includes: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device implements the privacy-preserving vertical federated learning method based on additive integration.
[0026] Compared with the prior art, the beneficial effects of the present disclosure are:
[0027] In the privacy-preserving vertical federated learning method based on additive integration of the present disclosure, each participant predicts the residuals of the previous underlying model in additive integration vertical federated learning (AE-VFL), rather than directly predicting the target labels. This design makes it more difficult for attackers to infer the target labels based on the gradient information transmitted from the server. The success rate of label inference attacks in traditional VFL on the Adult dataset is as high as 99.62%. This is mainly because the Adult dataset is a binary classification problem with fewer label types, making attacks more likely to succeed. In contrast, AE-VFL significantly reduces the success rate of label inference attacks to 46.44%, thus effectively enhancing security.
[0028] In the privacy-preserving vertical federated learning method based on additive integration of the present disclosure, AE-VFL even outperforms centralized training in some tasks. The fundamental reason for this advantage is that centralized training trains the model by simply aggregating data, while AE-VFL, although without centralized data, combines the ideas of ensemble learning and gradient boosting, forming a complementary effect among participant models and maximizing the potential of distributed participants. AE-VFL significantly improves the overall performance of the model by leveraging the knowledge of distributed participants while ensuring data privacy. On the CIFAR-10 dataset, the F-score of AE-VFL is 19.26% higher than that of centralized training (74.46% vs. 55.20%).
[0029] In the privacy-preserving vertical federated learning method based on additive integration of the present disclosure, AE-VFL is significantly superior to traditional VFL in mitigating label inference attacks based on gradient signs. Especially on the CIFAR-10 dataset, AE-VFL reduces the success rate of label inference attacks of traditional VFL from 50.22% to 16.61%. This result indicates that AE-VFL not only outperforms traditional VFL in terms of model performance but also shows significant advantages in terms of security. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The specification drawings forming a part of the present disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure.
[0031] Figure 1 It is the overall framework diagram of the privacy-preserving vertical federated learning method based on additive integration according to the embodiment of the present disclosure;
[0032] Figure 2 It is the structural framework diagram of traditional vertical federated learning. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The present disclosure will be further described below in conjunction with the drawings and embodiments.
[0034] It should be noted that the following detailed description is illustrative and aims to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure pertains.
[0035] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0036] Embodiment 1
[0037] In one embodiment of the present disclosure, a privacy-preserving vertical federated learning method based on additive integration is provided, including the following steps:
[0038] Step 1: Initialize the top-level model on the server side and the bottom-level models of each participant, and construct an additive integration vertical federated learning framework;
[0039] Step 2: In the additive integration vertical federated learning framework, the top-level model on the server side is assigned to each participant, and each participant performs sequential training of the bottom-level model through the gradient boosting mechanism;
[0040] Among them, the sequential training of the bottom-level model through the gradient boosting mechanism includes: after the participant trains the bottom-level model using local private data, the output of the bottom-level model is uploaded to its corresponding top-level model, the server calculates the loss of each top-level model, and propagates the gradient to the bottom-level model of the next participant. The next participant serializes the training of the bottom-level model and the top-level model of the next participant by considering the residuals of the previous participants. Through ensemble learning and gradient boosting, a complementary effect between the participant models is formed to achieve additive integration vertical federated learning.
[0041] As an embodiment, the specific implementation process of the privacy-preserving vertical federated learning method based on additive integration of the present disclosure is as follows:
[0042] Step 1: Initialize the top-level model on the server side and the bottom-level models of each participant, and construct an additive integration vertical federated learning framework;
[0043] Specifically, in the additive integrated vertical federated learning framework, the server assigns an independent top-level model to each participant, and the participants train in sequence. The training objective of subsequent participants becomes the pseudo-residuals of the previous participants. By training the bottom-level model for each participant, the training objectives of each participant no longer share the same label.
[0044] Problem Definition:
[0045] In a typical vertical federated learning (VFL) system, T data owners jointly train a machine learning model while keeping their respective private data D 1 , D 2 ,..., D T locally stored. Let D = D 1 ∪ D 2 ∪... ∪ D T denote their data sets. Further define the feature space as X, the label space as Y, and the sample ID space as I. Thus, D t = (I t , X t ) is the data set owned by party t. In VFL, a trusted third-party server holds the label Y, while each party only has partial features X t . VFL can be defined as follows:
[0046]
[0047] Different from horizontal federated learning (HFL) where data is partitioned by sample, vertical federated learning (VFL) assumes that data is partitioned by feature. The VFL system assumes there are N overlapping samples to train a joint machine learning model, where x i is the feature vector and y i is the label information. Each feature vector is assigned to T participants, where where d t is the feature dimension of party t. The goal of the VFL system is to collaboratively train a joint machine learning model using the data set D while maintaining the privacy and security of the local data and models.
[0048] The AE-VFL - additive integrated vertical learning method of the present disclosure aims to improve the performance and security of VFL. First, the AE-VFL framework is outlined.
[0049] The deployment process of AE-VFL mainly includes the following steps:
[0050] (1) Use the private data D t of party t to train the bottom-level model G t ;
[0051] (2) Upload the output of the underlying model G t to the corresponding top-level model H located on the server t ;
[0052] (3) Calculate the loss of each top-level model H t and propagate the gradient to the underlying model G t ;
[0053] (4) Serialize the training of the underlying model G t+1 and the top-level model H t+1 of the next participant by considering the residuals of the previous participant
[0054] This process is different from traditional VFL, in which the server side only contains one top-level model H, responsible for processing the outputs of all participants' underlying models. However, in AE-VFL, each participant has an independent top-level model H t on the server side, and the participants are trained sequentially. In addition, the training objective of the next participant is the pseudo-residuals of the previous participant. The advantages of the improved method of the present disclosure are as follows: First, by training a weak model for each participant, multiple weak models can significantly enhance the overall performance of vertical federated learning. In addition, each participant no longer shares the same label with the training objective, which greatly reduces the risk of label inference attacks in vertical federated learning
[0055] Step 2: In the additive ensemble longitudinal federated learning framework, the top-level model on the server side is assigned to each participant, and each participant performs sequential training of the underlying model through the gradient boosting mechanism
[0056] The AE-VFL framework follows the core idea of Boosting to improve the overall performance by sequentially generating models, and each participant is used to complement the deficiencies of the previous participants. Specifically, AE-VFL constructs an underlying model for each participant, and each underlying model is used to make up for the defects of the underlying models of the previous participants. The goal of AE-VFL is to train a stronger model by integrating the underlying models of multiple participants. Therefore, the proposed AE-VFL can effectively improve the overall performance of the model. On the other hand, each participant has an independent training objective (i.e., the residuals of the previous model rather than the original labels on the server side). Therefore, it is more difficult for malicious participants to infer the target labels through gradient information. The specific steps are as follows
[0057] (1) Use the private data D t to train the underlying model G t on participant t
[0058] (2) The underlying model G tThe output is uploaded to the corresponding top-level model H on the server side t ;
[0059] (3) Calculate the loss of each top-level model H t and backpropagate the gradients to the bottom-level model G t ;
[0060] (4) Serialize the training of the bottom-level model G t+1 and top-level model H t+1 of the next participant by considering the residuals of the previous participants
[0061] The local outputs from the participants are passed as inputs to their corresponding top-level models. In the additive-ensemble vertical federated learning framework, each participant has its own independent top-level model, which is implemented by another DNN model. Through the top-level model, the final output prediction of each participant's data is obtained. The final output of the participant for the multi-class task is a probability, and the final prediction represented takes into account the outputs of the previous bottom-level models according to the gradient boosting mechanism
[0062] Furthermore, the server performs backpropagation and calculates two gradients for the model parameters of the top-level model and the forward outputs of the bottom-level models from each participant. Using the gradients of the top-level model, the label owner calculates the average gradient of each private data sample and updates its top-level model
[0063] Each participant calculates the gradients of its bottom-level model and updates its bottom-level model. The process is repeated until convergence. Each participant calculates the gradients of the parameters of its bottom-level model based on the local private data and gradients of the forward outputs from the label owner, and then updates its bottom-level model
[0064] Specifically, the participants using the local private data to train the bottom-level model includes: dividing the public dataset D into different participants according to features. Among the data samples aligned among all participants, each participant will complete the forward propagation process using the local private data based on its bottom model, and obtain the latent representation of each private data sample using the deep neural network model as the bottom-level model. The specific implementation process includes
[0065] Step ①: Private set intersection
[0066] Private set intersection (PSI) is a secure multi-party computation cryptographic technique that allows two participants to hold sets and compare the encrypted versions of these sets to calculate the intersection. Therefore, PSI can be used to privately identify the intersection of the training samples of all participants in AE-VFL. Using PSI, no information other than the intersection elements will be leaked between participants. The present disclosure divides the public dataset D among different participants according to features
[0067] Step ②: Forward propagation of the underlying model.
[0068] After determining the alignment of data samples among all participants, each participant will use its local data D t to complete the forward propagation process based on its underlying (local) model. A deep neural network (DNN) model is used as the local underlying model G t to obtain the latent representation of each private data sample. The formula is as follows:
[0069]
[0070] where, represents the i-th sample in the t-th participant. is the hidden representation of this sample, and θ t represents the parameters of the underlying model of the t-th participant.
[0071] Specifically, the local architecture of the underlying model G t is not strictly fixed. The underlying model G t learns the representation of local data, and its structure can be dynamically adjusted according to specific scenarios.
[0072] Step ③: Forward output transmission. The forward output contains the intermediate results of the local neural network, which converts the original attributes into features.
[0073] Specifically, the local output from participant t t is passed to its corresponding top-level model H
[0074] as input.
[0075] Different from traditional VFL, in traditional VFL, there is only one shared top-level model. In AE-VFL, each participant has its own independent top-level model H t , which is implemented by another DNN model. Through the top-level model H t , the final output prediction of each participant's data can be obtained. The specific formula is as follows:
[0076]
[0077] For classification tasks with the number of classes C > 2, the label is defined by a 1-of-C vector, where y i,c = 1 indicates that the instance belongs to class c, otherwise y i,c = 0. Here, the cross-entropy loss l t calculated for participant t is as follows:
[0078]
[0079] wherein, is the probability that a given instance in participant t belongs to category c. More specifically, can be obtained by the following formula:
[0080]
[0081] Similarly, the above formula is valid for all participants, which provides the possibility of applying gradient boosting to other base learners. The final output of participant T in a multi-class task is the probability
[0082] Furthermore, gradient boosting is an important strategy in ensemble learning, which is extended additively based on different fitting objectives. The most significant feature of gradient boosting is that the t-th sub-model fits the residual of the ensemble model composed of the first t - 1 sub-models. The gradient boosting algorithm improves the learning accuracy by synthesizing the prediction results of multiple sub-models.
[0083] For a given training data set D with N samples, that is The goal of a machine learning algorithm is to find an approximate function F(x) that maps an instance x to its output y. Generally, this learning process can be regarded as an optimization problem, in which the goal of the model is to minimize the expected value of a given loss function l, that is, E[l(Y,F(X))], where X represents the input set and Y represents the label set. The expected value of the loss can be approximated using data-based estimation: In the specific scenario of gradient boosting, the model is created by additive expansion, and the gradient boosting calculation process is:
[0084] F t (x i ) = F t-1 (x i ) + ρ t g t (x i )
[0085] wherein, ρ t represents the weight of the t-th sub-model g t . When a new sub-model g t is created, it does not modify the model contained in F t -1 (x i ). First, the additive expansion is evaluated by a constant:
[0086]
[0087] wherein, θ 1Representing the parameters of the first sub-model g 1 Subsequently, the goal of creating the next sub-model is to minimize:
[0088]
[0089] Each sub-model g t is trained to learn the gradient vector of the loss function based on the data. To this end, each sub-model g t is trained on a new data set where r i t is the pseudo-residual term, representing the negative gradient of the loss function at F t-1 (x i ).
[0090]
[0091] The output value of the t-th sub-model g t is expected to be close to the pseudo-residual r of the given data point i t . This is consistent with the gradient of the loss function l at F t-1 (x i ).
[0092] According to the gradient boosting calculation process, the final predicted value is denoted as taking into account the outputs of the previous sub-models before t.
[0093]
[0094] For the regression task, the mean squared error (MSE) loss function is used for the regression task. Specifically, the formula is as follows:
[0095]
[0096] According to the gradient boosting algorithm calculation process, after adding the weak model of participant t (i.e., the combination of G and H t and H t ), the final output is denoted as
[0097]
[0098] Similar to the classification task, the final prediction takes into account the outputs of all sub-models before t.
[0099] Step ⑤: Top model backpropagation. The label owner, i.e., the server, performs backpropagation and calculates two gradients:
[0100] 1) Top model Ht The gradient of the model parameters;
[0101] 2) Each participant's bottom model G t The gradient of the forward output of .
[0102] Using the gradients of the top model, the label owner can calculate the average gradient for each batch of samples and update its model:
[0103]
[0104] In all After being calculated, the server calculates the training loss according to the task type. Then, the server calculates the gradient of its global module And update its top model:
[0105] ψ t =ψ t -γ h *ΔΨ t l t
[0106] Among them, γ h represents the learning rate of the top model.
[0107] Step ⑥: Backward output transmission. The gradients of the forward output are sent back to each participant. It can be noted that the required communication cost (number of bits transmitted) is usually much smaller than the communication cost in step ③ because they are gradients rather than intermediate outputs.
[0108] Step 7: Back propagation of the underlying model. Next, the server calculates the gradient of each participant and pass them back. Finally, each participant t calculates its local underlying model θ according to the following formula t The gradient of the local model is calculated based on the local data and the forward output gradient from the label owner, and then its underlying model is updated. This process is iterated until convergence. Each participant calculates the gradient of the local model parameters based on the local data and the forward output gradient from the label owner, and then updates its underlying model.
[0109]
[0110] Among them, γ g Represents the underlying model G t The learning rate.
[0111] Simulation experiment
[0112] 1. Dataset and evaluation metrics
[0113] The performance of AE-VFL is evaluated on two datasets dedicated to classification tasks and another two datasets for regression tasks. The statistics of the four datasets are listed in Table 1.
[0114] 1. The Adult dataset is used for the task of predicting whether an individual's income exceeds $50K per year, based on census data. To fairly evaluate the performance of the model, the training set is divided into a training set and a validation set at a ratio of 4:1. The final performance of the model is ultimately evaluated on an independent test set. The feature set of this dataset includes numerical features and categorical features, with a total of 107 features. To simulate the scenario of vertical federated learning, these 107 features are assigned to two different participants, with each participant having 50 features and 57 features respectively.
[0115] 2. The CIFAR-10 dataset contains 60,000 color images, each with a resolution of 32×32 pixels, distributed across 10 different classes. 50,000 training images are randomly divided into a training set (40,000 images) and a validation set (10,000 images), maintaining a 4:1 ratio. In the context of VFL, the features of the three channels in the CIFAR-10 dataset are assigned to two different institutions. In the VFL setting, each participant has data features from one or two channels respectively. This division method simulates the collaboration between participants while maintaining data privacy and security.
[0116] 3. The King County House Sales dataset (KCHouse) provides information on houses sold in King County between May 2014 and May 2015. To simulate the VFL scenario, the 13-dimensional features are randomly divided between two different institutions. One institution is assigned 6 features, while the other institution has 7 features. The dataset is further divided into a training set, a validation set, and a test set. Such a division ensures that each institution holds a subset of the overall feature space, enabling collaborative analysis while maintaining data privacy.
[0117] 4. The Song Popularity Prediction dataset (Song) is used for a regression problem, aiming to predict the popularity of songs. This dataset contains 14 features, including attributes such as energy, acousticness, instrumentalness, liveness, and danceability. The 14 features are randomly divided between two different participants, with each participant having 7 features from the dataset. Similarly, the dataset is divided into a training set and a validation set at a ratio of 4:1, and the remaining 3,767 samples are used as the test set.
[0118] Table 1 Statistical information of the four datasets
[0119]
[0120] 2. Evaluation metrics for classification tasks.
[0121] In the experiments of this disclosure, special attention is paid to the evaluation metrics of the following two classification tasks: Accuracy and F-score. Accuracy measures the proportion of correct predictions among all predictions. It represents the ability of the model to correctly classify instances, and a higher accuracy value indicates that the model is more accurate. On the other hand, the F-score is a metric that combines Precision and Recall into a single value. The F-score is obtained by calculating the harmonic mean of Precision and Recall. It provides a balanced assessment of the precision and recall performance of the model.
[0122] 3. Evaluation Metrics for Regression Tasks.
[0123] The evaluation of regression tasks focuses on two commonly used metrics: Root Mean Square Error (RMSE) and R 2 . RMSE measures the difference between the predicted value and the actual expected value. RMSE is obtained by calculating the square root of the Mean Square Error (MSE) or loss. R 2 , also known as the coefficient of determination, is a statistical measure in the regression model. It represents the proportion of the variance in the dependent variable that can be explained by the independent variable(s). R 2 values range from 0 to 1, where a higher R 2 value indicates a better fit of the model to the data.
[0124] The experiments were conducted on a Linux GPU server equipped with an Intel(R) Xeon(R) Silver 4214R CPU @ 2.40GHz, an Nvidia A100 GPU, and 80GB of RAM. The AE-VFL framework was developed using PyTorch 1.13.1, and the default number of participants in the experiments was set to 2.
[0125] All prediction functions added to the model are Multi-Layer Perceptrons (MLPs) with two hidden layers. The number of neurons in each hidden layer was set to 200. The hyperparameter λ for balancing the regularization term was set to 0.001. The boosting rate γh was initially set to 1 and was automatically adjusted during the correction step. For classification tasks, each prediction model was trained for 1000 epochs. The optimizer used was Adam, with a batch size of 64 and a learning rate of 0.00001. For regression tasks, the number of learning epochs for each task was set to 10000, the batch size was 64, the learning rate was 0.0001, and the optimizer used was also Adam.
[0126] To evaluate the effectiveness of the method proposed in this disclosure in classification and regression tasks, a comparative evaluation with baseline methods was conducted on four benchmark datasets. Table 2 shows the classification results on the Adult and CIFAR-10 datasets, and Table 3 shows the comparison results on the KCHouse and Song datasets, highlighting the performance comparison between our method and existing methods. The following notable observations can be made:
[0127] First, whether in classification tasks or regression tasks, AE-VFL consistently outperforms other methods in all evaluation metrics. These results clearly demonstrate that AE-VFL is highly efficient in effectively handling classification and regression tasks. For example, compared with the most similar traditional VFL method, AE-VFL improves by 18.74% and 33.09% respectively in terms of accuracy and F-score on the Adult dataset. Similarly, on the Song dataset, AE-VFL also shows significant improvement, with the RMSE reduced by 21.70, while the R 2 is increased by 0.98. The key factor leading to these improvements is that AE-VFL does not simply aggregate the outputs of multiple parties like traditional VFL. Instead, each party trains its own weak model, aiming to make up for the deficiencies of the previous model. Therefore, AE-VFL utilizes the advantages of ensemble learning and is able to train a model with stronger performance and generalization ability.
[0128] Second, it should be noted that the AE-VFL method outperforms the centralized training method. For example, on CIFAR-10, the AE-VFL method has an absolute advantage of 19.26% in terms of F-score over the centralized training method (74.46% vs. 55.20%). This is because the centralized training method simply aggregates the training data. In contrast, although the data of the AE-VFL method is not centralized, it can utilize the features and data of different parties during the training process. In addition, AE-VFL incorporates the concepts of ensemble learning and gradient boosting. Each party is used to make up for the deficiencies of the previous party's model. This architecture makes better use of the capabilities of distributed parties and demonstrates superior overall model performance compared to centralized training.
[0129] Table 2 Classification Results on Adult and CIFAR-10 Datasets
[0130]
[0131] Comparative evaluations show that both traditional VFL and centralized learning methods generally outperform the independent training of labeled participants. For example, on the Adult dataset, traditional VFL and centralized training methods improved the F-score by approximately 15.83% and 15.82% respectively compared to independent training. Similarly, on the CIFAR-10 dataset, the F-score of traditional VFL and centralized training methods increased by approximately 2.63% and 0.81% respectively compared to independent learning. The key difference lies in the way they achieve the improvement. Centralized learning directly benefits from the extended training data, while VFL and AE-VFL take data privacy into account and adopt distributed learning methods. By aggregating models, the collective knowledge of multiple participants is utilized while maintaining data privacy.
[0132] These findings highlight the advantages of collaborative learning methods such as VFL and AE-VFL, which strike a balance between data privacy and model performance and demonstrate their potential to advance machine learning in scenarios with limited data sharing.
[0133] Table 3 Comparison results on KCHouse and Song datasets
[0134]
[0135] Example 2
[0136] In one embodiment of the present disclosure, a privacy-preserving vertical federated learning system based on additive integration is provided, including:
[0137] An initialization module for initializing the top-level model on the server side and the bottom-level models of each participant, and constructing an additive integration longitudinal federated learning framework;
[0138] A federated learning module for, in the additive integration longitudinal federated learning framework, the top-level model on the server side is assigned to each participant, and each participant performs sequential training of the bottom-level model through a gradient boosting mechanism;
[0139] Among them, the sequential training of the bottom-level model through the gradient boosting mechanism includes: after a participant trains the bottom-level model using local private data, the output of the bottom model is uploaded to its corresponding top model, the server calculates the loss of each top model, and propagates the gradient to the bottom model of the next participant. The next participant serializes the training of the bottom model and the top model of the next participant by considering the residuals of the previous participant, and through ensemble learning and gradient boosting, a complementary effect between the participant models is formed to achieve additive integration vertical federated learning.
[0140] Example 3
[0141] In one embodiment of the present disclosure, there is provided a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the privacy-preserving vertical federated learning method based on additive integration.
[0142] Embodiment 4
[0143] In one embodiment of the present disclosure, there is provided a non-transitory computer-readable storage medium, characterized in that the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, it implements the privacy-preserving vertical federated learning method based on additive integration.
[0144] Embodiment 5
[0145] In one embodiment of the present disclosure, there is provided an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes and implements the privacy-preserving vertical federated learning method based on additive integration.
[0146] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified function in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0147] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the specified function in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0148] Although the specific implementation manners of the present disclosure have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present disclosure are still within the protection scope of the present disclosure.
Claims
1. A privacy-preserving vertical federated learning method based on additive integration, characterized in that Including: Initialize the top-level model on the server side and the bottom-level models of each participant, and construct an additive integrated vertical federated learning framework; In the additive integrated vertical federated learning framework, the top-level model on the server side is assigned to each participant, and each participant performs sequential training of the bottom-level model through a gradient boosting mechanism; Among them, the sequential training of the bottom-level model through the gradient boosting mechanism includes: after a participant uses local private data to train the bottom-level model, the output of the bottom model is uploaded to its corresponding top model, the server calculates the loss of each top model, and propagates the gradient to the bottom model of the next participant. The next participant serializes the training of the bottom model and the top model of the next participant by considering the residuals of the previous participant. Through ensemble learning and gradient boosting, a complementary effect between participant models is formed to achieve additive integrated vertical federated learning.
2. The privacy-preserving vertical federated learning method based on additive integration according to claim 1, wherein In the additive integrated vertical federated learning framework, the server side assigns an independent top-level model to each participant, and the participants perform sequential training. The training objective of subsequent participants becomes the pseudo-residuals of the previous participants. By training the bottom-level models for each participant, the training objectives of each participant no longer share the same label.
3. The privacy-preserving vertical federated learning method based on additive integration according to claim 1, wherein A participant uses local private data to train the bottom-level model, including: dividing the public dataset D into different participants according to features. Among the data samples aligned among all participants, each participant will complete the forward propagation process using local private data based on its bottom model, and use a deep neural network model as the bottom-level model to obtain the latent representation of each private data sample.
4. The privacy-preserving vertical federated learning method based on additive integration according to claim 1, wherein, The local output from the participants is passed as input to their corresponding top-level models. In the additive integrated vertical federated learning framework, each participant has its own independent top-level model, which is implemented through another DNN model. Through the top-level model, the final output prediction of each participant's data is obtained. The final output of the participant for a multi-class task is a probability. According to the gradient boosting mechanism, the final prediction represented considers the output of the previous bottom-level model.
5. The privacy-preserving vertical federated learning method based on additive integration according to claim 1, characterized in that, The server performs backpropagation and calculates two gradients for the model parameters of the top-level model and the forward output of the bottom-level models from each participant. Using the gradient of the top-level model, the label owner calculates the average gradient of each private data sample and updates its top-level model.
6. The privacy-preserving vertical federated learning method based on additive integration according to claim 1, wherein, Each participant calculates the gradient of its bottom-level model and updates its bottom-level model. The process is repeated until convergence. Each participant calculates the gradient of the parameters of its bottom-level model based on the local private data and the gradient of the forward output from the label owner, and then updates its bottom-level model.
7. A privacy-preserving vertical federated learning system based on additive integration, characterized in that, Including: An initialization module for initializing the top-level model on the server side and the bottom-level models of each participant, and constructing an additive integrated vertical federated learning framework; A federated learning module for, in the additive integrated vertical federated learning framework, the top-level model on the server side being assigned to each participant, and each participant performing sequential training of the bottom-level model through a gradient boosting mechanism; Among them, the sequential training of the underlying model through the gradient boosting mechanism includes: after the participant uses the local private data to train the underlying model, the output of the bottom model is uploaded to its corresponding top model, the server calculates the loss of each top model, and propagates the gradient to the bottom model of the next participant. The next participant serializes the training of the bottom model and the top model of the next participant by considering the residuals of the previous participant. Through ensemble learning and gradient boosting, a complementary effect between the participant models is formed to achieve additive integration of vertical federated learning.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the privacy-preserving vertical federated learning method based on additive integration according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, it implements the privacy-preserving vertical federated learning method based on additive integration according to any one of claims 1-6.
10. An electronic device, characterized in that, Including: A processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device implements the privacy-preserving vertical federated learning method based on additive integration according to any one of claims 1-6.