A Participant Selection Method Based on Model Accuracy Prediction in Vertical Federated Learning
Through the participant selection method based on model accuracy estimate in vertical federated learning, combined with the estimation of encrypted mutual information and the number of sample intersections, the problem of how to efficiently select the right participant to improve model performance is solved, which achieves privacy protection and data security, while maximizing the expected model performance of task publishers.
Patent Information
- Application Number
- CN202411520613.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-10-29
AI Technical Summary
In vertical federated learning, how to efficiently select the right participants for machine learning model training and assign prediction tasks, ensuring improved model performance while finding a balance between privacy protection and data security.
Participant selection method based on model accuracy estimate is used to estimate model performance by encrypting mutual information and the number of sample intersections, and fit the best prediction function of model accuracy. Combined feature value encoding is used to realize the privacy computing mutual information algorithm, combined with PCA dimensionality reduction and feature discretization, and is embedded in the encrypted sample alignment process of vertical federated learning to achieve privacy protection. Based on greedy strategies, participant selection algorithm is implemented to maximize the performance benefits of test samples, and to complete participants who complete the model prediction task.
It realizes the estimated contribution of participants to model performance in vertical federated learning, improves the efficiency and accuracy of model training, ensures the security of data privacy, and maximizes the expected model performance of task publishers.
Smart Images

Figure CN119026671B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of federated learning technology, specifically relates to the participant selection technology in vertical federated learning, and particularly relates to a method for selecting participants based on model accuracy prediction in vertical federated learning. Background Art
[0002] Traditional artificial intelligence model training usually requires a large amount of data. However, this data is often distributed in different institutions and regions, and restricted by regulations such as privacy protection laws, it cannot be directly centralized on a central server for processing. Federated learning, a privacy-preserving distributed machine learning method, can collaborate among multiple parties to jointly train a model on the premise of ensuring that the data does not leave the local device. It can not only use the data of each party for model training but also protect data privacy well. Different from the horizontal federated learning scenario where the data sets of each party have the same feature space and different sample spaces, vertical federated learning aims to enable each party sharing the same sample space and different feature spaces to collaborate in training by using the scattered features of their aligned samples, which has important research value in fields such as financial credit and healthcare. In the vertical federated learning mode, due to the different distributions of data features, it is crucial to efficiently select appropriate participants for machine learning model training and allocate prediction tasks, as this will directly affect the performance of the finally trained joint model. To maximize the inference performance of the federated model on new samples, the participant selection problem in vertical federated learning faces the following challenges:
[0003] 1) In vertical federated learning, each participant contributes to the federated machine learning model by sharing its features with the federated task publisher. In practical applications, there are many participants with feature dimensions that have little relevance to the training task or are even harmful to the model performance. How to quantify the contribution of each potential participant in vertical federated learning to the final model performance to help the task publisher screen participants, thereby improving the training efficiency, is a major challenge, which also has great significance for saving the task publisher and training costs and promoting healthy and sustainable federated learning.
[0004] 2) Currently, most of the calculation methods for model contribution and benefit distribution in the context of federated machine learning are based on the assumption that the participants are determined, and they study how to evaluate the contribution of each participant after the vertical federated learning training is completed. However, in real application scenarios, the task publisher needs to pay a reward for each cooperation with the participant. It is obviously unreasonable to evaluate the contribution of the participant and then make a selection after the cooperation ends. A method that can estimate the contribution of the participant before the vertical federated training starts is needed.
[0005] 3) It is not easy to develop a highly adaptable participant selection mechanism for vertical federated learning and find a balance between contribution measurement and data privacy for the following reasons:
[0006] a. The performance of the model has high randomness and is affected by various factors, such as the number of public samples, feature importance, task difficulty, the selection order of training samples, etc. It is very difficult to directly use a single metric to estimate to what extent each participant can improve the performance of a specific model, which makes problem modeling more complex.
[0007] b. Since federated learning requires data to stay local to the data owner and the task publisher cannot access the data, it is very difficult to obtain important factors such as the importance of features for the training task. In addition, in the scenario of vertical federation, it is usually assumed that the participants are honest but curious, and each participant will try to infer as much specific content as possible from the encrypted information obtained by other participants. Existing contribution measurement methods will inevitably expose the feature values of each participant, leading to the problem of privacy leakage. How to obtain the interaction information of each participant while ensuring privacy security is also a huge challenge.
[0008] c. After the vertical federated learning model training is completed, in order to obtain the best inference performance on more new samples, the task publisher also needs to select appropriate participants according to its own budget to cooperate in the corresponding inference tasks, which requires further solving a new constrained optimization problem based on the original participant selection results. Summary of the Invention
[0009] Object of the Invention: Aiming at the participant selection problem of vertical federated learning existing in the above-mentioned existing work, the present invention provides a participant selection method based on model accuracy prediction in vertical federated learning, which is a privacy-protected participant selection method based on model accuracy prediction, and this method can play a role before the start of vertical federated training.
[0010] Technical Solution: A participant selection method based on model accuracy prediction in vertical federated learning, the method is based on the vertical federated learning framework, estimates the performance of the vertical federated learning model according to the encrypted mutual information and the number of sample intersections, and fits to obtain the best prediction function of the model accuracy, uses the combined eigenvalue coding to implement the privacy calculation mutual information algorithm, and embeds it into the encrypted sample alignment process of vertical federated learning through PCA dimensionality reduction and feature discretization means to achieve the privacy protection task;
[0011] The method further includes implementing a participant selection algorithm based on a greedy strategy to maximize the performance gain of test samples, and accordingly completing the participant allocation for the model prediction task;
[0012] The method includes the following steps:
[0013] (1) Conduct a correlation analysis on the number of training samples and the mutual information between features and labels, including training both on a public dataset to verify the positive correlation between the accuracy of the vertical federated learning model, the number of training samples, and the mutual information.
[0014] (2) Based on the analysis results in step (1), build a model accuracy prediction function, and split and combine the data of the task publisher to simulate the vertical federated learning process, obtaining corresponding data points representing the relationship between mutual information and model accuracy, as well as data points representing the relationship between the number of training samples and the estimation error; then use a linear function to fit the relationship between model accuracy and mutual information , and further use a quadratic function to fit the relationship between the number of training samples and the accuracy estimation error , correct and fine-tune the accuracy estimation error based on mutual information, and then fit the joint effect of mutual information and the number of training samples on the accuracy of the vertical federated learning model.
[0015] The model accuracy prediction function is fitted using the following formula:
[0016]
[0017] represents the model accuracy estimated based on these two factors of mutual information and the number of training samples.
[0018] (3) Reduce the residual between the actual performance A and the estimated performance to correct the fitting error of the model performance prediction function. The steps include:
[0019] (31) Calculate the absolute error of each sample to obtain the value obtained after calculation, divide the error into intervals, count the number of times the error appears in each interval, and finally divide the number of error times in each interval by the total number of times to calculate the probability distribution.
[0020] (32) Repeat the model accuracy fitting operation in step (2) based on the data points obtained from the real federated training, and also obtain a probability distribution of the fitting error.
[0021] (33) Calculate the estimation error after each round of real vertical federated training process to obtain an increasingly accurate error mean offset err, and then correct the estimated performance according to the following formula to make the average estimation error close to 0:
[0022]
[0023] In the formula, That is the finally obtained model accuracy prediction function, which is fitted by the task publisher according to its own data and part of the historical training data before the start of vertical federated training;
[0024] (4)The mutual information between the features participating in the training and the task labels in privacy computing, and the calculation process includes:
[0025] (41)The task publisher and the federated participant k each discretize their features in each dimension, divide the data range of the feature values of each feature dimension into b equidistant intervals numbered from 1 to b. For each floating-point type feature, calculate the interval number it falls into according to its feature value, and then use the corresponding numbered value to replace the continuous floating-point value as the discretized feature value;
[0026] (42)Implement the encrypted sample alignment between the task publisher and the federated participant k based on the privacy intersection technology of the Diffie-Hellman encryption algorithm;
[0027] In step (42), the task publisher and the participant k first encrypt their own sample ids with their own private keys and send them to each other. Then the participant k encrypts the sample id encrypted by the task publisher with its own private key again and sends it to the task publisher. Based on the Diffie-Hellman encryption algorithm, a symmetric cipher algorithm is used for encryption and decryption. The task publisher only needs to compare whether the values of the sample id of the participant k encrypted by the private keys e and d in sequence are the same as the values of the sample id of the task publisher encrypted by the private keys d and e in sequence to obtain the sample intersection.
[0028] (43)During the encrypted sample alignment process, a combined feature value encoding is transmitted for the mutual information between the features and labels shared by the privacy computing participant and the task publisher. The calculation of this combined feature value encoding includes:
[0029] Suppose the dataset of the federated participant has dimensional features after PCA dimensionality reduction, and the feature values are discretized and scattered into intervals. The mapping function relationship between the discretized value and the original feature value is Suppose the sample record format is:
[0030] [ id , f 1 , f 2 ,..., f d k ]
[0031] The combined encoding of the feature values of the i-th aligned sample of the participant k is expressed by the following formula:
[0032]
[0033] It can be obtained therefrom that for any combination of characteristic values, the corresponding combined characteristic value coding is uniquely determined;
[0034] (5) The task publisher calculates each value in the following formula by combining the characteristic value coding sent by the participant, its own characteristic value coding, and the label column value, represents the number of data samples under the given variable value, and further calculates the mutual information estimation value between the combined characteristic and the label:
[0035] ;
[0036] (6) According to the training set D and the test set T of the task publisher, the goal is converted to maximizing the expected model performance obtained by all samples in the test set T, specifically including:
[0037] According to the model accuracy prediction function and the distribution of the estimation error obtained in step (3), the expected model maximization problem is represented by re-modeling as follows:
[0038] max ∑ i ∈ T E [ U i ] s . t . U i = max { k | i ∈ T k & k ∈ S ∪ { O } } Q k Q k = f ( I k , n k ) + e k , k ∈ S e k ~ N ( μ , σ 2 ) , k ∈ S n k ≤ | D ∩ D k | T k = T ∩ D k ∑ k n k ≤ B
[0039] Among them, the task publisher is O, its training budget is B, the set of participants is S, and for each participant k in S, its sample set The intersection with the test set of the task publisher is , The samples in have both the characteristic dimensions of k and the task publisher. For these samples, the model jointly trained by the task publisher and participant k can be used to perform inference prediction, The estimated performance of is The variable is a decision variable, corresponding to the number of training samples brought by participant k. When is larger, the predicted model performance is larger. At the same time, the samples of participant k will occupy more quotas in the budget B, and the best inference performance obtained by the test sample i in the test set T is the maximum value of all
[0040] Calculate the estimated accuracy of the model jointly trained by each participant k and the task publisher according to steps (1) to (5). Sort all participants according to the estimated accuracy from high to low. If there is The accuracy is lower than the model trained by the task publisher using only its own characteristic dimension dataIf its precision does not meet the requirements, it will be excluded from the selection range. Then, starting from the participant with the highest estimated precision, try to purchase training data from them in turn until the cumulative amount of purchased data exceeds the budget limit B or there are no more participants available for selection. Finally, according to the purchase situation, inform the selected participants of the amount of training data they need to provide and conduct vertical federated learning training.
[0041] (7) Considering that after selecting participants using the greedy strategy in step (6) before training, there will be a problem that several participants will obtain several vertical federated learning models. In the model prediction and inference stage, it is necessary to perform the best model task allocation. Assign the prediction tasks of the corresponding test set samples to each participant according to the priority ranking of the participants, and then analyze and compare the prediction results of the test data with the true labels, and calculate the proportion of the number of samples correctly classified by the model to the total number of samples to obtain the prediction accuracy of the vertical federated learning model.
[0042] Specifically, the prediction accuracy of the vertical federated learning model described in this method is calculated according to the following formula:
[0043]
[0044] Among them, TP represents the number of samples correctly predicted as the positive class by the model, TN represents the number of samples correctly predicted as the negative class by the model, FP represents the number of samples that the model wrongly predicts the negative class as the positive class, and FN represents the number of samples that the model wrongly predicts the positive class as the negative class.
[0045] Further, the correlation analysis process of the number of training samples in step (1) includes training it using three machine learning models including random forest, logistic regression, and support vector machine, and then recording the relationship between the number of training samples and the model prediction accuracy, and verifying that under their respective model and dataset configurations, the model accuracy and the number of training samples are both significantly positively correlated.
[0046] The correlation analysis process of the mutual information in this step includes:
[0047] The mutual information MI of two random variables is a measure of the mutual dependence between variables. The mutual information that measures the information shared by two random variables X and Y is represented by and the standard definition formula of mutual information is as follows:
[0048]
[0049] Assume that the feature set held by the task publisher of vertical federated learning is X, and the feature set held by the federal participant k is , then the mutual information between all features and labels for federated learning can be expressed as , and this mutual information can measure The mutual dependence between and, according to the definition of mutual information, is calculated as follows:
[0050]
[0051] where and y represent the possible values of and Y, is the probability that both are satisfied simultaneously, is the KL divergence between the two distributions. Since the true distribution of the data is difficult to obtain, we consider using the maximum likelihood estimation method to approximate them based on the feature distribution and label distribution in the training set, as shown in the following formula:
[0052]
[0053] In the formula, is the distribution of estimated from the training set, is the number of data samples for a given variable value, is the total number of common data samples between participant k and the task publisher;
[0054] For continuous features, they are converted into discrete ones, and then the probability distribution is calculated through maximum likelihood estimation. or The dimension of
[0055] includes converting it into a low-dimensional vector by using PCA dimensionality reduction.
[0056] The higher the mutual information between the features and the labels, the higher the contribution degree of the features to predicting the labels, which helps to improve the inference performance of the model, that is, the label prediction accuracy.
[0057] Suppose the task publisher T has a total of z-dimensional features, and the task publisher T selects a feature combination containing z' features , and the remaining features are used as the features of the federated participants. Then the task publisher randomly selects an integer within the range of , and then randomly selects a feature combination containing features from the remaining features ; then uses as the feature set of the task publisher, As a set of features of the federal participant, by partitioning the feature dimension data of the task publisher, the task publisher can separately simulate the execution of vertical federated learning and obtain the performance of the model before cooperating with other participants.
[0058] Furthermore, in step (4), when applying the model performance estimation function to the actual vertical federated training process, the mutual information and the number of aligned training samples need to be obtained. Among them, the number of samples is naturally obtained through data interaction, and the calculation of mutual information requires privacy processing.
[0059] In the process of calculating the mutual information between the features and task labels participating in the training by privacy computing, it includes dimensionality reduction processing on the data feature dimensions of the participants or the task publisher;
[0060] The dimensionality reduction processing mentioned above refers to PCA dimensionality reduction, and it is performed on data feature dimensions exceeding three dimensions to reduce the data transmission overhead and calculation overhead.
[0061] Furthermore, in step (7), assuming there are several federal participants, and they have different sample intersections with the task publisher T. In the model inference stage, the common samples between participant k and the task publisher can use the model jointly trained by the two of them ; for the samples shared by some participants and the task publisher, the model with relatively better estimated performance is selected for prediction; for the samples only held by the task publisher, the model M0 trained by the task publisher itself is used for prediction.
[0062] Step (7) includes using the model accuracy estimation function to substitute the encrypted mutual information and the number of sample intersections to obtain the accuracy estimation values of the vertical federated learning models obtained by each participant's cooperation with the task publisher. Then, by sorting the estimated values of the model accuracy, corresponding priorities are assigned to each participant, where the participant with a larger estimated model accuracy has a higher priority;
[0063] For each sample in the test set of the task publisher, the system will check which participants have records related to the sample and assign prediction tasks according to the priority; the specific steps are as follows:
[0064] For each sample in the test set of the task publisher, the system will check which participants have records related to the sample and assign prediction tasks according to the priority;
[0065] For each sample ID in the test set, the system will sequentially check each participant in the order of priority to see if they have a record corresponding to the sample ID. If a participant has a record corresponding to the sample ID, the system will mark the participant as "1", otherwise mark it as "0";
[0066] If only one participant has the record of the sample, the task publisher will choose to use the model trained in cooperation with this participant to predict the sample;
[0067] If multiple participants all have records corresponding to the sample ID, select the model trained in cooperation with the participant with the highest priority to be responsible for the prediction task of this sample.
[0068] Beneficial effects: Compared with the prior art, the substantial progress and remarkable effects of the present invention are as follows:
[0069] 1) By designing a simulation experiment of vertical federated learning, the present invention explores the factors affecting the accuracy of the vertical federated learning model, including feature importance and the number of samples that can be aligned. Inspired by the mutual information selection method in feature engineering, the mutual information value is used to reflect feature importance, and the joint influence of the above two factors on the performance of the vertical federated model is modeled, a model accuracy prediction function is fitted, and the prediction function is corrected based on the error probability distribution.
[0070] 2) The contribution measurement methods that have been proposed so far are not applicable before the start of vertical federated training, and most of the measurement criteria are irrelevant to the training task. The present invention innovatively proposes to split the feature data of the task publisher itself and use the results of historical training data to fit a low-error model accuracy prediction function, realizing the pre-quantitative estimation of the contribution of participants to the improvement of model performance in vertical federated learning for a specific training task.
[0071] 3) The method described in the present invention discretizes the continuous feature values after PCA dimensionality reduction and encodes them according to the designed rules, and completes the privacy calculation of mutual information in the data encryption state. This method can not only be efficiently embedded into the sample alignment process of vertical federated training, but also takes into account the data transmission efficiency and computational complexity during the training process.
[0072] 4) The present invention proposes a participant selection strategy applicable when multiple participants can improve the model performance. Through the model prediction task allocation based on the greedy algorithm, the model performance expected by the task publisher is maximized. The experimental results show that the test set accuracy obtained by vertical federated prediction according to the participant selection strategy proposed by the present invention is generally higher than the result of single federated model prediction without selection. In addition, the change of the priority of the participants will also affect the effect of the model, which further illustrates the effectiveness of the priority sorting strategy of the participants calculated according to the prediction accuracy of the present invention. Brief Description of the Drawings
[0073] Figure 1 is a schematic diagram of the training transaction service framework of the vertical federated learning model in the present invention;
[0074] Figure 2 The figure shown in this invention is a fitting example diagram of the relationship between mutual information and model accuracy;
[0075] Figure 3 It is a schematic diagram of the probability distribution of the error of the fitting function;
[0076] Figure 4 It is an image of the final model prediction function obtained based on the breast cancer dataset;
[0077] Figure 5 It is an example diagram of the model prediction task allocation strategy described in this invention. Specific implementation manners
[0078] In order to elaborate in detail the technical solution disclosed in this invention, the following further elaboration will be made in combination with the specification drawings and specific embodiments.
[0079] For vertical federated learning in the scenario of collaborative training of models among multiple organizations, it is assumed that there are K federated participants who can all cooperate with the federated task publisher to jointly train a machine learning model through the vertical federated learning mechanism. They first need to perform encrypted entity alignment with the task publisher to find common data samples, and then jointly train a vertical federated model based on richer feature information without revealing the original data. The vertical federated learning setting of this invention includes a task publisher and a federated participant.
[0080] The task publisher is an organization that hopes to build a better-performing model by combining more relevant features of other participants through vertical federated learning. Its goal is to improve the prediction accuracy of the trained model for all new customer samples. The task publisher itself holds a certain amount of feature dimension information and the label information of the samples, and there is an overlapping part between its sample id and that of the federated participant.
[0081] The federated participant has no label information, but has feature dimension information that the task publisher does not have. There is an overlapping part between its sample id and that of the task publisher, and it can provide model training services for the task publisher. The federated participant's cooperation with the task publisher in performing vertical federated training not only provides its own data value for the task publisher, but also consumes its own computing resources and communication bandwidth. Therefore, the task publisher often needs to pay equivalent compensation to it.
[0082] In fact, not all participants can bring significant performance improvement to the model. For the task publisher, since the profit brought by the model performance improvement is limited, it is very important to screen out suitable federated participants before the start of vertical federated training. To address this issue, the present invention provides a method for evaluating the performance of participants with low error. It preferentially selects those participants with excellent performance from K participants to conduct two-party vertical federated training with the task publisher respectively, ensuring that the prediction accuracy of the obtained model on new samples is higher than that of the model trained by the task publisher itself, and designs a suitable participant selection strategy to assign the prediction task to the appropriate participant to cover more test cases and further improve the accuracy of the model on new samples.
[0083] In the method described in the present invention, this method is a participant selection algorithm based on model accuracy prediction applicable to vertical federated learning. Different from related technologies such as the method for measuring the contribution of participants in existing federated learning, this method can quantify the potential contribution of each participant to the performance of the joint model before the start of vertical federated training and at the same time takes into account the privacy protection task of vertical federated learning. In order to maximize the expected model performance obtained for all samples in the test set of the task publisher, the method described in the present invention includes constructing a joint prediction function of the encrypted mutual information and the number of intersections of training samples on the model performance, and implementing the participant selection algorithm based on the greedy strategy and completing the participant assignment for the model prediction task.
[0084] Furthermore, as shown in Figure 1 the specific implementation steps of the technical solution provided by the present invention can be described as follows:
[0085] Step 1: Conduct a correlation analysis on the factors affecting the performance of the vertical federated learning model.
[0086] This step takes the number of training samples and the mutual information between the features and labels participating in the training as the key factors affecting the model performance for analysis.
[0087] (11) Correlation analysis of the number of training samples.
[0088] In vertical federated learning training, only aligned samples can be used as the training set for subsequent model training. When the number of intersections of samples between a federated participant and the task publisher is small, the model is likely to have a high variance, the data lacks diversity, and it is more likely to overfit, resulting in poor generalization performance of the model on the test data.
[0089] This invention conducts experiments on three machine learning models, namely Random Forest Classification (RFC), Logistic Regression (LR), and Support Vector Classification (SVC), using two classic publicly available datasets, the breast cancer dataset and the adult income dataset. It records the relationship between the number of training samples and the model prediction accuracy, and verifies that under the respective model and dataset configurations, there is an obvious positive correlation between the model accuracy and the number of training samples. Therefore, the number of common samples between the participating parties and the task publisher is an important reference indicator for estimating the accuracy of vertical federated models.
[0090] (12) Correlation analysis of mutual information.
[0091] The mutual information (MI) of two random variables is a measure of the mutual dependence between variables. Mutual information measures the information shared by two random variables X and Y, denoted by The standard definition formula of mutual information is as follows:
[0092]
[0093] Assume that the feature set held by the task publisher in vertical federated learning is X, and the feature set held by the federal participant k is , then the mutual information between all features and labels for federated learning can be expressed as , and this mutual information can measure and The mutual dependence between them. According to the definition of mutual information, the calculation formula of is as follows:
[0094]
[0095] where and y represent and the possible values of Y, is The probability that they are satisfied simultaneously, is the KL divergence between the two distributions. Since the true distribution of the data is difficult to obtain, it is considered to approximate them using the maximum likelihood estimation method based on the feature distribution and label distribution in the training set, as shown in the following formula:
[0096]
[0097] In the formula is the distribution of estimated from the training set, is the number of data samples under the given variable value. is the total number of common data samples between participant k and the task publisher. When some features are continuous, they are converted into discrete ones, and then the maximum likelihood estimation is applied to calculate the probability distribution. If or has a high dimension and a large amount of calculation, PCA dimensionality reduction can be performed to convert it into a low-dimensional vector.
[0098] The higher the mutual information between the feature and the label, the higher the contribution degree of the feature to predicting the label, which helps to improve the inference performance of the model (i.e., the label prediction accuracy).
[0099] According to our experimental results, under the above three machine learning models, as the mutual information value MI increases, the model accuracy improves rapidly. Therefore, the mutual information value between the features shared by the participating party and the task publisher and the label is also an important reference index for estimating the accuracy of the vertical federated model.
[0100] Step 2: Fitting of the model accuracy prediction function.
[0101] Fitting the model accuracy prediction function, it can be obtained from the correlation analysis results in Step 1 that the influence of mutual information on the model performance is more significant than the influence of the number of training samples on the model performance. Therefore, in this step, the relationship between the model performance and the mutual information is first fitted, and then the factor of the number of training samples is used for fine-tuning.
[0102] Suppose the task publisher T has a total of z-dimensional features, and the task publisher T can retain a feature combination containing z' features ( ), and the remaining features are used as the features of the federated participant. The task publisher randomly selects an integer within the range of , and then randomly selects a feature combination containing features from the remaining features . Use as the feature set of the task publisher, and as the feature set of the federated participant. In this way, by dividing the feature dimension data of the task publisher, the task publisher can separately simulate the execution of vertical federated learning and obtain the model performance before cooperating with other participants. as the feature set of the federated participant, so that by dividing the feature dimension data of the task publisher, the task publisher can separately simulate the execution of vertical federated learning and obtain the model performance before cooperating with other participants.
[0103] This embodiment uses breast cancer datasets and adult income datasets to simulate a real longitudinal federated learning scenario. Taking the adult income dataset as an example, this dataset has a total of 14-dimensional features. Randomly select 8-dimensional features as the features of the task publisher, and the remaining 6 features are the features available for each federated participant. The task publisher randomly selects m features from his 8-dimensional features. For the same m (m < 8), there are multiple feature combinations. Then sample n combinations from these feature combinations and calculate their mutual information with the label and the accuracy of the model trained only with this part of the feature dimensions. Note that these are the results that the task publisher can obtain by training only with his own data. In Figure 2 it is marked with circular dots. The present invention uses these simulated data points to preliminarily fit the relationship between mutual information and model performance. To verify the effect of fitting these data points using linear functions and quadratic functions, randomly select feature combinations with different numbers of features from the remaining 6 features, and combine them with the 8-dimensional features of the task publisher to calculate the mutual information with the label and test the accuracy of the jointly trained model. Using these data, the Figure 2 triangle dots in it can be plotted. Use a quadratic function to fit these points to obtain the corresponding line "Quadratic function fitting the true training accuracy". One of the objectives of this fitting step is to hope that the line fitted according to the circular dots can be as close as possible to the trend of the line fitted according to the triangle dots. After many repeated experiments, it is found that the relationship between mutual information and model performance fitted using a linear function is better than that fitted using a quadratic function in most cases. For the breast cancer dataset, the fitting function is 0.1263x + 0.8664. For the adult income dataset, the fitting function is: 0.1808x + 0.7977.
[0104] Then model the relationship between the number of training samples and the model performance estimation error based on mutual information. To further fine-tune the fitting effect using the factor of the number of training set samples, the experiment sets the ratio of the training samples participating in the calculation of mutual information to all samples: {0.2, 0.4, 0.6, 0.8, 1}, and calculates the mutual information and the corresponding model accuracy for each ratio. Taking the sampling rate as the independent variable and the difference (i.e., the error) between the predicted accuracy value estimated by the fitting line of the circular dots and the actual accuracy corresponding to the circular dots as the dependent variable, a quadratic function is fitted. For the breast cancer dataset, the fitting function is y = -0.04061x2 + 0.07673x - 0.02817. For the adult income dataset, the fitting function is: y = -0.01432x2 + 0.02825x - 0.01064. Based on the function fitted based on mutual information, adding the estimated error fitting amount based on the sampling rate can correct a more accurate model performance, that is: 。
[0105] Step 3: Correction of the fitting error of the model performance prediction function.
[0106] There will still be a certain error between the model performance estimated by the model accuracy prediction function obtained in Step 2 and the actual performance. On the one hand, the reason for this error is that there may be other influencing factors besides mutual information and the number of training samples. On the other hand, the method of calculating mutual information using the feature discretization method itself has a certain error. In order to adjust the average estimation error to a state close to 0, the present invention proposes to calculate the absolute error of each sample, that is, substitute the mutual information of the dots and the number of training samples fitted in Step 2 into the fitting function, calculate the error between the fitting accuracy and the accuracy marked by the dots, divide this error according to a certain interval, calculate the probability distribution of the error, corresponding to Figure 3 the "Err of F^s for SVFL" line in Figure 3 In the same way, calculate the error between the fitting function estimation result of the triangular points in Step 2 and the actual accuracy rate, obtain the probability distribution, corresponding to Figure 3 the solid line "Err of F^r for VFL" in Figure 3 It can be seen from
[0107] In the present invention, an estimated error is calculated and recorded after each new longitudinal federated training process is completed, thereby obtaining an increasingly accurate mean shift amount. In practical applications, there may not be enough historical longitudinal federated training results for further error correction. At this time, based on the idea of transfer learning, the task publisher can use its own data to simulate longitudinal federated training and obtain the performance results of the model. Since it is much easier to obtain the shift amount of the estimated error than to directly fit the model performance prediction function, the number of required pre-trained longitudinal federated training results is much smaller. Figure 3 "Corrected Err of F^s" in Figure 3 reflects the above-mentioned error correction effect. It can be seen that according to the method proposed in the present invention, the corrected error distribution image is closer to the real error distribution image. In addition, from the perspective of participant selection, only the relative advantages and disadvantages between participants need to be estimated, and a slight overall shift in the estimated error value will not greatly affect the final result of participant selection.
[0108] Step 4: Implement the privacy calculation of the mutual information between the features and labels of the federated participants.
[0109] Applying the model performance prediction function to the actual longitudinal federated training process requires obtaining the mutual information and the number of aligned training samples. Among them, the number of samples can be naturally obtained through data interaction, while the calculation of mutual information requires certain privacy processing.
[0110] It is further pointed out that according to the calculation formula in Step 1, the calculation of mutual information depends on , . How to obtain the value without accessing the original data of the federated participants is an important problem to be solved. The present invention proposes a mutual information calculation method with high computational efficiency and privacy protection. This algorithm can be naturally embedded in the encrypted sample alignment process in longitudinal federated learning. The specific process of this algorithm is introduced below.
[0111] In this embodiment, first, in order to reduce the data transmission overhead and computing overhead, the federated participants and the task publisher need to perform PCA dimensionality reduction on their own datasets in advance to convert high-dimensional features into low-dimensional features. Next, the task publisher and the federated participants each perform discretization processing on the features of each dimension after dimensionality reduction. The present invention uses an equal-width binning discretization method to divide the features of each dimension into b bins. Specifically, first calculate the minimum value "min_val" and the maximum value "max_val" of the data of each feature dimension, and then divide the data range of this feature dimension into b equidistant intervals. Then the width of each interval is "dis=(max_val-min_val) / 5". For each floating-point type feature, subtract the minimum value min_val from the feature value, then divide by the interval dis of this dimension feature, and then round down the obtained result to get the number of the interval (that is, which interval it falls into), and use this numbered value to replace the continuous floating-point value as the discretized feature value.
[0112] Then the task publisher and the federated participant k can start the normal encrypted sample alignment process. The difference is that in addition to encrypting and transmitting their own sample ids, they also need to transmit an additional combined feature value encoding, which is used to calculate the mutual information between the features and labels shared by the two. After the values of each feature dimension are mapped to their respective discretized intervals, the discretized feature combination of each aligned sample is encoded through the following calculation formula:
[0113]
[0114] where, is the mapping function relationship between the discretized value and the original feature value, is the number of discretized intervals of the feature value. The feature vector of a sample is denoted as where each feature The encoding in its corresponding discrete interval is 。The encoding process can be regarded as a weighted sum, where the weights corresponding to the eigenvalues of features in different dimensions are different, and the weights are based on the total number of features and the value range. Through this encoding method, any combination of eigenvalue can be mapped to a unique encoding value, thus avoiding confusion during the processing and facilitating further mutual information calculation. After obtaining this encoding, since the task publisher does not know how many discrete intervals the federal participants divide the eigenvalues into during eigenvalue discretization, and even if it knows, it can only deduce the interval numbers after discretization, so it cannot accurately know the actual eigenvalues, thus realizing the privacy calculation of mutual information. To further improve privacy, this encoding value can be further function-mapped and then sent to the task publisher, so that the task publisher can more difficultly infer the original data. The task publisher combines the eigenvalue encoding sent by the participating party, its own eigenvalue encoding, and the label column to calculate the value, and then further calculate the mutual information between the joint features and the label.
[0115] Step 5: Embed the mutual information privacy calculation into the encrypted sample alignment process of vertical federated learning.
[0116] a) Assume that at this time, the federal participant k wants to perform sample alignment with the task publisher. Participant k needs to first hash the ID of each sample it holds locally to obtain the hash value , and the task publisher also hashes the ID of each sample it holds to obtain the hash value .
[0117] b) The task publisher encrypts the hashed sample ID using the encryption key and sends the encrypted ID to participant k. After receiving the encrypted ID sent by the task publisher, participant k further encrypts it, and at the same time sends the combined eigenvalue encoding and the encrypted ID back to the task publisher.
[0118] c) The task publisher finds the matching sample pairs by comparing the received encrypted IDs and calculates the number of common samples , thus realizing the encrypted sample alignment process of vertical federated learning.
[0119] d) For each pair of matching samples, the task publisher counts according to the combined eigenvalue encoding of the matching samples and the combined eigenvalue encoding of its own data and the corresponding labels, records the frequencies of occurrence, and obtains the and values required for calculating the mutual information. represents the number of samples with feature encoding and and label y. .
[0120] e) Finally, the task publisher calculates the mutual information based on the statistical results .
[0121] Through this algorithm, the task publisher can calculate the mutual information between the feature combination and the label without disclosing the privacy data of the participants. This method ensures the security of the data while obtaining valuable statistical information.
[0122] Step 6: Implementation of the participant selection algorithm based on model accuracy prediction.
[0123] Before starting the vertical federated learning training, the task publisher needs to select cooperative participants based on the model accuracy prediction results obtained in the previous five steps and determine the amount of sample data that each participant needs to provide to maximize the performance of the expected model within the budget limit.
[0124] The task publisher first needs to train a model on its own local data and test the accuracy on the test set . Assuming the participant set is S and the total budget of the training samples is B, the task publisher needs to first calculate the estimated accuracy of the model obtained by cooperative training with each participant k , and then sort the priorities of the participants from high to low according to to obtain the sorted participant set .
[0125] The participant selection algorithm initializes two empty sets to store the selected participants and the training samples they provide. Then, the algorithm traverses the sorted participant set in turn . For each participant k, if its accuracy is greater than or equal to , the algorithm will calculate the number of samples needed to be obtained from this participant based on the greedy strategy. As long as the number of samples held by this participant is still within the remaining budget of the total training sample quantity of the task publisher, it will preferentially purchase its training data and add this participant and the corresponding number of training samples to the corresponding sets. If the total number of training samples exceeds the budget limit B after obtaining the samples, the algorithm will stop selecting more participants and terminate the loop. If the accuracy of the current participant is less than the threshold , the algorithm will skip this participant and continue to check the next participant. When the algorithm has traversed all participants or has reached the target sample number B, the algorithm returns the final selected participant set and the corresponding training sample set.
[0126] Combined withFigure 4 The image of the final model prediction function obtained based on the breast cancer dataset as shown. Through this step, the algorithm ensures that only those participants with high accuracy and significant sample contributions are selected, while avoiding exceeding the budget, thus achieving a balance between the optimal sample quantity and quality.
[0127] Step 7: Participant prediction task assignment in the model inference stage.
[0128] After selecting participants using the greedy strategy described in Step 6 before training, since multiple participants will result in multiple vertical federated models, the best model task assignment needs to be performed in the model prediction and inference stage. Taking Figure 5 as an example, we assume there are two federated participants, P1 and P2, which have different sample intersections with the task publisher T. In the model inference stage, for the common samples between participant P1 and T, the model M1 jointly trained by P1 and T can be used; for the common samples between participant P2 and T, the model M2 jointly trained by P2 and T can be used. For the samples that are common to P1, P2, and T, the one with better predicted performance among M1 and M2 is selected for prediction. For the samples that only T has and neither P1 nor P2 has, the model M0 trained by T itself is used for prediction. Note that the inference performance of M1 and M2 should be better than that of M0, otherwise the task publisher will not choose to jointly train models with them.
[0129] Specifically, after substituting the encrypted mutual information and the number of sample intersections into the model accuracy prediction function to obtain the accuracy estimation values of the vertical federated learning models obtained by each participant collaborating with the task publisher, the corresponding priorities can be assigned to each participant by sorting the estimated values of the model accuracy. Among them, the higher the predicted accuracy of the model, the higher the priority of the participant. For each sample in the test set of the task publisher, the system will check which participants have records related to the sample and assign prediction tasks according to the priority. The specific steps are as follows:
[0130] a) For each sample ID in the test set, the system will sequentially check whether each participant (in the order of priority) has a record corresponding to the sample ID. If a participant has a record corresponding to the sample ID, the system will mark the participant as "1", otherwise as "0".
[0131] b) If only one participant has the record of the sample, then this participant will be responsible for making predictions on the sample. Even if this participant is not the one with the highest priority, since other participants do not have the relevant record, this participant automatically undertakes this task. If multiple participants have records corresponding to the sample ID, the system will select the participant with the highest priority according to the priority order to be responsible for the prediction task of this sample.
[0132] Finally, according to the above steps, the system assigns a specific participant to each sample to perform the prediction task. This ensures that the task is assigned to the participant with the relevant record and the highest priority as much as possible.
[0133] After the task publisher cooperates with each participant to carry out model prediction according to the assigned tasks, the task publisher records the label prediction results of all its test samples, and counts the predicted labels and the true labels, and a confusion matrix of the classification model can be obtained. It is a tool for evaluating the performance of the classification model. In this matrix, there are four key concepts:
[0134] TP (True Positives): This is the number of samples that the model correctly predicts as the positive class, that is, the samples where both the model prediction result and the actual label are positive.
[0135] TN (True Negatives): This is the number of samples that the model correctly predicts as the negative class, that is, the samples where both the model prediction result and the actual label are negative.
[0136] FP (False Positives): This is the number of samples that the model wrongly predicts the negative class as the positive class.
[0137] FN (False Negatives): This is the number of samples that the model wrongly predicts the positive class as the negative class.
[0138] Finally, the prediction accuracy of the vertical federated learning model is calculated according to the following formula:
[0139]
[0140] Based on the result of this accuracy, the effectiveness of the participant selection algorithm and prediction task assignment strategy proposed in the present invention based on model accuracy estimation can be verified.
Claims
1. A method for selecting federated participants based on model accuracy estimation in vertical federated learning, characterized in that: The method is based on the framework of vertical federated learning. It estimates the performance of the vertical federated learning model according to the encrypted mutual information and the number of sample intersections, and fits the best estimation function for the model accuracy. It uses combined eigenvalue coding to implement the privacy calculation mutual information algorithm, and embeds it into the encrypted sample alignment process of vertical federated learning through PCA dimensionality reduction and feature discretization to achieve privacy protection tasks. The method also includes implementing a federation participant selection algorithm based on a greedy strategy to maximize the test sample performance benefit, thereby completing the allocation of federation participants for the model prediction task; The method specifically comprises the following steps: (1) Conduct correlation analysis on the number of training samples and the mutual information between features and labels, including training the two on public datasets to verify the positive correlation between the accuracy of the longitudinal federated learning model and the number of training samples and mutual information; (2) Based on the analysis results in step (1), a model accuracy estimation function is constructed, and the data of the task publisher is split and combined to simulate the longitudinal federated learning process, and the corresponding data points representing the relationship between mutual information and model accuracy and the relationship between the number of training samples and the estimation error are obtained; then a linear function is used to fit the relationship between model accuracy and mutual information I Then use the quadratic function to further fit the relationship between the number of training samples n and the accuracy estimation error Correct and fine-tune the mutual information-based precision estimation error, and then fit the joint effect of mutual information and the number of training samples on the accuracy of the longitudinal federated learning model; The model accuracy estimation function is fitted using the following formula: Represents the model accuracy estimated by the mutual information I between all training features and task labels and the number of training samples n; (3) Reduce the actual performance A and the estimated performance calculated by the model accuracy estimation function in step (2) The residual between is used to correct the fitting error of the model performance prediction function. The steps include: (31) Calculate the absolute error of each sample and get The error is divided into intervals based on the calculated value, and the number of errors in each interval is counted. Finally, the number of errors in each interval is divided by the total number to calculate the probability distribution. (32) Repeat the model accuracy fitting operation in step (2) based on the data points obtained from the real federated training, and also obtain a probability distribution of the fitting error; (33) After each round of the real longitudinal federation training process, the estimation error is calculated to obtain an increasingly accurate error mean offset err, and then the estimation performance is corrected according to the following formula so that the average estimation error is close to 0: In the formula, F r This is the final model accuracy estimation function, which is fitted by the task publisher based on its own data and part of the historical training data before the start of the vertical federated training; (4) Privacy calculation is the mutual information between the features involved in training and the task labels. The calculation process includes: (41) The task publisher and federated participant k each discretize the features of each dimension, divide the data range of the feature value of each feature dimension into b equally spaced intervals, numbered 1-b, and for each floating-point type feature, calculate the interval number in which it falls according to its feature value, and then use the corresponding number value to replace the continuous floating-point value as the discretized feature value; (42) Based on the privacy intersection technology of the Diffie-Hellman encryption algorithm, the task publisher and the federation participant k can align the encrypted samples; (43) During the encrypted sample alignment process, a combined feature value code is transmitted to calculate the mutual information of features and labels shared by the federation participants and the task publisher. The combined feature value code calculation includes: Suppose the dataset of federated participants has d after PCA dimensionality reduction k dimensional features, the eigenvalues are discretized and scattered to b k In the interval, the mapping function relationship between the discretized value and the original eigenvalue is bin(). Assume that the sample record format is: In the formula, Indicates samples 1 to d k The eigenvalue of the dimension feature, the combined encoding of the eigenvalue of the i-th aligned sample of federated participant k is expressed as follows: It can be concluded that for any combination of feature values, the corresponding combined feature value code is uniquely determined; (5) The task publisher combines the feature value codes sent by the participants with its own feature value codes and label columns to calculate the values of N(·) in the following formula, where N(·) represents the number of data samples under a given variable value, and further calculates the mutual information estimate of the joint feature and label: For each pair of matching samples, the task publisher encodes F according to the combined feature values of the matching samples. i k The combined eigenvalue encoding F of the data itself i And the corresponding label y is counted, the frequency of occurrence is recorded, and the N(F i k ,F i ,y) and N(F i k ,F i ), N(F i k ,F i ,y) indicates that the feature encoding is F i k and F i And the number of samples with label y, (6) According to the task publisher’s training set D and test set T, the goal is converted to maximizing the expected model performance [U i ], including: According to the model accuracy prediction function and the distribution of estimation errors obtained in step (3), the expected model maximization problem is re-modeled as follows: Q k =f(I k ,n k )+e k ,k∈S e k ~N(μ,σ 2 ),k∈S n k ≤|D∩D k | T k =T∩D k Among them, the task publisher is O, its training budget is B, the set of federated participants is S, and for each federated participant k in S, its sample set D k The intersection with the test set of the task publisher is T k =T∩D k , T k The samples in have both k and the feature dimensions of the task publisher. For these samples, the model M trained by the task publisher and the federated participant k can be used. k Perform inference prediction, M k The estimated performance is Q k , variable n k is a decision variable corresponding to the number of training samples brought by federation participant k. k When the larger the value, the estimated model performance f(I k ,n k ) is larger, and the sample of federated participant k will occupy more quota in budget B, and the best reasoning performance obtained by test sample i in test set T is the best performance of all Q k The maximum value in ; According to steps (1) to (5), the model M trained by each federation participant k in cooperation with the task publisher is calculated. k The estimated accuracy Q k , all federated participants are assigned an estimated accuracy Q k Sort from high to low, if there is model M k If the accuracy of the task publisher is lower than the accuracy of the model M0 trained by the task publisher using only its own feature dimension data, it will be removed from the selection range, and then starting from the federation participant with the highest estimated accuracy, try to purchase training data from them in turn until the cumulative amount of data purchased exceeds the budget limit B or there are no more federation participants to choose from; finally, based on the purchase situation, inform the selected federation participant of the amount of training data it needs to provide, and carry out vertical federated learning training; (7) Consider the problem that after the federation participants are selected using the greedy strategy in step (6) before training, there are several federation participants and several longitudinal federated learning models will be obtained. In the model prediction and reasoning stage, it is necessary to perform the best model task allocation. According to the priority ranking of the federation participants, each federation participant is assigned the prediction task of the corresponding test set sample. Then, the prediction results of the test data are analyzed and compared with the true labels, and the proportion of the number of samples correctly classified by the model to the total number of samples is calculated to obtain the prediction accuracy of the longitudinal federated learning model, which includes: After substituting the encrypted mutual information and the number of sample intersections into the model accuracy estimation function to obtain the accuracy estimation value of the longitudinal federated learning model obtained by the cooperation between each federation participant and the task publisher, the corresponding priority is assigned to each federation participant by sorting the estimated values of the model accuracy. The federation participant with a higher model estimation accuracy has a higher priority. For each sample in the task publisher's test set, the system checks which federated participants have records related to the sample and assigns prediction tasks based on priority; For each sample ID in the test set, the system will check whether each federation participant has a record corresponding to the sample ID in order of priority. If a federation participant has a record corresponding to the sample ID, the system will mark the federation participant as "1", otherwise it will be marked as "0"; If only one federation participant has a record of the sample, the task publisher will choose to use the model trained in cooperation with this federation participant to predict the sample; If multiple federation participants have records corresponding to the sample ID, the model trained in collaboration with the federation participant with the highest priority is selected to be responsible for the prediction task of the sample.
2. The method for selecting federated participants based on model accuracy estimation in vertical federated learning according to claim 1, characterized in that: Step S1: The correlation analysis process for the number of training samples includes training them using three machine learning models, including random forest, logistic regression and support vector machine, and then recording the relationship between the number of training samples and the model prediction accuracy, and verifying that under the respective model and data set configurations, the model accuracy and the number of training samples are significantly positively correlated; The correlation analysis process of mutual information in this step includes: The mutual information MI of two random variables is a measure of the mutual dependence between the variables. The mutual information measures the information shared by two random variables X and Y and is represented by I(X; Y). The standard definition of mutual information is as follows: Assume that the feature set held by the task publisher of vertical federated learning is X, and the feature set held by federated participant k is X k , then the mutual information between all features and labels for federated learning can be expressed as I(X∪X k ; Y), this mutual information can measure X∪X k The mutual dependence between and Y, according to the definition of mutual information, I(X∪X k ; Y) is calculated as follows: where x k ,x and y represent X k , possible values of X and Y, p(x k ,x,y) is X=x,X k =x k ,Y=The probability that y satisfies at the same time, D KL It is the KL divergence between the two distributions. Since the true distribution of the data is difficult to obtain, we consider using the maximum likelihood estimation method to approximate them based on the feature distribution and label distribution in the training set, as shown in the following formula: In the formula, is the distribution of p(·) estimated from the training set, N(·) is the number of data samples under a given variable value, and n k is the total number of public data samples between federation participant k and task publisher; For continuous features, convert them into discrete ones, and then calculate the probability distribution by maximum likelihood estimation. If X k Or the dimension of X is high, including using PCA to reduce its dimension and convert it into a low-dimensional vector; The higher the mutual information between features and labels, the higher the contribution of the features to predicting labels, which helps improve the reasoning performance of the model, that is, the accuracy of label prediction.
3. The method for selecting federated participants based on model accuracy estimation in vertical federated learning according to claim 1, characterized in that: In step (2), the relationship between model performance and mutual information is first fitted, and then fine-tuned using the number of training samples, expressed as: Assume that the task publisher T has a total of z dimensions of features, and the task publisher T chooses to retain the feature combination f containing z' features t , the remaining zz′ features are used as the features of the federated participants, and then the task publisher randomly selects an integer in the range [1,zz′] Then randomly select from the remaining zz′ features The feature combination f of features d ; Then use f t As the feature set of the task publisher, f d As a feature set of federated participants, by dividing the feature dimension data of the task publisher, the task publisher can simulate and perform longitudinal federated learning alone and obtain the performance of the model before cooperating with other federated participants.
4. The method for selecting federated participants based on model accuracy estimation in vertical federated learning according to claim 1, characterized in that: In step (4), the model performance estimation function is applied to the actual longitudinal federated training process, which requires obtaining the mutual information and the number of aligned training samples. The number of samples is naturally obtained through data interaction, and the calculation of mutual information needs to be privacy-processed.
5. The method for selecting federated participants based on model accuracy estimation in vertical federated learning according to claim 4, characterized in that: In the mutual information process between the features and task labels involved in privacy computing training, the dimension reduction of the data features of the federation participants or task publishers is performed; The dimensionality reduction processing refers to PCA dimensionality reduction, and processing is performed on data with feature dimensions exceeding three dimensions to reduce data transmission overhead and calculation overhead.
6. The method for selecting federated participants based on model accuracy estimation in vertical federated learning according to claim 1, characterized in that: In step (42), the task publisher and the federated participant k first encrypt their own sample IDs with their own private keys e and d and send them to each other. Then, the federated participant k encrypts the sample ID encrypted by the task publisher again with its own private key e and sends it to the task publisher. Based on the Diffie-Hellman encryption algorithm, a symmetric encryption algorithm is used for encryption and decryption. The task publisher only needs to compare the sample ID of the federated participant k encrypted with the private keys of e and d and the sample ID of the task publisher encrypted with the private keys of d and e to see if they are consistent to obtain the sample intersection.
7. The method for selecting federated participants based on model accuracy estimation in vertical federated learning according to claim 1, characterized in that: In step (7), it is assumed that there are several federated participants, which have different sample intersections with the task publisher T. In the model inference phase, the common samples between the federated participant k and the task publisher can use the model M trained by both of them. k ; For samples shared by some federation participants and task publishers, the estimated performance Q is selected k A relatively better model is used to make predictions; for some samples that are only held by the task publisher, the model M0 trained by the task publisher itself is used to make predictions.
8. The method for selecting federated participants based on model accuracy estimation in vertical federated learning according to claim 1 or 7, characterized in that: The prediction accuracy of the longitudinal federated learning model described in this method is calculated according to the following formula: Among them, TP represents the number of samples correctly predicted by the model as positive, TN represents the number of samples correctly predicted by the model as negative, FP represents the number of samples incorrectly predicted by the model as negative, and FN represents the number of samples incorrectly predicted by the model as positive.
Citation Information
Patent Citations
Federal learning client contribution evaluation method based on significant score
CN115905859A
Longitudinal federal learning participant selection method and device, equipment and storage medium
CN118468984A