Longitudinal federal prediction optimization method and device, equipment, storage medium and product
By using business label prediction models and residual prediction models in vertical federated learning systems and sharing data through sample alignment, the problem of privacy leakage risk in vertical federated learning is solved while improving prediction accuracy.
Patent Information
- Application Number
- CN202311725306.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-17
AI Technical Summary
In vertical federated learning, participants can infer the tag information of the test sample of the tag holder based on the local model obtained by training, thus bringing serious privacy leakage risks.
In the vertical federated learning system, a business tag prediction model is deployed on the first device and a residual prediction model is deployed on the second device. Through sample alignment, the first device shares data of the aligned sample with the second device, the second device uses its residual prediction model to make predictions and sends the prediction results back to the first device. The first device updates the model it deploys based on the received prediction results and the output of the local model, thereby improving prediction accuracy while avoiding the leakage of private data.
Through this method, only the training residual prediction results and model gradients are shared between the first device and the second device, avoiding privacy data leakage to the participants and the tag holders, while improving the accuracy of business tag predictions.
Smart Images

Figure CN120163206A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of artificial intelligence in financial technology (Fintech), and particularly relates to a vertical federated prediction optimization method, device, equipment, storage medium and product. Background Technique
[0002] With the continuous development of financial technology, especially Internet technology finance, more and more technologies (such as distributed, artificial intelligence, etc.) are applied in the financial field, but the financial industry also puts forward higher requirements for technologies.
[0003] The application of artificial intelligence in the financial industry is becoming more and more extensive. The training of models often requires a large amount of sample business data. Federated learning can effectively expand the scale of sample business data, thereby improving the performance of the model. Vertical federated learning is to take out the part of users and data with the same users but different sample business data characteristics for joint machine learning training when the data characteristics of the participants overlap less and the user overlap is more.
[0004] However, in the related technology, each participating party of vertical federated learning cooperates with the label holder to train a vertical federated model, and the vertical federated model is used to perform label prediction for the label holder. The label holder holds the labels, while the participating party only holds the auxiliary features. However, in this way, the participating party can infer the label information of the test samples of the label holder based on the locally trained model, thus bringing a serious risk of privacy leakage. At present, there are two categories of methods for protecting the privacy of the actively participating party's labels. One is the privacy protection method based on cryptography, which uses technologies such as homomorphic encryption / multi-party secure computing to protect data privacy. However, this type of solution has a huge computational overhead. The other is the privacy protection method based on perturbation, which protects privacy by adding noise to the data. However, this will cause a significant decline in the performance of the federated learning model. Summary of the Invention
[0005] The main purpose of the present application is to provide a vertical federated prediction optimization method, device, equipment, storage medium and product, aiming to solve the technical problem of relatively high privacy leakage risk in vertical federated learning in the related technology.
[0006] To achieve the above object, the present application provides a vertical federated prediction optimization method, which is applied to a first device in a vertical federated learning system, and a business label prediction model is deployed on the first device; the vertical federated learning system further includes a second device, and a residual prediction model is deployed on the second device; the vertical federated prediction optimization method includes the following steps:
[0007] Perform sample alignment with the second device to determine aligned samples;
[0008] Obtain the first-party training sample service data and the first sample service label of the aligned sample, and obtain the service label training prediction result obtained by the trained service label prediction model for service label prediction based on the first-party training sample service data;
[0009] Determine the service label prediction residual according to the first sample service label and the service label training prediction result;
[0010] Receive the training residual prediction result sent by the second device, where the training residual prediction result is obtained by the second device through the residual prediction model for residual prediction based on the second-party training sample service data of the aligned sample;
[0011] Based on the training residual prediction result and the service label prediction residual, determine the residual prediction model gradient corresponding to the second device, and send the residual prediction model gradient to the second device for the second device to update the locally deployed residual prediction model based on the received residual prediction model gradient.
[0012] This application also provides a vertical federated prediction optimization method, which is applied to the second device in a vertical federated learning system. A residual prediction model is deployed on the second device, and the vertical federated learning system further includes a first device on which a service label prediction model is deployed; the vertical federated prediction optimization method includes the following steps:
[0013] Perform sample alignment with the first device to determine the aligned sample, and obtain the second-party training sample service data of the aligned sample;
[0014] Input the second-party training sample service data into the residual prediction model for residual prediction to obtain the training residual prediction result;
[0015] Send the training residual prediction result to the first device for the first device to obtain the first-party training sample service data and the first sample service label of the aligned sample, and obtain the service label training prediction result obtained by the trained service label prediction model for service label prediction based on the first-party training sample service data. Determine the service label prediction residual according to the first sample service label and the service label training prediction result, and determine the residual prediction model gradient corresponding to the second device based on the training residual prediction result and the service label prediction residual;
[0016] Receive the residual prediction model gradient sent by the first device, and update the residual prediction model based on the residual prediction model gradient.
[0017] The present application also provides a vertical federated prediction optimization method, which is applied to a third device and includes the following steps:
[0018] Obtain the first-party business data of the sample to be predicted of the sample to be predicted, and obtain the local business label prediction result obtained by the trained business label prediction model for business label prediction based on the first-party business data of the sample to be predicted;
[0019] Perform sample alignment with at least one second device, so that the target second device containing the second-party business data of the sample to be predicted corresponding to the sample to be predicted in each second device performs residual prediction based on the second-party business data of the sample to be predicted through the residual prediction model deployed by itself, and obtain a residual prediction result, wherein the residual prediction model is trained by using the vertical federated prediction optimization method as described above;
[0020] Receive the residual prediction results sent by each target second device, and aggregate the local business label prediction result and the residual prediction results to obtain a federated business label prediction result.
[0021] The present application also provides a vertical federated prediction optimization device, which is applied to a first device in a vertical federated learning system, and a business label prediction model is deployed on the first device; the vertical federated learning system further includes a plurality of second devices, and a residual prediction model is deployed on each second device; the vertical federated prediction optimization device includes:
[0022] A first sample alignment module, configured to perform sample alignment with the second device to determine aligned samples;
[0023] An acquisition module, configured to acquire the first-party training sample business data and the first sample business label of the aligned samples, and acquire the business label training prediction result obtained by the trained business label prediction model for business label prediction based on the first-party training sample business data;
[0024] A business label prediction residual determination module, configured to determine a business label prediction residual according to the first sample business label and the business label training prediction result;
[0025] A receiving module, configured to receive the training residual prediction result sent by the second device, where the training residual prediction result is obtained by the second device performing residual prediction based on the second-party training sample business data of the aligned samples through the residual prediction model;
[0026] A gradient determination module, configured to determine a residual prediction model gradient corresponding to the second device based on the training residual prediction result and the service label prediction residual, and send the residual prediction model gradient to the second device for the second device to update the locally deployed residual prediction model based on the received residual prediction model gradient.
[0027] This application also provides a vertical federated prediction optimization device, which is applied to a second device in a vertical federated learning system. A residual prediction model is deployed on the second device, and the vertical federated learning system further includes a first device on which a service label prediction model is deployed. The vertical federated prediction optimization device includes:
[0028] A second sample alignment module, configured to perform sample alignment with the first device to determine aligned samples, and obtain second-party training sample service data of the aligned samples;
[0029] A residual prediction module, configured to input the second-party training sample service data into the residual prediction model to perform residual prediction, and obtain a training residual prediction result;
[0030] A sending module, configured to send the training residual prediction result to the first device for the first device to obtain first-party training sample service data and first sample service labels of the aligned samples, and obtain a service label training prediction result obtained by the trained service label prediction model performing service label prediction based on the first-party training sample service data. According to the first sample service label and the service label training prediction result, determine a service label prediction residual, and determine a residual prediction model gradient corresponding to the second device based on the training residual prediction result and the service label prediction residual;
[0031] An update module, configured to receive the residual prediction model gradient sent by the first device, and update the residual prediction model based on the residual prediction model gradient.
[0032] This application also provides a vertical federated prediction optimization device, which is applied to a third device. The vertical federated prediction optimization device includes:
[0033] A prediction module, configured to obtain first-party to-be-predicted sample service data of a to-be-predicted sample, and obtain a local service label prediction result obtained by the trained service label prediction model performing service label prediction based on the first-party to-be-predicted sample service data;
[0034] A third sample alignment module, configured to perform sample alignment with at least one second device, so that for a target second device among the second devices that contains second-party to-be-predicted sample service data corresponding to the to-be-predicted sample, a residual prediction is performed based on the second-party to-be-predicted sample service data through a residual prediction model deployed by itself, and a residual prediction result is obtained, where the residual prediction model is trained by using the vertical federated prediction optimization method as described above;
[0035] An aggregation module, configured to receive the residual prediction results sent by the target second devices, and aggregate the local service label prediction result and the residual prediction results to obtain a federated service label prediction result.
[0036] The present application further provides an electronic device, which is a physical device. The electronic device includes: a memory, a processor, and a program of the vertical federated prediction optimization method stored on the memory and executable on the processor. When the program of the vertical federated prediction optimization method is executed by the processor, the steps of the vertical federated prediction optimization method as described above can be implemented.
[0037] The present application further provides a storage medium, which is a computer-readable storage medium. A program for implementing the vertical federated prediction optimization method is stored on the computer-readable storage medium. When the program of the vertical federated prediction optimization method is executed by a processor, the steps of the vertical federated prediction optimization method as described above are implemented.
[0038] The present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the vertical federated prediction optimization method as described above are implemented.
[0039] The present application provides a vertical federated prediction optimization method, apparatus, device, storage medium and product. The vertical federated prediction optimization method is applied to a first device in a vertical federated learning system. A business label prediction model is deployed on the first device. The vertical federated learning system further includes a plurality of second devices, and a residual prediction model is deployed on each of the second devices. First, sample alignment is performed with the second devices to determine aligned samples, and first-party training sample business data and first sample business labels of the aligned samples are obtained. Also, a business label training prediction result obtained by the trained business label prediction model performing business label prediction based on the first-party training sample business data is obtained. According to the first sample business labels and the business label training prediction results, business label prediction residuals are determined. The business label prediction residuals thus determined are the true values of the residuals generated by the trained business label prediction model for model prediction, thus achieving the purpose of determining the true values of the residuals. Furthermore, training residual prediction results sent by the second devices are received, where the training residual prediction results are obtained by the second devices performing residual prediction based on second-party training sample business data of the aligned samples through the residual prediction model. In this way, the second-party training sample business data owned by the second devices can be used for residual prediction, achieving the purpose of determining the predicted values of the residuals. Furthermore, based on the training residual prediction results and the business label prediction residuals, a residual prediction model gradient corresponding to the second devices is determined, and the residual prediction model gradient is sent to the second devices for the second devices to update the residual prediction models deployed on their own sides based on the received residual prediction model gradient, achieving the update of the residual prediction models on the second devices. In this way, during the entire training process, only the training residual prediction results and the residual prediction model gradients are shared between the first device and the second devices. For the second devices, since the first device cannot know which sample features the second devices are based on to obtain the training residual prediction results, it is impossible to infer the second-party training sample business data owned by the second devices, so privacy protection for the second devices can be achieved. For the first device, since the residual is the difference between the first sample business label and the business label training prediction result and is an embodiment of the model performance, it is difficult for the second device to infer the label information or privacy data owned by the first device only based on the residual when the second device neither knows the first-party training sample business data nor the first sample business label nor the business label prediction model, so privacy protection for the first device can be achieved. Therefore, it overcomes the technical defect that each participating party in vertical federated learning cooperates with the label holder to train a vertical federated model, and then uses the vertical federated model to perform label prediction for the label holder. In this way, the participating parties can infer the labels of the test samples of the label holder based on the locally trained models, thus bringing a serious privacy leakage risk.Moreover, compared with the service label prediction model trained only based on local data, vertical federated learning can utilize diverse sample service data distributed across multiple second devices to more accurately predict the residual between the prediction result and the true result of the service label prediction model, thereby correcting the prediction result of the service label prediction model, that is, improving the accuracy of the finally output prediction result. Therefore, it can not only improve prediction accuracy but also avoid privacy leakage. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0041] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or related technologies. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 Schematic flowchart of the first embodiment of the vertical federated prediction optimization method of the present application;
[0043] Figure 2 Schematic scenario diagram of an implementable manner of the federated learning system in the embodiments of the present application;
[0044] Figure 3 Schematic flowchart of the second embodiment of the vertical federated prediction optimization method of the present application;
[0045] Figure 4 Schematic flowchart of the third embodiment of the vertical federated prediction optimization method of the present application;
[0046] Figure 5 Schematic structural diagram of the vertical federated prediction optimization device in the embodiments of the present application;
[0047] Figure 6 Schematic structural diagram of the device of the hardware operating environment involved in the vertical federated prediction optimization method in the embodiments of the present application.
[0048] The implementation, functional features, and advantages of the objectives of the present application will be further described in conjunction with the embodiments with reference to the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] To make the above objects, features, and advantages of the present invention more apparent and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0050] Artificial intelligence is increasingly widely used in the financial industry. The training of models often requires a large amount of sample business data. Each enterprise can obtain limited sample business data, and the features are relatively single. The user portraits depicted may be inaccurate. If each enterprise only trains a prediction model based on its own local sample business data, the performance of the obtained prediction model is limited, and the prediction accuracy is relatively low. Therefore, federated learning can be used to expand the scale of sample business data, thereby improving the model performance.
[0051] Vertical federated learning is to extract the part of users and data with the same users but different sample business data features when the data features of the participating parties overlap less and the user overlap is more. In a vertical federated learning system, it usually includes a participating party with label information and a participating party without label information. For the convenience of description, in the subsequent embodiments, the first device refers to the participating party with label information, and the first device needs to perform a business label prediction task. The second device refers to the participating party without label information, and the second device can assist the first device to complete the business label prediction task and improve the prediction accuracy of the business label prediction task.
[0052] Exemplarily, the service label prediction task may be the prediction of information click-through rate. The first device may be a device of an information recommendation platform, such as a video website, a short video platform, etc. The second device may be a device of an institution such as an e-commerce company, a financial institution, a bank, etc. The information recommendation platform may cooperate with institutions such as e-commerce companies, financial institutions, and banks to jointly predict the information click behavior of common users. Thus, the information recommendation platform can recommend information to users more accurately. The information recommendation platform has the click records of users and can use the click records as sample service labels for model training. The information recommendation platform may also have first-party sample service data such as users' information browsing behavior and page stay time. These first-party sample service data can be used for local training of the service label prediction model deployed on its own side, so as to better predict the information click-through rate. The second device may have other sample service data of the same users as the information recommendation platform, such as purchase behavior, purchase preferences, financial management information, etc. These second-party sample service data can provide auxiliary features for the service label prediction model deployed on the first device, so as to correct the information click-through rate predicted by the service label prediction model deployed on the first device and obtain a more accurate information click-through rate. In this process, neither the information recommendation platform nor institutions such as e-commerce companies, financial institutions, and banks want to disclose the privacy data of users.
[0053] In the related art, the first device and each second device cooperate to train a vertical federated model together. The vertical federated model is used to perform label prediction for the first device. The label holder holds the labels, while the participating party only holds the auxiliary features. However, in this way, the participating party can infer the label information of the test samples of the label holder based on the trained local model, thus bringing a serious privacy leakage risk. Currently, there are two categories of methods for protecting the privacy of the actively participating label. One is the privacy protection method based on cryptography, which uses technologies such as homomorphic encryption / multi-party secure computing to protect data privacy. However, this type of solution has a huge computational overhead. The other is the privacy protection method based on perturbation, which protects privacy by adding noise to the data. However, this will cause a significant decline in the performance of the federated learning model.
[0054] Currently, the methods for protecting the privacy of the first device's labels include privacy protection methods based on cryptography, that is, using technologies such as homomorphic encryption or multi-party secure computing to protect data privacy. However, this type of solution has a huge computational overhead. In addition, the methods for protecting the privacy of the first device's labels also include privacy protection methods based on perturbation, that is, protecting privacy by adding noise to the data. However, this will cause a significant decline in the performance of the vertical federated learning model.
[0055] In this application, a business label prediction model is deployed on a first device, and a residual prediction model is deployed on a second device. The entire business label prediction task is divided into two parts: local business label prediction and result correction based on residuals. In this way, the first device can train the business label prediction model only based on the sample business data it locally has. At the same time, due to the limited sample business data locally available, the performance of the business label prediction model is limited and the prediction accuracy is relatively low. There is still a large prediction residual between the prediction result obtained by using the business label prediction model to predict all aligned data and the true result. Therefore, the sample business data of the samples aligned between the second device and the first device is used, and through the residual prediction model on the second device, the prediction residual of the business label prediction model on the first device is predicted. By jointly correcting the prediction result of the business label prediction model with the prediction residual of the second device, a federated business label prediction result closer to the true result can be obtained, thereby achieving the purpose of improving the prediction effect. Moreover, since a residual prediction model is deployed on the second device, during the entire training process, only the training residual prediction result and the residual prediction model gradient are shared between the first device and the second device. For the second device, since the first device cannot know which sample features the second device used to obtain the training residual prediction result, it is impossible to infer the second-party training sample business data owned by the second device. Therefore, privacy protection for the second device can be achieved; for the first device, since the residual is the difference between the first sample business label and the business label training prediction result, which reflects the model performance, in the case where the second device neither knows the first-party training sample business data nor the first sample business label nor the business label prediction model, it is difficult for the second device to infer the label information or privacy data owned by the first device only based on the residual. Therefore, privacy protection for the first device can be achieved. Thus, it overcomes the technical defect that the participants in vertical federated learning and the label holder jointly cooperate to train a vertical federated model, and then use the vertical federated model for label prediction for the label holder. In this way, the participants can infer the labels of the test samples of the label holder based on the locally trained model, thus bringing a serious privacy leakage risk. Moreover, compared with the business label prediction model trained only based on local data, vertical federated learning can use a variety of sample business data distributed on the second device to more accurately predict the residual between the prediction result of the business label prediction model and the true result, thereby correcting the prediction result of the business label prediction model, that is, improving the accuracy of the finally output prediction result. Therefore, both the prediction accuracy can be improved and privacy leakage can be avoided.
[0056] Embodiment 1
[0057] The embodiments of the present application provide a vertical federated prediction optimization method. In the first embodiment of the vertical federated prediction optimization method of the present application, the vertical federated prediction optimization method is applied to a first device in a vertical federated learning system, and a business label prediction model is deployed on the first device; the vertical federated learning system further includes a second device, and a residual prediction model is deployed on the second device; referring to Figure 1 , the vertical federated prediction optimization method includes the following steps:
[0058] Step S10, perform sample alignment with the second device to determine aligned samples;
[0059] The execution subject of the method in this embodiment can be a vertical federated prediction optimization device, or a vertical federated prediction optimization terminal device or server. In this embodiment, a vertical federated prediction optimization device is used as an example, and this vertical federated prediction optimization device can be integrated on terminal devices such as smartphones and computers with data processing functions.
[0060] In this embodiment, it should be noted that the vertical federated learning system includes at least one first device and at least one second device corresponding to each first device. For the convenience of understanding, in the following, one first device and a second device corresponding to the first device are used for description. The vertical federated prediction optimization method includes the training process of the vertical federated learning model. The vertical federated learning model includes a business label prediction model deployed on the first device and a residual prediction model deployed on the second device. Among them, the business label prediction model can be determined according to the actual business label prediction task and the form of sample data. For example, a convolutional neural network model can be selected for image data, and a classification model can be selected for classification tasks, etc. This embodiment does not limit this; the residual prediction model can be an improved model based on residual learning, and the business label prediction residual of the business label prediction model on the first device is fitted by executing the forward propagation algorithm. Therefore, the training process of the vertical federated learning model needs to be jointly completed by the first device and the second device, and the vertical federated prediction optimization method provided in this embodiment is applied to the first device in the vertical federated learning system.
[0061] During the training process of the vertical federated learning model, the service label prediction model is trained first, and then the trained service label prediction model is used to train the residual prediction model. The training of the service label prediction model can be carried out on the first device or on other devices, and then the trained service label prediction model is deployed to the first device, and the residual prediction model is trained through the first device. This embodiment does not limit this. The service label prediction model is used to predict service labels based on the sample service data owned by the first device. After the service label prediction model is trained, the service label prediction residual between the service label training prediction result obtained by the service label prediction model performing the service label prediction task and the first sample service label owned by the first device can be used as the true value for the residual prediction model on the second device to perform residual prediction; the residual prediction model is used to predict the service label prediction residual of the service label prediction model on the first device to obtain a predicted value of the service label prediction residual; thus, the residual prediction model on the second device can be updated by the gradient descent method.
[0062] As an example, step S10 includes: the first device and the second device perform sample alignment to determine aligned samples. Among them, the sample alignment relationship between each first-party training sample service data and each second-party training sample service data can be determined by aligning sample IDs, or the sample alignment relationship between each first-party training sample service data and each second-party training sample service data can be determined by other identification information of the aligned samples. Specifically, it can be determined according to the actual situation. This embodiment does not limit this. For the convenience of description, the following will take the aligned sample ID as an example for description.
[0063] In an implementable manner, there may be multiple second devices. In this case, the aligned sample refers to the sample aligned by the first device and at least one second device, that is, the first device has the sample data of all the aligned samples, and each second device may have part or all of the sample data of the aligned samples.
[0064] Step S20, obtain the first-party training sample service data and the first sample service label of the aligned samples, and obtain the service label training prediction result obtained by the trained service label prediction model performing service label prediction based on the first-party training sample service data;
[0065] In this embodiment, it should be noted that the first device has first-party training sample service data of multiple aligned samples, and first sample service labels corresponding to each of the aligned samples; for each second device, it may have second-party training sample service data of some or all of the aligned samples, and through sample alignment, the second device can extract second-party training sample service data of the aligned samples from the sample service data it owns. For example, the first device has deposit data of aligned sample S1, aligned sample S2, aligned sample S3, and aligned sample S4, the first-party training sample service data is the deposit data of aligned sample S1, aligned sample S2, aligned sample S3, and aligned sample S4, the second device P1 has consumption data of S1, S2, S3, and non-aligned sample S5, and the second device P2 has credit investigation data of S1, S4, and non-aligned sample S6. Through sample alignment, it can be determined that the second-party training sample service data owned by P1 is the consumption data of S1, S2, and S3, and the second-party training sample service data owned by P2 is the credit investigation data of S1 and S4.
[0066] As an example, step S20 includes: obtaining a batch of first-party training sample service data, where the batch of first-party training sample task data includes first-party training sample service data of at least one aligned sample, and obtaining first sample service labels corresponding to each of the aligned samples; further, inputting each of the first-party training sample service data into the trained service label prediction model for service label prediction to obtain service label training prediction results corresponding to each of the aligned samples.
[0067] In an implementable manner, the service label prediction model may include a feature extractor, a fully connected layer, and an activation layer. For each training sample service data, after inputting the first-party training sample service data into the trained service label prediction model, the feature extractor can be used to extract features from the first-party training sample service data to obtain first-party training sample features. Further, each of the first-party training sample features can be concatenated into a first-party training sample feature vector or a first-party training sample feature matrix, and the fully connected layer is used to perform a full connection on the first-party training sample feature vector or the first-party training sample feature matrix. Then, the output of the fully connected layer is activated through the activation function preset by the activation layer to obtain the service label training prediction result.
[0068] Further, before the steps of obtaining the first-party training sample service data and the first sample service labels of the aligned samples, and obtaining the service label training prediction results obtained by the trained service label prediction model for service label prediction based on the first-party training sample service data, the following steps are also included:
[0069] Obtain the service label prediction model training sample service data and the second sample service label, and iteratively optimize the service label prediction model based on the service label prediction model training sample service data and the second sample service label to obtain a trained service label prediction model.
[0070] In this embodiment, it should be noted that since the true value of the residual prediction by the residual prediction model on the second device is determined by the service label prediction model on the first device, before iteratively optimizing the residual prediction model on the second device, it is necessary to first complete the training of the service label prediction model to ensure the correctness of the true value used for training the residual prediction model.
[0071] Exemplarily, before iteratively optimizing the residual prediction model on the second device, first initialize the service label prediction model; then obtain a batch of service label prediction model training sample service data and the corresponding second sample service label from the local sample data of the first device, input the service label prediction model training sample service data into the service label prediction model to perform service label prediction, obtain the service label training prediction result, calculate the service label prediction model loss based on the difference between the service label training prediction result and the second sample service label, determine whether the service label prediction model loss converges. If the service label prediction model loss converges, it is determined that the service label prediction model training is completed. If the service label prediction model loss does not converge, perform one round of update on the service label prediction model according to the service label prediction model gradient calculated from the service label prediction model loss, and return to execute the step of obtaining the service label prediction model training sample service data and the second sample service label until the service label prediction model converges to obtain a trained service label prediction model.
[0072] Step S30: Determine the service label prediction residual according to the first sample service label and the service label training prediction result;
[0073] In this embodiment, it should be noted that the first sample service label, the service label training prediction result, and the service label prediction residual are in one-to-one correspondence with each alignment sample. The service label prediction residual corresponding to each alignment sample is the true value of the residual generated by the service label prediction model for performing service label prediction on this alignment sample.
[0074] As an example, step S30 includes: calculating the service label prediction residual corresponding to each alignment sample according to the first sample service label corresponding to each alignment sample and the service label training prediction result corresponding to each alignment sample.
[0075] Step S40: Receive the training residual prediction result sent by the second device, where the training residual prediction result is obtained by the second device through residual prediction based on the second-party training sample service data of the aligned samples using the residual prediction model.
[0076] As an example, Step S40 includes: After sample alignment, the second device can, based on the aligned sample ID, search for the second-party training sample service data corresponding to the aligned sample ID from the sample service data it owns. Then, input each piece of the second-party training sample service data into the residual prediction model deployed on its own side for residual learning. Based on the forward propagation algorithm, at least one training residual prediction result is calculated. The training residual prediction result is the predicted value of the residual generated by the service label prediction model for the aligned samples. Then, send each of the training residual prediction results to the first device. Therefore, the first device can receive the training residual prediction results sent by the second device. It should be noted that the training residual prediction results also correspond one-to-one with the aligned sample IDs. Therefore, the training residual prediction results can be corresponded one-to-one with the service label prediction residuals determined on the first device based on the aligned sample IDs.
[0077] Step S50: Based on the training residual prediction result and the service label prediction residual, determine the residual prediction model gradient corresponding to the second device, and send the residual prediction model gradient to the second device for the second device to update the residual prediction model deployed on its own side based on the received residual prediction model gradient.
[0078] As an example, Step S50 includes: According to the aligned sample IDs, match each of the training residual prediction results with each of the service label prediction residuals, and calculate the loss function of the residual prediction loss of the residual prediction model on the second device based on the difference value between the mutually matched training residual prediction result and the service label prediction residual. Then, calculate the residual prediction model gradient of the loss function corresponding to the second device with respect to the training residual prediction result, and send the residual prediction model gradient to the second device for the second device to update the residual prediction model deployed on its own side using the gradient descent method based on the received residual prediction model gradient.
[0079] In an implementable manner, there may be multiple second devices. In this case, the loss functions corresponding to each of the second devices may be the same or different. That is, the method for calculating the residual prediction loss of the residual prediction model on each of the second devices based on the difference value between the mutually matched training residual prediction result and the service label prediction residual may be: calculating the loss function of the residual prediction loss of the residual prediction model on each second device respectively according to the difference value between the mutually matched training residual prediction result and the service label prediction residual; or determining the difference value between the mutually matched training residual prediction result and the service label prediction residual, and then averaging and aggregating the difference values corresponding to all the aligned samples to determine the federated difference value, and then calculating the loss function based on the federated difference value, so that the loss functions of each second device determined in this way are the same.
[0080] In this way, one round of iterative update of the residual prediction model can be completed. By performing multiple rounds of iterative update on the residual prediction model until the preset federated training end condition is met. Among them, the preset federated training end condition may be that among all the second devices participating in federated learning, the residual prediction models in more than the preset number of second devices converge, or reach the preset maximum number of iterations of federated learning, or reach the preset maximum training time of federated learning, etc. Specifically, it can be determined according to the actual situation, and this embodiment does not limit this.
[0081] Further, the step of determining the gradient of the residual prediction model corresponding to the second device based on the training residual prediction result and the service label prediction residual includes:
[0082] Step A10, determining the service label prediction residual corresponding to each of the training residual prediction results based on the sample alignment relationship between each first-party training sample service data and each second-party training sample service data;
[0083] Step A20, calculating the initial gradient of the residual prediction model corresponding to each second-party training sample service data based on the difference between each of the training residual prediction results and the service label prediction residual corresponding to each of the training residual prediction results;
[0084] Step A30, determining the average value of each of the initial gradients of the residual prediction models as the gradient of the residual prediction model corresponding to the second device.
[0085] In this embodiment, it should be noted that the number of alignment samples in which the second device is aligned with the first device can be multiple. In this case, there are also multiple training residual prediction results that the second device can determine. For each training residual prediction result, a corresponding business label prediction residual can be determined. Therefore, for each alignment sample of each second device, a set of training residual prediction results and business label prediction residuals can be determined. Therefore, for each alignment sample of each second device, a corresponding initial gradient of the residual prediction model can be determined, and then aggregated by taking the average value to determine the gradient of the residual prediction model corresponding to each second device.
[0086] As an example, the steps A10 - A30 include: receiving the training residual prediction results sent by the second device, matching each of the training residual prediction results with each of the business label prediction residuals according to the alignment sample ID, determining the business label prediction residuals corresponding to each of the training residual prediction results, calculating the loss function of the residual prediction loss of the residual prediction model on the second device according to the difference value between the mutually matched training residual prediction result and the business label prediction residual, and then calculating the initial gradient of the residual prediction model of each of the loss functions with respect to the corresponding training residual prediction result, calculating the average value of the initial gradients of the residual prediction models corresponding to each second device, and determining the average value of the initial gradients of the residual prediction models corresponding to each second device as the gradient of the residual prediction model corresponding to the second device.
[0087] In an implementable manner, in the case where the number of alignment samples in which the second device is aligned with the first device is multiple, the step of determining the gradient of the residual prediction model corresponding to the second device based on the training residual prediction result and the business label prediction residual may also include: receiving the training residual prediction results sent by the second device, matching each of the training residual prediction results with each of the business label prediction residuals according to the alignment sample ID, determining the business label prediction residuals corresponding to each of the training residual prediction results, and determining the difference value between the mutually matched training residual prediction result and the business label prediction residual, so as to calculate the average difference value corresponding to each second device and the average training residual prediction result corresponding to each second device, calculating the loss function of the residual prediction loss of the residual prediction model on the second device according to the average difference value, and then calculating the gradient of the residual prediction model of each second device with respect to the average training residual prediction result of each second device.
[0088] Further, the step of determining the gradient of the residual prediction model corresponding to the second device based on the training residual prediction result and the service label prediction residual, and sending the gradient of the residual prediction model to the second device for the second device to update the locally deployed residual prediction model based on the received gradient of the residual prediction model includes:
[0089] Step B10, determine whether the current condition meets the end condition of federated training;
[0090] Step B20, when it is determined that the current condition does not meet the end condition of federated training, determine the gradient of the residual prediction model corresponding to the second device based on each of the training residual prediction results and the service label prediction residual;
[0091] Step B30, send the gradient of the residual prediction model to the second device for the second device to update the locally deployed residual prediction model based on the received gradient of the residual prediction model, and return to execute the step of obtaining the first-party training sample service data and the first sample service label of the aligned sample.
[0092] In this embodiment, it should be noted that after receiving the training residual prediction result sent by the second device each time, it can be first determined whether the current condition meets the end condition of federated training. When the current condition does not meet the end condition of federated training, it is also necessary to further calculate the loss and gradient to iteratively optimize the residual prediction model. When the current condition meets the end condition of federated training, it means that the residual prediction model has been trained.
[0093] The preset end condition of federated training may be that among all the second devices participating in federated learning, the residual prediction models in more than the preset number of second devices converge, or reach the preset maximum number of iterations of federated learning, or reach the preset maximum training time of federated learning, etc. It can be specifically determined according to the actual situation, and this embodiment does not limit this.
[0094] As an example, the steps B10 - B30 include: receiving the training residual prediction results sent by the second device. Then, at least one of the current iteration number, the current training time, and the current model convergence condition can be used to determine whether the current condition for ending the federated training is met. When it is determined that the current condition for ending the federated training is not met, according to the aligned sample IDs, each of the training residual prediction results is matched with each of the service label prediction residuals, and based on the difference value between the mutually matched training residual prediction result and the service label prediction residual, a loss function for the residual prediction loss of the residual prediction model on the second device is calculated. Furthermore, the residual prediction model gradient of the loss function corresponding to the second device with respect to its corresponding training residual prediction result is calculated, and the residual prediction model gradient is sent to the second device for the second device to update the locally deployed residual prediction model using the gradient descent method based on the received residual prediction model gradient, and then return to execute the step of obtaining the first - party training sample service data and the first sample service label of the aligned sample, and perform the next round of iterative update until the condition for ending the federated training is met, thus completing the training of the residual prediction model.
[0095] Further, there are multiple second devices, and the step of determining whether the current condition for ending the federated training is met includes:
[0096] Step B11, judging whether the residual prediction models deployed on each of the second devices converge according to the differences between the training residual prediction results corresponding to each of the second devices and the corresponding service label prediction residuals;
[0097] Step B12, when the number of detected converged residual prediction models exceeds a preset number threshold, determining that the current condition for ending the federated training is met.
[0098] As an example, the steps B11 - B12 include: after receiving the training residual prediction results sent by each of the second devices, the differences between the training residual prediction results corresponding to each second device and the service label prediction residuals corresponding to each of the training residual prediction results can be used to judge whether the residual prediction models deployed on each of the second devices converge, and count the number of converged residual prediction models. When the number of detected converged residual prediction models exceeds the preset number threshold, it is determined that the current condition for ending the federated training is met. When the number of detected converged residual prediction models does not exceed the preset number threshold, it is determined that the current condition for ending the federated training is not met.
[0099] In an implementable manner, referring to Figure 2, the vertical federated learning system includes a first device and N second devices. Each of the second devices is communicatively connected to the first device. After sample alignment, each of the second devices can perform residual prediction on the second training sample service data of the aligned federated training samples through a residual prediction model, obtaining training residual prediction results R1, R2, R3, ……, R N , and sending the training residual prediction results R1, R2, R3, ……, R N to the first device. After the first device receives the training residual prediction results R1, R2, R3, ……, R N , based on the training residual prediction results R1, R2, R3, ……, R N and each of the service label prediction residuals, determine the residual prediction model gradients G1, G2, G3, ……, G N corresponding to each of the second devices respectively, and send the residual prediction model gradients G1, G2, G3, ……, G N to their respective corresponding second devices respectively, for each of the second devices to update the residual prediction models M1, M2, M3, ……, M N deployed by themselves based on the received residual prediction model gradients.
[0100] In this embodiment, the vertical federated prediction optimization method is applied to a first device in a vertical federated learning system. A business label prediction model is deployed on the first device. The vertical federated learning system further includes a plurality of second devices, and a residual prediction model is deployed on each of the second devices. First, sample alignment is performed with the second devices to determine aligned samples, and the first-party training sample business data and the first sample business label of the aligned samples are obtained. Also, the business label training prediction result obtained by the trained business label prediction model through business label prediction based on the first-party training sample business data is obtained. According to the first sample business label and the business label training prediction result, the business label prediction residual is determined. The business label prediction residual determined in this way is the true value of the residual generated by the trained business label prediction model for model prediction. Therefore, the purpose of determining the true value of the residual is achieved. Furthermore, the training residual prediction result sent by the second device is received. The training residual prediction result is obtained by the second device through residual prediction based on the second-party training sample business data of the aligned samples by the residual prediction model. In this way, the second-party training sample business data owned by the second device can be used for residual prediction, and the purpose of determining the predicted value of the residual is achieved. Furthermore, based on the training residual prediction result and the business label prediction residual, the residual prediction model gradient corresponding to the second device is determined, and the residual prediction model gradient is sent to the second device for the second device to update the locally deployed residual prediction model based on the received residual prediction model gradient, thus achieving the update of the residual prediction model on the second device. In this way, during the entire federated training process, only the training residual prediction result and the residual prediction model gradient are shared between the first device and the second device. For the second device, since the first device cannot know which sample features the second device is based on to obtain the training residual prediction result, it is impossible to infer the second-party training sample business data owned by the second device. Therefore, privacy protection for the second device can be achieved. For the first device, since the residual is the difference between the first sample business label and the business label training prediction result and is an embodiment of the model performance, it is difficult for the second device to infer the label information or privacy data owned by the first device only based on the residual when the second device neither knows the first-party training sample business data, nor the first sample business label, nor the business label prediction model. Therefore, privacy protection for the first device can be achieved. Thus, it overcomes the technical defect that each participating party in vertical federated learning cooperates with the label holder to train a vertical federated model, and then uses the vertical federated model for label prediction for the label holder. In this way, the participating party can infer the label of the test sample of the label holder based on the locally trained model, thus bringing a serious privacy leakage risk.Moreover, compared with the service label prediction model trained only based on local data, vertical federated learning can utilize diverse sample service data distributed across multiple second devices to more accurately predict the residual between the prediction result and the true result of the service label prediction model, thereby correcting the prediction result of the service label prediction model, that is, improving the accuracy of the finally output prediction result. Therefore, it can not only improve prediction accuracy but also avoid privacy leakage.
[0101] Embodiment 2
[0102] Furthermore, the present application also provides a vertical federated prediction optimization method. In the second embodiment of the present application, for the same or similar content as the above embodiment, reference can be made to the above introduction and will not be elaborated hereinafter. On this basis, the vertical federated prediction optimization method is applied to the second device in the vertical federated learning system. A residual prediction model is deployed on the second device, and the vertical federated learning system further includes a first device on which a service label prediction model is deployed; referring to Figure 3 , the vertical federated prediction optimization method includes the following steps:
[0103] Step C10: Align samples with the first device to determine the aligned samples, and obtain the second-party training sample service data of the aligned samples;
[0104] The execution subject of the method in this embodiment can be a vertical federated prediction optimization device, or a vertical federated prediction optimization terminal device or server. In this embodiment, a vertical federated prediction optimization device is taken as an example. The vertical federated prediction optimization device can be integrated on terminal devices such as smart phones and computers with data processing functions.
[0105] In this embodiment, it should be noted that the vertical federated prediction optimization method provided in this embodiment is applied to the second device in the vertical federated learning system.
[0106] As an example, the step C10 includes: the residual prediction model can be initialized first, and then samples are aligned with the first device to determine the alignment sample IDs of at least one aligned sample, and the second-party training sample service data corresponding to the alignment sample IDs are found from the sample service data owned by itself based on each of the alignment sample IDs.
[0107] Step C20: Input the second-party training sample service data into the residual prediction model for residual prediction to obtain a training residual prediction result;
[0108] As an example, step C20 includes: inputting each piece of the second-party training sample service data into a residual prediction model deployed by itself for residual learning, and based on the forward propagation algorithm, calculating at least one training residual prediction result, where the training residual prediction result is the predicted value of the residual generated by the service label prediction model for predicting the service label of the aligned sample.
[0109] In an implementable manner, a boosting model can be initialized as the residual prediction model. The boosting model can be a three-layer fully connected neural network. The first layer is the input layer, and the dimension of the input layer is the same as the dimension of the data features. The third layer is the output layer, and the dimension of the output layer is the same as the dimension of the residuals. When performing residual prediction, first input each piece of the second-party training sample service data into the first layer, perform an inner product operation with the model parameters of the first layer to obtain the output result vector of the first layer; then input the output result vector of the first layer into the second layer, perform an inner product operation with the model parameters of the second layer to obtain the output result vector of the second layer; then input the output result vector of the second layer into the third layer, perform an inner product operation with the model parameters of the third layer to obtain the training residual prediction result. In this way, the training residual prediction result can be calculated based on the forward propagation algorithm by the residual prediction model.
[0110] Step C30: Send the training residual prediction result to the first device, so that the first device can obtain the first-party training sample service data and the first sample service label of the aligned sample, and obtain the service label training prediction result obtained by the trained service label prediction model for predicting the service label based on the first-party training sample service data. Determine the service label prediction residual according to the first sample service label and the service label training prediction result, and determine the residual prediction model gradient corresponding to the second device based on the training residual prediction result and the service label prediction residual.
[0111] As an example, step C30 includes: sending the training residual prediction result to the first device, so that the first device can obtain the first-party training sample service data and the first sample service label of the aligned sample, and obtain the service label training prediction result obtained by the trained service label prediction model for predicting the service label based on the first-party training sample service data. Determine the service label prediction residual according to the first sample service label and the service label training prediction result, and determine the residual prediction model gradient corresponding to the second device based on the training residual prediction result and the service label prediction residual. After the first device determines the residual prediction model gradient corresponding to the second device, it sends the residual prediction model gradient to the second device.
[0112] Step C40: Receive the residual prediction model gradient sent by the first device, and update the residual prediction model based on the residual prediction model gradient.
[0113] As an example, step C40 includes: receiving the residual prediction model gradient sent by the first device, and based on the received residual prediction model gradient, using the gradient descent method to update the residual prediction model deployed by itself.
[0114] In this embodiment, for the second device, the second-party training sample service data also belongs to the user's privacy information. By performing residual prediction through the residual prediction model, only the residual information can be sent to the first device. Without the first device knowing the second-party training sample service data, it is difficult for the second device to infer the privacy data owned by the second device only based on the residual information. Therefore, privacy protection for the second device can be achieved.
[0115] Embodiment Three
[0116] Furthermore, the present application also provides a vertical federated prediction optimization method. In the third embodiment of the present application, for the same or similar content as the above embodiments, reference can be made to the above introduction and will not be elaborated hereinafter. On this basis, the vertical federated prediction optimization method is applied to the third device. Referring to Figure 4 , the vertical federated prediction optimization method includes the following steps:
[0117] Step D10: Obtain the first-party to-be-predicted sample service data of the to-be-predicted sample, and obtain the local service label prediction result obtained by the trained service label prediction model for performing service label prediction based on the first-party to-be-predicted sample service data;
[0118] In this embodiment, it should be noted that the vertical federated prediction optimization method is applied to the third device. The third device can be the first device. After completing the model training of the service label prediction model and the residual prediction model, service label prediction can be directly performed locally; the third device can also be other electronic devices other than the first device. After using the first device to complete the model training of the service label prediction model and the residual prediction model, the trained service label prediction model is deployed on other electronic devices for service label prediction. It can be specifically determined according to the actual situation, and this embodiment does not limit this.
[0119] As an example, step D10 includes: obtaining the first-party to-be-predicted sample service data of the to-be-predicted sample, and then inputting the first-party to-be-predicted sample service data into the trained service label prediction model for service label prediction to obtain the local service label prediction result.
[0120] In an implementable manner, the service label prediction model may include a feature extractor, a fully connected layer, and an activation layer. After obtaining the first-party service data of the sample to be predicted, the first-party service data of the sample to be predicted may be first subjected to feature extraction based on the feature extractor to obtain first-party sample features. Furthermore, the first-party sample features may be concatenated into a first-party sample feature vector or a first-party sample feature matrix, and the first-party sample feature vector or the first-party sample feature matrix may be fully connected through the fully connected layer. Then, the output of the fully connected layer may be activated through the activation function preset by the activation layer to obtain the local service label prediction result.
[0121] Step D20: Align samples with at least one second device, so that for each target second device among the second devices that contains the second-party service data of the sample to be predicted corresponding to the sample to be predicted, perform residual prediction based on the second-party service data of the sample to be predicted through the residual prediction model deployed by itself, and obtain a residual prediction result, where the residual prediction model is trained by using the vertical federated prediction optimization method described above;
[0122] As an example, step D20 includes: The third device may send the ID of the sample to be predicted of the sample to be predicted to at least one second device, so that each second device may search for the second-party service data of the sample to be predicted corresponding to the ID of the sample to be predicted from the sample service data it owns based on the received ID of the sample to be predicted, thereby achieving sample alignment with each second device. The second device that can perform sample alignment with the third device may be determined as the target second device. When the target second device finds the second-party service data of the sample to be predicted, it may input each second-party service data of the sample to be predicted into the trained residual prediction model deployed by itself to perform residual prediction, obtain a residual prediction result, and then send each training residual prediction result to the third device. When the second device cannot find the second-party service data of the sample to be predicted corresponding to the ID of the sample to be predicted, it may not need to perform the residual prediction task, or may send a null value to the third device, or may send a prompt message of no aligned sample to the third device. Among them, the training method of the residual prediction model may refer to the specific content corresponding to steps S10-S40 above, and will not be elaborated here.
[0123] Step D30: Receive the residual prediction results sent by each target second device, and aggregate the local service label prediction result and the residual prediction results to obtain a federated service label prediction result.
[0124] As an example, step D30 includes: receiving the residual prediction results sent by each of the target second devices, aggregating the local service label prediction result and the residual prediction results to obtain a federated service label prediction result, where the method of aggregating the local service label prediction result and the residual prediction results may be addition operation, average operation, weighted summation operation, weighted average operation, etc.
[0125] In this embodiment, the model performance of the service label prediction model trained based on the local sample service data is limited. Therefore, the residual between the obtained local service label prediction result and the true value will be relatively large. By using the residual prediction models distributed on different second devices, the prediction result of the service label prediction model can be corrected by using the more diverse sample service data owned by different second devices while ensuring the privacy of each party. This can effectively reduce the residual between the local service label prediction result and the true value, thereby obtaining a federated service label prediction result closer to the true value and improving the accuracy of service label prediction.
[0126] Embodiment 4
[0127] Furthermore, the embodiment of the present application further provides a vertical federated prediction optimization device. Referring to Figure 5 , the vertical federated prediction optimization device is applied to the first device in the vertical federated learning system, and the service label prediction model is deployed on the first device; the vertical federated learning system further includes multiple second devices, and the residual prediction models are deployed on the second devices; the vertical federated prediction optimization device includes:
[0128] The first sample alignment module 10 is used to perform sample alignment with the second device to determine the aligned samples;
[0129] The acquisition module 20 is used to acquire the first-party training sample service data and the first sample service label of the aligned samples, and acquire the service label training prediction result obtained by the trained service label prediction model for service label prediction based on the first-party training sample service data;
[0130] The service label prediction residual determination module 30 is used to determine the service label prediction residual according to the first sample service label and the service label training prediction result;
[0131] The receiving module 40 is used to receive the training residual prediction result sent by the second device, where the training residual prediction result is obtained by the second device through the residual prediction model for residual prediction based on the second-party training sample service data of the aligned samples;
[0132] A gradient determination module 50, configured to determine a residual prediction model gradient corresponding to the second device based on the training residual prediction result and the service label prediction residual, and send the residual prediction model gradient to the second device for the second device to update the locally deployed residual prediction model based on the received residual prediction model gradient.
[0133] Further, the gradient determination module 50 is further configured to:
[0134] Determine a service label prediction residual corresponding to each training residual prediction result based on the sample alignment relationship between each first-party training sample service data and each second-party training sample service data;
[0135] Calculate an initial gradient of the residual prediction model corresponding to each second-party training sample service data based on the difference between each training residual prediction result and the service label prediction residual corresponding to each training residual prediction result;
[0136] Determine the average value of each initial gradient of the residual prediction model as the residual prediction model gradient corresponding to the second device.
[0137] Further, the gradient determination module 50 is further configured to:
[0138] Determine whether the current condition for ending the federated training is satisfied;
[0139] When it is determined that the current condition for ending the federated training is not satisfied, determine the residual prediction model gradient corresponding to the second device based on each training residual prediction result and the service label prediction residual;
[0140] Send the residual prediction model gradient to the second device for the second device to update the locally deployed residual prediction model based on the received residual prediction model gradient, and return to execute the step of obtaining the first-party training sample service data and the first sample service label of the aligned sample.
[0141] Further, the gradient determination module 50 is further configured to:
[0142] Judge whether the residual prediction models deployed on each second device converge according to the difference between each training residual prediction result corresponding to each second device and the corresponding service label prediction residual;
[0143] When the number of detected converged residual prediction models exceeds a preset number threshold, determine that the current condition for ending the federated training is satisfied.
[0144] Further, before the operations of obtaining the first-party training sample service data and the first sample service label of the alignment sample, and obtaining the service label training prediction result obtained by the trained service label prediction model for service label prediction based on the first-party training sample service data, the vertical federated prediction optimization device further includes a local training module, and the local training module is used for:
[0145] Obtain the service label prediction model training sample service data and the second sample service label, and iteratively optimize the service label prediction model based on the service label prediction model training sample service data and the second sample service label to obtain a trained service label prediction model.
[0146] The vertical federated prediction optimization device provided by the present invention adopts the vertical federated prediction optimization method in the above embodiment, and solves the technical problem of relatively high privacy leakage risk in vertical federated learning in the related art. Compared with the related art, the benefits of the vertical federated prediction optimization device provided by the embodiment of the present invention are the same as those of the vertical federated prediction optimization method provided by the above embodiment, and other technical features in the vertical federated prediction optimization device are the same as the features disclosed in the method of the above embodiment, and will not be elaborated here.
[0147] Embodiment Five
[0148] Further, an embodiment of the present application further provides a vertical federated prediction optimization device. The vertical federated prediction optimization device is applied to a second device in a vertical federated learning system. A residual prediction model is deployed on the second device. The vertical federated learning system further includes a first device, and a service label prediction model is deployed on the first device. The vertical federated prediction optimization device includes:
[0149] A second sample alignment module, configured to perform sample alignment with the first device, determine the alignment sample, and obtain the second-party training sample service data of the alignment sample;
[0150] A residual prediction module, configured to input the second-party training sample service data into the residual prediction model for residual prediction to obtain a training residual prediction result;
[0151] A sending module, configured to send the training residual prediction result to the first device, so that the first device obtains the first-party training sample service data and the first sample service label of the alignment sample, and obtains the service label training prediction result obtained by the trained service label prediction model for service label prediction based on the first-party training sample service data. Determine the service label prediction residual according to the first sample service label and the service label training prediction result, and determine the residual prediction model gradient corresponding to the second device based on the training residual prediction result and the service label prediction residual;
[0152] An update module, configured to receive the residual prediction model gradient sent by the first device, and update the residual prediction model based on the residual prediction model gradient.
[0153] Embodiment Six
[0154] Furthermore, an embodiment of the present application further provides a vertical federated prediction optimization device, which is applied to a third device. The vertical federated prediction optimization device includes:
[0155] A prediction module, configured to obtain the first-party to-be-predicted sample service data of the to-be-predicted sample, and obtain the local service label prediction result obtained by performing service label prediction on the first-party to-be-predicted sample service data based on the trained service label prediction model;
[0156] A third sample alignment module, configured to perform sample alignment with at least one second device, so that a target second device among the second devices that includes the second-party to-be-predicted sample service data corresponding to the to-be-predicted sample performs residual prediction on the second-party to-be-predicted sample service data through the residual prediction model deployed by itself, and obtain a residual prediction result, where the residual prediction model is trained by using the vertical federated prediction optimization method as described above;
[0157] An aggregation module, configured to receive the residual prediction results sent by the target second devices, and aggregate the local service label prediction result and the residual prediction results to obtain a federated service label prediction result.
[0158] Embodiment Seven
[0159] Furthermore, an embodiment of the present invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the vertical federated prediction optimization method in the above embodiments.
[0160] Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as Bluetooth headsets, mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0161] As shown Figure 6 in the figure, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and arrays required for the operation of the electronic device are also stored. The processing device, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0162] Generally, the following systems may be connected to the I / O interface: input devices including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices including, for example, a magnetic tape, a hard disk, etc.; and a communication device. The communication device may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange arrays. Although the figure shows an electronic device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.
[0163] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from a storage device, or installed from the ROM. When the computer program is executed by the processing device, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0164] The electronic device provided by the present invention adopts the vertical federated prediction optimization method in the above embodiment, and solves the technical problem of relatively high privacy leakage risk in vertical federated learning in the related art. Compared with the related art, the advantages of the electronic device provided by the embodiment of the present invention are the same as those of the vertical federated prediction optimization method provided by the above embodiment, and other technical features in the electronic device are the same as the features disclosed in the method of the above embodiment, and will not be elaborated here.
[0165] It should be understood that each part of the present disclosure may be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in a suitable manner in any one or more embodiments or examples.
[0166] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims described above.
[0167] Embodiment VIII
[0168] Furthermore, this embodiment provides a computer-readable storage medium having computer-readable program instructions stored thereon, and the computer-readable program instructions are used to execute the vertical federated prediction optimization method in the above embodiments.
[0169] The computer-readable storage medium provided by the embodiments of the present invention can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0170] The above computer-readable storage medium can be included in an electronic device; or it can exist alone without being assembled into the electronic device.
[0171] The above computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to: perform sample alignment with the second device to determine aligned samples; obtain first-party training sample service data and first sample service labels of the aligned samples, and obtain service label training prediction results obtained by the trained service label prediction model based on the first-party training sample service data for service label prediction; determine service label prediction residuals according to the first sample service labels and the service label training prediction results; receive the training residual prediction results sent by the second device, where the training residual prediction results are obtained by the second device through the residual prediction model based on second-party training sample service data of the aligned samples; determine the residual prediction model gradient corresponding to the second device based on the training residual prediction results and the service label prediction residuals, and send the residual prediction model gradient to the second device for the second device to update the residual prediction model deployed by itself based on the received residual prediction model gradient.
[0172] Alternatively, the above computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to: perform sample alignment with the first device to determine aligned samples and obtain second-party training sample service data of the aligned samples; input the second-party training sample service data into the residual prediction model for residual prediction to obtain training residual prediction results; send the training residual prediction results to the first device for the first device to obtain first-party training sample service data and first sample service labels of the aligned samples, and obtain service label training prediction results obtained by the trained service label prediction model based on the first-party training sample service data for service label prediction, determine service label prediction residuals according to the first sample service labels and the service label training prediction results, determine the residual prediction model gradient corresponding to the second device based on the training residual prediction results and the service label prediction residuals; receive the residual prediction model gradient sent by the first device and update the residual prediction model based on the residual prediction model gradient.
[0173] Alternatively, the above computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to: obtain the first-party business data of the sample to be predicted, and obtain the local business label prediction result obtained by the trained business label prediction model for predicting the business label based on the first-party business data of the sample to be predicted; align the samples with at least one second device, so that the target second device containing the second-party business data of the sample to be predicted in each second device performs residual prediction based on the second-party business data of the sample to be predicted through the residual prediction model deployed by itself, and obtain a residual prediction result, where the residual prediction model is trained by using the vertical federated prediction optimization method as described above; receive the residual prediction results sent by each target second device, and aggregate the local business label prediction result and the residual prediction results to obtain a federated business label prediction result.
[0174] Computer program code for carrying out operations for aspects of the present disclosure may be written in one or more programming languages, or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0175] The flowcharts and block diagrams in the figures illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the figures. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0176] The modules involved in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0177] The computer-readable storage medium provided by the present invention stores computer-readable program instructions for executing the above-mentioned vertical federated prediction optimization method, and solves the technical problem of relatively high privacy leakage risk in vertical federated learning in the related art. Compared with the related art, the advantages of the computer-readable storage medium provided by the embodiments of the present invention are the same as those of the vertical federated prediction optimization method provided by the above embodiments, and will not be elaborated here.
[0178] Embodiment Nine
[0179] Furthermore, the present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned vertical federated prediction optimization method are implemented.
[0180] The computer program product provided by the present application solves the technical problem of relatively high privacy leakage risk in vertical federated learning in the related art. Compared with the related art, the advantages of the computer program product provided by the embodiments of the present invention are the same as those of the vertical federated prediction optimization method provided by the above embodiments, and will not be elaborated here.
[0181] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent scope of the present application by the same token.
Claims
1. A vertical federated prediction optimization method, characterized in that, The vertical federated prediction optimization method is applied to a first device in a vertical federated learning system. A business label prediction model is deployed on the first device. The vertical federated learning system further includes a second device on which a residual prediction model is deployed. The vertical federated prediction optimization method includes the following steps: Perform sample alignment with the second device to determine aligned samples; Obtain the first-party training sample business data and the first sample business label of the aligned samples, and obtain the business label training prediction results obtained by the trained business label prediction model for business label prediction based on the first-party training sample business data; Determine the business label prediction residuals according to the first sample business label and the business label training prediction results; Receive the training residual prediction results sent by the second device, where the training residual prediction results are obtained by the second device through residual prediction based on the second-party training sample business data of the aligned samples; Based on the training residual prediction results and the business label prediction residuals, determine the residual prediction model gradient corresponding to the second device, and send the residual prediction model gradient to the second device for the second device to update the locally deployed residual prediction model based on the received residual prediction model gradient.
2. The vertical federated prediction optimization method according to claim 1, characterized in that, The step of determining the residual prediction model gradient corresponding to the second device based on the training residual prediction results and the business label prediction residuals includes: Based on the sample alignment relationship between each first-party training sample business data and each second-party training sample business data, determine the business label prediction residuals corresponding to each of the training residual prediction results; Based on the differences between each of the training residual prediction results and the business label prediction residuals corresponding to each of the training residual prediction results, calculate the initial gradient of the residual prediction model corresponding to each second-party training sample business data; Determine the average value of each of the initial gradients of the residual prediction models as the residual prediction model gradient corresponding to the second device.
3. The vertical federated prediction optimization method according to claim 1, characterized in that, The step of determining the residual prediction model gradient corresponding to the second device based on the training residual prediction results and the business label prediction residuals, and sending the residual prediction model gradient to the second device for the second device to update the locally deployed residual prediction model based on the received residual prediction model gradient includes: Judge whether the current condition meets the end condition of federated training; In the case of determining that the current condition does not meet the end condition of federated training, determine the residual prediction model gradient corresponding to the second device based on each of the training residual prediction results and the business label prediction residuals; Send the residual prediction model gradient to the second device for the second device to update the locally deployed residual prediction model based on the received residual prediction model gradient, and return to execute the step of obtaining the first-party training sample business data and the first sample business label of the aligned samples.
4. The vertical federated prediction optimization method according to claim 3, characterized in that, There are multiple second devices. The step of judging whether the current condition meets the end condition of federated training includes: Based on the differences between the respective training residual prediction results corresponding to each of the second devices and the corresponding service label prediction residuals, determine whether the residual prediction models deployed on each of the second devices converge; In the case where the number of converged residual prediction models detected exceeds a preset number threshold, determine that the current condition for the end of federated training is satisfied.
5. The vertical federated prediction optimization method according to claim 1, characterized in that, Before the step of obtaining the first-party training sample service data and the first sample service label of the aligned sample, and obtaining the service label training prediction result obtained by the trained service label prediction model for service label prediction based on the first-party training sample service data, it further includes: Obtain service label prediction model training sample service data and second sample service labels, and iteratively optimize the service label prediction model based on the service label prediction model training sample service data and the second sample service labels to obtain a trained service label prediction model.
6. A vertical federated prediction optimization method, characterized in that, The vertical federated prediction optimization method is applied to a second device in a vertical federated learning system. A residual prediction model is deployed on the second device. The vertical federated learning system further includes a first device on which a service label prediction model is deployed. The vertical federated prediction optimization method includes the following steps: Perform sample alignment with the first device to determine the aligned sample, and obtain the second-party training sample service data of the aligned sample; Input the second-party training sample service data into the residual prediction model for residual prediction to obtain a training residual prediction result; Send the training residual prediction result to the first device for the first device to obtain the first-party training sample service data and the first sample service label of the aligned sample, and obtain the service label training prediction result obtained by the trained service label prediction model for service label prediction based on the first-party training sample service data. Determine the service label prediction residual according to the first sample service label and the service label training prediction result, and determine the residual prediction model gradient corresponding to the second device based on the training residual prediction result and the service label prediction residual; Receive the residual prediction model gradient sent by the first device, and update the residual prediction model based on the residual prediction model gradient.
7. A vertical federated prediction optimization method, characterized in that, The vertical federated prediction optimization method is applied to a third device and includes the following steps: Obtain the first-party to-be-predicted sample service data of the to-be-predicted sample, and obtain the local service label prediction result obtained by the trained service label prediction model for service label prediction based on the first-party to-be-predicted sample service data; Perform sample alignment with at least one second device for a target second device among the second devices that contains the second-party to-be-predicted sample service data corresponding to the to-be-predicted sample to perform residual prediction based on the second-party to-be-predicted sample service data through the residual prediction model deployed by itself to obtain a residual prediction result, where the residual prediction model is trained by using the vertical federated prediction optimization method described in any one of claims 1-6; Receive the residual prediction results sent by each of the target second devices, aggregate the local service label prediction results and the residual prediction results to obtain a federated service label prediction result.
8. A vertical federated prediction optimization device, characterized in that, The vertical federated prediction optimization device is applied to a first device in a vertical federated learning system, and a service label prediction model is deployed on the first device; the vertical federated learning system further includes a plurality of second devices, and a residual prediction model is deployed on each of the second devices; the vertical federated prediction optimization device includes: A first sample alignment module, configured to perform sample alignment with the second device to determine aligned samples; An acquisition module, configured to acquire the first-party training sample service data and the first sample service label of the aligned samples, and acquire the service label training prediction result obtained by the trained service label prediction model performing service label prediction based on the first-party training sample service data; A service label prediction residual determination module, configured to determine a service label prediction residual according to the first sample service label and the service label training prediction result; A receiving module, configured to receive the training residual prediction result sent by the second device, where the training residual prediction result is obtained by the second device performing residual prediction on the second-party training sample service data of the aligned samples through the residual prediction model; A gradient determination module, configured to determine the residual prediction model gradient corresponding to the second device based on the training residual prediction result and the service label prediction residual, and send the residual prediction model gradient to the second device for the second device to update the locally deployed residual prediction model based on the received residual prediction model gradient.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by at least one of the processors, and the instructions are executed by at least one of the processors so that at least one of the processors can execute the steps of the vertical federated prediction optimization method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a program for implementing the vertical federated prediction optimization method is stored on the computer-readable storage medium. The program for implementing the vertical federated prediction optimization method is executed by a processor to implement the steps of the vertical federated prediction optimization method according to any one of claims 1 to 7.
11. A product, the product being a computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the vertical federated prediction optimization method according to any one of claims 1 to 7.