Model construction optimization method, device, storage medium, and program product

CN115114862BActive Publication Date: 2026-09-25WEBANK (CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210890498.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2026-09-25
Estimated Expiration
2042-07-27

AI Technical Summary

Technical Problem

一个典型场景是,各方通过用户匹配发现重叠的用户比例较高,然而各方拥有的相应的标签比例较低,因此可以实际用于训练的样本较少,如何能够利用无标签的数据是一个具有重要实际意义的问题

Benefits of technology

[0033]本发明中,通过在参与纵向联邦学习的第一参与方设备部署第一编码器和解码器,在第二参与方设备部署第二编码器,通过第一参与方设备将第一样本数据输入至第一编码器进行编码得到第一编码特征,并获取第二参与方设备将第二样本数据输入至第二编码器进行编码得到的第二编码特征,将第一编码特征和第二编码特征进行融合后输入至解码器进行解码得到重构样本数据,基于重构样本数据与第一样本数据之间的误差更新第一编码器以及计算得到用于更新第二编码器的中间结果,将中间结果发送给第二参与方设备以使得第二参与方设备对第二编码器进行更新,在对第一编码器和第二编码器进行至少一轮迭代更新后,基于更新后的第一编码器和第二编码器与第二参与方设备进行纵向联邦学习得到目标模型,实现了采用参与纵向联邦学习的各参与方的无标签的样本数据来对第一编码器和第二编码器进行预训练,以通过预训练来优化模型构建过程。并且本发明中的该预训练过程使得第一编码器和第二编码器能够学习到如何挖掘各参与方样本数据之间的关联性。相比于采用各参与方中有标签的样本数据对随机初始化或根据人工经验初始化的第一编码器、第二编码器和预测器进行纵向联邦学习,本发明中采用预训练后的第一编码器和第二编码器作为纵向联邦学习的基础,由于第一编码器和第二编码器学习到了如何挖掘各参与方样本数据之间的联系,使得在纵向联邦学习阶段,通过第一编码器和第二编码器能够编码得到更利于预测器得出准确预测结果的编码特征,进而能够帮助提高训练得到的目标模型的预测准确度,也能够帮助缩短纵向联邦学习的时长,进而能够减少纵向联邦学习阶段的计算资源消耗。并且,也实现了将无标签的样本数据应用于纵向联邦学习场景参与模型训练,进而实现了基于无标签的样本数据来提升纵向联邦学习场景中的模型预测准确度。另外,在预训练过程中,各参与方设备并没有直接发送原始的样本数据,所以也保证了各参与方样本数据的隐私安全。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115114862B_ABST
    Figure CN115114862B_ABST
Patent Text Reader

Abstract

The application discloses a model construction optimization method and device, a storage medium and a program product. The method comprises the following steps: inputting first sample data into a first encoder to obtain first encoding features; obtaining second encoding features obtained by inputting second sample data into a second encoder by a second participant device; inputting the first encoding features and the second encoding features into a decoder after fusion to obtain reconstructed sample data; updating the first encoder based on an error between the reconstructed sample data and the first sample data and calculating an intermediate result; sending the intermediate result to the second participant device to update the second encoder; and performing vertical federated learning based on the updated first encoder and the second encoder to obtain a target model. The application realizes model pre-training by using unlabeled data, so that the unlabeled data can be used for participating in vertical federated learning, thereby helping to improve the prediction accuracy of the model obtained by vertical federated learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a model building optimization method, device, storage medium, and program product. Background Technology

[0002] Vertical federated learning is used to address the data silo problem in the financial sector, where multiple parties collaborate to build business models. Its application scenarios involve situations where participating parties have significant user overlap but limited feature overlap, and different parties possess information about users from different domains and perspectives. A typical scenario is that parties discover a high proportion of overlapping users through user matching, but each party possesses a low proportion of corresponding labels, resulting in a limited number of samples suitable for training. Therefore, how to utilize unlabeled data is a problem of significant practical importance. Summary of the Invention

[0003] The main objective of this invention is to provide a model building optimization method, device, storage medium, and program product. It aims to propose a method for model pre-training using unlabeled data in a vertical federated learning scenario, so that unlabeled data can be used to participate in vertical federated learning, thereby helping to improve the prediction accuracy of the model obtained by vertical federated learning.

[0004] To achieve the above objectives, the present invention provides a model building optimization method. This method is applied to a first participating device in longitudinal federated learning, where the first participating device deploys a first encoder and a decoder, and a second participating device in longitudinal federated learning deploys a second encoder. The method includes the following steps:

[0005] The first sample data is input into the first encoder for encoding to obtain the first encoded feature;

[0006] The second encoded feature is obtained by the second participating device inputting the second sample data into the second encoder for encoding;

[0007] The first encoded feature and the second encoded feature are fused and then input into the decoder for decoding to obtain reconstructed sample data;

[0008] The first encoder is updated based on the error between the reconstructed sample data and the first sample data, and an intermediate result for updating the second encoder is calculated. The intermediate result is then sent to the second participating device for the second participating device to update the second encoder.

[0009] After at least one round of iterative updates to the first encoder and the second encoder, the target model is obtained by longitudinal federated learning based on the updated first encoder and the second encoder and the second participating device.

[0010] Optionally, the step of inputting the first sample data into the first encoder for encoding to obtain the first encoded feature includes:

[0011] A portion of the data in the first sample data is transformed, and the transformed first sample data is input into the first encoder for encoding to obtain the first encoded feature.

[0012] Optionally, the reconstructed sample data is reconstructed data of the transformed portion of the first sample data, and the steps of updating the first encoder based on the error between the reconstructed sample data and the first sample data and calculating intermediate results for updating the second encoder include:

[0013] Calculate the self-loss function that characterizes the error between the reconstructed sample data and the transformed portion of the first sample data;

[0014] Calculate the first gradient value of the self-loss function with respect to the parameters in the first encoder, and calculate the second gradient value of the self-loss function with respect to the second encoded feature;

[0015] Update the parameters in the first encoder according to the first gradient value to update the first encoder;

[0016] The second gradient value is used as an intermediate result for updating the second encoder.

[0017] Optionally, the step of transforming a portion of the data in the first sample data includes:

[0018] When the first sample data includes multiple attribute values, some attribute values ​​in the first sample data are transformed into preset values ​​or random noise is added;

[0019] When the first sample data is image data, the pixel values ​​of some pixels in the image data are transformed into preset pixel values;

[0020] When the first sample data is text data, some words in the text data are transformed into preset words.

[0021] Optionally, the first participating device further deploys a projection model, and the step of fusing the first encoded features and the second encoded features and inputting them into the decoder for decoding to obtain reconstructed sample data includes:

[0022] The first encoded feature and the second encoded feature are input into the projection model for feature cross-processing, and the fused feature obtained after feature cross-processing is input into the decoder for decoding to obtain reconstructed sample data.

[0023] Optionally, the first participating device further deploys a predictor, and the step of obtaining the target model by performing longitudinal federated learning based on the updated first encoder and second encoder and the second participating device after at least one round of iterative updates to the first encoder and the second encoder includes:

[0024] Based on the labeled third sample data and the fourth sample data in the second participant device that is aligned with the third sample data, longitudinal federated learning is performed on the predictor and the updated first encoder and second encoder to obtain a target model including the trained first encoder, second encoder and predictor.

[0025] To achieve the above objectives, the present invention also provides a model building optimization method, which is applied to a second participating device in longitudinal federated learning. The first participating device in longitudinal federated learning deploys a first encoder and a decoder, and the second participating device deploys a second encoder. The method includes the following steps:

[0026] The second sample data is input into the second encoder for encoding to obtain the second encoded feature;

[0027] The second encoded feature is sent to the first participating device, so that the first participating device can fuse the first encoded feature and the second encoded feature and input them into the decoder to decode and obtain reconstructed sample data. The first encoder is updated based on the error between the reconstructed sample data and the first sample data, and intermediate results for updating the second encoder are calculated. The first encoded feature is obtained by the first participating device inputting the first sample data into the first encoder for encoding.

[0028] Obtain the intermediate result and update the second encoder based on the intermediate result;

[0029] After at least one round of iterative updates to the first encoder and the second encoder, a target model is obtained by longitudinal federated learning based on the updated first encoder and the second encoder and the first participating device.

[0030] To achieve the above objectives, the present invention also provides a model building optimization device, the model building optimization device comprising: a memory, a processor, and a model building optimization program stored in the memory and executable on the processor, wherein the model building optimization program, when executed by the processor, implements the steps of the model building optimization method as described above.

[0031] Furthermore, to achieve the above objectives, the present invention also proposes a computer-readable storage medium storing a model building optimization program, which, when executed by a processor, implements the steps of the model building optimization method as described above.

[0032] Furthermore, to achieve the above objectives, the present invention also proposes a computer program product, comprising a computer program that, when executed by a processor, implements the steps of the model building optimization method as described above.

[0033] In this invention, a first encoder and decoder are deployed on a first participating device in vertical federated learning, and a second encoder is deployed on a second participating device. The first participating device inputs first sample data into the first encoder for encoding to obtain first encoded features, and the second participating device obtains second encoded features by inputting second sample data into the second encoder. The first and second encoded features are fused and input into the decoder to obtain reconstructed sample data. The first encoder is updated based on the error between the reconstructed sample data and the first sample data, and intermediate results for updating the second encoder are calculated. These intermediate results are sent to the second participating device to update the second encoder. After at least one round of iterative updates to the first and second encoders, vertical federated learning is performed using the updated first and second encoders and the second participating device to obtain the target model. This achieves pre-training of the first and second encoders using unlabeled sample data from each participating party in vertical federated learning, thereby optimizing the model construction process. Furthermore, this pre-training process enables the first and second encoders to learn how to mine the correlations between the sample data from each participating party. Compared to using labeled sample data from each participant to perform longitudinal federated learning on randomly initialized or manually initialized first encoders, second encoders, and predictors, this invention uses pre-trained first and second encoders as the foundation for longitudinal federated learning. Because the first and second encoders learn how to mine the relationships between the sample data from each participant, they can encode features that are more conducive to the predictor's accurate predictions during the longitudinal federated learning phase. This helps improve the prediction accuracy of the trained target model and shortens the duration of longitudinal federated learning, thus reducing computational resource consumption. Furthermore, it also enables the application of unlabeled sample data to participate in model training in the longitudinal federated learning scenario, thereby improving the model's prediction accuracy based on unlabeled sample data. Additionally, during pre-training, the participating devices do not directly send raw sample data, thus ensuring the privacy and security of the sample data from each participant. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the hardware operating environment involved in the embodiments of the present invention;

[0035] Figure 2 This is a flowchart illustrating the first embodiment of the model construction optimization method of the present invention;

[0036] Figure 3 This is a flowchart illustrating the third embodiment of the model construction optimization method of the present invention;

[0037] Figure 4 This is a schematic diagram of a pre-training architecture according to an embodiment of the present invention;

[0038] Figure 5 This is a schematic diagram of a model fine-tuning architecture according to an embodiment of the present invention.

[0039] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0040] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0041] like Figure 1 As shown, Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention.

[0042] It should be noted that the model building and optimization device in this embodiment of the invention can be a smartphone, a personal computer, a server, or other such device, and no specific limitations are imposed here.

[0043] like Figure 1 As shown, the optimized device for this model construction may include: a processor 1001, such as a CPU; a network interface 1004; a user interface 1003; a memory 1005; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or stable non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0044] Those skilled in the art will understand that Figure 1 The device structure shown does not constitute a limitation on the model building optimization device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0045] like Figure 1As shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a model building and optimization program. The operating system is a program that manages and controls the device's hardware and software resources, supporting the operation of the model building and optimization program and other software or programs. Figure 1 In the device shown, the user interface 1003 is mainly used for data communication with the client; the network interface 1004 is mainly used for the server to establish a communication connection; and the processor 1001 can be used to call the model building optimization program stored in the memory 1005 and execute the operations described in the following embodiments of the model building optimization method of the present invention.

[0046] Based on the above structure, various embodiments of the model building optimization method are proposed.

[0047] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the model construction optimization method of the present invention.

[0048] This invention provides an embodiment of a model building optimization method. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order. In this embodiment, the model building optimization method is applied to a first participant device participating in longitudinal federated learning. The first participant device is deployed in the first participant of longitudinal federated learning, and devices deployed in other participants for longitudinal federated learning are referred to as second participant devices. The first and second participant devices can be devices such as smartphones, personal computers, and servers, and are not limited in this embodiment. In this embodiment, the model building optimization method includes:

[0049] Step S10: Input the first sample data into the first encoder for encoding to obtain the first encoded feature;

[0050] Participants in vertical federated learning typically include one participant providing labeled data (referred to as the data user) and at least one participant providing feature data (referred to as the data provider). The data user provides both labeled and feature data. The feature spaces of the data user and the data provider are different. The purpose of vertical federated learning is to combine the feature data of the same sample from each participant across different feature spaces to perform modeling, thereby improving the model's prediction accuracy.

[0051] In specific application scenarios, in addition to data providers and data users, third parties can also participate to help them transmit intermediate data that needs to be exchanged during the modeling process in an encrypted manner. This embodiment can be applied to scenarios with third-party participation or without it, without limitation. It is understood that when a third party is involved, the data sent by each party to the other is encrypted. The data sent back after processing the encrypted data locally is decrypted by the third party before being sent back. This ensures that each party cannot know the other party's original data and can only perform calculations based on the other party's encrypted data, thus protecting the privacy and security of the sample data of each party.

[0052] In this embodiment, one of the participants in the vertical federated learning is referred to as the first participant, and the other participants are referred to as the second participants. The device deployed on the first participant for vertical federated learning is referred to as the first participant device, and the device deployed on the second participant for vertical federated learning is referred to as the second participant device. This embodiment does not restrict whether the data application party is the first or the second participant.

[0053] In specific application scenarios, model tasks can be set as needed, and models to be trained for each participant can be designed based on these tasks. The models to be trained are then deployed on the devices of each participant. A model task refers to the intended use of the model after modeling is completed, such as risk prediction or advertising recommendation. In this embodiment, the model task is not limited.

[0054] In this embodiment, the model to be trained may include an encoder deployed on a first participant device (hereinafter referred to as the first encoder), an encoder deployed on a second participant device (the second encoder), and a predictor deployed on a data application device. It is understood that the data application device can be either the first participant device or the second participant device, and this embodiment does not impose any limitation. The encoder types selected for the first and second encoders, as well as the designed data structure of the input data, can be set according to the type of sample data required by the model task. For example, if the sample data is image data, then the first and second encoders can use convolutional neural networks, and the data structure of the input data can be fixed or variable-sized image data. Alternatively, if the sample data is text data, then the first and second encoders can use recurrent neural networks, and the data structure of the input data can be word IDs. Furthermore, if the sample data is tabular data, then the first and second encoders can use fully connected networks, and the data structure of the input data can be a vector composed of the feature values ​​in the tabular data. The type of predictor selected can be set according to the prediction results required by the model task; for example, if the model task is binary classification, then the predictor can use a binary classifier; or if the model task is multi-class classification, then the predictor can use a multi-class classifier. The category to which a sample belongs can be determined based on the results of the predictor output.

[0055] In this embodiment, unlabeled sample data from the first and second participants are used to pre-train the first and second encoders. Pre-training refers to the pre-training of the first and second encoders before the formal longitudinal federated learning. That is, compared to using labeled sample data from each participant to perform longitudinal federated learning on randomly initialized or manually initialized first and second encoders and predictors, this embodiment uses pre-trained first and second encoders as the foundation for longitudinal federated learning. Compared to manual or experience-based initialization, pre-trained first and second encoders can learn how to mine the relationships between the sample data from each participant. This allows the first and second encoders to encode features that are more conducive to accurate predictions by the predictor during the longitudinal federated learning stage, thereby improving the prediction accuracy of the trained target model, shortening the duration of longitudinal federated learning, and reducing computational resource consumption during this stage. Furthermore, it also enables the application of unlabeled sample data to participate in model training in the longitudinal federated learning scenario, thus achieving improved model prediction accuracy based on unlabeled sample data in the longitudinal federated learning scenario.

[0056] Specifically, the pre-training of the first encoder and the second encoder may include one or more rounds of parameter updates. The following description uses a one-round update process as an example to illustrate the pre-training process. Before starting pre-training, the parameters in the first encoder and the second encoder may be randomly initialized or initialized based on experience.

[0057] The first participating party's device can acquire sample data (hereinafter referred to as first sample data) owned by the first participating party from local or remote sources. A single piece of sample data for a given sample is distributed across various participating parties. The sample data owned by each participating party includes the feature data of that sample within that participating party's feature space. That is, the sample data for a given sample belongs to different feature spaces across different participating parties, and can be seen as describing the sample from different perspectives. For example, in one embodiment, one participating party is a bank, and another is an e-commerce institution. The bank possesses the financial activity data of users (samples), and the e-commerce institution possesses the users' purchase record data. The model task can be to use the users' financial activity data and purchase record data to predict the users' loan repayment risk.

[0058] In a specific implementation, if all participants' samples are shared samples, then the sample data of the shared samples in each participant's sample data can be directly used for pre-training. If each participant's samples include both shared and non-shared samples, then the first participant's device and the second participant's device can perform sample alignment beforehand, that is, determine the shared samples of each participant, and then each participant selects the sample data of the shared samples in its own sample data for pre-training. It should be noted that there may be multiple shared samples, and a large number of shared samples can be used for pre-training. However, for ease of description, the following uses one sample data point of a shared sample for pre-training as an example, that is, the first sample data represents the sample data of a sample in the first participant's sample data. Since the pre-training process does not require sample label data, unlabeled sample data can be used for pre-training. There are various methods for sample alignment, and this embodiment does not impose any limitations.

[0059] After acquiring the first sample data, the first participating device can input the first sample data into the first encoder for encoding to obtain encoded features (hereinafter referred to as the first encoded features for distinction). The encoding process using the encoder is the process of applying the parameters in the encoder to the input data to be encoded. Depending on the type of encoder, the types and number of parameters in the encoder are different, and the specific way in which the parameters are applied to the data to be encoded is also different. For example, when the encoder uses a convolutional neural network, the encoder parameters include the convolution kernel parameters of each convolutional layer in the convolutional neural network, and applying the parameters to the data to be encoded specifically means performing convolution on the data to be encoded.

[0060] In a specific implementation, the first participating device can directly convert the first sample data into the data structure form of the input data of the first encoder and then input it into the first encoder, or it can first transform a portion of the data in the first sample data before converting the data structure and inputting it into the first encoder. This embodiment does not impose any limitations. Hereinafter, the portion of the first sample data before transformation is referred to as the transformed portion of the first sample data. For example, if the first sample data includes attribute value a1 of attribute A and attribute value b2 of attribute B, and a1 is transformed into a1', then a1 is the transformed portion of the first sample data, and b2 is the untransformed portion of the first sample data.

[0061] Step S20: Obtain the second encoded feature, wherein the second encoded feature is obtained by the second participant device inputting the second sample data into the second encoder for encoding;

[0062] The second participant device can acquire sample data (hereinafter referred to as second sample data) owned by the second participant, either locally or remotely. The second sample data and the first sample data belong to the same sample object, and the first sample data and the second sample data are the feature data of the sample in different feature spaces. The second participant device inputs the second sample data into the second encoder for encoding to obtain coded features (hereinafter referred to as "second coded features" for distinction).

[0063] The first participating device can obtain the second encoded feature obtained by the second participating device. In specific implementations, the first participating device can obtain it directly from the second participating device, or it can be forwarded by a third party. When there are high requirements for data privacy protection, the second encoded feature obtained by the first participating device from the second participating device can be encrypted. The encryption algorithm can be homomorphic encryption, differential privacy, etc., and there are no restrictions here.

[0064] Step S30: The first encoded feature and the second encoded feature are fused and then input into the decoder for decoding to obtain reconstructed sample data;

[0065] To enable pre-training of the first encoder and the second encoder using unlabeled sample data, in this embodiment, a decoder is also deployed in the first participating device to decode the encoded features output by the first encoder and the second encoder.

[0066] The first participating device fuses the first and second encoded features and inputs the fused data into the decoder for decoding. The decoded data is called reconstructed sample data. The reconstructed sample data can be reconstructed data based on the first sample data, or reconstructed data based on the transformed portion of the first sample data. It should be noted that reconstruction can also be called restoration, that is, the aim is to obtain data identical to the first sample data, or data identical to the transformed portion of the first sample data, by the decoder based on the first and second encoded features. The data structure of the decoder's output data can be configured to obtain reconstructed sample data based on the first sample data or the transformed portion of the first sample data. For example, in one embodiment, the first sample data is image data, then the data structure of the decoder's output data can be image data of the same size as the first sample data, or image data of the same size as the transformed portion of the first sample data. As another example, in another embodiment, the first sample data is text data, then the data structure of the decoder's output data can be text data of the same length as the first sample data, or text data of the same length as the transformed portion of the first sample data.

[0067] There are many ways to fuse the first coding feature and the second coding feature. For example, it can be a vector or matrix concatenation, or a direct addition, etc. There are no specific limitations in this embodiment.

[0068] It should be noted that before the first and second encoders are pre-trained, they encode based on initialized parameters, which are generally initialized based on human experience. Therefore, the resulting first and second encoded features cannot accurately represent the latent features of the sample data, leading to a large error between the reconstructed sample data obtained by the decoder based on these features and the original sample data or the transformed portion thereof. Pre-training the first and second encoders aims to update their parameters, enabling them to encode based on the updated parameters. This results in first and second encoded features that more accurately represent the latent features of the sample data. This ensures that the reconstructed sample data obtained by the decoder based on the first and second coding features is as close as possible to the original first sample data or the transformed part of the first sample data (so that the decoder output is as close as possible to the original first sample data). Since the decoder needs to use the first coding features obtained by the first encoder to encode the first sample data when decoding the reconstructed sample data, and also needs to use the second coding features obtained by the second encoder to encode the second sample data, the first and second coding features reflect the correlation features between the first and second sample data when reconstructing the sample data. Therefore, after achieving the above pre-training objectives, it can be considered that the first encoder and the second encoder have learned how to mine the correlation between the first and second sample data.

[0069] Step S40: Update the first encoder based on the error between the reconstructed sample data and the first sample data, and calculate the intermediate result for updating the second encoder. Send the intermediate result to the second participant device so that the second participant device can update the second encoder.

[0070] After acquiring the reconstructed sample data, the first participating device can calculate the error between the reconstructed sample data and the first sample data. To reduce this error, it calculates the updated parameters of the first encoder and intermediate results for updating the second encoder. Specifically, the method for calculating the updated parameters and intermediate results for updating the second encoder can employ a gradient descent algorithm. This involves calculating a loss function that represents the error between the reconstructed sample data and the first sample data, or the transformed portion of the first sample data. The loss function can be a root mean square error or other self-loss function. The gradient of the loss function relative to each parameter in the first encoder is calculated, and this gradient is used to update the parameters in the first encoder. The gradient of the loss function relative to the second encoded feature is then used as an intermediate result for updating the second encoder, or the second encoded feature is updated based on its gradient, and the updated second encoded feature is used as an intermediate result for updating the second encoder.

[0071] The first participating device sends the intermediate results to the second participating device, either directly or via a third party. When data privacy is critical, the intermediate results sent from the first participating device to the second participating device can be encrypted. The encryption algorithm can be homomorphic encryption, differential privacy, etc., without restriction. After receiving the intermediate results, the second participating device can update the second encoder. Specifically, when the intermediate result is the gradient value of the second encoded feature, the second participating device can calculate the gradient values ​​of each parameter in the second encoder using the gradient backpropagation algorithm, and then use the gradient values ​​to update the parameters in the second encoder. When the intermediate result is the updated second encoded feature, the second participating device can calculate the loss function between the updated and unupdated second encoded features, calculate the gradient value of this loss function with respect to each parameter in the second encoder, and then use the gradient values ​​to update the parameters in the second encoder.

[0072] Furthermore, in one embodiment, the first participating device may also update the decoder based on the error between the reconstructed sample data and the first sample data. Specifically, the first participating device may calculate the gradient value of the loss function with respect to each parameter in the decoder, and then use the gradient value to update the parameters in the decoder to update the decoder.

[0073] Understandably, since the first and second participating devices interact with encoded features and intermediate results used to update the encoders during the pre-training process of the first and second encoders, and do not directly send the original sample data, the privacy and security of the sample data of each participating device are guaranteed.

[0074] Step S50: After at least one round of iterative updates to the first encoder and the second encoder, a target model is obtained by longitudinal federated learning based on the updated first encoder and the second encoder and the second participating device.

[0075] After at least one round of iterative updates to the first and second encoders, the first and second participating devices can perform longitudinal federated learning based on the updated first and second encoders. The trained model will be referred to as the "target model" for distinction. The process of performing longitudinal federated learning based on the updated first and second encoders can refer to the conventional longitudinal federated learning process, and will not be elaborated in this embodiment.

[0076] In this embodiment, a first encoder and decoder are deployed on the first participating device in the vertical federated learning, and a second encoder is deployed on the second participating device. The first participating device inputs first sample data into the first encoder for encoding to obtain first encoded features, and the second participating device obtains second encoded features by inputting second sample data into the second encoder. The first and second encoded features are fused and input into the decoder to obtain reconstructed sample data. The first encoder is updated based on the error between the reconstructed sample data and the first sample data, and intermediate results for updating the second encoder are calculated. These intermediate results are sent to the second participating device to update the second encoder. This achieves pre-training of the first and second encoders using unlabeled sample data from each participating party in the vertical federated learning. Furthermore, this pre-training process in this embodiment enables the first and second encoders to learn how to mine the correlations between the sample data of each participating party. Compared to using labeled sample data from each participant to perform longitudinal federated learning on randomly initialized or manually initialized first encoders, second encoders, and predictors, this embodiment uses pre-trained first and second encoders as the foundation for longitudinal federated learning. Because the first and second encoders learn how to mine the relationships between the sample data from each participant, they can encode features that are more conducive to the predictor's accurate predictions during the longitudinal federated learning phase. This helps improve the prediction accuracy of the trained target model and shortens the duration of longitudinal federated learning, thus reducing computational resource consumption. Furthermore, it also enables the application of unlabeled sample data to participate in model training in the longitudinal federated learning scenario, thereby improving the model's prediction accuracy based on unlabeled sample data. Additionally, during pre-training, the participating devices do not directly send raw sample data, thus ensuring the privacy and security of the sample data from each participant.

[0077] Furthermore, it should be noted that since the model building optimization method in this embodiment enables the first encoder and the second encoder to learn the ability to mine the correlation between the first sample data and the second sample data, the first sample data and the second sample data do not necessarily need to have a strong correlation. Therefore, compared with the model building optimization method based on contrastive learning, which requires the data of each participant to have a strong correlation, the model building optimization method in this embodiment has a wider range of applications.

[0078] Furthermore, in one embodiment, the first participating device also deploys a predictor, and step S50 includes:

[0079] Step S501: After at least one round of iterative updates to the first encoder and the second encoder, based on the labeled third sample data and the fourth sample data in the second participant device that is aligned with the third sample data, longitudinal federated learning is performed on the predictor and the updated first encoder and the second encoder to obtain a target model including the trained first encoder, the second encoder and the predictor.

[0080] In this embodiment, the first participant may be the party that possesses the sample label data. A predictor is deployed in the first participant's device, and longitudinal federated learning is performed on the predictor and the pre-trained first encoder and second encoder.

[0081] Specifically, after at least one round of iterative updates in conjunction with the second participating device, the first participating device can perform longitudinal federated learning on the predictor, the updated first encoder, and the second encoder based on labeled third sample data and fourth sample data aligned with the third sample data from the second participating device, to obtain a target model including the trained first encoder, second encoder, and predictor. In a specific implementation, the first and second participating devices can first align the labeled samples and use shared labeled sample data for longitudinal federated learning; the first participating device can input the third sample data into the pre-trained first encoder to obtain encoded features (hereinafter referred to as third encoded features), and obtain the encoded features obtained by the second participating device by inputting the fourth sample data into the pre-trained second encoder (hereinafter referred to as fourth encoded features); the first participating device fuses the third and fourth encoded features and inputs them into the predictor for prediction to obtain a prediction result, obtains the label data of the samples corresponding to the third and fourth sample data, and calculates a loss function based on the label data and the prediction result. The loss function can be a supervised loss function such as the cross-entropy loss function; the first participating device... The first participating device calculates the gradient of the loss function relative to the parameters in the first encoder and predictor, and updates the parameters in the first encoder and predictor according to the gradient value to perform one round of updates for the first encoder and predictor. The first participating device calculates the gradient of the loss function relative to the fourth encoded feature, and sends the gradient value to the second participating device, so that the second participating device can calculate the gradient value of the parameters in the second encoder according to the gradient value of the fourth encoded feature, and update the parameters in the second encoder according to the gradient value to perform one round of updates for the second encoder. The first participating device and the second participating device perform at least one round of iterative updates for the first encoder, the second encoder and predictor. Training can end when the stopping condition of the iterative update is met, and the finally updated first encoder, second encoder and predictor are used as the target model to complete the model task.

[0082] It should be noted that since the first encoder and the second encoder have already been pre-trained, the longitudinal federated learning can be regarded as a fine-tuning process of the first encoder and the second encoder. Compared with directly performing longitudinal federated learning on the first encoder and the second encoder that have not been pre-trained, the scheme in this embodiment can improve the convergence speed of the model in the longitudinal federated learning process, save training time, and thus save computing resources.

[0083] Furthermore, in a specific implementation, after the target model is trained, the first encoder and predictor in the target model can still be deployed on the first participant device, and the second encoder can still be deployed on the second participant device, with the first participant device and the second participant device jointly completing the model task. Alternatively, the target model can be deployed on the same device, with that device independently completing the model task.

[0084] Furthermore, based on the first embodiment described above, a second embodiment of the model construction optimization method of the present invention is proposed. In this embodiment, step S10 includes:

[0085] Step S101: Transform a portion of the data in the first sample data, and input the transformed first sample data into the first encoder for encoding to obtain the first encoded feature.

[0086] In this embodiment, the first participating device can transform a portion of the data in the first sample data while leaving the rest unchanged. The transformed first sample data is then input to the first encoder for encoding to obtain a first encoded feature. It is understood that the transformed first sample data includes both the untransformed portion and the transformed portion. Transforming a portion of the first sample data and inputting it to the first encoder to obtain the first encoded feature, then inputting the second sample data to the second encoder to obtain the second encoded feature, and finally inputting both the first and second encoded features to the decoder to obtain reconstructed sample data, aims to allow the second sample data to help reconstruct the transformed portion of the first sample data, thereby further improving the ability of the first and second encoders to discover the correlation between the first and second sample data. In this context, the second sample data helps reconstruct the transformed portion of the first sample data. This means that, for the encoder and decoder, the transformation of a portion of the first sample data is equivalent to the loss of a portion. The decoder decodes the transformed first sample data based on the first coding features of the transformed first sample data and the second coding features of the second sample data to obtain the reconstructed sample data. The reconstructed sample data is then pre-trained to make it closer to the original first sample data and the transformed portion of the first sample data. In other words, the second sample data helps reconstruct the transformed portion of the first sample data.

[0087] In specific implementations, different transformation methods are adopted depending on the type of sample data. This embodiment does not limit the transformation method. For example, a portion of the data in the first sample data can be replaced with a preset value, such as replacing it with 0.

[0088] Furthermore, in a specific implementation, when the first participating device transforms a portion of the data in the first sample data and then inputs it to the first encoder for encoding, the reconstructed sample data can be reconstructed data for the first sample data, or it can be reconstructed data for only the transformed portion of the first sample data.

[0089] Further, in one embodiment, the reconstructed sample data is reconstructed data for the transformed portion of the first sample data, and the steps in step S40 of updating the first encoder based on the error between the reconstructed sample data and the first sample data and calculating the intermediate result for updating the second encoder include:

[0090] Step S401: Calculate the self-loss function that characterizes the error between the reconstructed sample data and the transformed portion of the first sample data;

[0091] In this embodiment, when the reconstructed sample data is the reconstructed data of the transformed portion of the first sample data, the first participating device can calculate a loss function (hereinafter referred to as the self-loss function for distinction) that characterizes the error between the reconstructed sample data and the transformed portion of the first sample data. The self-loss function can be a loss function commonly used in autoencoders, such as the root mean square error, and is not limited in this embodiment.

[0092] Step S402: Calculate the first gradient value of the self-loss function with respect to the parameters in the first encoder, and calculate the second gradient value of the self-loss function with respect to the second encoded feature;

[0093] After calculating the self-loss function, the first participating device calculates the gradient value of the self-loss function relative to the parameters in the first encoder (hereinafter referred to as the first gradient value). It should be noted that the first encoder includes multiple parameters, so the first gradient value includes the gradient value corresponding to each parameter. The first participating device also calculates the gradient value of the self-loss function relative to the second encoded feature (hereinafter referred to as the second gradient value).

[0094] Step S403: Update the parameters in the first encoder according to the first gradient value to update the first encoder;

[0095] After obtaining the first gradient value, the first participating device can update the parameters in the first encoder based on the first gradient value. After updating the parameters, the first encoder has completed one round of updates. The specific process of updating parameters based on the gradient value will not be elaborated here.

[0096] Step S404: Use the second gradient value as an intermediate result for updating the second encoder.

[0097] The first participating device uses the calculated second gradient value as an intermediate result to update the second encoder. Further, the first participating device sends the intermediate result to the second participating device. Upon receiving the intermediate result, i.e., the second gradient value, the second participating device can calculate the gradient values ​​of each parameter in the second encoder based on the second gradient value, and update the parameters in the second encoder using the calculated gradient values, thereby completing one round of updating the second encoder.

[0098] Furthermore, in one embodiment, the step of transforming a portion of the data in the first sample data in step S101 includes:

[0099] Step S1011: When the first sample data includes multiple attribute values, some attribute values ​​in the first sample data are transformed into preset values ​​or random noise is added.

[0100] In this embodiment, when the first sample data is tabular data including multiple attribute values, some attribute values ​​in the first sample data can be transformed to preset values ​​or random noise can be added to transform a portion of the data. The preset values ​​can be set as needed, for example, to 0. Adding random noise can involve adding a random number to the attribute values; this random number can be generated using a random algorithm, for example, adding normally distributed noise. Which attribute values ​​in the first sample data are transformed can be predetermined, or several attribute values ​​can be randomly selected according to a certain proportion. For example, 2 attribute values ​​can be randomly selected from 10 attribute values ​​at a 20% ratio and set to 0 or have noise added.

[0101] Accordingly, the reconstructed sample data can be either the transformed attribute values ​​or all the attribute values ​​of the reconstructed first sample data.

[0102] Step S1012: When the first sample data is image data, the pixel values ​​of some pixels in the image data are transformed into preset pixel values;

[0103] When the first sample data is image data, the pixel values ​​of some pixels in the image data can be transformed to preset pixel values ​​to transform a portion of the data in the first sample data. The preset pixel values ​​can be set as needed to achieve the effect of occluding certain areas of the image data.

[0104] Accordingly, the reconstructed sample data can be either the part of the original image data that was occluded, or the entire original image data.

[0105] Step S1013: When the first sample data is text data, some words in the text data are transformed into preset words.

[0106] When the first sample data is text data, some words in the text data can be transformed into preset words to transform a portion of the data in the first sample data. The preset words can be set as needed.

[0107] Accordingly, reconstructing sample data can be either reconstructing the replaced words in the text data or reconstructing the entire text data.

[0108] Further, in one embodiment, step S30 includes:

[0109] Step S301: Input the first encoded feature and the second encoded feature into the projection model for feature cross-processing, and input the fused feature obtained after feature cross-processing into the decoder for decoding to obtain reconstructed sample data.

[0110] The first participating device can also deploy a projection model. This projection model can be implemented using any model structure that allows multiple input features to undergo cross-fusion, such as a one- or two-layer fully connected model. After obtaining the first and second encoded features, the first participating device can input these features into the projection model for feature cross-processing. This process involves applying parameters from the projection model to the input data to be projected. Depending on the type of projection model, the types and number of parameters vary, as does the specific method of applying these parameters to the data to be projected. For example, when the projection model is fully connected, applying the parameters to the data to be projected involves matrix multiplication of the parameter matrices of the fully connected model with the data to be projected.

[0111] In a specific implementation, the first participating device can add or concatenate the first and second encoded features and then input them into the projection model. This can be achieved by setting the data structure of the input data of the projection model. For example, if the data structure of the input data of the projection model is set to a 256-dimensional vector, and the first and second encoded features are each 128-dimensional vectors, then the first and second encoded features can be concatenated to obtain a 256-dimensional vector, which is then input into the projection model.

[0112] Furthermore, in one embodiment, when the first participating device updates the first encoder and decoder, it can also update the projection model at the same time. Specifically, it can calculate the gradient value of the loss function with respect to each parameter in the projection model, and update each parameter in the projection model according to the gradient value to update the projection model.

[0113] It should be noted that by deploying a projection model in the first participating device to perform feature cross-processing on the first and second encoded features before inputting them into the decoder, the first and second encoded features can be further cross-combined. As mentioned earlier, the purpose of pre-training the first and second encoders is to make the reconstructed sample data obtained by the decoder based on the first and second encoded features as close as possible to the original first sample data or the transformed part of the first sample data. The first and second encoded features reflect the correlation between the first and second sample data during the reconstruction of the sample data. The encoded features obtained by cross-combining the first and second encoded features can better reflect the correlation between the first and second sample data during the reconstruction of the sample data. Therefore, when the cross-combined encoded features are input into the decoder for decoding, it is more conducive to the first and second encoders learning the ability to mine the correlation between the first and second sample data through the pre-training process.

[0114] Furthermore, based on the first and / or second embodiments described above, a third embodiment of the model construction optimization method of the present invention is proposed. In this embodiment, the method is applied to the second participating party's equipment, referring to... Figure 3 As shown, the method includes the following steps:

[0115] Step A10: Input the second sample data into the second encoder for encoding to obtain the second encoded feature;

[0116] Step A20: The second encoded feature is sent to the first participating device so that the first participating device can fuse the first encoded feature and the second encoded feature and input them into the decoder to obtain reconstructed sample data. The first encoder is updated based on the error between the reconstructed sample data and the first sample data, and intermediate results for updating the second encoder are calculated. The first encoded feature is obtained by the first participating device inputting the first sample data into the first encoder for encoding.

[0117] Step A30: Obtain the intermediate result and update the second encoder based on the intermediate result;

[0118] Without A40, after at least one round of iterative updates to the first encoder and the second encoder, the target model is obtained by longitudinal federated learning based on the updated first encoder and the second encoder and the first participating device.

[0119] The specific implementation of steps A10 to A40 in this embodiment can be referred to the specific implementation of steps S10 to S50 in the above embodiment, and will not be repeated here.

[0120] Further, in one embodiment, the first participant can be a bank, possessing business data generated when users (samples) conduct business with the bank, specifically including deposit and withdrawal records, loan records, repayment records, etc., and may also possess some user loan repayment risk labels. The second participant can be an e-commerce institution, possessing purchase record data generated when users purchase goods on the e-commerce platform, specifically including purchase amount, payment method, number of returns and exchanges, etc. The first and second participants can set up a model task to predict user loan repayment risk based on user business data and purchase record data. A first encoder and predictor are deployed in the first participant's device, and a second encoder is deployed in the second participant's device. By performing vertical federated learning on the first encoder, second encoder, and predictor using labeled shared user business data and purchase record data, a loan repayment risk prediction model is obtained to predict user loan repayment risk. Before performing vertical federated learning on the first encoder, second encoder, and predictor, the first and second participant devices use unlabeled shared user business data (first sample data) and purchase record data (second sample data) to pre-train the first encoder and second encoder. One round of pre-training process may include:

[0121] The first participating device inputs the business data of the shared users into the first encoder for encoding to obtain the first encoded feature;

[0122] The first participating device acquires the second coding feature, wherein the second coding feature is obtained by the second participating device inputting the purchase record data of the common users into the second encoder for encoding;

[0123] The first participating device fuses the first and second encoded features and inputs them into the decoder to obtain the reconstructed sample data;

[0124] The first encoder is updated based on the error between the reconstructed sample data and the business data, and an intermediate result for updating the second encoder is calculated. The intermediate result is then sent to the second participant device for updating the second encoder.

[0125] After at least one round of iterative updates to the first encoder and the second encoder, the first participating device can obtain a loan repayment risk prediction model by performing longitudinal federated learning with the second participating device based on the updated first encoder and second encoder.

[0126] This embodiment utilizes unlabeled shared user business data and purchase record data from banks and e-commerce institutions to pre-train the first and second encoders. This pre-training process enables the first and second encoders to learn how to mine the correlation between user business data and purchase record data. Compared to using labeled sample data for longitudinal federated learning on randomly initialized or manually initialized first and second encoders and predictors, this embodiment uses pre-trained first and second encoders as the foundation for longitudinal federated learning. Because the first and second encoders have learned how to mine the relationship between user business data and purchase record data, during the longitudinal federated learning stage, they can encode features that are more conducive to the predictor obtaining accurate loan repayment risk prediction results. This helps improve the prediction accuracy of the trained loan repayment risk prediction model and shortens the duration of longitudinal federated learning, thereby reducing computational resource consumption during the longitudinal federated learning stage. Furthermore, it also enables the application of unlabeled business data and purchase record data to the longitudinal federated learning scenario for model training, thus achieving improved prediction accuracy of the loan repayment risk model based on unlabeled business data and purchase record data. In addition, during the pre-training process, the devices of each participating party did not directly send raw business data and purchase record data, thus ensuring the privacy and security of users.

[0127] Furthermore, the first participating device can transform a portion of the business data, and input the transformed business data into the first encoder for encoding to obtain the first encoded feature. Specifically, transforming a portion of the business data can involve replacing some attribute values ​​with preset values ​​or adding random noise. For example, user deposit and withdrawal records in the business data can be transformed.

[0128] Furthermore, in one embodiment, such as Figure 4 The diagram illustrates a pre-training framework based on data reconstruction. Party A has feature data and labels, while Party B only has feature data. Party A and Party B are either the first participating device or the second participating device. Party A and Party B each deploy their own encoding model (also called an encoder). A and Encoder B Output the encoded features f respectively A and f B In addition, A has a projection model P.A Decoder (also known as a decoding model) A Basic pre-training process:

[0129] Initialization: Each participant initializes its respective model Encoder. A Encoder B P A and Decoder A Party A and Party B align their sample data using encryption.

[0130] ① Party A will provide sample data X A Through data processing steps, X is obtained. A,prossed The purpose is to transform the original sample data. Common methods include randomly perturbing some features, setting them to predefined values, such as 0; or adding a random perturbation, such as noise from a normally distributed system. For example, randomly perturbing 20% ​​of the values ​​in a 10-dimensional feature by setting them to 0 or adding noise. For image or text data, commonly used perturbation schemes can be used. For example, for images, a portion of the image area can be obscured; for text, some words can be replaced. Accordingly, in this case, a flag can be added to the text to indicate the replaced words, so that the loss function can be calculated between the original text data and the reconstructed text data indicating the words in the flag.

[0131] ② Parties A and B input their respective sample data into their respective encoding models (Encoders). A and Encoder B Encoding is performed to obtain the encoded feature f A and f B Party B will f B Before transmission, f can be encrypted using various methods and sent to party A. B Encryption can be used to further enhance data security; for example, differential privacy methods can be employed for encryption.

[0132] ③ This step is the data reconstruction process, the purpose of which is to allow the model to model the relationship between sample data from side A and sample data from side B, so that the features of sample data from side B can help reconstruct sample data from side A. Side A will f A and f B To achieve fusion, either splicing or addition methods can be used, and then the result can be input into the projection model P. A , to obtain z A Projection model P A A one- or two-layer fully connected model can be used to handle feature interactions. Then, Z... A Input to the decoding model Decoder AThe output h is obtained with the same shape as the sample data input by A. A .

[0133] For X A and h A To calculate the loss, a commonly used loss function is the root mean square error, which takes the following form:

[0134] L=Σ(X A,i -h A,i ) 2

[0135] Where i represents the sample number, when calculating the loss function, we iterate through each sample in the current training batch of both A and B sides and add up the root mean square error calculated for each sample.

[0136] The structure of the decoding model can be determined based on the specific format of the sample data. For example, for tabular data, the decoding model can use a fully connected network, and the data structure of its last layer output is the same as X. A The data structure is the same as X, or it is the same as X. A Detailed data structure diagram of the transformed portion of the data; for image data, the decoding model can use a model containing a deconvolution module to transform the input z-axis... A The image is converted to the same size as the original image input by A and then fed into the decoding model. For text data, the decoding model can use a transformer or BERT model, and correspondingly, the loss function can be designed to characterize the error between the text output by the decoding model and the transformed text in the original input text of A.

[0137] ④ Party A updates the Encoder model based on the loss function. A P A and Decoder A A calculates the value of f. B partial derivative g B It is then sent back to Party B, and the transmission process can be encrypted using methods such as differential privacy.

[0138] ⑤B based on the obtained partial derivative g B Update your own encoding model Encoder B .

[0139] like Figure 5 As shown, for Encoder A and Encoder B After pre-training, the Encoder can be... A and Encoder B Fine-tuning is performed, a process similar to regular longitudinal federated learning, except that A and B use the Encoder obtained from the previous pre-training step.A and Encoder B Each model is initialized with weights.

[0140] First, A and B respectively use the Encoder obtained from the model pre-training step. A and Encoder B Initialization. Party A also has a classification model (also called a prediction model) named Classifier. A And initialized randomly.

[0141] Next, follow these steps to perform federated learning:

[0142] ① Party A and Party B are based on aligned and labeled data (X) A Y A ) and X B Through the model Encoder A and Encoder B Encode to obtain f A and f B Party B will f B Pass it to Party A.

[0143] ②Party A will f A and f B After fusion (usually by concatenation, but addition can also be used), it is input into the classification model Classifier. A The predicted result Y is obtained pred Based on the prediction results and label Y A Calculate the cross-entropy loss function LCE.

[0144] ③ Party A updates the Encoder based on the cross-entropy loss function LCE. A and Classifier A and for f B Calculate the partial derivative g B Return it to Party B.

[0145] Based on the partial derivative g obtained by Party B B Update your own encoding model Encoder B .

[0146] Furthermore, embodiments of the present invention also propose a computer-readable storage medium storing a model building optimization program, wherein when the model building optimization program is executed by a processor, it implements the steps of the model building optimization method described below.

[0147] The present invention also proposes a computer program product, including a computer program that, when executed by a processor, implements the steps of the model building optimization method as described above.

[0148] The various embodiments of the model building and optimization device, computer-readable storage medium, and computer program product of the present invention can all refer to the various embodiments of the model building and optimization method of the present invention, and will not be repeated here.

[0149] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0150] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0151] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0152] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A model construction optimization method, characterized in that, The method is applied to a first participating device in longitudinal federated learning, wherein the first participating device deploys a first encoder and a decoder, and a second participating device in longitudinal federated learning deploys a second encoder. The method includes the following steps: A portion of the data in the first sample data is transformed, and the transformed first sample data is input into the first encoder for encoding to obtain a first encoded feature; wherein, when the first sample data includes multiple attribute values, some attribute values ​​in the first sample data are transformed into preset values ​​or random noise is added; when the first sample data is image data, the pixel values ​​of some pixels in the image data are transformed into preset pixel values; when the first sample data is text data, some words in the text data are transformed into preset words; A second encoding feature is obtained, wherein the second encoding feature is obtained by the second participating device inputting the second sample data into the second encoder for encoding; wherein the second sample data and the first sample data belong to the same sample object, and the first sample data and the second sample data are feature data of the same sample object in different feature spaces; The first coding feature and the second coding feature are fused and then input into the decoder for decoding to obtain reconstructed sample data. The reconstructed sample data is the reconstructed data of the transformed part of the first sample data. The first encoder is updated based on the error between the reconstructed sample data and the first sample data, and an intermediate result for updating the second encoder is calculated. The intermediate result is then sent to the second participating device for the second participating device to update the second encoder. After at least one round of iterative updates to the first encoder and the second encoder, the target model is obtained by longitudinal federated learning based on the updated first encoder and the second encoder and the second participating device. The steps of updating the first encoder based on the error between the reconstructed sample data and the first sample data, and calculating the intermediate result for updating the second encoder, include: Calculate the self-loss function that characterizes the error between the reconstructed sample data and the transformed portion of the first sample data; Calculate the first gradient value of the self-loss function with respect to the parameters in the first encoder, and calculate the second gradient value of the self-loss function with respect to the second encoded feature; Update the parameters in the first encoder according to the first gradient value to update the first encoder; The second gradient value is used as an intermediate result for updating the second encoder.

2. The model construction and optimization method as described in claim 1, characterized in that, The first participating device also deploys a projection model, and the step of fusing the first encoded features and the second encoded features and inputting them into the decoder for decoding to obtain reconstructed sample data includes: The first encoded feature and the second encoded feature are input into the projection model for feature cross-processing, and the fused feature obtained after feature cross-processing is input into the decoder for decoding to obtain reconstructed sample data.

3. The model construction optimization method as described in any one of claims 1 to 2, characterized in that, The first participating device also deploys a predictor, and the step of obtaining the target model by performing longitudinal federated learning with the second participating device based on the updated first encoder and second encoder includes: Based on the labeled third sample data and the fourth sample data in the second participant device that is aligned with the third sample data, longitudinal federated learning is performed on the predictor and the updated first encoder and second encoder to obtain a target model including the trained first encoder, second encoder and predictor.

4. A model construction optimization method, characterized in that, The method is applied to a second participating device in longitudinal federated learning, wherein a first participating device deploys a first encoder and a decoder, and the second participating device deploys a second encoder. The method includes the following steps: The second sample data is input into the second encoder for encoding to obtain the second encoded feature; The second encoded feature is sent to the first participating device, which then fuses the first and second encoded features and inputs the fused features into the decoder to obtain reconstructed sample data. The first participating device calculates a self-loss function representing the error between the reconstructed sample data and the transformed portion of the first sample data. It also calculates a first gradient value of the self-loss function relative to the parameters in the first encoder and a second gradient value of the self-loss function relative to the second encoded feature. The parameters in the first encoder are updated based on the first gradient value to update the first encoder. The second gradient value is used as an intermediate result for updating the second encoder. The first encoded feature is derived by the first participating device from a portion of the first sample data. The transformation is performed, and the transformed first sample data is input into the first encoder for encoding. The reconstructed sample data is the reconstructed data of the transformed portion of the first sample data. Specifically, when the first sample data includes multiple attribute values, some attribute values ​​in the first sample data are transformed to preset values ​​or random noise is added; when the first sample data is image data, the pixel values ​​of some pixels in the image data are transformed to preset pixel values; when the first sample data is text data, some words in the text data are transformed to preset words; the second sample data and the first sample data belong to the same sample object, and the first sample data and the second sample data are feature data of the same sample object in different feature spaces. Obtain the intermediate result and update the second encoder based on the intermediate result; After at least one round of iterative updates to the first encoder and the second encoder, a target model is obtained by longitudinal federated learning based on the updated first encoder and the second encoder and the first participating device.

5. A model building optimization device, characterized in that, The model building optimization device includes: a memory, a processor, and a model building optimization program stored in the memory and executable on the processor. When the model building optimization program is executed by the processor, it implements the steps of the model building optimization method as described in any one of claims 1-3, or the steps of the model building optimization method as described in claim 4.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a model building optimization program, which, when executed by a processor, implements the steps of the model building optimization method as described in any one of claims 1-3, or the steps of the model building optimization method as described in claim 4.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the model building optimization method as described in any one of claims 1-3, or the steps of the model building optimization method as described in claim 4.