Data Processing Method and Apparatus, Storage Medium, and Electronic Device
Through the combination of vertical federated learning and segmented learning, data fusion and decoding and dimensionality reduction of the bottom model and the top model are used to solve the problem of low data security in cross-platform user preference prediction, and improve prediction accuracy and privacy protection without leaking user data.
Patent Information
- Application Number
- CN202210571451.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-24
AI Technical Summary
In cross-platform user preference prediction, it is difficult for the existing technology to effectively protect user data privacy and security, especially when multi-platform cooperation poses a risk of data leakage.
The vertical federated learning and segmented learning method are adopted to combine the bottom model and the top model of the first participant and the second participant, and the data is fused using split layers, and the K-dimensional inference results are processed by decoding the dimension reduction module to obtain the M-dimensional inference results to improve data security.
Without leaking user data, the accuracy and data security of user preference prediction are improved and user privacy is protected.
Smart Images

Figure CN115130548B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular, to a data processing method, apparatus, storage medium, and electronic device. Background Art
[0002] In the related art, predicting user preferences can recommend media resources that the user is interested in. For example, video recommendation, product recommendation, and so on.
[0003] Based on the user's historical behavior data, a neural network model obtained through machine learning can predict the user's behavior. The larger the amount of data in the training samples, the higher the prediction accuracy of the neural network model. Therefore, the amount of data in the training sample data has always been an important factor affecting the prediction accuracy of user preferences. The user's historical behavior data accumulated on a single platform is limited. If multiple different platforms provide the user's historical behavior data together, the amount of data in the training samples can be increased, and thus the prediction accuracy of the neural network model can be improved. However, in the way of multi-platform cooperation, there is a problem of data leakage, and it is difficult to ensure the privacy and security of user data.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present application provide a data processing method, apparatus, storage medium, and electronic device to at least solve the technical problem of low data security.
[0006] According to one aspect of the embodiments of the present application, a data processing method is provided, including: obtaining a first forward output and a second forward output, where the first forward output is an output obtained by a first bottom model in a first participant processing first feature data, and the second forward output is an output obtained by a second bottom model in a second participant processing second feature data; processing the first forward output and the second forward output by a top model in the second participant to obtain a K-dimensional inference result, where K is a positive integer greater than 2; decoding and dimension-reducing the K-dimensional inference result by a decoding and dimension-reducing module in the second participant to obtain an M-dimensional inference result, where M is less than K and greater than or equal to 2 and is a positive integer.
[0007] Optionally, after obtaining the first forward output and the second forward output, the method further includes: fusing the first forward output and the second forward output through a splitting layer in the second participant to obtain a fusion result; the processing the first forward output and the second forward output through a top model in the second participant to obtain a K-dimensional inference result includes: processing the fusion result through the top model in the second participant to obtain the K-dimensional inference result.
[0008] Optionally, the fusing the first forward output and the second forward output through a splitting layer in the second participant to obtain a fusion result includes: concatenating the first forward output and the second forward output through the splitting layer in the second participant to obtain the fusion result; or, performing an averaging process on the first forward output and the second forward output through the splitting layer in the second participant to obtain the fusion result.
[0009] Optionally, before obtaining the first forward output and the second forward output, the method further includes: obtaining a first set of sample feature data, a second set of sample feature data, and an M-dimensional set of known sample labels with corresponding relationships, where each M-dimensional known sample label in the M-dimensional set of known sample labels is used to represent an actual sample label; performing dimensionality-raising encoding on the M-dimensional set of known sample labels through a dimensionality-raising encoding module to obtain a K-dimensional set of known encoded sample labels; using the first set of sample feature data, the second set of sample feature data, and the K-dimensional set of known encoded sample labels to jointly train a first bottom model to be trained in the first participant, a second bottom model to be trained in the second participant, and a top model to be trained until a first end condition is met, ending the training, and obtaining the first bottom model, the second bottom model, and the top model.
[0010] Optionally, jointly training the to-be-trained first bottom model in the first party, the to-be-trained second bottom model in the second party, and the to-be-trained top model by using the first sample feature data set, the second sample feature data set, and the K-dimensional known encoded sample label set includes: performing the j-th round of joint training on the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model through the following steps, where j is a positive integer greater than or equal to 1, and the first bottom model, the second bottom model, and the top model obtained in the 0-th round of training are the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model that have not been trained, including: inputting S first sample feature data used in the j-th round in the first sample feature data set into the first bottom model obtained in the (j - 1)-th round of training to obtain S first sample forward outputs output in the j-th round of training, where S is a positive integer greater than or equal to 1; inputting S second sample feature data used in the j-th round in the second sample feature data set into the second bottom model obtained in the (j - 1)-th round of training to obtain S second sample forward outputs output in the j-th round of training; inputting the S first sample forward outputs output in the j-th round of training and the S second sample forward outputs output in the j-th round of training into the top model obtained in the (j - 1)-th round of training to obtain S K-dimensional sample inference results output in the j-th round of training; ending the training to obtain the first bottom model, the second bottom model, and the top model when the S K-dimensional sample inference results output in the j-th round of training and the S K-dimensional known encoded sample labels used in the j-th round satisfy the first end condition, and the K-dimensional known encoded sample labels used in the j-th round have the corresponding relationship with the first sample feature data used in the j-th round and the second sample feature data used in the j-th round.
[0011] Optionally, the method further includes: when the S K-dimensional sample inference results output in the j-th round of training and the S K-dimensional known encoded sample labels used in the j-th round do not satisfy the first end condition, adjusting the model parameters of the first bottom model, the second bottom model, and the top model obtained in the (j - 1)-th round of training through backpropagation.
[0012] Optionally, adjusting the model parameters of the top model obtained in the (j - 1)-th round through direction propagation includes: the second party obtains S first gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S K-dimensional sample inference results, and adjusts the model parameters of the top model obtained in the (j - 1)-th round through the S first gradients;
[0013] Adjusting the model parameters of the second bottom model obtained from the (j - 1)-th round of training through forward propagation includes: the second party obtains S second gradients based on the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the model parameters of the second bottom model obtained from the (j - 1)-th round of training, and adjusts the model parameters of the second bottom model obtained from the (j - 1)-th round of training through the S second gradients;
[0014] Adjusting the first bottom model obtained from the (j - 1)-th round of training through backpropagation includes: the second party obtains S third gradients based on the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S first sample forward outputs; the second party normalizes the S third gradients through the second norm to obtain S normalized gradients; the second party sends the S normalized gradients to the first party, and the first party adjusts the first bottom model obtained from the (j - 1)-th round of training through the S normalized gradients.
[0015] Optionally, the method further includes: obtaining an M-dimensional known sample label set, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label; performing dimension elevation on the M-dimensional known sample labels in the M-dimensional known sample label set through a dimension elevation module to obtain a K-dimensional known sample label set; jointly training the encoding module to be trained and the decoding module to be trained using the K-dimensional known sample label set until the second end condition is satisfied among the first loss function, the second loss function, and the third loss function, ending the training, and obtaining the encoding module and the decoding module, where the decoding and dimension reduction module includes the decoding module and the dimension reduction module, the dimension elevation and encoding module includes the encoding module and the dimension elevation module, and the dimension reduction module corresponds to the dimension elevation module in the dimension elevation and encoding module; where the first loss function is the loss function between the K-dimensional decoding output of the decoder to be trained and the corresponding K-dimensional known sample label, the second loss function is the loss function between the K-dimensional encoding output of the encoder to be trained and the corresponding K-dimensional known sample label, and the third loss function is the information entropy function of the K-dimensional encoding output of the encoder to be trained.
[0016] Optionally, the joint training of the to-be-trained encoding module and the to-be-trained decoding module using the set of K-dimensional known sample labels includes: performing the i-th round of joint training on the to-be-trained encoding module and the to-be-trained decoding module through the following steps, where i is a positive integer greater than or equal to 1, and the encoding module and the decoding module obtained in the 0-th round of training are the untrained to-be-trained encoding module and the to-be-trained decoding module: inputting the K-dimensional known sample labels used in the i-th round in the set of K-dimensional known sample labels into the encoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional encoded output obtained in the i-th round of training; inputting the K-dimensional encoded output obtained in the i-th round of training into the decoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional decoded output obtained in the i-th round of training; wherein, when the first loss function between the K-dimensional decoded output obtained in the i-th round of training and the K-dimensional known sample labels, the second loss function between the K-dimensional encoded output obtained in the i-th round of training and the K-dimensional known sample labels, and the third loss function of the K-dimensional encoded output obtained in the i-th round of training do not satisfy the second end condition, adjusting the parameters in the to-be-trained encoding module and the to-be-trained decoding module; otherwise, ending the training to obtain the encoding module and the decoding module.
[0017] Optionally, the performing the i-th round of joint training on the to-be-trained encoding module and the to-be-trained decoding module further includes: inputting the K-dimensional decoded output obtained in the i-th round of training and the K-dimensional known sample labels into the first loss function to obtain a first loss value; inputting the K-dimensional encoded output obtained in the i-th round of training and the K-dimensional known sample labels into the second loss function to obtain a second loss value; inputting the K-dimensional encoded output obtained in the i-th round of training into the third loss function to obtain a third loss value; determining that the second end condition is satisfied when the difference between the first loss value, the second loss value, and the third loss value is less than or equal to a preset threshold.
[0018] Optionally, the decoding and dimension reduction of the K-dimensional inference result by the decoding and dimension reduction module in the second party to obtain an M-dimensional inference result includes: decoding the K-dimensional inference result through the decoding module in the decoding and dimension reduction module to obtain a K-dimensional decoded inference result; performing dimension reduction on the K-dimensional decoded inference result through the dimension reduction module in the decoding and dimension reduction module to obtain the M-dimensional inference result.
[0019] Optionally, dimensionality elevation of the M-dimensional known sample labels in the dimensionality elevation coding module includes: obtaining K K-dimensional one-hot coding vectors, randomly dividing the K K-dimensional one-hot coding vectors into M groups, or dividing them into M groups according to a preset rule, to obtain M groups of K-dimensional one-hot coding vectors; mapping the M-dimensional known sample labels to the M groups of K-dimensional one-hot coding vectors respectively to obtain K-dimensional known sample labels corresponding to the M-dimensional known sample labels; dimensionality reduction of the K-dimensional decoded inference result by a dimensionality reduction module includes: performing an inner product operation on the K-dimensional inference result and M K-dimensional decoding vectors to obtain the M-dimensional inference result, where the M K-dimensional decoding vectors are decoding vectors obtained according to the M groups of K-dimensional one-hot coding.
[0020] According to another aspect of the embodiments of the present application, there is also provided a data processing method, including: obtaining a first sample feature data set, a second sample feature data set, and an M-dimensional known sample label set with corresponding relationships, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label; performing dimensionality elevation coding on the M-dimensional known sample label set through a dimensionality elevation coding module to obtain a K-dimensional known coded sample label set; using the first sample feature data set, the second sample feature data set, and the K-dimensional known coded sample label set to jointly train a first bottom model to be trained in a first party, a second bottom model to be trained in a second party, and a top model to be trained until a first end condition is satisfied, ending the training, and obtaining the first bottom model in the first party, the second bottom model in the second party, and the top model.
[0021] Optionally, jointly training the to-be-trained first bottom model in the first party, the to-be-trained second bottom model in the second party, and the to-be-trained top model by using the first sample feature data set, the second sample feature data set, and the K-dimensional known encoded sample label set includes: performing the j-th round of joint training on the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model through the following steps, where j is a positive integer greater than or equal to 1, and the first bottom model, the second bottom model, and the top model obtained in the 0-th round of training are the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model that have not been trained, including: inputting S first sample feature data used in the j-th round in the first sample feature data set into the first bottom model obtained in the (j - 1)-th round of training to obtain S first sample forward outputs output in the j-th round of training, where S is a positive integer greater than or equal to 1; inputting S second sample feature data used in the j-th round in the second sample feature data set into the second bottom model obtained in the (j - 1)-th round of training to obtain S second sample forward outputs output in the j-th round of training; inputting the S first sample forward outputs output in the j-th round of training and the S second sample forward outputs output in the j-th round of training into the top model obtained in the (j - 1)-th round of training to obtain S K-dimensional sample inference results output in the j-th round of training; ending the training to obtain the first bottom model, the second bottom model, and the top model when the S K-dimensional sample inference results output in the j-th round of training satisfy the first ending condition with respect to the S K-dimensional known encoded sample labels used in the j-th round, and the K-dimensional known encoded sample labels used in the j-th round have the corresponding relationship with the first sample feature data used in the j-th round and the second sample feature data used in the j-th round.
[0022] Optionally, the method further includes: when the S K-dimensional sample inference results output in the j-th round of training do not satisfy the first ending condition with respect to the S K-dimensional known encoded sample labels used in the j-th round, adjusting the model parameters of the first bottom model obtained in the (j - 1)-th round of training, the second bottom model, and the top model obtained in the (j - 1)-th round of training through backpropagation.
[0023] Optionally, adjusting the model parameters of the top model obtained in the (j - 1)-th round of training through direction propagation includes: the second party obtaining S first gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S K-dimensional sample inference results, and adjusting the model parameters of the top model obtained in the (j - 1)-th round of training through the S first gradients;
[0024] Adjusting the model parameters of the second bottom model obtained in the (j-1)-th round of training through forward propagation includes: the second party obtaining S second gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the model parameters of the second bottom model obtained in the (j-1)-th round of training, and adjusting the model parameters of the second bottom model obtained in the (j-1)-th round of training through the S second gradients;
[0025] Adjusting the first bottom model obtained in the (j-1)-th round of training through backpropagation includes: the second party obtaining S third gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S first sample forward outputs; the second party normalizing the S third gradients through the second norm to obtain S normalized gradients; the second party sending the S normalized gradients to the first party, and the first party adjusting the first bottom model obtained in the (j-1)-th round of training through the S normalized gradients.
[0026] Optionally, the method further includes: obtaining an M-dimensional known sample label set, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label; performing dimension elevation on the M-dimensional known sample labels in the M-dimensional known sample label set through a dimension elevation module to obtain a K-dimensional known sample label set; using the K-dimensional known sample label set to jointly train a to-be-trained encoding module and a to-be-trained decoding module until a second end condition is satisfied among a first loss function, a second loss function, and a third loss function, ending the training, and obtaining the encoding module in the dimension elevation encoding module and the decoding module in the decoding dimension reduction module, where the decoding dimension reduction module includes the decoding module and a dimension reduction module, the dimension elevation encoding module includes the encoding module and the dimension elevation module, and the dimension reduction module corresponds to the dimension elevation module in the dimension elevation encoding module; where the first loss function is the loss function between the K-dimensional decoding output output by the to-be-trained decoder and the corresponding K-dimensional known sample label, the second loss function is the loss function between the K-dimensional encoding output output by the to-be-trained encoder and the corresponding K-dimensional known sample label, and the third loss function is the information entropy function of the K-dimensional encoding output output by the to-be-trained encoder.
[0027] Optionally, jointly training the to-be-trained encoding module and the to-be-trained decoding module using the K-dimensional known sample label set includes: performing the i-th round of joint training on the to-be-trained encoding module and the to-be-trained decoding module through the following steps, where i is a positive integer greater than or equal to 1, and the encoding module and the decoding module obtained in the 0-th round of training are the untrained to-be-trained encoding module and the to-be-trained decoding module: inputting the K-dimensional known sample labels used in the i-th round in the K-dimensional known sample label set into the encoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional encoded output obtained in the i-th round of training; inputting the K-dimensional encoded output obtained in the i-th round of training into the decoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional decoded output obtained in the i-th round of training; where, in the case that the first loss function between the K-dimensional decoded output obtained in the i-th round of training and the K-dimensional known sample labels, the second loss function between the K-dimensional encoded output obtained in the i-th round of training and the K-dimensional known sample labels, and the third loss function of the K-dimensional encoded output obtained in the i-th round of training do not satisfy the second end condition, adjusting the parameters in the to-be-trained encoding module and the to-be-trained decoding module; otherwise, ending the training to obtain the encoding module and the decoding module.
[0028] Optionally, the i-th round of joint training on the to-be-trained encoding module and the to-be-trained decoding module further includes: inputting the K-dimensional decoded output obtained in the i-th round of training and the K-dimensional known sample labels into the first loss function to obtain a first loss value; inputting the K-dimensional encoded output obtained in the i-th round of training and the K-dimensional known sample labels into the second loss function to obtain a second loss value; inputting the K-dimensional encoded output obtained in the i-th round of training into the third loss function to obtain a third loss value; determining that the second end condition is satisfied in the case that the difference between the first loss value, the second loss value, and the third loss value is less than or equal to a preset threshold.
[0029] Optionally, dimensionally ascending the M-dimensional known sample labels in the M-dimensional known sample label set through the dimension-ascending module in the dimension-ascending encoding module includes: obtaining K K-dimensional one-hot encoding vectors, randomly dividing the K K-dimensional one-hot encoding vectors into M groups, or dividing them into M groups according to a preset rule to obtain M groups of K-dimensional one-hot encoding vectors; mapping the M-dimensional known sample labels to the M groups of K-dimensional one-hot encoding vectors respectively to obtain the K-dimensional known sample labels corresponding to the M-dimensional known sample labels.
[0030] According to another aspect of the embodiments of the present application, there is also provided a data processing device, including: a first acquisition module, configured to acquire a first forward output and a second forward output, where the first forward output is an output obtained by processing first feature data by a first bottom model in a first participating party, and the second forward output is an output obtained by processing second feature data by a second bottom model in a second participating party; a processing module, configured to process the first forward output and the second forward output through a top model in the second participating party to obtain a K-dimensional inference result, where K is a positive integer greater than 2; a decoding and dimensionality reduction module, configured to perform decoding and dimensionality reduction on the K-dimensional inference result through a decoding and dimensionality reduction module in the second participating party to obtain an M-dimensional inference result, where M is a positive integer less than K and greater than or equal to 2.
[0031] According to another aspect of the embodiments of the present application, there is also provided a data processing device, including: a second acquisition module, configured to acquire a set of first sample feature data, a set of second sample feature data, and a set of M-dimensional known sample labels with a corresponding relationship, where each M-dimensional known sample label in the set of M-dimensional known sample labels is used to represent an actual sample label; a dimensionality increase and encoding module, configured to perform dimensionality increase and encoding on the set of M-dimensional known sample labels through a dimensionality increase and encoding module to obtain a set of K-dimensional known encoded sample labels; a training module, configured to jointly train a first bottom model to be trained in a first participating party, a second bottom model to be trained in a second participating party, and a top model to be trained by using the set of first sample feature data, the set of second sample feature data, and the set of K-dimensional known encoded sample labels until a first end condition is satisfied, and end the training to obtain the first bottom model in the first participating party, the second bottom model in the second participating party, and the top model.
[0032] According to yet another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the above data processing method when running.
[0033] According to yet another aspect of the embodiments of the present application, there is provided a computer program product or a computer program, where the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method as described above.
[0034] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the above data processing method through the computer program.
[0035] In the embodiments of the present application, the first participating party provides first feature data, the second participating party provides second feature data, the first bottom model is located in the first participating party, and the second bottom model is located in the second participating party. The first forward output is obtained by processing the first feature data through the first bottom model, the second forward output is obtained by processing the second feature data through the second bottom model, and the K-dimensional inference result is obtained by processing the first forward output and the second forward output through the top model in the second participating party. The M-dimensional inference result is obtained by decoding and dimension reduction of the K-dimensional inference result through the decoding and dimension reduction module in the second participating party. Since the M-dimensional inference result can only be obtained after decoding and dimension reduction, the technical effect of improving data security is achieved, and thus the technical problem of low data security is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0037] Figure 1 is a schematic diagram of an application environment of an optional data processing method according to the embodiments of the present application;
[0038] Figure 2 is a schematic flowchart of an optional data processing method according to the embodiments of the present application;
[0039] Figure 3 is a schematic diagram of an optional vertical federated deep learning method according to the embodiments of the present application;
[0040] Figure 4 is another schematic diagram of an optional inference stage according to the embodiments of the present application;
[0041] Figure 5 is another schematic diagram of an optional training stage according to the embodiments of the present application;
[0042] Figure 6 is another schematic flowchart of an optional federated deep model training process according to the embodiments of the present application;
[0043] Figure 7 is another optional. Schematic diagram of the federated deep model usage process;
[0044] Figure 8It is a schematic structural diagram of an optional data processing device according to an embodiment of the present application;
[0045] Figure 9 It is a block diagram of a computer system structure of an optional electronic device according to an embodiment of the present application;
[0046] Figure 10 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0047] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0048] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0049] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:
[0050] Federated Learning (FL for short);
[0051] Vertical Federated Learning (VFL for short);
[0052] Split Learning (SL for short);
[0053] Label Inference Attack (LIA for short);
[0054] Autoencoder (AE for short)
[0055] The present application will be described below in conjunction with embodiments:
[0056] According to one aspect of an embodiment of the present invention, a data processing method is provided. Optionally, as an alternative implementation, the above data processing method may be applied, but is not limited to, an application environment such as Figure 1 shown in the figure. The application environment may include: user device 101, server 102, server 103, and user device 104. Among them, user device 101 and server 102 are located at the second participating party, and user device 104 and server 103 are located at the first participating party. Among them, the first participating party may also be referred to as the Host party, and the second participating party may be referred to as the Guest party.
[0057] Optionally, in this embodiment, the above user device 101 and user device 104 may include, but are not limited to, at least one of the following: mobile phone (such as Android mobile phone, iOS mobile phone, etc.), laptop computer, tablet computer, handheld computer, MID (Mobile Internet Devices), PAD, desktop computer, smart TV, medical device, etc. The user device may be configured with a target client, and the target client may be a game client, instant messaging client, browser client, video client, shopping client, etc.
[0058] Optionally, the above network may include, but is not limited to: wired network, wireless network. Among them, the wired network includes: local area network, metropolitan area network, and wide area network, and the wireless network includes: Bluetooth, WIFI, and other networks that implement wireless communication.
[0059] Optionally, the above server 102 and server 103 may be a single server, or a server cluster composed of multiple servers, or a cloud server. Server 102 and server 103 may include, but are not limited to: a database and a processing engine. The database can be used to store data. For example, the database in server 102 can be used to store first feature data, and the database in server 103 can be used to store target media resource features. The processing engine processes data. For example, the processing engine in server 102 can be used to obtain a first forward output by processing the first feature data through a first bottom model, and the processing engine in server 103 can be used to obtain a second forward output by processing the second feature data through a second bottom model. The first forward output and the second forward output are processed through a top model to obtain a K-dimensional inference result, and the K-dimensional inference result is decoded and dimension-reduced through a decoding and dimension-reduction module to obtain an M-dimensional inference result.
[0060] Optionally, in this embodiment, the above data processing method may also be implemented through a server. For example,Figure 1 implemented in the server 103 of the second participating party shown; or jointly implemented by the user device 104 and the server 103 of the second participating party.
[0061] The above is only an example, and this embodiment is not specifically limited.
[0062] It can be understood that in the specific implementation of this application, when it comes to data related to user information (for example, first feature data, second feature data), etc., when the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0063] Optionally, as an alternative implementation, as Figure 2 shown, the above data processing method includes:
[0064] Step S202, obtaining a first forward output and a second forward output, where the first forward output is the output obtained by processing first feature data by a first bottom model in a first participating party, and the second forward output is the output obtained by processing second feature data by a second bottom model in a second participating party;
[0065] Among them, the above first feature data may be feature data of a target object, the target object may be a target user, the first feature data may be portrait features of the target user (for example, gender, age, education level, location, etc.), and the above second feature data may be feature data of a target media resource, and the target media resource includes but is not limited to media resources such as videos, audios, and commodities. Taking a video as an example, the second feature data may be data such as the playing duration of the video, the content of the video, and the creation location of the video.
[0066] Step S204, processing the first forward output and the second forward output through a top model in the second participating party to obtain a K-dimensional inference result, where K is a positive integer greater than 2;
[0067] Step S206, decoding and dimension reduction of the K-dimensional inference result through a decoding and dimension reduction module in the second participating party to obtain an M-dimensional inference result, where M is a positive integer less than K and greater than or equal to 2.
[0068] Among them, the above M-dimensional inference result is used to represent a predicted operation result, and the operation result is used to represent whether the target object performs a target operation on the target media resource. For example, whether the target user clicks to play the video, or represents whether the target user will purchase a commodity, etc.
[0069] As an alternative implementation, Federated Learning (FL) technology trains an effective machine learning model through the interaction of models or intermediate results, which can avoid the direct transmission of data and provides a new solution for cross-department, cross-platform, cross-organization, and cross-industry data cooperation. According to the distribution of data among different participants, federated learning can be divided into Vertical Federated Learning (VFL), Horizontal Federated Learning (HFL), and Federated Transfer Learning (FTL). Split Learning (SL), as an emerging vertical federated learning paradigm, realizes the training of a federated deep learning model without data exchange through the message interaction of a split layer (also known as a cut layer or interactive layer). This application can be applied to vertical federated learning (VFL) and split learning (SL) scenarios. Vertical federated learning increases the amount of features involved in training by combining multiple participants who have the same sample space and different feature space datasets, thereby improving the training effect of the model. In this scenario, the Host party (i.e., the first participant) and the Guest party (i.e., the second participant) have the same sample space. To train an inference model, the Guest party hopes to collaboratively apply the feature data of the Host party without exposing privacy information while ensuring that its own labels are not leaked.
[0070] As Figure 3 shown, the core idea of vertical federated deep learning is to let each participant train a bottom model locally using the features it owns. During the training process, the Guest party fuses (can be concatenated or calculate the mean, etc.) the output features of all bottom models and then inputs them into the top model. The loss function and backpropagation gradient are calculated using the data labels and the output of the top model, realizing the update of the bottom model parameters. Split learning can be regarded as a special form of vertical federated deep learning. Whether in the vertical federated learning or split learning scenario, the data of the Host party and the Guest party do not leave the domain and are not shared, and the data is available but invisible, thereby strengthening the protection of data privacy.
[0071] However, during the training process, the Host can implement a Label Inference Attack (LIA) against the Guest by analyzing the backpropagated gradients of the segmentation layer. Since this attack is launched by the Host and is independent of the Guest's data characteristics and the existence of the underlying model, both vertical federated learning and split learning may be threatened by the label inference attack. The label protection schemes implemented on the Guest add isotropic Gaussian noise to the backpropagated gradients, or add optimized perturbation noise to the backpropagated gradients, or train an Autoencoder (AE) to encode the Guest's label information. The scheme based on isotropic Gaussian noise degrades the performance of the model on the main task because the noise and the original model gradients are in different distributions. The defense mechanism based on optimized perturbation noise requires different parameters to be selected for different datasets to balance model performance and privacy protection, and often a better defense effect will correspondingly lead to a significant degradation in the performance of the model on the training task. In this implementation, a label protection scheme for federated deep learning based on dimensional transformation and autoencoder is proposed, which can be applied to the vertical federated deep learning scenario of one Host and one Guest.
[0072] As Figure 4 shown in the schematic diagram of the inference stage, during the federated model inference stage, at the first participant, the first feature data X(A) is input into the first underlying model V(A), and the forward output of the first feature data is calculated through the first underlying model V(A) to obtain the first forward output Z(A), and the first forward output Z(A) is sent to the second participant. The second participant uses the locally trained second underlying model V(B) to calculate the forward output of the second feature data X(B) to obtain the second forward output Z(B). The second participant fuses the second forward output Z(B) and the first forward output Z(A) output by the first participant through the split layer to obtain a fusion result. The second participant processes the fusion result using the trained top model T(B) to obtain the K-dimensional inference result p (K) , and then decodes the K-dimensional inference result using the decoding module D(·), and uses the dimensionality reduction module (M K,M, M can be determined according to the actual situation, taking M = 2 as an example in the figure. The dimensionality reduction module static mapping maps the decoded K-dimensional inference result to an M-dimensional inference result. Finally, the predicted operation result is obtained based on the argmax(·) function. This operation result is used to indicate whether the target object performs the target operation on the target media resource. argmax(·) returns the index of the maximum value of the input vector. For example, whether the target user clicks to play the video, argmax([0.1, 0.9]) = 1 indicates that the predicted operation result is that the target user clicks to play the video. argmax([0.8, 0.2]) = 0 indicates that the predicted operation result is that the target user does not click to play the video.
[0073] Optionally, after obtaining the first forward output and the second forward output, the method further includes: fusing the first forward output and the second forward output through the splitting layer in the second party to obtain a fusion result; the process of obtaining the K-dimensional inference result by processing the first forward output and the second forward output through the top model in the second party includes: inputting the fusion result into the top model to obtain the K-dimensional inference result output by the top model.
[0074] As an optional implementation manner, there are various ways for the above splitting layer to fuse the first forward output and the second forward output, including but not limited to direct splicing, or calculating the mean value, etc., to obtain a fusion result. The fusion result is input into the top model T(B) to obtain the K-dimensional inference result p output by the top model T(B). (K) .
[0075] Optionally, the process of fusing the first forward output and the second forward output through the splitting layer in the second party to obtain a fusion result includes: splicing the first forward output and the second forward output through the splitting layer in the second party to obtain the fusion result; or, performing mean processing on the first forward output and the second forward output through the splitting layer in the second party to obtain the fusion result.
[0076] As an optional implementation manner, the above splicing can be directly splicing the first forward output and the second forward output, and the spliced fusion result R = [Z(A), Z(B)]. Or, the above splicing can be performing mean processing on the first forward output and the second forward output, and the fusion result R after mean processing = (Z(A) + Z(B)) / 2.
[0077] Optionally, before obtaining the first forward output and the second forward output, the method further includes: obtaining a first set of sample feature data, a second set of sample feature data, and an M-dimensional set of known sample labels with a corresponding relationship, where each M-dimensional known sample label in the M-dimensional set of known sample labels is used to represent an actual label, and the first sample feature data in the first set of sample feature data has the corresponding relationship with the second sample feature data in the second set of sample feature data; performing dimensionality increase encoding on the M-dimensional set of known sample labels through a dimensionality increase encoding module to obtain a K-dimensional set of known encoded sample labels; using the first set of sample feature data, the second set of sample feature data, and the K-dimensional set of known encoded sample labels to jointly train a first bottom model to be trained in the first party, a second bottom model to be trained in the second party, and a top model to be trained until a first end condition is satisfied, ending the training, and obtaining the first bottom model, the second bottom model, and the top model; where the K-dimensional sample inference result is the output obtained by the top model to be trained processing the first sample forward output and the second sample forward output, the first sample forward output is the output obtained by the first bottom model to be trained processing the first sample feature data in the first set of sample feature data, and the second sample forward output is the output obtained by the second bottom model to be trained processing the second sample feature data in the second set of sample feature data.
[0078] As an optional implementation, taking M = 2 as an example above, in the vertical federated deep learning scenario of the first participant and the second participant, the M-dimensional known sample labels provided by the second participant are binary classifications after one-hot encoding (for example, [0, 1] or [1, 0]). The M-dimensional known sample labels can be used to represent the actual sample operation results, and the sample operation results are used to represent whether the sample object represented by the sample object feature performs the target operation on the sample media resource represented by the sample resource feature. The first sample feature data in the above first sample feature data set is the feature of the sample object. For example, the sample object is a sample user, and the first sample feature data is the gender, age, education level, etc. of the sample user. The second sample feature data in the above second sample feature data set can be the feature of the sample media resource. For example, the video type, product type, etc. The first sample feature data, the second sample feature data, and the M-dimensional known sample labels in the first sample feature data set, the second sample feature data set, and the M-dimensional known sample label set are corresponding. For example, the first sample feature data of the sample object is female, 15 years old, and a middle school student. The second sample feature data of the sample media resource is that the product type is a Barbie doll, and the corresponding M-dimensional known sample label is [0, 1] (indicating that the actual sample operation result is purchase). Another example is that the first sample feature data of the sample object is male, 55 years old, and the second sample feature data of the sample media resource is that the video content is a beauty blogger video, and the corresponding M-dimensional known sample label is [1, 0], and the corresponding M-dimensional known sample label is [1, 0] (indicating that the actual sample operation result is not to click to play). In this embodiment, there are two participants in the vertical federated deep learning scenario. The first participant provides the first sample feature data, and the second participant provides the second sample feature data and the M-dimensional known sample labels (the second participant can also only provide the M-dimensional known sample labels). And the first participant and the second participant do not perform any sample feature and label data transmission. This scenario ensures that the sample data and label data only remain local and do not leave the local, providing privacy protection at the sample feature data and label data levels.
[0079] As an optional implementation, the above-mentioned dimensionality-raising encoding module is a module that has been trained by the second participant. The M-dimensional known sample labels in the M-dimensional known sample label set are subjected to dimensionality-raising encoding through the trained dimensionality-raising encoding module to obtain a K-dimensional known encoded sample label set. After the second participant completes the dimensionality-raising encoding of the original M-dimensional known sample labels, the first participant initializes a bottom model V (A) (·), and the second participant initializes a bottom model V (B) (·) and a top model T (B) (·).
[0080] For a given K-dimensional known encoded sample label, the sample set can be expressed as Among them, represents the first sample feature data, represents S pieces of second sample feature data, indicates that the first sample feature data is the second sample feature data is the corresponding M-dimensional known sample label, and can also represent the corresponding sample object pair the actual sample operation result of the corresponding sample media resource, for example, purchasing a commodity, clicking to play a video, etc.
[0081] Input the first sample feature data into the first bottom model to be trained of the first participant. The first bottom model to be trained calculates the forward output of the first sample forward output z (A) , input the S pieces of second sample feature data into the second bottom model to be trained of the second participant. The second bottom model to be trained calculates the forward output of the second sample forward output z(B). Through the split layer of the second participant, z (A) and z (B) are fused (which can be direct splicing or calculating the mean, etc.) to obtain a fusion result. The fusion result is input into the top model T (B) (·) to obtain the K-dimensional sample inference result p (K) output by the top model to be trained.
[0082] Then calculate the loss L = CE(y (K,enc) , p (K) ), where CE can be the cross-entropy loss function, and y (K,enc) is the K-dimensional known encoded sample label, and p (K) is the K-dimensional sample inference result output by the top model to be trained. If the output value of the cross-entropy loss function between y( K,enc ), p (K) is less than or equal to a preset threshold (the preset threshold can be determined according to the actual situation, such as 0.01, 0.02, etc.), then the loss function converges.
[0083] The above first end condition includes but is not limited to: the loss function converges, the model parameters converge, reaching a preset number of iterations, reaching a preset training duration, etc. After satisfying the first end condition, stop training to obtain the first bottom model, the second bottom model, and the top model. If the above first end condition is not satisfied, then backpropagate to modify the model parameters of the first bottom model to be trained, the second bottom model to be trained, and the top model to be trained.
[0084] Optionally, jointly training the to-be-trained first bottom model in the first party, the to-be-trained second bottom model in the second party, and the to-be-trained top model by using the first sample feature data set, the second sample feature data set, and the K-dimensional known encoded sample label set includes: performing the j-th round of joint training on the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model through the following steps, where j is a positive integer greater than or equal to 1, and the first bottom model, the second bottom model, and the top model obtained in the 0-th round of training are the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model that have not been trained, including: inputting the S first sample feature data used in the j-th round in the first sample feature data set into the first bottom model obtained in the (j - 1)-th round of training to obtain S first sample forward outputs output in the j-th round of training, where S is a positive integer greater than or equal to 1; inputting the S second sample feature data used in the j-th round in the second sample feature data set into the second bottom model obtained in the (j - 1)-th round of training to obtain S second sample forward outputs output in the j-th round of training; inputting the S first sample forward outputs output in the j-th round of training and the S second sample forward outputs output in the j-th round of training into the top model obtained in the (j - 1)-th round of training to obtain S K-dimensional sample inference results output in the j-th round of training; ending the training to obtain the first bottom model, the second bottom model, and the top model when the S K-dimensional sample inference results output in the j-th round of training satisfy the first end condition with respect to the S K-dimensional known encoded sample labels used in the j-th round, and the K-dimensional known encoded sample labels used in the j-th round have the corresponding relationship with the first sample feature data used in the j-th round and the second sample feature data used in the j-th round.
[0085] As an alternative implementation, in each round of training, mini-batch (the smallest sample set for model training) is used to jointly train the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model. Assume that the mini-batch (the smallest sample set for model training) contains S first sample feature data provided by the Host party, S second sample feature data provided by the Guest party, and S K-dimensional known encoded sample labels (the S K-dimensional known encoded sample labels are obtained by performing dimensionality increase encoding on the corresponding S M-dimensional known sample labels through a dimensionality increase encoding module). The S first sample feature data, the S second sample feature data, and the S K-dimensional known encoded sample labels have a one-to-one correspondence relationship and can be used to represent the actual operation result of the sample object corresponding to the first sample feature data on the sample media resource corresponding to the second sample feature data.
[0086] As Figure 5In the schematic diagram of the training phase shown, assume that X(A) is the first sample feature data used in the j-th round above, and X(B) is the second sample feature data used in the j-th round. Input X(A) into the first bottom model V(A) obtained from the (j - 1)-th round of training to obtain the first sample forward output Z(A) of the j-th round of training. Input X(B) into the second bottom model V(B) obtained from the (j - 1)-th round of training to obtain the second sample forward output Z(B) of the j-th round of training. Fuse Z(A) and Z(B) through a split layer to obtain the fusion result of the j-th round of training output, and input the fusion result of the j-th round of training output into the top model T(B) obtained from the (j - 1)-th round of training to obtain the K-dimensional sample inference result of the j-th round of training. When the loss value between the K-dimensional sample inference result of the j-th round of training output and the K-dimensional known encoded sample label used in the j-th round satisfies the above L = CE(y (K,enc) , p (K) ), the loss function converges. When the above first end condition is satisfied, end the training. The first bottom model obtained from the (j - 1)-th round of training is determined as the first bottom model, the second bottom model obtained from the (j - 1)-th round of training is determined as the second bottom model, and the top model obtained from the (j - 1)-th round of training is determined as the top model. Among them, CE can be the cross-entropy loss function, y (K ,enc) is the K-dimensional known encoded sample label used in the j-th round, and p (K) is the K-dimensional sample inference result of the j-th round of training output. In addition, the above K-dimensional known encoded sample label used in the j-th round is obtained by performing dimensionality increase encoding on the M-dimensional known sample label used in the j-th round through a dimensionality increase encoding module (including the dimensionality increase module and the encoding module in the figure).
[0087] In the federated training phase of the bottom model (the first bottom model) of the first party above, the bottom model (the second bottom model) of the second party, and the top model, the first party uses the bottom model to calculate the forward output of the local first sample feature data and sends it to the second party. The second party also uses its own bottom model to calculate the forward output of the local second sample feature data, fuses it with the forward output of the first party, and then uses the top model to calculate the K-dimensional sample inference result. Finally, use the corresponding K-dimensional known encoded sample label to calculate the loss function and the gradient of backpropagation, and update the parameters of the top model and the bottom model. The protection mechanism in this embodiment is implemented on the second party. In the split learning scenario, the second party may not have any feature data of the sample, so only one top model needs to be trained, which has good scalability and can also be applied to the split learning scenario.
[0088] The above embodiments are applicable to the vertical federated deep learning scenario and can cope with the label inference attack that the first party may implement in this scenario. It can protect the model training in the binary classification task, that is, the labels of the second party can be binary classification after one-hot encoding ([0,1] or [1,0]). The functional performance on the product side of vertical federated deep learning can include being a model training module for federated learning tasks, improving the usability of the federated learning system platform, and adding functional modules to the federated learning system, that is, adding a label protection scheme based on label dimension conversion and autoencoders to improve the privacy and security characteristics of the federated learning model. The main product form can provide federated learning services externally in the form of a federated learning platform on a public cloud or a private cloud.
[0089] Optionally, the method further includes: in the case that the S K-dimensional sample inference results output in the j-th round of training do not satisfy the first end condition with the S K-dimensional known encoded sample labels used in the j-th round of training, adjusting the model parameters of the first bottom model, the second bottom model, and the top model obtained in the (j - 1)-th round of training through backpropagation.
[0090] Optionally, the adjusting the model parameters of the top model obtained in the (j - 1)-th round of training through backpropagation includes: the second party obtains S first gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S K-dimensional sample inference results, and adjusts the model parameters of the top model obtained in the (j - 1)-th round of training through the S first gradients; the adjusting the model parameters of the second bottom model obtained in the (j - 1)-th round of training through backpropagation includes: the second party obtains S second gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the model parameters of the second bottom model obtained in the (j - 1)-th round of training, and adjusts the model parameters of the second bottom model obtained in the (j - 1)-th round of training through the S second gradients; the adjusting the first bottom model obtained in the (j - 1)-th round of training through backpropagation includes: the second party obtains S third gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S first sample forward outputs; the second party normalizes the S third gradients through the second norm to obtain S normalized gradients; the second party sends the S normalized gradients to the first party, and the first party adjusts the first bottom model obtained in the (j - 1)-th round of training through the S normalized gradients.
[0091] As an optional implementation, during the backpropagation process, the second party first calculates the first gradient:
[0092]
[0093]
[0094] where L u is the u-th loss function among S loss functions, and p(K) u is the u-th K-dimensional sample inference result among S K-dimensional sample inference results. The loss function can be the above-mentioned cross-entropy loss function, thus obtaining S first gradients, and using the Chain Rule to update the parameters of the top model T (B) (·).
[0095] Then calculate the second gradient:
[0096]
[0097] where w(B) refers to the parameters of the model V (B) (·) (the model parameters of the second bottom model obtained from the (j - 1)-th round of training). Update the second bottom model V (B) (·) to be trained through S second gradients.
[0098] The second party calculates the third gradient:
[0099]
[0100] Since each mini-batch (the smallest sample set for model training) contains S samples (here s is the defined batch size), for the set of S third gradients corresponding to the current mini-batch perform normalization processing through the second-order norm, and the specific operation is as follows:
[0101] First, perform a mean processing on the S third gradients to obtain the mean processing result
[0102]
[0103] Then, for the mean processing result obtain S normalized gradients:
[0104]
[0105] where ||·|| is the function for calculating the second-order norm. Subsequently, for each sample, the Guest party sends to the Host party for the first bottom model V (A)(·) Parameter update. The above steps are performed once for each sample in each Epoch (iteration round), and repeated η Epochs in total until the model converges on the training sample set, obtaining the first bottom model, the second bottom model, and the top model.
[0106] Optionally, the method further includes: obtaining an M-dimensional known sample label set, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample operation result, and the sample operation result is used to represent whether the sample object represented by the first sample feature data performs the target operation on the sample media resources represented by S second sample feature data, and the first sample feature data has a corresponding relationship with the S second sample feature data; performing dimension elevation on the M-dimensional known sample labels in the M-dimensional known sample label set through a dimension elevation module to obtain a K-dimensional known sample label set; using the K-dimensional known sample label set to jointly train the to-be-trained encoding module and the to-be-trained decoding module until the second end condition is satisfied among the first loss function, the second loss function, and the third loss function, ending the training, and obtaining the encoding module and the decoding module, where the decoding dimension reduction module includes the decoding module and the dimension reduction module, the dimension elevation encoding module includes the encoding module and the dimension elevation module, and the dimension reduction module corresponds to the dimension elevation module in the dimension elevation encoding module; where the first loss function is the loss function between the K-dimensional decoding output output by the to-be-trained decoder and the corresponding K-dimensional known sample label in the K-dimensional known sample label set, the second loss function is the loss function between the K-dimensional encoding output output by the to-be-trained encoder and the corresponding K-dimensional known sample label in the K-dimensional known sample label set, and the third loss function is the information entropy function of the K-dimensional encoding output output by the to-be-trained encoder.
[0107] As an optional implementation manner, the second party can custom-decide a dimension K of the elevated label and define a pair of static mappings between M dimensions and K dimensions, denoted as M M,K (from M dimensions to K dimensions) and M K,M (from K dimensions to M dimensions), where M M,K is implemented through the dimension elevation module in the dimension elevation encoding module, and M K,M is implemented through the dimension reduction module in the decoding dimension reduction module.
[0108] The second participating party uses the dimensionality - raising module to raise the dimensionality of the M - dimensional known sample labels to obtain K - dimensional known sample labels, and trains the encoding module E(·) and the decoding module D(·) based on the raised K - dimensional known sample labels. Here, the input and output of the encoder are both K - dimensional. Then, the second participating party calculates the corresponding K - dimensional known sample labels for the M - dimensional known sample labels of each sample for federated model training (performing federated training on the first bottom model, the second bottom model, and the top model to be trained). The label dimension in the federated model training step is K - dimensional, and the dimension mapping method is unknown to the first participating party compared to the encoding method of the encoding module, improving data security.
[0109] As an alternative implementation, the second participating party can construct the dimensionality - raising encoded labels locally. Assume a given set containing N samples where is the first sample feature data set of the second participating party, is the second sample feature data set of the second participating party, is the M - dimensional known sample label set of the second participating party. X (A) and X (B) can be used to represent the features of a certain sample for the first and second participating parties respectively, and y is used to represent the M - dimensional known sample label of the sample.
[0110] Taking the above - mentioned M - dimensional known sample labels as binary - classification labels as an example, the second participating party defines a dimensionality - raising module and a dimensionality - lowering module for the M - dimensional known sample labels. Among them, the dimensionality - raising module is a static mapping M from 2 - D to K - D 2,K (from 2 - D to K - D), and the dimensionality - lowering module is a static mapping M from K - D to 2 - D K,2 (from K - D to 2 - D). Generally, K is taken as a multiple of M.
[0111] For example, when K = 4 and M = 2, the dimensionality - raising module, which is a static mapping M from 2 - D to K - D 2,K can be:
[0112] M 2,4 : ([1, 0], [0, 1]) → (m 4,0 , m 4,1 )
[0113] where m 4,0 ∈{[1, 0, 0, 0], [0, 0, 1, 0]}, m 4,1 ∈{[0, 1, 0, 0], [0, 0, 0, 1]}. In the case where [1, 0] represents that the sample object performs a target operation on the sample media resource, m 4,0Indicates that the sample object performs a target operation on the sample media resource. Similarly, in the case where [0, 1] indicates that the sample object does not perform the target operation on the sample media resource, m 4,1 Indicates that the sample object does not perform the target operation on the sample media resource.
[0114] The dimensionality reduction module is a mapping from K dimensions to 2 dimensions. Since the K-dimensional data is not necessarily one-hot encoded, when K = 4, based on the above-given M 2,4 , assuming that the decoded inference result is p (4,dec) , the two-dimensional vector can be obtained through the following mapping:
[0115] M 4,2 : y (4,dec) → [dot(p (4,dec ), [1, 0, 1, 0]), dot(p (4,dec) , [0, 1, 0, 1])]
[0116] where dot is the inner product operation of vectors. The conventional processing method here is to set the dimension with the largest value in p (4,dec) to 1 and the rest to 0, and then map back to 2 dimensions.
[0117] Suppose p (4,dec) is [1, 0, 0, 0], M 4,2 : y (4,dec) = dot([1, 0, 0, 0], [1, 0, 1, 0]), dot([1, 0, 0, 0], [0, 1, 0, 1])] = [1, 0].
[0118] When K = 6, M = 2, the dimensionality increase module is a static mapping from 2 dimensions to K dimensions M 2,K can be:
[0119] M 2,6 : ([1, 0], [0, 1]) → (m 6,0 , m 6,1 ),
[0120] m 6,0 ∈ {[1, 0, 0, 0, 0, 0], [0, 0, 1, 0, 0, 0], [0, 0, 0, 1, 0, 0]}
[0121] m 6,1 ∈ {[0, 1, 0, 0, 0, 0], [0, 0, 0, 0, 1, 0][0, 0, 0, 0, 0, 1]}
[0122] In the case where [1, 0] indicates that the sample object performs a target operation on the sample media resource, m 6,0Indicates that the sample object performs a target operation on the sample media resource. Similarly, in the case where [0, 1] indicates that the sample object does not perform the target operation on the sample media resource, m 6,1 Indicates that the sample object does not perform the target operation on the sample media resource.
[0123] Regarding the mapping from K dimensions to 2 dimensions, since the data in K dimensions is not necessarily one-hot encoded, when K = 6, based on the above-given M 2,6 , assuming that the decoded inference result is p (6,dec ), a two-dimensional vector can be obtained through the following mapping
[0124] M 6,2 : y (6,dec) → [dot(p (6,dec) , [1, 0, 1, 1, 0, 0]), dot(p (6,dec) , [0, 1, 0, 0, 1, 1])]
[0125] where dot is the inner product operation of vectors. The conventional processing method here is to set the dimension with the largest value in p (6,dec) to 1 and the rest to 0, and then map back to 2 dimensions.
[0126] Suppose p (6,dec) = [1, 0, 0, 0, 0, 0],
[0127] M 6,2 : y (6,dec) = [dot([1, 0, 0, 0, 0, 0], [1, 0, 1, 1, 0, 0]), dot([1, 0, 0, 0, 0, 0], [0, 1, 0, 0, 1, 1])] = [1, 0]
[0128] As an alternative implementation, the second participating party trains the encoding module E(·) and the decoder D(·) module based on the K-dimensional known sample labels after dimensionality increase. The input and output of both are K-dimensional. The network model structures of the encoding module and the decoding module are the same. For example, it can be FC((6K) 2 )-ReLU-FC(K)-SoftMax, where FC is the fully connected layer (FullyConnectedLayer), and ReLU and SoftMax are common activation functions. For the input K-dimensional known sample labels, through the encoding module and the decoding module, we can obtain:
[0129] y (K,enc) = E(y (K) )
[0130] y (K,dec) = D(y (K,enc) )
[0131] Among them, y (K) is the K-dimensional known sample label, y (K,enc) is the K-dimensional encoded sample label output by the encoding module, and y (K,dec) is the K-dimensional decoded sample label output by the decoding module.
[0132] Define the above convergence function:
[0133] L = CE(y (K) , y (K,dec) ) - α·CE(y (K) , y (K,enc) ) - β·Entropy(y (K,enc) )
[0134] Among them, CE(·) is the cross-entropy loss function (Cross-Entropy Loss), CE(y (K) , y (K,dec) ) is the above first loss function, α·CE(y (K) , y (K,enc) ) is the above second loss function, β·Entropy(y (K,enc) ) is the above third loss function, Entropy(·) is the information entropy, and α and β are two custom parameters that can be set according to the actual situation, such as 0.3, 0.5, etc., used to adjust the weights of different parts of the loss function. When the above convergence function satisfies the second end condition (which can be that the value of L is less than or equal to a preset threshold, and the preset threshold can be set according to the actual situation), the training is ended to obtain the encoding module and the decoding module.
[0135] In the above embodiment, a dimension conversion scheme based on a static mapping mechanism is adopted to expand the encoding space of the sample label, and the trained encoding module is used to encode the upsampled label, without adding random or optimized noise to the backpropagation gradient. After dimension conversion and label encoding, the second party can directly transmit the backpropagation gradient obtained from the training to the first party without the risk of leaking the true label of the sample. The label protection method based on dimension conversion and autoencoder can be applied to any binary classification dataset and can address the problem of poor protection effect caused by the small encoding space of the traditional autoencoder scheme.
[0136] Optionally, the joint training of the to-be-trained encoding module and the to-be-trained decoding module using the K-dimensional known sample label set includes: performing the i-th round of joint training on the to-be-trained encoding module and the to-be-trained decoding module through the following steps, where i is a positive integer greater than or equal to 1, and the encoding module and the decoding module obtained from the 0-th round of training are the untrained to-be-trained encoding module and the to-be-trained decoding module: inputting the K-dimensional known sample label used in the i-th round in the K-dimensional known sample label set into the encoding module obtained from the (i - 1)-th round of training to obtain the K-dimensional encoded output obtained from the i-th round of training; inputting the K-dimensional encoded output obtained from the i-th round of training into the decoding module obtained from the (i - 1)-th round of training to obtain the K-dimensional decoded output obtained from the i-th round of training; where, in the case that the first loss function between the K-dimensional decoded output obtained from the i-th round of training and the K-dimensional known sample label, the second loss function between the K-dimensional encoded output obtained from the i-th round of training and the K-dimensional known sample label, and the third loss function of the K-dimensional encoded output obtained from the i-th round of training do not satisfy the second end condition, adjusting the parameters in the to-be-trained encoding module and the to-be-trained decoding module; otherwise, ending the training to obtain the encoding module and the decoding module.
[0137] As an optional implementation manner, during the i-th round of joint training of the to-be-trained encoding module and the to-be-trained decoding module, for:
[0138] yi (K,enc) = E i-1 (yi (K) )
[0139] yi (K,dec) = D i-1 (yi (K,enc) )
[0140] where, yi (K) is the K-dimensional known sample label used in the i-th round, E i-1 () is the encoding module obtained from the (i - 1)-th round of training, D i-1 () is the decoding module obtained from the (i - 1)-th round of training, yi (K,enc) is the K-dimensional encoded output obtained from the i-th round of training, yi( K ,dec) is the K-dimensional decoded output obtained from the i-th round of training, input the following convergence function:
[0141] Li = CE(yi (K) , yi (K,dec) ) - α·CE(yi (K) , yi (K,enc) ) - β·Entropy(yi (K,enc) )
[0142] If Li is less than or equal to a preset threshold (which can be set according to the actual situation, such as 0.01, 0.02, etc.), it is determined that the above second end condition is satisfied, and the training is ended to obtain the encoding module and the decoding module.
[0143] Optionally, the joint training of the to-be-trained encoding module and the to-be-trained decoding module in the i-th round further includes: inputting the K-dimensional decoding output obtained in the i-th round and the K-dimensional known sample label into the first loss function to obtain a first loss value; inputting the K-dimensional encoding output obtained in the i-th round and the K-dimensional known sample label into the second loss function to obtain a second loss value; inputting the K-dimensional encoding output obtained in the i-th round into the third loss function to obtain a third loss value; and determining that the second end condition is satisfied when the difference between the first loss value, the second loss value, and the third loss value is less than or equal to a preset threshold.
[0144] As an optional implementation manner, for the convergence function: L = CE(y (K) , y (K,dec) ) - α·CE(y (K) , y (K,enc) ) - β·Entropy(y (K,enc) ), the first loss function is CE(y (K) , y (K,dec) ), the second loss function is -α·CE(y (K) , y (K ,enc) ), and the third loss function is -β·Entropy(y (K,enc) ). y (K) is the K-dimensional known sample label, y (K,dec) is the K-dimensional decoding output obtained in the i-th round, and y (K,enc) is the K-dimensional encoding output obtained in the i-th round. The second end condition is that the L value is less than or equal to a preset threshold (which can be set according to the actual situation).
[0145] Optionally, the decoding and dimensionality reduction of the K-dimensional inference result by the decoding and dimensionality reduction module in the second party to obtain an M-dimensional inference result includes: decoding the K-dimensional inference result through the decoding module in the decoding and dimensionality reduction module to obtain a K-dimensional decoded inference result; and reducing the dimensionality of the K-dimensional decoded inference result through the dimensionality reduction module in the decoding and dimensionality reduction module to obtain the M-dimensional inference result.
[0146] As an optional implementation manner, the above decoding and dimensionality reduction module may include a decoding module and a dimensionality reduction module. The input and output of the decoding module are both K-dimensional. The input is the K-dimensional inference result (such as y (K,enc) ), and the output is the K-dimensional decoded inference result (such as, y(K,dec) ), that is, y (K,dec) = D(y (K,enc) ). The K-dimensional decoding inference result is passed through the dimensionality reduction module in the above embodiments. For example:
[0147] M 4,2 : y (4,dec) → [dot(p (4,dec) , [1, 0, 1, 0]), dot(p (4,dec) , [0, 1, 0, 1])]
[0148] M 6,2 : y (6,dec) → [dot(p (6,dec) , [1, 0, 1, 1, 0, 0]), dot(p (6,dec) , [0, 1, 0, 0, 1, 1])]
[0149] Decode the K-dimensional decoding inference result to obtain the above M-dimensional inference result.
[0150] In the above embodiments, the main process in the inference stage is
a. The model outputs an inference set based on the inference sample set
b. The decoder decodes the inference results in the inference set
c. Map and reduce the dimensionality of the decoded inference results
[0151] For example, define the structure of the encoder E(·) as FC((6K) 2 )-ReLU-FC(K)-SoftMax, with both input and output being K-dimensional; define the structure of the decoder D(·) as FC((6K) 2 )-ReLU-FC(2)-SoftMax, with the input being K-dimensional and the output being 2-dimensional.
[0152] Given the sample set and the M-dimensional known sample label set There is:
[0153] y (K,enc) = E(y (K) ),
[0154] y (2,dec) = D(y (K,enc) ).
[0155] Define the loss function as:
[0156] L = CE(y (K) , y (2,dec) ) - α·CE(y (K) , y (K,enc) ) - β·Entropy(y (K,enc) ).
[0157] Among them, different from the above embodiments, the output y (2,dec) of the decoding module is two-dimensional, CE(·) is the cross-entropy loss function, Entropy(·) is the information entropy, and α and β are two custom parameters used to adjust the weights of different parts of the loss function in the training of the autoencoder. In this embodiment, the main process in the inference stage is simplified to [a. The model outputs an inference set based on the inference sample set] -> [b. The decoder decodes the inference results in the inference set], reducing information loss, and the remaining steps are the same as the technical solution in 3.
[0158] Optionally, the M-dimensional known sample labels in the M-dimensional known sample label set are dimensionally upscaled by the dimension upscaling module in the dimension upscaling and encoding module, including: obtaining K K-dimensional one-hot encoding vectors, randomly dividing the K K-dimensional one-hot encoding vectors into M groups, or dividing them into M groups according to a preset rule, to obtain M groups of K-dimensional one-hot encoding vectors; mapping the M-dimensional known sample labels to the M groups of K-dimensional one-hot encoding vectors respectively to obtain K-dimensional known sample labels corresponding to the M-dimensional known sample labels; the K-dimensional decoded inference result is dimensionally downscaled by the dimension downscaling module, including: performing an inner product operation on the K-dimensional inference result and M K-dimensional decoding vectors to obtain the M-dimensional inference result, where the M K-dimensional decoding vectors are decoding vectors obtained according to the M groups of K-dimensional one-hot encodings.
[0159] As an optional implementation manner, taking the above M = 2 as an example, [1, 0], [0, 1] are one-hot encodings that can correspond to M classifications, or can represent M + 1 classifications, including a all-zero vector (which can be selected or not used), and can be used to represent whether an object performs a target operation on a media resource, such as a video click play operation, a commodity purchase operation, etc. K can be an integer multiple of M or not an integer multiple of M.
[0160] In the dimension upscaling stage, assuming K = 4, M = 2, the dimension upscaling module is a static mapping M from 2D to KD 2,K can be:
[0161] M 2,4 : ([1, 0], [0, 1]) → (m 4,0 , m 4,1 )
[0162] First, generate K K-dimensional one-hot encoding vectors. In this embodiment, generate 4 four-dimensional one-hot encoding vectors: [1, 0, 0, 0], [0, 0, 1, 0], [0, 1, 0, 0], [0, 0, 0, 1]. Randomly or according to a preset rule, divide the 4 four-dimensional one-hot encoding vectors into M = 2 groups. Among them, "randomly" can be randomly selecting any one vector from the 4 four-dimensional one-hot encoding vectors as one group, and the remaining three vectors as another group. The preset rule can be an average distribution, with two vectors from the 4 four-dimensional one-hot encoding vectors as one group and the remaining two vectors as another group.
[0163] Taking the example of dividing the 4 four-dimensional one-hot encoding vectors [1, 0, 0, 0], [0, 0, 1, 0], [0, 1, 0, 0], [0, 0, 0, 1] into [1, 0, 0, 0], [0, 0, 1, 0] and [0, 1, 0, 0], [0, 0, 0, 1] these two groups. Among them, m 4,0 ∈ {[1, 0, 0, 0], [0, 0, 1, 0]}, that is, adding at the last two bits. m 4,1 ∈ {[0, 1, 0, 0], [0, 0, 0, 1]}. Thus, obtain the K-dimensional known sample labels m 4,0 and m 4,1 .
[0164] In the dimensionality reduction stage, based on the above-given M 2,4 , assuming the decoded inference result is p (4,dec) , a two-dimensional vector can be obtained through the following mapping:
[0165] M 4,2 : y (4,dec) → [dot(p (4,dec) , [1, 0, 1, 0]), dot(p (4,dec) , [0, 1, 0, 1])]
[0166] Among them, dot is the inner product operation of vectors. The conventional processing method here is to set the dimension with the largest value in p (4,dec) to 1, and the rest to 0, and then map back to 2D. The above [1, 0, 1, 0] and [0, 1, 0, 1] are the corresponding decoding vectors.
[0167] When K = 6, M = 2, and the dimensionality increase module is a static mapping M 2,K from 2D to K-dimensional, it can be:
[0168] M 2,6 : ([1, 0], [0, 1]) → (m 6,0 , m 6,1 ),
[0169] First, generate K K-dimensional one-hot encoding vectors. In this embodiment, generate 6 six-dimensional one-hot encoding vectors:
[0170] [1, 0, 0, 0, 0, 0], [0, 0, 1, 0, 0, 0], [0, 0, 0, 1, 0, 0][0, 1, 0, 0, 0, 0], [0, 0, 0, 0, 1, 0][0, 0, 0, 0, 0, 1]
[0171] Randomly or according to a preset rule, divide six 6 - dimensional one - hot encoded vectors into M = 2 groups. Among them, random division can be randomly selecting any one or two vectors from the six 6 - dimensional one - hot encoded vectors as one group, and the remaining vectors as another group. The preset rule can be an even distribution, that is, three vectors from the six 6 - dimensional one - hot encoded vectors form one group, and the remaining three vectors form another group.
[0172] Take the example of evenly dividing six 6 - dimensional one - hot encoded vectors into two groups.
[0173] m 6,0 ∈ {[1, 0, 0, 0, 0, 0], [0, 0, 1, 0, 0, 0], [0, 0, 0, 1, 0, 0]}
[0174] m 6,1 ∈ {[0, 1, 0, 0, 0, 0], [0, 0, 0, 0, 1, 0][0, 0, 0, 0, 0, 1]}
[0175] When [1, 0] represents that the sample object performs a target operation on the sample media resource, m 6,0 represents that the sample object performs a target operation on the sample media resource. Similarly, when [0, 1] represents that the sample object does not perform a target operation on the sample media resource, m 6,1 represents that the sample object does not perform a target operation on the sample media resource.
[0176] Regarding the mapping from K - dimensional to 2 - dimensional, since the K - dimensional data is not necessarily one - hot encoded. When K = 6, based on the above - given M 2,6 , assuming that the decoded inference result is p (6,dec) , a two - dimensional vector can be obtained through the following mapping
[0177] M 6,2 : y (6,dec) → [dot(p (6,dec) [1, 0, 1, 1, 0, 0]), dot(p (6,dec) [0, 1, 0, 0, 1, 1])]
[0178] where dot is the inner - product operation of vectors. The conventional processing method here is to set the dimension with the largest value in p (6,dec) to 1, and the rest to 0, and then map it back to 2 - dimensional.
[0179] Suppose p(6,dec) = [1, 0, 0, 0, 0, 0],
[0180] M 6,2 : y (6,dec) = [dot([1, 0, 0, 0, 0, 0], [1, 0, 1, 1, 0, 0]), dot([1, 0, 0, 0, 0, 0], [0, 1, 0, 0, 1, 1])] = [1, 0]
[0181] In the above embodiments, the sample feature data of the first participant and the second participant do not leave the local, and the first participant realizes the label confusion of binary classification through the combination of label dimension conversion and autoencoder, and the backpropagation gradient is normalized, reducing the risk of leaking sample label information in the gradient sent back to the first participant. It should be noted that if in the traditional vertical federated learning scenario, that is, the second participant has a certain number of features, this solution can also be directly executed, because the label confusion module does not affect the performance of the training task, and the normalization of the backpropagation gradient is carried out based on the gradient received by the first participant, which has no impact on the local gradient calculation of the second participant.
[0182] As an optional implementation manner, the training process of the federated deep model refers to Figure 6 and the usage process of the federated deep model refers to Figure 7 .
[0183] In Figure 6 the training process of the federated deep model includes the following steps:
[0184] The following steps are performed for the first participant:
[0185] Step S601, obtain the first sample feature data in the first sample feature data set, and the first sample feature data can be the features of the target object;
[0186] Step S602, input the first sample feature data into the first bottom model to be trained by the first participant to obtain the first sample forward output;
[0187] The following steps are performed for the second participant:
[0188] Step S603, obtain the second sample feature data and the M-dimensional known sample label in the second sample feature data set, and the second sample feature data can be the features of the target media resource;
[0189] Step S604, input the second sample feature data into the second bottom model to be trained by the second participant to obtain the second sample forward output;
[0190] Step S605: Input the M - dimensional known sample labels into the dimensionality - increasing encoding module of the second party to obtain the K - dimensional known encoded sample labels;
[0191] Step S606: Fuse the first sample forward output and the second sample forward output through the split layer of the second party to obtain a fusion result;
[0192] Step S607: Process the fusion result through the top model of the second party to obtain the K - dimensional sample inference result;
[0193] Step S608: Determine whether the K - dimensional sample inference result and the K - dimensional known encoded sample labels satisfy the first end condition. If satisfied, end the training to obtain the first bottom model, the second bottom model, and the top model. If the first end condition is not satisfied, adjust the model parameters and continue training.
[0194] In Figure 7 The process of using the federated deep model includes the following steps:
[0195] For the first party, perform the following steps:
[0196] Step S701: Obtain the first feature data, such as the gender, age, etc. of the target object;
[0197] Step S702: Input the first feature data into the first bottom model of the first party to obtain the first forward output;
[0198] For the second party, perform the following steps:
[0199] Step S703: Obtain the second feature data, such as video content, product type, etc.;
[0200] Step S704: Input the second feature data into the second bottom model of the second party to obtain the second forward output;
[0201] Step S705: Fuse the first forward output and the second forward output through the split layer of the second party to obtain a fusion result;
[0202] Step S706: Process the fusion result through the top model of the second party to obtain the K - dimensional inference result;
[0203] Step S707: Decode and reduce the dimension of the K - dimensional inference result through the decoding and dimensionality - reduction module to obtain the M - dimensional preset label.
[0204] The above embodiments can be applied to vertical federated deep learning. The second party in vertical federated learning performs dimensionality conversion of labels and training of autoencoders locally. In the training phase, the second party uses the trained encoder to encode the upsampled labels (changing binary classification to K-classification), and conducts federated deep learning model training based on the encoding results. In the inference phase, the second party uses the decoder to decode the inference results, and uses static mapping to downsample the decoded labels to a binary classification task to obtain the final inference result. In the proposed method for label-protected vertical federated deep learning.
[0205] According to another aspect of the embodiments of the present application, there is also provided a data processing method, including: obtaining a first set of sample feature data, a second set of sample feature data, and an M-dimensional set of known sample labels with corresponding relationships, where each M-dimensional known sample label in the M-dimensional set of known sample labels is used to represent an actual sample label; performing upsampling encoding on the M-dimensional set of known sample labels through an upsampling encoding module to obtain a K-dimensional set of known encoded sample labels; using the first set of sample feature data, the second set of sample feature data, and the K-dimensional set of known encoded sample labels to jointly train a first bottom model to be trained in the first party, a second bottom model to be trained in the second party, and a top model to be trained until a first end condition is satisfied, ending the training, and obtaining the first bottom model in the first party, the second bottom model in the second party, and the top model.
[0206] Optionally, jointly training the first bottom model to be trained in the first party, the second bottom model to be trained in the second party, and the top model to be trained by using the first sample feature data set, the second sample feature data set, and the K-dimensional known encoded sample label set includes: performing the j-th round of joint training on the first bottom model to be trained, the second bottom model to be trained, and the top model to be trained through the following steps, where j is a positive integer greater than or equal to 1, and the first bottom model, the second bottom model, and the top model obtained in the 0-th round of training are the first bottom model to be trained, the second bottom model to be trained, and the top model to be trained that are not trained, including: inputting S first sample feature data used in the j-th round in the first sample feature data set into the first bottom model obtained in the (j - 1)-th round of training to obtain S first sample forward outputs output in the j-th round of training, where S is a positive integer greater than or equal to 1; inputting S second sample feature data used in the j-th round in the second sample feature data set into the second bottom model obtained in the (j - 1)-th round of training to obtain S second sample forward outputs output in the j-th round of training; inputting the S first sample forward outputs output in the j-th round of training and the S second sample forward outputs output in the j-th round of training into the top model obtained in the (j - 1)-th round of training to obtain S K-dimensional sample inference results output in the j-th round of training; ending the training to obtain the first bottom model, the second bottom model, and the top model when the S K-dimensional sample inference results output in the j-th round of training satisfy the first end condition with respect to the S K-dimensional known encoded sample labels used in the j-th round, and the K-dimensional known encoded sample labels used in the j-th round have the corresponding relationship with the first sample feature data used in the j-th round and the second sample feature data used in the j-th round.
[0207] Optionally, the method further includes: when the S K-dimensional sample inference results output in the j-th round of training do not satisfy the first end condition with respect to the S K-dimensional known encoded sample labels used in the j-th round, adjusting the model parameters of the first bottom model, the second bottom model, and the top model obtained in the (j - 1)-th round of training through backpropagation.
[0208] Optionally, adjusting the model parameters of the top model obtained from the (j-1)-th round of training through forward propagation includes: the second party obtaining S first gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S K-dimensional sample inference results, and adjusting the model parameters of the top model obtained from the (j-1)-th round of training through the S first gradients; adjusting the model parameters of the second bottom model obtained from the (j-1)-th round of training through forward propagation includes: the second party obtaining S second gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the model parameters of the second bottom model obtained from the (j-1)-th round of training, and adjusting the model parameters of the second bottom model obtained from the (j-1)-th round of training through the S second gradients; adjusting the first bottom model obtained from the (j-1)-th round of training through backpropagation includes: the second party obtaining S third gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S first sample forward outputs; the second party normalizing the S third gradients through the second norm to obtain S normalized gradients; the second party sending the S normalized gradients to the first party, and the first party adjusting the first bottom model obtained from the (j-1)-th round of training through the S normalized gradients.
[0209] Optionally, the method further includes: obtaining an M-dimensional known sample label set, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label; dimensionally ascending the M-dimensional known sample labels in the M-dimensional known sample label set through a dimension-ascending module to obtain a K-dimensional known sample label set; using the K-dimensional known sample label set to jointly train a to-be-trained encoding module and a to-be-trained decoding module until a second end condition is satisfied among a first loss function, a second loss function, and a third loss function, ending the training, and obtaining the encoding module in the dimension-ascending encoding module and the decoding module in the decoding dimension-reduction module, where the decoding dimension-reduction module includes the decoding module and a dimension-reduction module, the dimension-ascending encoding module includes the encoding module and the dimension-ascending module, and the dimension-reduction module corresponds to the dimension-ascending module in the dimension-ascending encoding module; where the first loss function is the loss function between the K-dimensional decoding output output by the to-be-trained decoder and the corresponding K-dimensional known sample label, the second loss function is the loss function between the K-dimensional encoding output output by the to-be-trained encoder and the corresponding K-dimensional known sample label, and the third loss function is the information entropy function of the K-dimensional encoding output output by the to-be-trained encoder.
[0210] Optionally, the joint training of the to-be-trained encoding module and the to-be-trained decoding module using the K-dimensional known sample label set includes: performing the i-th round of joint training on the to-be-trained encoding module and the to-be-trained decoding module through the following steps, where i is a positive integer greater than or equal to 1, and the encoding module and the decoding module obtained in the 0-th round of training are the untrained to-be-trained encoding module and the to-be-trained decoding module: inputting the K-dimensional known sample label used in the i-th round in the K-dimensional known sample label set into the encoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional encoding output obtained in the i-th round of training; inputting the K-dimensional encoding output obtained in the i-th round of training into the decoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional decoding output obtained in the i-th round of training; where, in the case that the first loss function between the K-dimensional decoding output obtained in the i-th round of training and the K-dimensional known sample label, the second loss function between the K-dimensional encoding output obtained in the i-th round of training and the K-dimensional known sample label, and the third loss function of the K-dimensional encoding output obtained in the i-th round of training do not satisfy the second end condition, adjusting the parameters in the to-be-trained encoding module and the to-be-trained decoding module; otherwise, ending the training to obtain the encoding module and the decoding module.
[0211] Optionally, the performing the i-th round of joint training on the to-be-trained encoding module and the to-be-trained decoding module further includes: inputting the K-dimensional decoding output obtained in the i-th round of training and the K-dimensional known sample label into the first loss function to obtain a first loss value; inputting the K-dimensional encoding output obtained in the i-th round of training and the K-dimensional known sample label into the second loss function to obtain a second loss value; inputting the K-dimensional encoding output obtained in the i-th round of training into the third loss function to obtain a third loss value; in the case that the difference between the first loss value, the second loss value, and the third loss value is less than or equal to a preset threshold, determining that the second end condition is satisfied.
[0212] Optionally, the dimension elevation of the M-dimensional known sample label in the M-dimensional known sample label set by the dimension elevation module in the dimension elevation encoding module includes: obtaining K K-dimensional one-hot encoding vectors, randomly dividing the K K-dimensional one-hot encoding vectors into M groups, or dividing them into M groups according to a preset rule to obtain M groups of K-dimensional one-hot encoding vectors; mapping the M-dimensional known sample labels to the M groups of K-dimensional one-hot encoding vectors respectively to obtain the K-dimensional known sample labels corresponding to the M-dimensional known sample labels.
[0213] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0214] According to another aspect of the embodiments of the present application, there is also provided a data processing device for implementing the above data processing method. As Figure 8 shown, the device includes: a first acquisition module 82, configured to acquire a first forward output and a second forward output, where the first forward output is an output obtained by processing first feature data by a first bottom model in a first party, and the second forward output is an output obtained by processing second feature data by a second bottom model in a second party; a processing module 84, configured to process the first forward output and the second forward output through a top model in the second party to obtain a K-dimensional inference result, where K is a positive integer greater than 2; a decoding and dimensionality reduction module 86, configured to decode and reduce the dimensionality of the K-dimensional inference result through a decoding and dimensionality reduction module in the second party to obtain an M-dimensional inference result, where M is a positive integer less than K and greater than or equal to 2.
[0215] Optionally, the above device is further configured to, after acquiring the first forward output and the second forward output, fuse the first forward output and the second forward output through a split layer in the second party to obtain a fusion result; and process the fusion result through a top model in the second party to obtain the K-dimensional inference result.
[0216] Optionally, the above device is further configured to splice the first forward output and the second forward output through the split layer in the second party to obtain the fusion result; or perform an average processing on the first forward output and the second forward output through the split layer in the second party to obtain the fusion result.
[0217] Optionally, the above device is further configured to obtain a first set of sample feature data, a second set of sample feature data, and an M-dimensional known sample label set with a corresponding relationship before obtaining the first forward output and the second forward output, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label; perform dimensionality increase encoding on the M-dimensional known sample label set through a dimensionality increase encoding module to obtain a K-dimensional known encoded sample label set; use the first set of sample feature data, the second set of sample feature data, and the K-dimensional known encoded sample label set to jointly train the to-be-trained first bottom model in the first party, the to-be-trained second bottom model in the second party, and the to-be-trained top model until a first end condition is satisfied, and end the training to obtain the first bottom model, the second bottom model, and the top model.
[0218] Optionally, the above device is further configured to perform the j-th round of joint training on the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model through the following steps, where j is a positive integer greater than or equal to 1, and the first bottom model, the second bottom model, and the top model obtained in the 0-th round of training are the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model that have not been trained, and include: inputting S first sample feature data used in the j-th round in the first set of sample feature data into the first bottom model obtained in the (j - 1)-th round of training to obtain S first sample forward outputs output in the j-th round of training, where S is a positive integer greater than or equal to 1; inputting S second sample feature data used in the j-th round in the second set of sample feature data into the second bottom model obtained in the (j - 1)-th round of training to obtain S second sample forward outputs output in the j-th round of training; inputting the S first sample forward outputs output in the j-th round of training and the S second sample forward outputs output in the j-th round of training into the top model obtained in the (j - 1)-th round of training to obtain S K-dimensional sample inference results output in the j-th round of training; ending the training to obtain the first bottom model, the second bottom model, and the top model when the first end condition is satisfied between the S K-dimensional sample inference results output in the j-th round of training and the S K-dimensional known encoded sample labels used in the j-th round, and the K-dimensional known encoded sample labels used in the j-th round have the corresponding relationship with the first sample feature data used in the j-th round and the second sample feature data used in the j-th round.
[0219] Optionally, the above device is further configured to, when the S K-dimensional sample inference results output in the j-th round of training do not satisfy the first end condition with respect to the S K-dimensional known encoded sample labels used in the j-th round, adjust the model parameters of the first bottom model, the second bottom model, and the top model obtained in the (j - 1)-th round of training through backpropagation.
[0220] Optionally, the above device is further configured to enable the second party to obtain S first gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S K-dimensional sample inference results, and adjust the model parameters of the top model obtained in the (j - 1)-th round of training through the S first gradients; the second party obtains S second gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the model parameters of the second bottom model obtained in the (j - 1)-th round of training, and adjusts the model parameters of the second bottom model obtained in the (j - 1)-th round of training through the S second gradients; the second party obtains S third gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S first sample forward outputs; the second party normalizes the S third gradients through the second norm to obtain S normalized gradients; the second party sends the S normalized gradients to the first party, and the first party adjusts the first bottom model obtained in the (j - 1)-th round of training through the S normalized gradients.
[0221] Optionally, the above device is further configured to obtain an M-dimensional known sample label set, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label; perform dimensionality elevation on the M-dimensional known sample labels in the M-dimensional known sample label set through a dimensionality elevation module to obtain a K-dimensional known sample label set; use the K-dimensional known sample label set to jointly train the encoding module and the decoding module to be trained until the second end condition is satisfied among the first loss function, the second loss function, and the third loss function, and end the training to obtain the encoding module and the decoding module, where the decoding dimensionality reduction module includes the decoding module and the dimensionality reduction module, the dimensionality elevation encoding module includes the encoding module and the dimensionality elevation module, and the dimensionality reduction module corresponds to the dimensionality elevation module in the dimensionality elevation encoding module; where the first loss function is the loss function between the K-dimensional decoding output of the decoder to be trained and the corresponding K-dimensional known sample label, the second loss function is the loss function between the K-dimensional encoding output of the encoder to be trained and the corresponding K-dimensional known sample label, and the third loss function is the information entropy function of the K-dimensional encoding output of the encoder to be trained.
[0222] Optionally, the above device is further configured to perform the i-th round of joint training on the to-be-trained encoding module and the to-be-trained decoding module through the following steps, where i is a positive integer greater than or equal to 1, and the encoding module and the decoding module obtained in the 0-th round of training are the to-be-trained encoding module and the to-be-trained decoding module that have not been trained: input the K-dimensional known sample label used in the i-th round in the K-dimensional known sample label set into the encoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional encoding output obtained in the i-th round of training; input the K-dimensional encoding output obtained in the i-th round of training into the decoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional decoding output obtained in the i-th round of training; wherein, when the first loss function between the K-dimensional decoding output obtained in the i-th round of training and the K-dimensional known sample label, the second loss function between the K-dimensional encoding output obtained in the i-th round of training and the K-dimensional known sample label, and the third loss function of the K-dimensional encoding output obtained in the i-th round of training do not satisfy the second end condition, adjust the parameters in the to-be-trained encoding module and the to-be-trained decoding module; otherwise, end the training to obtain the encoding module and the decoding module.
[0223] Optionally, the above device is further configured to input the K-dimensional decoding output obtained in the i-th round of training and the K-dimensional known sample label into the first loss function to obtain a first loss value; input the K-dimensional encoding output obtained in the i-th round of training and the K-dimensional known sample label into the second loss function to obtain a second loss value; input the K-dimensional encoding output obtained in the i-th round of training into the third loss function to obtain a third loss value; when the difference between the first loss value, the second loss value, and the third loss value is less than or equal to a preset threshold, it is determined that the second end condition is satisfied.
[0224] Optionally, the above device is further configured to decode the K-dimensional inference result through the decoding module in the decoding and dimensionality reduction module to obtain a K-dimensional decoded inference result; reduce the dimensionality of the K-dimensional decoded inference result through the dimensionality reduction module in the decoding and dimensionality reduction module to obtain the M-dimensional inference result.
[0225] Optionally, the above device is further configured to obtain K K-dimensional one-hot encoding vectors, randomly divide the K K-dimensional one-hot encoding vectors into M groups, or divide them into M groups according to a preset rule to obtain M groups of K-dimensional one-hot encoding vectors; map the M-dimensional known sample labels to the M groups of K-dimensional one-hot encoding vectors respectively to obtain the K-dimensional known sample labels corresponding to the M-dimensional known sample labels; perform an inner product operation on the K-dimensional inference result and M K-dimensional decoding vectors to obtain the M-dimensional inference result, where the M K-dimensional decoding vectors are decoding vectors obtained according to the M groups of K-dimensional one-hot encodings.
[0226] According to another aspect of the embodiments of the present application, a data processing device is further provided. The device includes: a second acquisition module, configured to acquire a first set of sample feature data, a second set of sample feature data, and an M-dimensional known sample label set with a corresponding relationship, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label; a dimensionality increase encoding module, configured to perform dimensionality increase encoding on the M-dimensional known sample label set through the dimensionality increase encoding module to obtain a K-dimensional known encoded sample label set; a training module, configured to jointly train a first bottom model to be trained in a first party, a second bottom model to be trained in a second party, and a top model to be trained using the first set of sample feature data, the second set of sample feature data, and the K-dimensional known encoded sample label set until a first end condition is satisfied, and end the training to obtain the first bottom model in the first party, the second bottom model in the second party, and the top model.
[0227] Optionally, the above device is further configured to perform the j-th round of joint training on the first bottom model to be trained, the second bottom model to be trained, and the top model to be trained through the following steps, where j is a positive integer greater than or equal to 1, and the first bottom model, the second bottom model, and the top model obtained in the 0-th round of training are the first bottom model to be trained, the second bottom model to be trained, and the top model to be trained that have not been trained, including: inputting S first sample feature data used in the j-th round in the first set of sample feature data into the first bottom model obtained in the (j - 1)-th round of training to obtain S first sample forward outputs output in the j-th round of training, where S is a positive integer greater than or equal to 1; inputting S second sample feature data used in the j-th round in the second set of sample feature data into the second bottom model obtained in the (j - 1)-th round of training to obtain S second sample forward outputs output in the j-th round of training; inputting the S first sample forward outputs output in the j-th round of training and the S second sample forward outputs output in the j-th round of training into the top model obtained in the (j - 1)-th round of training to obtain S K-dimensional sample inference results output in the j-th round of training; ending the training to obtain the first bottom model, the second bottom model, and the top model when the first end condition is satisfied between the S K-dimensional sample inference results output in the j-th round of training and the S K-dimensional known encoded sample labels used in the j-th round, and the K-dimensional known encoded sample labels used in the j-th round have the corresponding relationship with the first sample feature data used in the j-th round and the second sample feature data used in the j-th round.
[0228] Optionally, the above-mentioned device is further configured to, when the S K-dimensional sample inference results output in the j-th round of training do not satisfy the first end condition with respect to the S K-dimensional known encoded sample labels used in the j-th round of training, adjust the model parameters of the first bottom model, the second bottom model, and the top model obtained in the (j - 1)-th round of training through backpropagation.
[0229] Optionally, the above-mentioned device is further configured to obtain S first gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S K-dimensional sample inference results, and adjust the model parameters of the top model obtained in the (j - 1)-th round of training through the S first gradients; obtain S second gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the model parameters of the second bottom model obtained in the (j - 1)-th round of training, and adjust the model parameters of the second bottom model obtained in the (j - 1)-th round of training through the S second gradients; obtain S third gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S first sample forward outputs; the second party normalizes the S third gradients through the second norm to obtain S normalized gradients; the second party sends the S normalized gradients to the first party, and the first party adjusts the first bottom model obtained in the (j - 1)-th round of training through the S normalized gradients.
[0230] Optionally, the above-mentioned device is further configured to obtain an M-dimensional known sample label set, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label; perform dimensionality increase on the M-dimensional known sample labels in the M-dimensional known sample label set through a dimensionality increase module to obtain a K-dimensional known sample label set; use the K-dimensional known sample label set to jointly train the to-be-trained encoding module and the to-be-trained decoding module until the second end condition is satisfied among the first loss function, the second loss function, and the third loss function, end the training, and obtain the encoding module in the dimensionality increase encoding module and the decoding module in the decoding dimensionality reduction module, where the decoding dimensionality reduction module includes the decoding module and the dimensionality reduction module, the dimensionality increase encoding module includes the encoding module and the dimensionality increase module, and the dimensionality reduction module corresponds to the dimensionality increase module in the dimensionality increase encoding module; where the first loss function is the loss function between the K-dimensional decoding output output by the to-be-trained decoder and the corresponding K-dimensional known sample label, the second loss function is the loss function between the K-dimensional encoding output output by the to-be-trained encoder and the corresponding K-dimensional known sample label, and the third loss function is the information entropy function of the K-dimensional encoding output output by the to-be-trained encoder.
[0231] Optionally, the above-mentioned device is further configured to perform the i-th round of joint training on the to-be-trained encoding module and the to-be-trained decoding module through the following steps, where i is a positive integer greater than or equal to 1, and the encoding module and the decoding module obtained in the 0-th round of training are the to-be-trained encoding module and the to-be-trained decoding module that have not been trained: input the K-dimensional known sample label used in the i-th round in the K-dimensional known sample label set into the encoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional encoding output obtained in the i-th round of training; input the K-dimensional encoding output obtained in the i-th round of training into the decoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional decoding output obtained in the i-th round of training; where, in the case that the first loss function between the K-dimensional decoding output obtained in the i-th round of training and the K-dimensional known sample label, the second loss function between the K-dimensional encoding output obtained in the i-th round of training and the K-dimensional known sample label, and the third loss function of the K-dimensional encoding output obtained in the i-th round of training do not satisfy the second end condition, adjust the parameters in the to-be-trained encoding module and the to-be-trained decoding module; otherwise, end the training to obtain the encoding module and the decoding module.
[0232] Optionally, the above-mentioned device is further configured to input the K-dimensional decoding output obtained in the i-th round of training and the K-dimensional known sample label into the first loss function to obtain a first loss value; input the K-dimensional encoding output obtained in the i-th round of training and the K-dimensional known sample label into the second loss function to obtain a second loss value; input the K-dimensional encoding output obtained in the i-th round of training into the third loss function to obtain a third loss value; in the case that the difference between the first loss value, the second loss value, and the third loss value is less than or equal to a preset threshold, it is determined that the second end condition is satisfied.
[0233] Optionally, the above-mentioned device is further configured to obtain K K-dimensional one-hot encoding vectors, randomly divide the K K-dimensional one-hot encoding vectors into M groups, or divide them into M groups according to a preset rule to obtain M groups of K-dimensional one-hot encoding vectors; map the M-dimensional known sample labels to the M groups of K-dimensional one-hot encoding vectors respectively to obtain K-dimensional known sample labels corresponding to the M-dimensional known sample labels.
[0234] According to one aspect of the present application, there is provided a computer program product, which includes computer programs / instructions, and the computer programs / instructions include program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit 901, it executes various functions provided in the embodiments of the present application.
[0235] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0236] Figure 9 Schematically shown is a block diagram of a computer system of an electronic device for implementing the embodiments of the present application.
[0237] It should be noted that Figure 9 The computer system 900 of the shown electronic device is only an example and should not impose any limitation on the functions and the scope of use of the embodiments of the present application.
[0238] As Figure 9 shown, the computer system 900 includes a central processing unit 901 (Central Processing Unit, CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 902 (Read-Only Memory, ROM) or the program loaded from the storage section 908 into the random access memory 903 (Random Access Memory, RAM). In the random access memory 903, various programs and data required for system operation are also stored. The central processing unit 901, the read-only memory 902, and the random access memory 903 are connected to each other via a bus 904. The input / output interface 905 (Input / Output interface, i.e., I / O interface) is also connected to the bus 904.
[0239] The following components are connected to the input / output interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including such as a cathode ray tube (Cathode Ray Tube, CRT), a liquid crystal display (Liquid Crystal Display, LCD), etc. and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a local area network card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output interface 905 as required. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as required so that the computer program read from it can be installed into the storage section 908 as required.
[0240] In particular, according to an embodiment of the present application, the processes described in each method flowchart can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium. The computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit 901, various functions defined in the system of the present application are executed.
[0241] According to another aspect of the embodiments of the present application, an electronic device for implementing the above data processing method is further provided. The electronic device may be Figure 1 the terminal device or server shown. This embodiment takes the electronic device as the terminal device as an example for illustration. As Figure 10 shown, the electronic device includes a memory 1002 and a processor 1004. A computer program is stored in the memory 1002, and the processor 1004 is configured to execute the steps in any of the above method embodiments through the computer program.
[0242] Optionally, in this embodiment, the above electronic device may be at least one network device among multiple network devices in a computer network.
[0243] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0244] S1, obtain a first forward output and a second forward output, where the first forward output is the output obtained by processing first feature data by a first bottom model in a first participating party, and the second forward output is the output obtained by processing second feature data by a second bottom model in a second participating party;
[0245] S2, process the first forward output and the second forward output through a top model in the second participating party to obtain a K-dimensional inference result, where K is a positive integer greater than 2;
[0246] S3, decode and reduce the dimension of the K-dimensional inference result through a decoding and dimension reduction module in the second participating party to obtain an M-dimensional inference result, where M is a positive integer less than K and greater than or equal to 2.
[0247] Optionally, in this embodiment, the above processor may also be configured to execute the following steps through a computer program:
[0248] S1. Obtain a first set of sample feature data, a second set of sample feature data, and an M-dimensional known sample label set with corresponding relationships, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label;
[0249] S2. Perform dimensionality increase encoding on the M-dimensional known sample label set through a dimensionality increase encoding module to obtain a K-dimensional known encoded sample label set;
[0250] S3. Use the first set of sample feature data, the second set of sample feature data, and the K-dimensional known encoded sample label set to jointly train a to-be-trained first bottom model in the first party, a to-be-trained second bottom model in the second party, and a to-be-trained top model until a first end condition is met, end the training, and obtain the first bottom model in the first party, the second bottom model in the second party, and the top model.
[0251] Optionally, those of ordinary skill in the art can understand that Figure 10 the structure shown is only schematic, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 10 It does not limit the structure of the above electronic device. For example, the electronic device may further include Figure 10 more or fewer components (such as a network interface, etc.) than those shown, or have a different configuration from Figure 10 that shown.
[0252] Among them, the memory 1002 can be used to store software programs and modules, such as the program instructions / modules corresponding to the data processing method and device in this embodiment of the present application. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, that is, implements the above data processing method. The memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 1002 may further include a memory remotely set relative to the processor 1004, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. Among them, the memory 1002 can specifically but not limitedly be used to store information such as first feature data and second feature data. As an example, Figure 10As shown, the above-mentioned memory 1002 may but is not limited to include the first acquisition module 82, the processing module 84, and the decoding and dimensionality reduction module 86 in the above-mentioned data processing device. In addition, it may also include but is not limited to other module units in the above-mentioned data processing device, which will not be elaborated in this example.
[0253] Optionally, the above-mentioned transmission device 1006 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one example, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one example, the transmission device 1006 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0254] In addition, the above-mentioned electronic device further includes: a display 1008, which is used to display the above-mentioned M-dimensional inference result; and a connection bus 1010, which is used to connect each module component in the above-mentioned electronic device.
[0255] In other embodiments, the above-mentioned terminal device or server may be a node in a distributed system, where the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. Among them, the nodes can form a peer-to-peer (P2P, Peer To Peer) network, and any form of computing device, such as servers, terminals and other electronic devices, can become a node in the blockchain system by joining the peer-to-peer network.
[0256] According to one aspect of the present application, there is provided a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing methods provided in the above various optional implementation manners.
[0257] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be set to store a computer program for executing the following steps:
[0258] S1, obtain a first forward output and a second forward output, where the first forward output is the output obtained by processing first feature data by a first bottom model in a first party, and the second forward output is the output obtained by processing second feature data by a second bottom model in a second party;
[0259] S2. Process the first forward output and the second forward output through the top model in the second participating party to obtain a K-dimensional inference result, where K is a positive integer greater than 2;
[0260] S3. Decode and reduce the dimension of the K-dimensional inference result through the decoding and dimensionality reduction module in the second participating party to obtain an M-dimensional inference result, where M is a positive integer less than K and greater than or equal to 2.
[0261] Optionally, in this embodiment, the above computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0262] S1. Obtain a first set of sample feature data, a second set of sample feature data, and an M-dimensional set of known sample labels with corresponding relationships, where each M-dimensional known sample label in the M-dimensional set of known sample labels is used to represent an actual sample label;
[0263] S2. Perform dimensionality increase encoding on the M-dimensional set of known sample labels through the dimensionality increase encoding module to obtain a K-dimensional set of known encoded sample labels;
[0264] S3. Jointly train the first bottom model to be trained in the first participating party, the second bottom model to be trained in the second participating party, and the top model to be trained using the first set of sample feature data, the second set of sample feature data, and the K-dimensional set of known encoded sample labels until a first end condition is met, end the training, and obtain the first bottom model in the first participating party, the second bottom model in the second participating party, and the top model.
[0265] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructing the relevant hardware of the terminal device through a program, and this program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0266] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.
[0267] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage media. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in various embodiments of this application.
[0268] In the above embodiments of this application, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0269] In the several embodiments provided by this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0270] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0271] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit exists physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0272] The above is only the preferred embodiment of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A data processing method, characterized in that, Including: Obtain a first forward output and a second forward output, where the first forward output is an output obtained by processing first feature data by a first bottom model in a first participant, and the second forward output is an output obtained by processing second feature data by a second bottom model in a second participant; Process the first forward output and the second forward output through a top model in the second participant to obtain a K-dimensional inference result, where K is a positive integer greater than 2; Decode the K-dimensional inference result through a decoding module included in a decoding and dimensionality reduction module in the second participant to obtain a K-dimensional decoded inference result, where both the decoding module and an encoding module are trained using a K-dimensional known sample label set, and the K-dimensional known sample label set is obtained by dimensionality increasing the M-dimensional known sample labels in an M-dimensional known sample label set through a dimensionality increasing module, and both the encoding module and the dimensionality increasing module are located in a dimensionality increasing and encoding module in the second participant; In a dimensionality reduction module included in the decoding and dimensionality reduction module, perform an inner product operation on the K-dimensional decoded inference result and M K-dimensional decoding vectors to obtain an M-dimensional inference result, where the M K-dimensional decoding vectors are decoding vectors based on M groups of K-dimensional one-hot encoding vectors, and the M groups of K-dimensional one-hot encoding vectors are used to perform dimensionality increasing on the M-dimensional known sample labels in the M-dimensional known sample label set through the dimensionality increasing module, and M is a positive integer less than K and greater than or equal to 2.
2. The method according to claim 1, wherein: After obtaining the first forward output and the second forward output, the method further includes: fusing the first forward output and the second forward output through a splitting layer in the second participant to obtain a fusion result; The processing the first forward output and the second forward output through a top model in the second participant to obtain a K-dimensional inference result includes: processing the fusion result through a top model in the second participant to obtain the K-dimensional inference result.
3. The method according to claim 2, wherein The fusing the first forward output and the second forward output through a splitting layer in the second participant to obtain a fusion result includes: Concatenating the first forward output and the second forward output through the splitting layer in the second participant to obtain the fusion result; or Performing an averaging process on the first forward output and the second forward output through the splitting layer in the second participant to obtain the fusion result.
4. The method according to claim 1, wherein Before obtaining the first forward output and the second forward output, the method further includes: Obtain a first sample feature data set, a second sample feature data set with a corresponding relationship, and the M-dimensional known sample label set, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label; Perform dimensionality increasing and encoding on the M-dimensional known sample label set through the dimensionality increasing and encoding module to obtain a K-dimensional known encoded sample label set; Use the first sample feature data set, the second sample feature data set, and the K-dimensional known encoded sample label set to jointly train the to-be-trained first bottom model in the first party, the to-be-trained second bottom model in the second party, and the to-be-trained top model until the first end condition is satisfied, end the training, and obtain the first bottom model, the second bottom model, and the top model.
5. The method according to claim 4, wherein The joint training of the to-be-trained first bottom model in the first party, the to-be-trained second bottom model in the second party, and the to-be-trained top model using the first sample feature data set, the second sample feature data set, and the K-dimensional known encoded sample label set includes: Perform the j-th round of joint training on the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model through the following steps, where j is a positive integer greater than or equal to 1, and the first bottom model, second bottom model, and top model obtained in the 0-th round of training are the to-be-trained first bottom model, the to-be-trained second bottom model, and the to-be-trained top model that have not been trained, and include: Input the S first sample feature data used in the j-th round in the first sample feature data set into the first bottom model obtained in the (j - 1)-th round of training to obtain S first sample forward outputs output in the j-th round of training, where S is a positive integer greater than or equal to 1; Input the S second sample feature data used in the j-th round in the second sample feature data set into the second bottom model obtained in the (j - 1)-th round of training to obtain S second sample forward outputs output in the j-th round of training; Input the S first sample forward outputs output in the j-th round of training and the S second sample forward outputs output in the j-th round of training into the top model obtained in the (j - 1)-th round of training to obtain S K-dimensional sample inference results output in the j-th round of training; When the first end condition is satisfied between the S K-dimensional sample inference results output in the j-th round of training and the S K-dimensional known encoded sample labels used in the j-th round, end the training to obtain the first bottom model, the second bottom model, and the top model, and the K-dimensional known encoded sample labels used in the j-th round have the corresponding relationship with the first sample feature data used in the j-th round and the second sample feature data used in the j-th round.
6. The method according to claim 5, characterized in that, The method further includes: When the first end condition is not satisfied between the S K-dimensional sample inference results output in the j-th round of training and the S K-dimensional known encoded sample labels used in the j-th round, adjust the model parameters of the first bottom model, the second bottom model, and the top model obtained in the (j - 1)-th round of training through backpropagation.
7. The method according to claim 6, wherein Adjusting the model parameters of the top model obtained from the (j - 1)-th round of training through backpropagation includes: the second party obtains S first gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S K-dimensional sample inference results, and adjusts the model parameters of the top model obtained from the (j - 1)-th round of training through the S first gradients; Adjusting the model parameters of the second bottom model obtained from the (j - 1)-th round of training through backpropagation includes: the second party obtains S second gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the model parameters of the second bottom model obtained from the (j - 1)-th round of training, and adjusts the model parameters of the second bottom model obtained from the (j - 1)-th round of training through the S second gradients; Adjusting the first bottom model obtained from the (j - 1)-th round of training through backpropagation includes: the second party obtains S third gradients through the S loss functions between the S K-dimensional sample inference results and the S K-dimensional known encoded sample labels and the S first sample forward outputs; the second party normalizes the S third gradients through the second norm to obtain S normalized gradients; the second party sends the S normalized gradients to the first party, and the first party adjusts the first bottom model obtained from the (j - 1)-th round of training through the S normalized gradients.
8. The method according to claim 1, wherein The method further includes: Obtaining the set of M-dimensional known sample labels, where each M-dimensional known sample label in the set of M-dimensional known sample labels is used to represent an actual sample label; Dimensionality-increasing the M-dimensional known sample labels in the set of M-dimensional known sample labels through the dimensionality-increasing module to obtain the set of K-dimensional known sample labels; Using the set of K-dimensional known sample labels to jointly train the encoding module to be trained and the decoding module to be trained until the second termination condition is satisfied among the first loss function, the second loss function, and the third loss function, terminating the training, and obtaining the encoding module and the decoding module, where the decoding and dimensionality-reducing module includes the decoding module and the dimensionality-reducing module, the dimensionality-increasing and encoding module includes the encoding module and the dimensionality-increasing module, and the dimensionality-reducing module corresponds to the dimensionality-increasing module in the dimensionality-increasing and encoding module; Wherein, the first loss function is the loss function between the K-dimensional decoding output of the decoder to be trained and the corresponding K-dimensional known sample label, the second loss function is the loss function between the K-dimensional encoding output of the encoder to be trained and the corresponding K-dimensional known sample label, and the third loss function is the information entropy function of the K-dimensional encoding output of the encoder to be trained.
9. According to the method of claim 8, the using the set of K-dimensional known sample labels to jointly train the encoding module to be trained and the decoding module to be trained includes: The i-th round of joint training for the to-be-trained encoding module and the to-be-trained decoding module is performed through the following steps, where i is a positive integer greater than or equal to 1, and the encoding module and the decoding module obtained in the 0-th round of training are the untrained to-be-trained encoding module and the to-be-trained decoding module: Input the K-dimensional known sample label used in the i-th round in the K-dimensional known sample label set into the encoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional encoding output obtained in the i-th round of training; Input the K-dimensional encoding output obtained in the i-th round of training into the decoding module obtained in the (i - 1)-th round of training to obtain the K-dimensional decoding output obtained in the i-th round of training; Among them, in the case where the first loss function between the K-dimensional decoding output obtained in the i-th round of training and the K-dimensional known sample label, the second loss function between the K-dimensional encoding output obtained in the i-th round of training and the K-dimensional known sample label, and the third loss function of the K-dimensional encoding output obtained in the i-th round of training do not satisfy the second end condition, adjust the parameters in the to-be-trained encoding module and the to-be-trained decoding module; otherwise, end the training to obtain the encoding module and the decoding module.
10. The method according to claim 9, wherein the i-th round of joint training for the to-be-trained encoding module and the to-be-trained decoding module further includes: Input the K-dimensional decoding output obtained in the i-th round of training and the K-dimensional known sample label into the first loss function to obtain a first loss value; Input the K-dimensional encoding output obtained in the i-th round of training and the K-dimensional known sample label into the second loss function to obtain a second loss value; Input the K-dimensional encoding output obtained in the i-th round of training into the third loss function to obtain a third loss value; In the case where the difference between the first loss value, the second loss value, and the third loss value is less than or equal to a preset threshold, it is determined that the second end condition is satisfied.
11. The method according to claim 8, wherein The dimensionality increase of the M-dimensional known sample labels in the M-dimensional known sample label set by the dimensionality increase module in the dimensionality increase encoding module includes: obtaining K K-dimensional one-hot encoding vectors, randomly dividing the K K-dimensional one-hot encoding vectors into M groups, or dividing them into M groups according to a preset rule to obtain the M groups of K-dimensional one-hot encoding vectors; in the dimensionality increase module, map the M-dimensional known sample labels to the M groups of K-dimensional one-hot encoding vectors respectively to obtain the K-dimensional known sample labels corresponding to the M-dimensional known sample labels.
12. A data processing method, characterized in that including: Obtain a first sample feature data set, a second sample feature data set, and an M-dimensional known sample label set with corresponding relationships, where each M-dimensional known sample label in the M-dimensional known sample label set is used to represent an actual sample label; Obtain M groups of K-dimensional one-hot encoding vectors; In the dimensionality-raising encoding module, the M-dimensional known sample label set is encoded with the M groups of K-dimensional one-hot encoding vectors to obtain a K-dimensional known encoded sample label set. Among them, the dimensionality-raising encoding module includes an encoding module and a dimensionality-raising module. The dimensionality-raising module corresponds to the dimensionality-reduction module included in the decoding and dimensionality-reduction module. The M K-dimensional decoding vectors for performing dimensionality-reduction operations in the dimensionality-reduction module are decoding vectors obtained based on the M groups of K-dimensional one-hot encoding vectors; The first bottom model to be trained in the first party, the second bottom model to be trained in the second party, and the top model to be trained are jointly trained using the first sample feature data set, the second sample feature data set, and the K-dimensional known encoded sample label set until the first end condition is met, and then the training ends to obtain the first bottom model in the first party, the second bottom model in the second party, and the top model.
13. The method according to claim 12, characterized in that, The joint training of the first bottom model to be trained in the first party, the second bottom model to be trained in the second party, and the top model to be trained using the first sample feature data set, the second sample feature data set, and the K-dimensional known encoded sample label set includes: The following steps are used to perform the j-th round of joint training on the first bottom model to be trained, the second bottom model to be trained, and the top model to be trained, where j is a positive integer greater than or equal to 1. The first bottom model, the second bottom model, and the top model obtained in the 0-th round of training are the first bottom model to be trained, the second bottom model to be trained, and the top model to be trained that have not been trained, and it includes: The S first sample features used in the j-th round in the first sample feature data set are input into the first bottom model obtained in the (j - 1)-th round of training to obtain the S first sample forward outputs output in the j-th round of training, where S is a positive integer greater than or equal to 1; The S second sample features used in the j-th round in the second sample feature data set are input into the second bottom model obtained in the (j - 1)-th round of training to obtain the S second sample forward outputs output in the j-th round of training; The S first sample forward outputs output in the j-th round of training and the S second sample forward outputs output in the j-th round of training are input into the top model obtained in the (j - 1)-th round of training to obtain the S K-dimensional sample inference results output in the j-th round of training; When the first end condition is met between the S K-dimensional sample inference results output in the j-th round of training and the S K-dimensional known encoded sample labels used in the j-th round, the training ends to obtain the first bottom model, the second bottom model, and the top model. The K-dimensional known encoded sample labels used in the j-th round have the corresponding relationship with the first sample features used in the j-th round and the second sample features used in the j-th round.
14. The method according to claim 12, characterized in that Dimensionality-raising the M-dimensional known sample labels in the M-dimensional known sample label set through the dimensionality-raising module in the dimensionality-raising encoding module includes: Obtain K K-dimensional one-hot encoding vectors, randomly divide the K K-dimensional one-hot encoding vectors into M groups, or divide them into M groups according to a preset rule, to obtain the M groups of K-dimensional one-hot encoding vectors; map the M-dimensional known sample labels to the M groups of K-dimensional one-hot encoding vectors respectively, to obtain the K-dimensional known sample labels corresponding to the M-dimensional known sample labels.
15. A data processing device, characterized in that, It includes: A first acquisition module, configured to acquire a first forward output and a second forward output, where the first forward output is an output obtained by processing first feature data by a first bottom model in a first party, and the second forward output is an output obtained by processing second feature data by a second bottom model in a second party; A processing module, configured to process the first forward output and the second forward output through a top model in the second party to obtain a K-dimensional inference result, where K is a positive integer greater than 2; A decoding and dimensionality reduction module, configured to decode the K-dimensional inference result through a decoding module included in the decoding and dimensionality reduction module in the second party to obtain a K-dimensional decoded inference result, where both the decoding module and the encoding module are trained using a set of K-dimensional known sample labels, and the set of K-dimensional known sample labels is obtained by dimensionality increasing of the M-dimensional known sample labels in a set of M-dimensional known sample labels through a dimensionality increasing module, and both the encoding module and the dimensionality increasing module are located in the dimensionality increasing and encoding module in the second party; in a dimensionality reduction module included in the decoding and dimensionality reduction module, perform an inner product operation on the K-dimensional decoded inference result and M K-dimensional decoding vectors to obtain an M-dimensional inference result, where the M K-dimensional decoding vectors are decoding vectors based on M groups of K-dimensional one-hot encoding vectors, and the M groups of K-dimensional one-hot encoding vectors are used to perform dimensionality increasing on the M-dimensional known sample labels in the set of M-dimensional known sample labels through the dimensionality increasing module, and M is a positive integer less than K and greater than or equal to 2.
16. A data processing device, characterized in that, It includes: A second acquisition module, configured to acquire a set of first sample feature data, a set of second sample feature data, and a set of M-dimensional known sample labels with corresponding relationships, where each M-dimensional known sample label in the set of M-dimensional known sample labels is used to represent an actual sample label; A dimensionality increasing and encoding module, configured to acquire M groups of K-dimensional one-hot encoding vectors; use the M groups of K-dimensional one-hot encoding vectors in the dimensionality increasing and encoding module to perform dimensionality increasing encoding on the set of M-dimensional known sample labels to obtain a set of K-dimensional known encoded sample labels, where the dimensionality increasing and encoding module includes an encoding module and a dimensionality increasing module, the dimensionality increasing module corresponds to the dimensionality reduction module included in the decoding and dimensionality reduction module, and the M K-dimensional decoding vectors for performing dimensionality reduction operations in the dimensionality reduction module are decoding vectors based on the M groups of K-dimensional one-hot encoding vectors; A training module, configured to jointly train a first bottom model to be trained in a first party, a second bottom model to be trained in a second party, and a top model to be trained by using the first sample feature data set, the second sample feature data set, and the K-dimensional known encoded sample label set until a first end condition is satisfied, terminate the training, and obtain the first bottom model in the first party, the second bottom model in the second party, and the top model.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed by a terminal device or a computer, executes the method described in any one of claims 1 to 11 or 12 to 14.
18. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method described in any one of claims 1 to 11 or 12 to 14 are implemented.
19. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 11 or 12 to 14 through the computer program.
Citation Information
Patent Citations
Data dimension reduction method and device and related equipment
CN113240045A
Federal neural network model-based data processing method, related equipment and medium
CN113505882A
Targeted attack defense method and device
CN114139147A
Training method and device of longitudinal federal learning model and computer equipment
CN114239820A