Longitudinal federated learning privacy protection method and system considering inter-federated data leakage
Through semantic similar image set alignment and conditional diffusion model reconstruction technology, combined with perceived hash and pixel matching, the privacy protection problem of data leakage among allies in vertical federated learning is solved, and the privacy risk is accurately evaluated and reinforced, and the model performance is optimized.
Patent Information
- Application Number
- CN202510775759.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Existing vertical federated learning programs fail to effectively evaluate and prevent the risk of data leakage among allies, especially when both servers and clients have unknown labels, intermediate results and top-level models, the risk of privacy leakage is severe and difficult to detect.
By constructing semantic similar image set alignment and conditional diffusion model reconstruction technology, combining perceived hash and pixel matching, a three-level evaluation mechanism is established, privacy risks are evaluated, and privacy protection is achieved through gradient noise addition and compression strategies, monitoring top-level model accuracy fluctuations, and optimizing privacy protection and model performance.
The privacy leakage risk assessment is achieved in the case of unknown labels and intermediate results, the privacy protection capabilities of vertical federated learning systems are enhanced, the risk of data theft among allies is accurately quantified, and the balance between model performance and privacy protection is reached.
Smart Images

Figure CN120296798A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vertical federated learning privacy protection, and in particular to a vertical federated learning privacy protection method and system that takes into account data leakage between allies. Background Art
[0002] As privacy protection requirements become increasingly stringent, data islands have been created to avoid privacy leakage, which has promoted the development of federated learning. Vertical Federated Learning (VFL) guides multiple clients to conduct joint training through a coordinator (usually a third-party server). As a distributed learning paradigm that balances privacy and model availability, it provides basic privacy guarantees that data does not leave the local area and is "available but invisible", that is, the data of each participant will not be known by other participants or servers. Its main application scenario is that clients share the same sample space and different feature spaces. For example, traffic cameras belonging to different institutions shoot the same object from different angles and jointly predict the object category to implement autonomous driving algorithm training. Among them, their common sample space is the same object, and the different feature spaces are that the high-dimensional features are different due to different shooting angles; as allies participating in federated learning, the two can be used to jointly predict image categories; but based on the regulations of the Data Security Protection Law and the basic privacy guarantee of vertical federated learning, the data of the two should not leave the local area and should not be known to each other.
[0003] In order to achieve the purpose of data protection, on the one hand, from the perspective of training architecture, clients in vertical federated learning can only access some features related to prediction, cannot obtain the overall features and labels required for prediction, and cannot understand the feature information owned by other clients; on the other hand, the vertical federated learning architecture performs noise protection on the content of communication between the server and the client to combat potential data theft attacks between allies.
[0004] There is no existing privacy protection scheme that considers data leakage between allies in the presence of a server. That is, the existing privacy risk assessment framework mostly relies on privacy leakage caused by two of the three entities: known labels, intermediate results, and top-level models, while ignoring the risk when all three are unknown. Privacy protection schemes that only consider data leakage risk assessment of two of the three known entities have been fully implemented, by introducing a trusted third party to protect labels and top-level models, and encrypting intermediate results; the risk of privacy leakage caused by allies whose three are unknown is very serious, and because the privacy leakage process does not have any impact on the vertical federated learning training process, it is not easy to detect. How to assess the risk of data leakage between allies and strengthen privacy protection in vertical federated learning is a technical problem that needs to be solved in this field. Summary of the invention
[0005] In view of the above problems, the present invention provides a privacy protection method and system for vertical federated learning considering data leakage between allies, which can effectively enhance the privacy protection ability of the vertical federated learning system against the risk of data leakage between allies.
[0006] The technical solution adopted by the present invention is as follows:
[0007] In a first aspect, the present invention proposes a privacy protection method for vertical federated learning considering data leakage between allies, including:
[0008] Record the initial accuracy of the top-level model on the server side of the vertical federated learning system. The server maintains a data subset sampled from the local data of each client.
[0009] Treat each client as the second client in turn, and the remaining clients as the first clients. Calculate the risk that each first client steals the data of the second client. Take the maximum value of the risks as the overall privacy leakage risk of the second client. If the overall privacy leakage risk exceeds the risk threshold, add noise to the updated gradient sent by the server to the second client, and recalculate the overall privacy leakage risk based on the noisy result until the privacy protection condition is met. Traverse all clients.
[0010] When calculating the risk that the first client steals the data of the second client, train a label reconstructor based on the data subset of the first client maintained by the server; obtain the first semantic similarity image set of the data subset of the first client and the second semantic similarity image set of the data subset of the second client, and align the sample spaces of the two semantic similarity image sets; use the label reconstructor to generate pseudo-labels for the first samples from the first semantic similarity image set, and use the first samples and their pseudo-labels as conditions to guide the reconstruction of a candidate set of the second samples from the second semantic similarity image set using a conditional diffusion model. Screen the reconstructed samples from the candidate set, and obtain the risk that the first client steals the data of the second client according to the peak signal-to-noise ratio between the reconstructed samples and the second samples.
[0011] Further, in the vertical federated learning system, each client maintains a local bottom-level model. The intermediate features generated by each local bottom-level model are spliced on the server and the target category is predicted by the top-level model. The updated gradients of each local bottom-level model are calculated on the server.
[0012] Further, in the vertical federated learning system, the local data maintained by the client is object images from different perspectives, and the object category is predicted by fusing the intermediate features of the object images from different perspectives on the server.
[0013] Further, training a label reconstructor based on the data subset of the first client maintained by the server includes:
[0014] Input the data subset of the first client maintained by the server into the local underlying model of the first client, reduce the dimension and cluster the intermediate features generated by the local underlying model, and generate clustering labels.
[0015] Use the data subset of the first client and its clustering labels to train a fully connected layer so that the fully connected layer can predict the clustering labels based on the intermediate features generated by the local underlying model of the first client, and use the prediction result as the pseudo label.
[0016] Further, the process of obtaining the first semantic similarity image set and the second semantic similarity image set includes:
[0017] Sample from the data subsets of the first and second clients respectively, and use a search engine to collect image sets that are semantically similar to the sampling results as the first and second semantic similarity image sets.
[0018] Align the sample spaces of the first and second semantic similarity image sets to ensure that the corresponding images in the two similar image sets are different perspective images of the same object.
[0019] Further, the process of using the first sample and its pseudo label as conditional guidance to reconstruct the candidate set of the second sample from the second semantic similarity image set includes:
[0020] Perform forward diffusion on the second sample from the second semantic similarity image set, gradually add noise to generate noise images; then use the first sample from the first semantic similarity image set and its pseudo label as conditional inputs, and gradually denoise the noise images under conditional guidance to restore the second sample, and train a conditional diffusion model; the second sample and the first sample are different perspective images of the same object.
[0021] Use a pure noise image as the input, and use the trained conditional diffusion model to generate a predicted image under the conditional guidance of the first sample and its pseudo label. The predicted image is used as a candidate image of the second sample; run the conditional diffusion model multiple times to generate the candidate set of the second sample.
[0022] Further, the process of screening the reconstructed samples from the candidate set includes:
[0023] Calculate the fidelity score according to the distribution of the predicted labels of each sample in the candidate set and the first sample, and eliminate the candidate samples with lower fidelity scores.
[0024] Calculate the perceptual similarity according to the normalized Hamming distance between the perceptual hash sequences of each sample in the candidate set and the first sample, and eliminate the candidate samples with lower perceptual similarity.
[0025] For each sample in the candidate set, the pixel values of the same position area as the first sample are randomly sampled, and the L2 distance is calculated element by element to obtain the matching score. The sample with the highest matching score from the candidate set is used as the reconstructed sample.
[0026] Furthermore, the process of adding noise to the updated gradient sent by the server to the second client includes:
[0027] Gradually adding small-scale noise to the gradient calculated by the server and sent to the second client to update the local bottom-level model, obtaining the noisy gradient, recording the accuracy of the top-level model on the server after the noise addition, and calculating the difference with the initial accuracy;
[0028] If the accuracy difference does not exceed the accuracy fluctuation threshold, the noise addition result is retained, and the overall privacy leakage risk of the second client after the noise addition is further evaluated;
[0029] If the accuracy difference exceeds the accuracy fluctuation threshold, the newly added noise is cleared, and the elements with smaller absolute values in the gradient are gradually replaced with zero, and the overall privacy leakage risk of the second client after the noise is added is continuously evaluated.
[0030] Furthermore, the privacy protection condition means that the accuracy difference of the top-level model on the server side does not exceed the accuracy fluctuation threshold, and the overall privacy leakage risk of each client does not exceed the risk threshold.
[0031] In the second aspect, the present invention proposes a vertical federated learning privacy protection system that takes into account data leakage between allies, which is used to implement the above-mentioned vertical federated learning privacy protection method that takes into account data leakage between allies.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] The present invention constructs a privacy risk assessment system for unknown labels / intermediate results / top-level models. Through the alignment of semantically similar image sets and conditional diffusion model reconstruction technology, it pioneers a three-level assessment mechanism of "label reconstruction-candidate generation-optimal screening", uses clustered label reconstruction to achieve semantic guidance, and combines multi-dimensional screening with perceptual hashing and pixel matching. It breaks through the dependence of traditional methods on intermediate results and label information, and for the first time realizes privacy leakage risk assessment in scenarios where all three are unknown. It fills the technical gap in the field of privacy protection in vertical federated learning that considers data leakage between allies, and accurately quantifies the risk of data theft between allies in vertical federated learning.
[0034] The present invention achieves the optimal balance between privacy protection strength and model performance through the dual strategies of gradient noise addition and compression, combined with real-time monitoring of top-level model accuracy fluctuations, and effectively solves the game problem between privacy reinforcement and model availability. The present invention can comprehensively evaluate the privacy leakage risk of vertical federated learning and carry out targeted reinforcement. Brief Description of the Drawings
[0035] Figure 1 Figure 1 is a schematic diagram of the overall process for enhancing the privacy protection ability of a single client.
[0036] Figure 2 Figure 2 is a schematic diagram of a vertical federated learning architecture (taking enhancing the privacy protection ability of client B among three clients as an example).
[0037] Figure 3 Figure 3 is a method for evaluating the risk of an ally stealing client data (taking client A stealing data from client B as an example).
[0038] Figure 4 Figure 4 shows the experimental results for evaluating the effectiveness of the optimal selection module. Detailed Implementation Manner
[0039] The present invention will be further described and explained below in conjunction with the detailed implementation manner. The embodiments are only exemplary of the present disclosure content and do not delimit the scope of limitation. The technical features of each implementation manner in the present invention can be combined correspondingly without conflict.
[0040] The drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0041] The flowcharts shown in the drawings are only illustrative and do not necessarily include all steps. For example, some steps can be decomposed, while some steps can be combined or partially combined, so the actual execution order may be changed according to the actual situation.
[0042] As Figure 1 shown, the main steps of the vertical federated learning privacy protection method considering data leakage between allies include:
[0043] (1) The server adds noise to the gradients returned to the clients for local model update or performs gradient compression;
[0044] (2) Based on the steps of "label reconstruction - candidate generation - optimal selection", evaluate the privacy leakage risk between allies;
[0045] (3) According to the evaluation results, adjust the scale of the noise, and repeat step (2) until the data leakage risk is effectively avoided.
[0046] For step (1), the server adds noise to the gradients returned to the client for local model update or performs gradient compression. An alternative implementation process is as follows:
[0047] (1-1) For K clients jointly participating in vertical federated learning, the top-level model with a classification accuracy of is finally trained; for each client, the risk of the remaining K-1 clients stealing its data is calculated , where the risk of the k-th client stealing its data is ; the overall data leakage risk faced by the client is .
[0048] For ease of description, as Figure 2 shown, the vertical federated learning system consists of three clients and a third-party server. Taking the server's evaluation of the data leakage risks of clients A and C for client B and privacy reinforcement of B as an example, the inter-alliance privacy leakage risk faced by B is calculated , and the overall data leakage risk is .
[0049] (1-2) The gradient calculated by the server for sending to client B to update the local underlying model is , and small-scale noise is gradually added to it to obtain , thereby realizing privacy protection for this client. Record the classification accuracy of the top-level model at this time after adding noise as , and calculate the accuracy difference ; when and , the privacy protection for this client is completed; note that the small-scale addition needs to be gradual because excessive noise will affect the performance of the vertical federated learning system.
[0050] (1-3) When and , the noise is cleared, and small-scale gradient compression is gradually implemented until when and . The implementation range is for gradient compression method, and the specific implementation is: calculate the absolute value of each element in and sort them from largest to smallest to determine the indices of the first elements; create a zero tensor with the same shape as , and copy the elements with the largest absolute values to according to their indices.
[0051] For step (2), based on the steps of "label reconstruction - candidate generation - optimal selection", the privacy leakage risk among allies is evaluated. Here, the server evaluates the data leakage risk caused by client A to client B as an example for introduction.
[0052] In an embodiment provided by the present invention, the data leakage risk assessment system includes three modules: a label reconstruction module, a candidate generation module, and an optimal selection module. Its overall architecture is as Figure 3 shown, and the server retains a small-scale verification dataset for the implementation verification link, where , and are respectively verification data subsets randomly sampled from the local datasets , and of the three clients; the specific modules and the functions of each module are as follows:
[0053] (2-1) The label reconstruction module is mainly used to generate pseudo-labels for each sample in the verification data subset , achieving two purposes: on the one hand, the pseudo-labels serve as a guide for the candidate generation module to ensure the semantic consistency between the generated data and the target data; on the other hand, align the high-level features between the training data of the candidate generation module (from the semantic similar image set of external data: ) and the vertical federated learning dataset ( and ), thereby improving the generation quality. Therefore, it is necessary to use the label reconstruction module to generate pseudo-labels for each sample.
[0054] The specific process of generating pseudo-labels is as follows:
[0055] The server splices the fully connected layer of the underlying model as a label reconstructor, performs t-SNE dimensionality reduction processing on the intermediate high-dimensional feature vectors generated by the simulated attacker's underlying model , and performs a clustering operation to assign pseudo-labels to each sample, where represents the pseudo-label of the i-th sample in.
[0056] In the label reconstructor, is input , and the output is an intermediate high-dimensional feature vector. Its dimensionality is reduced - pseudo-labels are generated by clustering; then, and the pseudo-labels are used to train the fully connected layer, enabling to predict the pseudo-labels of each input sample.
[0057] (2-2) The candidate generation module uses a small-scale validation data subset and the underlying model , and generates multiple candidate reconstructed images. The specific process is as follows:
[0058] (2-2-1) Training dataset preparation: Randomly sample data from the local dataset of the simulated attacker , and use a search engine to collect an image set that is semantically similar to . At the same time, use a search engine to collect an image set that is semantically similar to ; and need to satisfy sample space alignment, that is the images in and the images in are different perspectives of the same object taken at the same time and place, and the sample index is i.
[0059] (2-2-2) Implement the forward diffusion process with a noise manager, and add Gaussian noise to . This process can be modeled as a Markov process, and Gaussian noise is added to iteratively through T rounds:
[0060]
[0061] where q represents the probability distribution in the forward diffusion process, represents the data at the t-th time step, is a hyperparameter of the noise schedule, N represents the Gaussian distribution, represents the identity matrix, represents the total amount of noise from the 1st time step to the t-th time step. Each step can be marginalized as:
[0062]
[0063] where, represents the current noise level, the image without added noise, that is .
[0064] (2-2-3) Design a denoising network to gradually remove the noise, using as the conditional guidance to achieve the restoration of the original input . In a specific implementation of the present invention, class information and sine time step embeddings are spliced in the embedding layer to enhance semantic similarity during the generation process, so as to better capture potential structures; subsequently, residual blocks are stacked to extract multi-scale features; skip connections are constructed to retain spatial information; attention blocks are used at the bottleneck part connecting the downsampling step and the upsampling step to dynamically focus on key features; the denoising network can be designed according to the well-known techniques in the art, and the present invention does not make limitations. Through conventional conditional guidance training, the conditional diffusion model (denoising network) can take a pure noise image as input and generate a predicted image under the condition of Generate a predicted image under the conditional guidance of
[0065] (2-2-4) Use the trained candidate generation module to generate data M times, with the same input each time as the condition, and the generation result of the mth time is expressed as:
[0066]
[0067] Finally, obtain the candidate set as the candidate generation result of
[0068] (2-3) The purpose of the optimal selection module is to determine the reconstruction result from the candidate set .
[0069] For each candidate picture , it has a certain degree of difference from other pictures in the candidate set. Based on Select the optimal candidate picture at the perception and feature point levels, and process the candidate set through coarse-grained and fine-grained filters in sequence. The coarse-grained filter calculates the fidelity score (FS) of each candidate after being spliced with , and then calculates the perceptual similarity (PS) between and . In this embodiment, after discarding the candidate pictures with the FS value ranking in the bottom 50%, the PS value of the remaining pictures is calculated, and the candidate pictures with the PS value ranking in the bottom 50% are further discarded.
[0070] The fine-grained filter randomly samples and the regions at the same positions of each , calculates the matching score (MS) between the two based on the pixel values of the regions, and finally selects the picture with the highest MS value as the reconstruction result . The specific process of screening the optimal data is as follows:
[0071] The calculation process of the Fidelity Score (FS) (2-3-1) is as follows: Use the general open-source model Inception-V3 network to predict the labels of each candidate image to obtain , where represents the probability distribution of all labels. Inception-V3 is trained on a large-scale real-world dataset to effectively capture high-level semantic information and evaluate the authenticity of images; based on this, high-fidelity images will produce a distribution concentrated on specific labels, and the "concentration degree" is quantified by calculating the Shannon entropy. The FS of candidate is shown in the following formula:
[0072]
[0073] where represents the j-th element of, and a smaller FS value indicates higher fidelity.
[0074] (2-3-2) The Perceptual Similarity (PS) uses perceptual hashing to evaluate the perceptual similarity between and . Because perceptual hashing pays more attention to low-frequency information, which can capture overall color changes, smooth regions, contours, and shapes, while ignoring high-frequency details. This method is particularly suitable for evaluating the global perceptual features of images. To calculate the perceptual similarity, first calculate the pHash sequences of the same length from and , that is and , and then calculate the normalized Hamming distance between them. The calculation formula is as follows:
[0075]
[0076] where represents the j-th element of, represents the length of the pHash sequence, represents the "exclusive OR" operation (that is, the two element values are the same as 1, different as 0), and a smaller PS value indicates higher similarity.
[0077] (2-3-3) The Matching Score (MS) is used as a fine-grained filter to evaluate the similarity of pixel values in the regions of randomly sampled same positions and each ; and respectively represent and the pixels in the sampling region, and the element-by-element Evaluate their matching degree based on the distance. The calculation formula is as follows:
[0078]
[0079] where represents the j-th element of and represents the number of pixel points in this area. A smaller MS value indicates a higher similarity.
[0080] Select the picture with the highest MS as the final reconstruction result (2-3-4) .
[0081] Evaluate the reconstruction result (2-4) for the privacy leakage risk indicated . Specifically, calculate the peak signal-to-noise ratio between the reconstruction result and the original picture to judge the degree of privacy leakage:
[0082]
[0083] where MAX is the maximum pixel value and MSE is the mean square error of and pixel by pixel; The larger the value, the greater the privacy leakage risk.
[0084] For step (3), evaluate the privacy leakage risk caused by each ally to client B, that is and , and take the maximum value as the privacy leakage risk between allies faced by client B ; repeat until the privacy protection condition in step (1) is met: and .
[0085] Traverse all clients to implement the privacy leakage risk assessment for each client. According to the risk assessment results, add noise to the gradients sent by the server to this client until the privacy protection condition is met, thereby obtaining a vertically federated learning system with enhanced privacy protection ability against data leakage risk between allies.
[0086] As shown in Table 1, the method proposed in the present invention shows excellent reconstruction quality in all tasks, even when the attacker only has 1 / 3 of the features. The case where the attacker holds 1 / 2 of the features is specifically analyzed. Specifically, the mean square error (MSE) values of the four tasks are all very low, all within 10 -3The order of magnitude, while the peak signal-to-noise ratio (PSNR) values are 32, 22, 28, and 42 respectively. In addition, the structural similarity (SSIM) values are close to 0.9 in the three tasks, and all these metrics indicate excellent reconstruction results.
[0087] In the ImageNet task, the performance of the present invention is particularly prominent, with an extremely low MSE (1.17e-4) and a high PSNR of 41.67. Although the performance in the medical task is relatively weak, it still achieves a satisfactory reconstruction effect, with an MSE of 7.94e-3 and a PSNR of 21.89.
[0088] The reasons for this performance difference are mainly the following three points:
[0089] The ImageNet dataset contains all the categories in Tiny-ImageNet, so the generated quality is higher;
[0090] The specific patterns related to human tissues in medical images will weaken the guiding role of generating semantically meaningful outputs;
[0091] The current design of perceptual-level evaluation metrics does not fully consider such data with special patterns.
[0092] Table 1: Experimental data on the performance of privacy leakage assessment methods
[0093]
[0094] Verify the effectiveness of the label reconstruction module: As shown in Table 2, the label reconstruction module shows high effectiveness in label reconstruction, even achieving 100% accuracy in the traffic task. At the same time, it also shows satisfactory label alignment effects when preprocessing the training dataset of the candidate generation module, thus helping to generate semantically consistent candidate samples. From the performance in the two tasks, as the proportion of features owned by the attacker decreases, the performance of the label reconstruction module does not show significant fluctuations, indicating that the module can maintain strong stability even under limited feature conditions.
[0095] Table 2: Experimental data on the effectiveness of the label reconstruction module
[0096]
[0097] The effectiveness of the candidate generation module: As shown in Table 3, the effectiveness of the candidate generation module is evaluated by measuring its performance on its design goals, which include: the integration of the guiding mechanism, the fidelity of the generated results, and the appropriate diversity among candidate samples.
[0098] To evaluate the guidance and fidelity, a classifier was trained using the VFL dataset, and the prediction accuracy of candidate samples was used as an index of their similarity to the corresponding real data. Therefore, the higher the classification accuracy of candidate samples, the better their fidelity. Specifically, for a binary classification task (medicine), outputs with a prediction probability lower than 0.75 (the average confidence of the real data task) were regarded as 0 because any random input can only achieve an accuracy of 0.5, which is equivalent to random guessing. As shown in the "Guidance and Fidelity" column of Table 3, the effectiveness of the guidance mechanism was evaluated by comparing the classifier prediction accuracies of candidate samples generated with and without guidance. A higher accuracy means better fidelity. The results show that candidate samples generated using the guidance mechanism perform significantly better. For example, the accuracy increased by more than 90% in the CIFAR task. In contrast, without guidance, the generation effect is extremely poor, with an accuracy even lower than 2% in the ImageNet task and close to 0 in the medical task, mainly because there are a very large number of categories in ImageNet. This comparison highlights the important role of the guidance mechanism.
[0099] Furthermore, the candidate samples generated with guidance were compared with the classification accuracy of the real data to evaluate the generation fidelity. As shown in Table 3, the average gap between the candidate samples generated with guidance and the real data does not exceed 6.14%. Using 1 / 2 features in the CIFAR task even exceeded the real data. This fully demonstrates that the candidate generation module has strong generation fidelity.
[0100] To evaluate the diversity among candidate samples, the standard deviation (std.) within the candidate sample set was compared with the standard deviation of the same-class samples in the target dataset. As shown in the "Diversity" column of Table 3, the standard deviation of the candidate sample set is approximately in the to order of magnitude, and the average reduction compared to the standard deviation of the within-class samples is 75.19%, indicating that there is a moderate degree of differentiation among the generated candidate samples.
[0101] Table 3: Experimental data on the effectiveness of the candidate generation module
[0102]
[0103] Effectiveness of the optimal selection module: To evaluate the effectiveness of the optimal selection module, the changing trends of the average PSNR, SSIM, MSE, MAE, and LPIPS of candidate samples after being filtered at each layer were shown, as Figure 4As shown. As the candidate samples are screened through each layer, their average PSNR and SSIM continue to rise, while MSE, MAE, and LPIPS continue to decline. These trends indicate that each layer of the filter can effectively screen out the poorer samples, thus converging to the optimal reconstruction result. Even when the overall quality of all candidate samples has reached a relatively good level (e.g., the average PSNR is about 20), the coarse-grained filter can still bring an average performance improvement of 9.23%, and the fine-grained filter further improves it by an average of 4.89% on this basis. Overall, the application of the optimal selection module has achieved an average performance improvement of 14.62%, proving its important value in screening candidate reconstruction results.
[0104] Based on the same inventive concept, in this embodiment, a vertical federated learning privacy protection system considering data leakage between allies is further provided, including:
[0105] The server side is used to record the initial accuracy of the top-level model on the server side of the vertical federated learning system. The server maintains a data subset sampled from the local data of each client;
[0106] The privacy leakage risk assessment module is used to regard each client as the second client in turn, and the remaining clients as the first clients, calculate the risk of the first clients stealing the data of the second client, and take the maximum value of the risks as the overall privacy leakage risk of the second client;
[0107] When calculating the risk of the first client stealing the data of the second client, a label reconstructor is trained based on the data subset of the first client maintained by the server; obtain the first semantic similarity image set of the data subset of the first client and the second semantic similarity image set of the data subset of the second client, and align the sample spaces of the two semantic similarity image sets; use the label reconstructor to generate the pseudo-labels of the first samples from the first semantic similarity image set, use the first samples and their pseudo-labels as conditions to guide, use the conditional diffusion model to reconstruct the candidate set of the second samples from the second semantic similarity image set, screen the reconstructed samples from the candidate set, and obtain the risk of the first client stealing the data of the second client according to the peak signal-to-noise ratio of the reconstructed samples and the second samples;
[0108] The gradient noise addition module is used to add noise to the updated gradient sent by the server to the second client. If the overall privacy leakage risk exceeds the risk threshold, the updated gradient sent by the server to the second client is added with noise.
[0109] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments, and the implementation methods of the remaining modules will not be elaborated here. The system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. A person of ordinary skill in the art can understand and implement it without creative work.
[0110] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.
[0111] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as a limitation on the scope of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A vertical federated learning privacy protection method considering data leakage among allies, characterized in that, Including: Record the initial accuracy of the top-level model on the server side of the vertical federated learning system. The server maintains a subset of data sampled from the local data of each client. Regarding each client as the second client in turn, and the remaining clients as the first clients, calculate the risk of the first clients stealing the data of the second client. Take the maximum value of the risks as the overall privacy leakage risk of the second client. If the overall privacy leakage risk exceeds the risk threshold, add noise to the updated gradient sent by the server to the second client, and recalculate the overall privacy leakage risk based on the noisy result until the privacy protection condition is met. Traverse all clients. When calculating the risk of the first client stealing the data of the second client, train a label reconstructor based on the data subset of the first client maintained by the server; obtain the first semantic similarity image set of the data subset of the first client and the second semantic similarity image set of the data subset of the second client, and align the sample spaces of the two semantic similarity image sets; use the label reconstructor to generate the pseudo-labels of the first samples from the first semantic similarity image set, and use the first samples and their pseudo-labels as conditional guidance to reconstruct the candidate set of the second samples from the second semantic similarity image set using a conditional diffusion model, screen the reconstructed samples from the candidate set, and obtain the risk of the first client stealing the data of the second client according to the peak signal-to-noise ratio between the reconstructed samples and the second samples.
2. The vertical federated learning privacy protection method considering data leakage among allies according to claim 1, characterized in that In the vertical federated learning system, each client maintains a local bottom-level model. The intermediate features generated by each local bottom-level model are spliced on the server and the target category is predicted by the top-level model. The updated gradients of each local bottom-level model are calculated on the server.
3. The vertical federated learning privacy protection method considering data leakage among allies according to claim 2, wherein In the vertical federated learning system, the local data maintained by the client is object images from different perspectives, and the intermediate features of the object images from different perspectives are fused on the server to predict the object category.
4. The vertical federated learning privacy protection method considering data leakage among allies according to claim 1, characterized in that, Training a label reconstructor based on the data subset of the first client maintained by the server includes: Input the data subset of the first client maintained by the server into the local bottom-level model of the first client, reduce the dimension and cluster the intermediate features generated by the local bottom-level model to generate cluster labels. Train a fully connected layer using the data subset of the first client and its cluster labels, so that the fully connected layer can predict the cluster labels according to the intermediate features generated by the local bottom-level model of the first client, and take the prediction result as the pseudo-label.
5. The vertical federated learning privacy protection method considering data leakage among allies according to claim 1, characterized in that, The process of obtaining the first semantic similarity image set and the second semantic similarity image set includes: Sample from the data subsets of the first and second clients respectively, and use a search engine to collect the image sets that are semantically similar to the sampling results as the first and second semantic similarity image sets. Align the sample spaces of the first and second semantic similarity image sets to ensure that the corresponding images in the two similar image sets are different perspective images of the same object.
6. The vertical federated learning privacy protection method considering data leakage among allies according to claim 1, characterized in that Reconstructing the candidate set of the second samples from the second semantic similarity image set using a conditional diffusion model with the first samples and their pseudo-labels as conditional guidance includes: Perform forward diffusion on the second samples from the second semantically similar image set, gradually adding noise to generate noisy images; then, using the first samples from the first semantically similar image set and their pseudo-labels as conditional inputs, gradually denoise the noisy images under conditional guidance to recover the second samples, and train a conditional diffusion model; the second samples and the first samples are images of the same object from different perspectives. Using a pure noise image as the input, use the trained conditional diffusion model to generate a predicted image under the conditional guidance of the first samples and their pseudo-labels. The predicted image is used as a candidate image for the second sample. Run the conditional diffusion model multiple times to generate a candidate set for the second sample.
7. The vertical federated learning privacy protection method considering data leakage among allies according to claim 1 or 6, characterized in that The process of screening the reconstructed samples from the candidate set includes: Calculate the fidelity score based on the distribution of the predicted labels of each sample in the candidate set and the first sample, and eliminate the candidate samples with lower fidelity scores. Calculate the perceptual similarity based on the normalized Hamming distance between the perceptual hash sequences of each sample in the candidate set and the first sample, and eliminate the candidate samples with lower perceptual similarity. Randomly sample the pixel values of the same position regions for each sample in the candidate set and the first sample, calculate the L2 distance element by element to obtain the matching score, and use the sample with the highest matching score from the candidate set as the reconstructed sample.
8. The vertical federated learning privacy protection method considering data leakage among allies according to claim 1, characterized in that, The process of adding noise to the updated gradients sent by the server to the second client includes: Gradually add small-scale noise to the gradients calculated by the server for updating the local underlying model to be sent to the second client to obtain the noisy gradients. Record the accuracy of the server-side top-level model after adding noise, and calculate the difference from the initial accuracy. If the accuracy difference does not exceed the accuracy fluctuation threshold, retain the result of adding noise and continue to evaluate the overall privacy leakage risk of the second client after adding noise. If the accuracy difference exceeds the accuracy fluctuation threshold, clear the newly added noise, gradually replace the elements with smaller absolute values in the gradients with zeros, and continue to evaluate the overall privacy leakage risk of the second client after adding noise.
9. The vertical federated learning privacy protection method considering data leakage among allies according to claim 8, characterized in that, The privacy protection condition refers to that the accuracy difference of the server-side top-level model does not exceed the accuracy fluctuation threshold, and the overall privacy leakage risk of each client does not exceed the risk threshold.
10. A vertical federated learning privacy protection system considering data leakage among allies, which is used to implement the vertical federated learning privacy protection method described in claim 1, and is characterized in that, The system includes: The server side is used to record the initial accuracy of the server-side top-level model of the vertical federated learning system. The server maintains a data subset sampled from the local data of each client. The privacy leakage risk assessment module is used to regard each client as the second client in turn, and the remaining clients as the first client, calculate the risk of the first clients stealing the data of the second client, and use the maximum value of the risks as the overall privacy leakage risk of the second client. When calculating the risk that the first client steals the data of the second client, a label reconstructor is trained based on the data subset of the first client maintained by the server; the first semantic similarity image set of the first client data subset and the second semantic similarity image set of the second client data subset are obtained, and the sample spaces of the two semantic similarity image sets are aligned; the label reconstructor is used to generate the pseudo-labels of the first samples from the first semantic similarity image set, and with the first samples and their pseudo-labels as conditional guidance, the conditional diffusion model is used to reconstruct the candidate set of the second samples from the second semantic similarity image set, the reconstructed samples are screened from the candidate set, and the risk that the first client steals the data of the second client is obtained according to the peak signal-to-noise ratio of the reconstructed samples and the second samples; The gradient noise addition module is used to add noise to the updated gradient sent by the server to the second client, and if the overall privacy leakage risk exceeds the risk threshold, the updated gradient sent by the server to the second client is added with noise.
Citation Information
Patent Citations
Privacy protection method based on federal learning of alliance chain
CN115952532A
Self-adaptive privacy protection federal learning method
CN116739079A
Method and system for providing differential privacy using federated learning
EP4149134A1
Federated-learning-based personal qualification evaluation method, apparatus and system, and storage medium
WO2022057108A1
Discrete variable preprocessing method in vertical federated learning
WO2024060400A1
Cited By
Power big data privacy protection method and system based on federated learning
CN120822242A
A federated learning-based power big data privacy protection method and system
CN120822242B