Diffusion Model Training and Sampling Method and System Based on Personalized Federated Learning
Through personalized federated learning and time segmentation sampling strategies, the problem of category information leakage in diffusion model training is solved, the generation quality is improved and privacy is protected, and efficient image generation is achieved in the federated learning environment.
Patent Information
- Application Number
- CN202510536167.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing diffusion model training methods have the risk of privacy leakage in federated learning, especially category information leakage problems, and the generation quality is insufficient.
Using the method of personalized federated learning, by locally training the second half of the diffusion model on the client side and introducing a personalized embedding layer, combined with the time segmentation sampling strategy, only images that conform to the local data distribution are generated locally, limiting the denoising ability of the global model and preventing category information leakage.
On the premise of ensuring data security and category-level privacy, the quality of image generation is significantly improved, the risk of privacy leakage is reduced, and higher quality image generation is achieved.
Smart Images

Figure CN120046048B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and particularly relates to a diffusion model training and sampling method and system based on personalized federated learning. Background Art
[0002] In recent years, diffusion models have attracted much attention due to their excellent performance in tasks such as generating high-quality images and audio. However, traditional diffusion model training usually needs to be carried out on a centralized dataset, which inevitably brings the risk of privacy leakage. Especially in sensitive fields such as medical images and financial data, there are significant security hazards in data sharing.
[0003] Federated learning, as a distributed training method, can collaboratively train a model without sharing the original data, thus effectively protecting data privacy. Although federated learning can avoid the leakage of the original data, there is still a risk of privacy leakage in the training of diffusion models. Specifically, the diffusion model may leak the category information of the client through the generated images. For example, the model of a certain client may generate images related to the data categories of other clients, resulting in privacy leakage.
[0004] To address this challenge, personalized federated learning has become a more suitable solution. Different from traditional federated learning, personalized federated learning aims to customize a personalized model for each client to solve the problem of data distribution differences. However, although existing personalized federated learning methods can ensure sample-level privacy protection, for diffusion models, they may still leak privacy, such as the category information of the data. Therefore, how to prevent category-level privacy leakage as much as possible while improving the generation quality has become an urgent challenge to be solved. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies in privacy protection and generation quality of the existing diffusion model training and sampling methods based on federated learning, and provide an innovative method, device and medium that can improve the generation quality of the diffusion model on the premise of ensuring data security and protecting category-level privacy.
[0006] To achieve the purpose of the present invention, the following technical solutions are provided:
[0007] In the first aspect, the present invention provides a diffusion model training and sampling method based on personalized federated learning, which includes the following steps:
[0008] S1: Uniformly set a time step segmentation point globally according to the time segmentation sampling strategy, divide all time steps of the denoising process into a first part and a second part. The time steps in the first part are responsible for gradually denoising the input noisy image to an intermediate state image, and the time steps in the second part are responsible for gradually denoising the intermediate state image to a clear image;
[0009] S2: For each client participating in federated learning, locally initialize the diffusion model and randomly sample time steps from the full time step range using local image data for training. After reaching the termination condition, each client saves the local optimal model parameters for sampling in the first part of the time steps;
[0010] S3: For each client participating in federated learning, re-initialize the diffusion model and set a personalized embedding layer for privacy protection, then randomly sample time steps from the second part and perform local personalized training on the diffusion model again. In each sampled time step, the input image of the diffusion model needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model;
[0011] S4: After each client completes the local personalized training of the diffusion model, keep the personalized embedding layer locally at the client, and only upload the remaining model parameters to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters and sends them back to each client for further optimization. Continuously loop the process of local personalized training and server weighted aggregation until the termination condition is reached, and each client obtains the global optimal model parameters;
[0012] S5: In the sampling stage, each client uses the diffusion model combined with the locally saved local optimal model parameters to gradually execute each time step of the first part, gradually denoise the input noisy image to an intermediate state image, and then use the diffusion model combined with the global optimal model parameters to gradually execute each time step of the second part and superimpose the locally saved personalized embedding layer on the input image of each time step, so as to gradually denoise the intermediate state image to a clear image.
[0013] As a preference of the first aspect above, the time step segmentation point adopts the median value of all time steps in the denoising process.
[0014] Preferably, in the above first aspect, in the step S3, when iteratively training the diffusion model by reusing the local image data, the training method for each round is as follows: first, randomly sample a time step from the second part, then calculate the noise-added image corresponding to the sampled time step through the forward diffusion process of the diffusion model, and then expand the personalized embedding layer to the same dimension as the noise-added image by means of repeated splicing and superimpose the two to obtain a superimposed image, and input the superimposed image into the denoising network to calculate the noise prediction value corresponding to the sampled time step, calculate the mean square error between the noise prediction value and the actual noise and use it as the loss function to perform gradient descent optimization on the denoising network and the personalized embedding layer in the diffusion model.
[0015] Preferably, in the above first aspect, in the step S4, after the server receives the model parameters uploaded by each client, it weights and aggregates the model parameters uploaded by all clients with the proportion of the training data volume used by each client in the total training data volume globally to obtain the updated global model parameters.
[0016] Preferably, the termination condition is to reach the set maximum number of training rounds or the performance index of the diffusion model converges.
[0017] Preferably, the denoising network in the diffusion model adopts U-Net.
[0018] Preferably, the personalized embedding layer adopts a learnable embedding vector with a dimension of 512.
[0019] In a second aspect, the present invention provides a diffusion model training and sampling system based on personalized federated learning, which includes:
[0020] A time step segmentation module, configured to uniformly set a time step segmentation point globally according to the time segmentation sampling strategy, divide all time steps of the denoising process into a first part and a second part, where the time steps in the first part are responsible for gradually denoising the input noise image to an intermediate state image, and the time steps in the second part are responsible for gradually denoising the intermediate state image to a clear image;
[0021] A complete time step client local training module, configured to, for each client participating in the federated learning, initialize the diffusion model locally and use the local image data to randomly sample time steps from the complete time step range for training. After reaching the termination condition, each client saves the local optimal model parameters for sampling in the first part of the sampling stage;
[0022] In the second half of the time step, the client local personalized training module is used to re-initialize the diffusion model for each client participating in federated learning, set the personalized embedding layer for privacy protection, and then randomly sample time steps from the second part and perform local personalized training on the diffusion model using the local image data again. In each sampled time step, the input image of the diffusion model needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model;
[0023] The global optimization module is used to, after each client completes the local personalized training of the diffusion model, retain the personalized embedding layer locally on the client, and only upload the remaining model parameters to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters, and sends them back to each client for further optimization. The process of local personalized training and server weighted aggregation is continuously looped until the termination condition is reached, and the global optimal model parameters are obtained on each client;
[0024] The phased image generation module is used to, in the sampling stage, each client uses the diffusion model combined with the locally saved local optimal model parameters to gradually execute each time step of the first part, gradually denoise the input noise image to an intermediate state image, and then use the diffusion model combined with the global optimal model parameters to gradually execute each time step of the second part and superimpose the locally saved personalized embedding layer on the input image of each time step, so as to gradually denoise the intermediate state image to a clear image.
[0025] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the diffusion model training and sampling method based on personalized federated learning as described in any one of the above first aspect solutions is implemented.
[0026] In a fourth aspect, the present invention provides a computer electronic device, characterized in that it includes a memory and a processor;
[0027] The memory is used to store a computer program;
[0028] The processor is used to, when executing the computer program, implement the diffusion model training and sampling method based on personalized federated learning as described in any one of the above first aspect solutions.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] 1. Improve the generation quality: The present invention adopts the method of personalized federated learning, allowing multiple data owners to jointly train a personalized federated diffusion model on the premise of ensuring data security and protecting class-level privacy. By sharing parameters and updates, they learn the common features and knowledge of other clients, and finally improve the generation quality of images.
[0031] 2. Protect class-level data privacy: With the improvement of privacy protection awareness, many data owners are reluctant to directly share their private data with third parties for training. Traditional federated learning allows the model to be trained on local devices without sharing local data, which can protect instance-level privacy. However, different from discriminative models, the diffusion model trained by federated learning may generate images from the distributions of other clients, which still has potential privacy leakage problems, such as leaking the class information of local data. The present invention effectively prevents other clients from generating images conforming to the local data class through the global model by introducing a personalized embedding layer that is only stored locally and not uploaded to the server side in the architecture of the federated diffusion model, significantly enhancing the privacy protection ability. On this basis, the present invention also proposes a time segmentation strategy, only training a part of the time steps during the joint training process, so as to prevent the global model from obtaining the knowledge of denoising all time steps of local data, further strengthening the privacy protection. In the sampling stage, the present invention adopts a time segmentation sampling strategy, and the missing time steps of the personalized federated model are supplemented by the local model, giving full play to the advantage of the good generation quality of the personalized federated diffusion model on the premise of ensuring data security and protecting class-level privacy, and obtaining better-quality pictures. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flowchart of the steps of a method for training and sampling a diffusion model based on personalized federated learning.
[0033] Figure 2 It is a block diagram of the components of a system for training and sampling a diffusion model based on personalized federated learning.
[0034] Figure 3 It is a schematic diagram of the structure of a computer electronic device. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given in conjunction with the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined correspondingly without conflict.
[0036] Before the specific description, several concepts mentioned in the present invention are defined as follows:
[0037] Federated learning: A distributed machine learning method that allows multiple clients to jointly train a global model without sharing the original data. Each client uses its local data to train the model locally. After several rounds of training, the client uploads the model parameter updates to the central server, and knowledge sharing and model optimization are achieved through parameter aggregation, thereby constructing a global model with better performance.
[0038] Diffusion Models: A generative model based on gradually adding and removing noise. It generates new data similar to the real data distribution by learning and simulating the denoising process. Diffusion Models define a Markov diffusion chain. The Markov diffusion chain gradually adds noise to the image data through the forward diffusion process until the data becomes approximately Gaussian distribution, and then gradually removes the noise through the reverse generation process to generate clear image data. In the present invention, the reverse generation process is called the denoising process. The denoising process consists of a series of time steps, and each time step needs to predict the noise that should be removed at the current time step based on the denoising network.
[0039] Personalized embedding layer: A model component unique to the client, used to protect data category-level privacy. The personalized embedding layer is randomly initialized locally by the client, ensuring that the model only generates images that conform to the local data distribution, while preventing other clients from obtaining the data distribution of the local data.
[0040] Time-sliced sampling strategy: An optimization strategy for the denoising process of diffusion models. The denoising process is divided into two stages: the first part of the time steps are completed by the local model of the client, gradually restoring the random noise to an intermediate state; the second part of the time steps are completed by the personalized federated diffusion model, generating the final clear image in combination with the personalized embedding layer. This strategy gives full play to the advantages of the personalized model, improves the generation quality and balances privacy protection.
[0041] Privacy leakage rate: An indicator used to quantitatively evaluate whether the generated image contains information about the data categories of other clients. The classifier is used to classify the generated image, and the proportion of the generated image classified as the category of other clients is statistically calculated to measure the category-level privacy protection effect.
[0042] Diffusion models have shown remarkable ability to generate high-quality images. However, considering data privacy, security, and accessibility, training generative models on large centralized datasets is challenging. Federated learning is a distributed training paradigm that allows models to be trained on local devices without sharing local data, effectively protecting the privacy of all parties involved. However, different from discriminative models, diffusion models trained through federated learning may generate images of other client distributions, which still poses potential privacy leakage problems, e.g., revealing the class information of client local data. The present invention proposes a training and sampling method for diffusion models based on personalized federated learning, aiming to more strictly protect the privacy of local data, i.e., achieve class-level privacy protection. This method aims to enhance the generation quality while protecting data privacy and preventing class information leakage by introducing a personalized embedding layer, a time-split sampling strategy, and a federated learning framework. First, in the model training stage, each client uses local data and the personalized embedding layer for training to ensure that the generated images conform to the local data distribution, and effectively avoids the server and other clients from generating pictures that conform to the local data classes through the personalized embedding layer that is only stored locally. Each client only trains the latter half of the time steps of the diffusion model denoising process, thus restricting the expansion of the global denoising ability and further ensuring that the server and other clients cannot completely generate pictures that conform to the local data classes. Then, in the model aggregation and update stage, after the clients complete training locally, they upload the parameters of the denoising network (excluding the personalized embedding layer) to the central server. The server performs weighted aggregation on the received model parameters to generate a new global model and distributes it to each client. The client combines the updated global model with the local personalized embedding layer and uses local data to further optimize the model to improve the generation quality. Through the iterative federated learning training process, the clients and the server gradually optimize the model performance until the set training objective or the maximum number of training epochs is reached. In the sampling stage, the present invention proposes a time-split sampling strategy adapted to the federated model that only trains some time steps, dividing the denoising process into two parts. The first half of the denoising is completed by the client local model, which is responsible for gradually restoring the random noise to an intermediate state; the second half of the denoising is completed by the personalized federated diffusion model to generate the final clear image. Through this strategy, the ability of the personalized diffusion model that only trains a part of the time steps is fully utilized, and the generation quality is improved under the premise of meeting privacy protection. The present invention can effectively improve the generation quality of the diffusion model, while ensuring data security and protecting class-level privacy by combining the personalized embedding layer and the time-split sampling strategy under the federated learning framework.
[0043] As Figure 1 shown, in a preferred embodiment of the present invention, a training and sampling method for a diffusion model based on personalized federated learning is provided, and the steps are as follows:
[0044] S1: Set a time-step segmentation point globally according to the time-slicing sampling strategy, and divide all time steps of the denoising process into a first part and a second part. The time steps in the first part are responsible for gradually denoising the input noisy image to an intermediate-state image, and the time steps in the second part are responsible for gradually denoising the intermediate-state image to a clear image.
[0045] It should be noted that the above time-step segmentation point is set for all clients globally. Based on this time-step segmentation point, all time steps of the denoising process are divided into a first part before and a second part after. This time-step segmentation point can be adjusted according to task requirements. For example, the midpoint or other points of the denoising time steps can be selected to balance the generation quality and privacy protection. Moreover, the time-step segmentation points of all clients need to be consistent during the training process and the sampling process. Assuming that the total number of steps in the denoising process is T, the set time-step segmentation point is , the starting time step of the denoising process is T, and the ending time step is 0.
[0046] S2: For each client participating in federated learning, initialize the diffusion model locally and randomly sample time steps from the full time-step range using local image data for training. After reaching the termination condition, each client saves the local optimal model parameters for sampling in the first part of the time steps during the sampling phase.
[0047] Each client of the present invention needs to perform non-federated training on the diffusion model using local image data in the conventional training manner of the diffusion model. During training, the aforementioned time-slicing sampling strategy does not need to be executed, the training is for the full time steps, and there is no need to introduce a personalized embedding layer. The model parameters of the diffusion model preliminarily trained by each client need to be saved as local optimal model parameters for subsequent calls during the sampling phase.
[0048] In an embodiment of the present invention, the denoising network in the diffusion model adopts U-Net.
[0049] In an embodiment of the present invention, the specific implementation method of the above step S2 is as follows:
[0050] Each client performs local training of the diffusion model with full time steps using its local image data. The termination condition for training can be any of the following conditions: reaching the set maximum number of training epochs, or the Frechet Inception Distance (FID) metric of the generated images converges on the validation dataset, that is, the FID does not decrease significantly in several consecutive epochs. After training ends, save the final model parameters of the diffusion model as local optimal model parameters. The local optimal model parameters saved on the th client are denoted as .
[0051] In addition, in addition to performing the training in step S2, each client also needs to re-initialize and perform personalized federated training according to steps S3 and S4. During this training process, a time-slicing sampling strategy needs to be considered. All time steps of the denoising process are divided into a first part and a second part, but only the time steps of the second part are trained using local data and the personalized embedding layer, and the time steps of the first part do not participate in the training.
[0052] S3: For each client participating in federated learning, after re-initializing the diffusion model and setting the personalized embedding layer for privacy protection, randomly sample time steps from the second part using the local image data again and perform local personalized training on the diffusion model. For each sampled time step, the input image of the diffusion model needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model.
[0053] In the embodiment of the present invention, the specific implementation method of the above step S3 is as follows:
[0054] S31. Model initialization: For each client participating in federated learning, each client initializes its diffusion model parameters, including the global parameters of the denoising network, and at the same time, a client-specific personalized embedding layer needs to be added. Among them, the personalized embedding layer is a learnable one-dimensional embedding vector randomly initialized by the client locally, which is used for privacy protection to ensure that the model only generates images that conform to the local data distribution, and at the same time prevents other clients from obtaining the data distribution of the local data. The dimension of the one-dimensional embedding vector can be adjusted according to actual needs. In the embodiment of the present invention, the dimension of the one-dimensional embedding vector is 512.
[0055] S32. Training of the second half of the denoising process: Re-train the diffusion model using the local image data of the client locally. Each client only trains the second part of the time steps of the diffusion model's denoising process (i.e., the process from the intermediate state image to the clear image). During each training, the client randomly samples a time step from the time step range and randomly samples a time step , and according to the forward diffusion process of the diffusion model, calculates the noise image corresponding to the time step through the following formula:
[0056]
[0057] Among them, is the real image data, is the random Gaussian noise, is the predefined single-step noise scheduling parameter. is 's cumulative version, The above-mentioned is the input image corresponding to the time step in the denoising process of a conventional diffusion model.
[0058] S33. Superposition of personalized embedding layers: During training, the personalized embedding layer is used as part of the input and embedded into the denoising network in a superposed manner:
[0059]
[0060] Among them, is the personalized embedding layer of the th client, is the noisy image obtained after superposition.
[0061] It should be noted that the dimension of the personalized embedding layer and may be different. Therefore, if the two dimensions are different, needs to be continuously copied and concatenated until its dimension is the same as the dimension of unrolled into a one-dimensional vector, thus realizing the superposition between the personalized embedding layer and the input image of the time step . The final input to the denoising network is the above-mentioned noisy image , rather than the original . The way the denoising network calculates the internal noise is the same as that of the traditional diffusion model. It predicts the output noise based on the input image and the corresponding time step . Therefore, the predicted noise value output can be expressed as .
[0062] S34. Calculation of the loss function and backpropagation update of model parameters: After the denoising network calculates the predicted noise value, the standard denoising error of the diffusion model is used as the loss function :
[0063]
[0064] Among them, is the predicted noise value output by the model, represents the model parameters, is the actually added random Gaussian noise, is the local image data distribution of the th client.
[0065] The training objective of the diffusion model is to minimize the difference between the predicted noise and the actual noise, that is, to minimize the loss function Moreover, during the training process, the personalized embedding layer and the diffusion model jointly participate in gradient descent optimization, but the personalized embedding layer is only retained locally subsequently and not uploaded to the server.
[0066] It should be noted that: during the training process, each client updates the original model parameters using the conventional gradient descent strategy. All client uploads, server aggregations, and global model distributions are based on the original model parameters without EMA processing. At the same time, to improve the stability of the training process and the generation effect, and following the mature practices in the current diffusion model field, a set of model parameters generated through the exponential moving average (EMA) mechanism is also maintained synchronously during the training process. The EMA mechanism is mainly used to smooth the update process of model parameters, reduce training fluctuations, and thus effectively improve the stability and overall quality of the generated images. In all comparative experiments, to make the evaluation of the generated images more in line with the mature practices in the current diffusion model field and obtain a more stable and smooth generation effect, each model uses the model parameters after EMA processing for sampling to evaluate the image generation quality; while the evaluation of the privacy leakage rate is based on the original model parameters without EMA smoothing processing, so as to more directly reflect the actual performance of each model in terms of privacy protection. This can not only ensure that the advantages of the EMA mechanism in improving model stability and effect are fully utilized in the image generation quality test, but also objectively evaluate the true performance of the model in privacy protection.
[0067] S4: Model aggregation and update: After each client completes the local personalized training of the diffusion model, the personalized embedding layer is retained locally on the client, and only the remaining model parameters are uploaded to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters, and sends them back to each client for further optimization. The process of local personalized training (note that only the local personalized training in step S2 is looped, and the training in step S1 does not need to be looped) and server weighted aggregation is continuously looped until the termination condition is reached, and the globally optimal model parameters are obtained on each client.
[0068] In the embodiment of the present invention, the specific implementation method of the above step S4 is as follows:
[0069] S41. Model parameter upload: After each client locally personalizes the training of the diffusion model to a preset number of rounds, the personalized embedding layer is retained locally on the client, and only the remaining model parameters are uploaded to the server
[0070] S42. Parameter aggregation: After the server receives the model parameters from each client, it uses the proportion of the training data volume used by each client in the total training data volume globally as the weight to perform weighted aggregation on the model parameters uploaded by all clients to obtain the updated global model parameters. The specific weighted aggregation formula is as follows:
[0071]
[0072] Among them, is the aggregated global model parameter, is the model parameter uploaded by the i-th client, and N is the total number of clients participating in federated learning. is the th amount of training data used by the client to train the diffusion model, is the total amount of training data of all clients.
[0073] S43. Global model distribution: The server sends the aggregated global model parameter to each client. After each client receives the global model parameter , it loads it into the denoising network of the diffusion model and combines it with the locally stored personalized embedding layer to form a model for the next round of training.
[0074] The above steps S3 and S4 need to be continuously repeated and iterated. After each client completes the preset number of local iterations (such as every 10,000 iterations), the current model parameters, including the denoising network parameters and the personalized embedding layer, need to be saved for subsequent model generation and analysis. At the same time, the model parameters of the denoising network are uploaded to the server for global weighted aggregation. The termination condition of the final federated training loop can be any of the following conditions: reaching the set maximum number of training rounds, or the Frechet Inception Distance (FID) metric of the generated images converges on the validation dataset, that is, the FID does not decrease significantly in several consecutive rounds.
[0075] Finally, after the federated training ends, the model parameters with the best performance on each client (or the model parameters of the maximum number of training rounds) are saved, which are called the global optimal parameters. The global optimal model parameters on the th client are denoted as . Therefore, the diffusion model based on the global optimal model parameter can be regarded as a federated diffusion model after federated learning. In addition, the personalized embedding layer of each client also retains the finally optimized result, denoted as the personalized embedding vector . This optimized personalized embedding vector needs to be combined into the diffusion model to change the input of the denoising network. Therefore, after introducing the personalized embedding vector into the federated diffusion model on each client, it is equivalent to constructing a personalized federated diffusion model for gradually performing the denoising process in the second part. In addition, each client also saves the local optimal model parameters saved after the complete time step client local training in step S2 , based on the local optimal model parameters The diffusion model can be regarded as a local diffusion model and is used to gradually execute the denoising process in the first part. It should be noted that the model structures of the federated diffusion model and the local diffusion model are the same, and the only difference lies in the model parameters.
[0076] S5: Time-slicing sampling strategy: In the sampling stage, each client uses the diffusion model in combination with the locally saved local optimal model parameters to gradually execute each time step of the first part, gradually denoising the input noisy image to an intermediate state image, and then uses the diffusion model in combination with the global optimal model parameters to gradually execute each time step of the second part and stack the locally saved personalized embedding layer into the input image of each time step, so as to gradually denoise the intermediate state image to a clear image.
[0077] It should be noted that in each time step of the first part above, each client uses the diffusion model in combination with the locally saved local optimal model parameters to gradually execute the denoising process, that is, uses the aforementioned local diffusion model to gradually execute the denoising process. And in each time step of the second part above, each client uses the diffusion model in combination with the global optimal model parameters and stacks the locally saved personalized embedding layer into the input image of each time step, so as to gradually execute the denoising process, that is, uses the aforementioned personalized federated diffusion model to gradually execute the denoising process.
[0078] In the embodiment of the present invention, for any i-th client, the specific implementation method of the above step S5 is as follows:
[0079] S51. Time-slicing setting: In the sampling stage, according to the preset time step segmentation point , the denoising process of the diffusion model is divided into two parts.
[0080] S52. In the first part (time step ): It is completed by the local diffusion model of the client and is responsible for gradually denoising the input noise to an intermediate state.
[0081] In the embodiment of the present invention, in each time step of the first part, the intermediate state can be denoised according to the following formula:
[0082]
[0083] where is the image at time step , is the local optimal model parameter of client . is the noise predicted by the local diffusion model. is a preset hyperparameter for controlling the degree of added noise, is the cumulative version of. represents the added noise term, where is the standard deviation of the noise, is the random noise sampled from the standard normal distribution, which is to maintain an appropriate amount of noise added during the denoising process and ensure the randomness of the generation process.
[0084] S53. In the second part (time step ): It is completed by the personalized federated diffusion model to complete the denoising from the intermediate state to the clear image.
[0085] In the embodiment of the present invention, for each time step in the second part, the image of the intermediate state can be denoised to generate the final clear image according to the following formula: The formula is as follows:
[0086]
[0087] where, represents the global optimal model parameters of the i-th client after the training ends, represents the personalized embedding vector of the i-th client after the training ends, represents the image at the time step, is the noise predicted by the personalized federated diffusion model.
[0088] S54. Generation result output: When the denoising process is completed (time step ), the client outputs the finally generated image .
[0089] It should be particularly emphasized that in the diffusion model training and sampling method based on personalized federated learning provided by the present invention, the design and implementation of each step are of great significance. Among them, the model training step enables each client to use local data and a personalized embedding layer for training, and only optimizes the latter half of the time steps in the denoising process of the diffusion model, realizing privacy protection and personalization of the generated images. The model aggregation and update step enables the client to upload model parameters and be weighted and aggregated by the server, realizing cross-client knowledge sharing on the premise of ensuring data security and protecting class-level privacy, and improving the generation ability of the model. The iterative step ensures the continuous improvement of the model generation ability and meets the user requirements. The time-slicing sampling strategy divides the denoising process into two parts. The first half is completed by the local model, and the second half is controlled by the personalized federated diffusion model, giving play to the ability of the personalized diffusion model that has only been trained for a part of the time steps, and realizing the balance between the improvement of generation quality and privacy protection.
[0090] Based on the diffusion model training and sampling method based on personalized federated learning shown in S1 - S5 in the above embodiments, it is applied to a specific example to demonstrate its effect. The specific process is as described above and will not be elaborated here. Below, the specific parameter settings and implementation effects are mainly shown.
[0091] Embodiment
[0092] Taking the training and sampling process of the diffusion model based on personalized federated learning shown in S1 - S5 in the above embodiments applied to a specific dataset as an example, the present invention is specifically described as follows:
[0093] 1) According to the aforementioned step S1, uniformly set a time step split point globally , where T is the total number of steps in the denoising process.
[0094] 2) According to the aforementioned step S2, divide the data in the dataset by category among multiple clients, and the data of each client has the characteristic of non - independent and identically distributed (Non - IID). Build a diffusion model based on the U - Net denoising network in each client, and use the local image data allocated to itself to complete the local training of the diffusion model. After iterating until the termination condition is met, save the local optimal model parameters to form a local diffusion model.
[0095] 3) According to the aforementioned step S3, each client re - initializes the diffusion model, and at the same time initializes a 512 - dimensional personalized embedding layer, and uses the local data and the personalized embedding layer to perform local personalized training on the diffusion model, only optimizing the latter half of the time steps in the denoising process of the diffusion model . In this process, each client only trains the local image data on its own device, and these image data will not be uploaded or shared, protecting the privacy at the instance level of user data. The personalized embedding layer ensures that the generated images conform to the local data distribution and prevents the generation of images of other client data categories, protecting the privacy at the category level of user data. Only jointly training the latter half of the time steps limits the complete denoising ability of the federated model, further protecting user privacy.
[0096] 4) According to the aforementioned step S4, after a certain number of rounds of training, each client only uploads the model parameter updates of the denoising network to the central server, and the personalized embedding layer always remains on the client side. Subsequently, the central server receives the model parameter updates from each client. Once the parameter collection is completed, the server will initiate the model parameter aggregation process. In the aggregation phase, the central server uses a widely applied federated averaging strategy to perform weighted averaging according to the weight ratio of the data volume contribution of each client, thereby aggregating and generating a unified global model. After the aggregation of the global model parameters is completed, they are distributed to each client. After each client receives the global model parameters, it combines them with the local personalized embedding layer to further optimize the local model. In this way, the model realizes cross-client knowledge sharing while ensuring data security and protecting class-level privacy, and improves the model's generation ability.
[0097] After each aggregation is completed and the global model parameters are distributed to each client, the federated learning training process needs to be repeated. The client and the central server continuously optimize the model performance and gradually improve the model's generation quality and privacy protection effect through the iterative process of training and aggregation. The loop process of local personalized training on the client side and global aggregation on the central server needs to be executed until either of the following stopping conditions A and B is met: A. The generation quality of the model on the validation set (such as the FID metric) converges; B. The training reaches the preset maximum number of iterations.
[0098] 5) According to the aforementioned step S5, in the sampling stage, a time-slicing sampling strategy is adopted, and the denoising process of the diffusion model is divided into two parts: the first half (time step ): The denoising is completed by the local diffusion model of the client, and the input random noise is gradually restored to the intermediate state; the second half (time step ): The denoising is completed by the personalized federated diffusion model to generate the final clear image. The design of this time-slicing sampling strategy effectively gives play to the advantage of good generation quality of the personalized federated diffusion model on the premise of ensuring data security and protecting class-level privacy, and ensures the improvement of the generation quality.
[0099] 6) Privacy evaluation: The privacy evaluation step judges whether the generated image contains the class information of other clients by classifying the generated image with a classifier, and quantitatively evaluates the class-level privacy leakage of various federated diffusion models, so as to verify the privacy protection effect. The specific implementation method is as follows:
[0100] 6.1) Definition of privacy leakage rate: The privacy leakage rate is used to measure the proportion of the class information of other clients' data contained in the generated image. Its formula is:
[0101]
[0102] Where: is the dataset of the th client, and its distribution; is the set of images generated by the model of the client; is the joint data distribution of all clients.
[0103] 6.2) Classification model construction: Use a pre-trained classifier to identify the categories of the generated images. The classifier should have high accuracy to ensure the credibility of the evaluation results. For example: For the CIFAR-10 dataset, the VGG13 classification model is used, and the classification accuracy is 94.2%; for the CIFAR-100 dataset, the SEResNet152 classification model is used, and the classification accuracy is 79.34%; for the Food-101 dataset, the ResNet50 classification model is used, and the classification accuracy is 62.79%.
[0104] 6.3) Classification result screening: Classify the generated images, and only retain the results with a classification confidence higher than 90% for privacy leakage evaluation to reduce the impact of classification errors.
[0105] 6.4) Privacy leakage judgment: If the classification result of the generated image does not belong to the data category of its corresponding client, it is regarded as privacy leakage. Count the number of images whose classification results do not belong to the data category of their corresponding clients, and calculate the privacy leakage rate accordingly.
[0106] 6.5) Output of privacy evaluation results: Output the privacy leakage rates of each client, and compare and analyze them with other methods to evaluate the privacy protection effect.
[0107] In this embodiment, multiple comparison methods are tested, including the local update algorithm (local), the federated learning algorithm with personalized layers (fedper), the federated proximal algorithm (fedprox), the federated graph learning algorithm (pfedgraph), and the federated average algorithm (fedavg). Among them, further taking fedavg as the baseline method, two ablation experiments are set up. The first ablation experiment is the baseline method + time-slicing sampling strategy (τ = T / 2), and the second ablation experiment is the baseline method + personalized embedding layer. Since the method of the present invention is equivalent to the baseline method + time-slicing sampling strategy (τ = T / 2) + personalized embedding layer, the first ablation experiment is equivalent to removing the personalized embedding layer of the present invention, and the second ablation experiment is equivalent to removing the time-slicing sampling strategy of the present invention.
[0108] As Figure 1As shown, it presents the improvement effects of the present invention and each comparative method on the image generation quality (FID(train) and FID(test)) and class-level privacy protection effect on the dataset. For all three metrics, the lower the better.
[0109] Table 1
[0110]
[0111] The experimental results shown in Table 1 indicate that the diffusion model training and sampling method based on personalized federated learning provided by the present invention is significantly superior to the existing methods in terms of both image generation quality and class-level privacy protection. The specific performance is as follows:
[0112] In terms of image generation quality: For the two metrics of FID(train) and FID(test), the results of the method of the present invention are 12.02 and 24.49 respectively, which are better than all other baseline methods (such as FedAvg, FedPer, and pFedGraph, etc.). In contrast, the generation quality of the local training model (local) is poor (FID(train)=14.08, FID(test)=26.48) because it cannot perform cross-client knowledge sharing and the model generalization ability is limited. Therefore, the present invention effectively utilizes the advantages of federated learning and improves the image generation quality by introducing the time-sliced sampling strategy (TSS) and the personalized embedding layer (Client-Condition).
[0113] In terms of class-level privacy protection: In terms of the privacy leakage rate metric, the leakage rate of the method of the present invention is 16.43%, which is much lower than other federated learning methods (such as 80.00% for FedAvg and 78.74% for FedPer). This result shows that the present invention ensures that the generated images strictly conform to the local data distribution of the client through the introduction of the personalized embedding layer, effectively preventing class-level information leakage. On this basis, the time-sliced strategy prevents the global model from obtaining the knowledge of denoising all time steps of the local data, further strengthening the privacy protection.
[0114] In summary, by introducing the personalized embedding layer and the time-sliced sampling strategy, the present invention significantly improves the image generation quality on the premise of ensuring data security and protecting class-level privacy, surpassing the existing federated learning baseline methods, and verifying its effectiveness and practical value in practical applications.
[0115] In another embodiment of the present invention, based on the same inventive concept, a diffusion model training and sampling system based on personalized federated learning is provided, as Figure 2 shown, which includes:
[0116] A time step segmentation module, configured to uniformly set a time step segmentation point globally according to a time segmentation sampling strategy, divide all time steps of the denoising process into a first part and a second part, where the time steps in the first part are responsible for gradually denoising the input noisy image to an intermediate state image, and the time steps in the second part are responsible for gradually denoising the intermediate state image to a clear image;
[0117] A complete time step client local training module, configured to, for each client participating in federated learning, initialize the diffusion model locally and train it using local image data. After reaching the termination condition, each client saves the local optimal model parameters for sampling in the first part of the sampling phase;
[0118] A second half time step client local personalized training module, configured to, for each client participating in federated learning, re-initialize the diffusion model and set a personalized embedding layer for privacy protection, and then randomly sample time steps from the second part and perform local training on the diffusion model using local image data again. In each sampled time step, the input image of the diffusion model needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model;
[0119] A global optimization module, configured to, after each client completes the local training of the diffusion model, retain the personalized embedding layer locally at the client, and only upload the remaining model parameters to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters and sends them back to each client for further optimization. The process of local training and server weighted aggregation is continuously looped until the termination condition is reached, and each client obtains the global optimal model parameters;
[0120] A phased image generation module, configured to, in the sampling phase, each client uses the diffusion model combined with the locally saved local optimal model parameters to gradually execute each time step of the first part, gradually denoise the input noisy image to an intermediate state image, and then use the diffusion model combined with the global optimal model parameters to gradually execute each time step of the second part and superimpose the locally saved personalized embedding layer on the input image of each time step, so as to gradually denoise the intermediate state image to a clear image.
[0121] Each module in the above diffusion model training and sampling system based on personalized federated learning corresponds to S1~S5 of the foregoing embodiments respectively. Therefore, the specific implementation methods can also be referred to the foregoing embodiments, and will not be elaborated here.
[0122] It should be noted that according to the embodiments disclosed in the present invention, the specific implementation functions of various modules in the above diffusion model training and sampling system based on personalized federated learning can be realized by writing computer software programs, and the computer programs contain program codes for executing corresponding methods.
[0123] In another embodiment of the present invention, based on the same inventive concept, a computer-readable storage medium is provided. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the method for training and sampling a diffusion model based on personalized federated learning as described in S1~S5 above is realized.
[0124] In another embodiment of the present invention, based on the same inventive concept, a computer device is provided, as Figure 3 shown, which includes a memory and a processor;
[0125] The memory is used to store a computer program;
[0126] The processor is used to realize the method for training and sampling a diffusion model based on personalized federated learning as described in S1~S5 above when executing the computer program.
[0127] It can be understood that the above storage medium may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0128] The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0129] It should be noted that the computer device may be any physical machine with a GPU, a CPU, and an intelligent network card slot, including personal computers (PCs) and servers.
[0130] The embodiments described above are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A method for training and sampling a diffusion model based on personalized federated learning, characterized in that It includes the following steps: S1: Set a time step segmentation point globally according to the time segmentation sampling strategy, divide all time steps of the denoising process into a first part and a second part. The time steps in the first part are responsible for gradually denoising the input noisy image to an intermediate state image, and the time steps in the second part are responsible for gradually denoising the intermediate state image to a clear image; S2: For each client participating in federated learning, initialize the diffusion model locally and randomly sample time steps from the full time step range using local image data for training. After reaching the termination condition, each client saves the local optimal model parameters for sampling in the first part of the sampling phase; S3: For each client participating in federated learning, after re-initializing the diffusion model and setting a personalized embedding layer for privacy protection, randomly sample time steps from the second part again using local image data and perform local personalized training on the diffusion model. In each sampled time step, the input image of the diffusion model needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model. When re-using local image data to perform iterative training on the diffusion model, the training method for each round is as follows: first randomly sample a time step from the second part, then calculate the noisy image corresponding to the sampled time step through the forward diffusion process of the diffusion model, and then expand the personalized embedding layer to the same dimension as the noisy image by repeated splicing and superimpose the two to obtain a superimposed image. The obtained superimposed image is input into the denoising network to calculate the noise prediction value corresponding to the sampled time step, calculate the mean square error between the noise prediction value and the actual noise and use it as the loss function to perform gradient descent optimization on the denoising network and the personalized embedding layer in the diffusion model; S4: After each client completes the local personalized training of the diffusion model, keep the personalized embedding layer locally on the client, and only upload the remaining model parameters to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters and sends them back to each client for further optimization. Continuously loop the process of local personalized training and server weighted aggregation until the termination condition is reached, and each client obtains the global optimal model parameters; S5: In the sampling phase, each client uses the diffusion model combined with the locally saved local optimal model parameters to gradually execute each time step of the first part, gradually denoise the input noisy image to an intermediate state image, and then use the diffusion model combined with the global optimal model parameters to gradually execute each time step of the second part and superimpose the locally saved personalized embedding layer on the input image of each time step, so as to gradually denoise the intermediate state image to a clear image.
2. The diffusion model training and sampling method based on personalized federated learning according to claim 1, wherein The time step segmentation point adopts the median value of all time steps in the denoising process.
3. The diffusion model training and sampling method based on personalized federated learning according to claim 1, characterized in that, In S4, after the server receives the model parameters uploaded by each client, it uses the proportion of the training data volume used by each client in the total training data volume globally as the weight to perform weighted aggregation on the model parameters uploaded by all clients to obtain the updated global model parameters.
4. The diffusion model training and sampling method based on personalized federated learning according to claim 1, wherein: The termination condition is reaching the set maximum number of training epochs or the performance metrics of the diffusion model converging.
5. The diffusion model training and sampling method based on personalized federated learning according to claim 1, characterized in that: The denoising network in the diffusion model uses U-Net.
6. The diffusion model training and sampling method based on personalized federated learning according to claim 1, characterized in that: The personalized embedding layer uses a learnable embedding vector of 512 dimensions.
7. A diffusion model training and sampling system based on personalized federated learning, characterized in that, It includes: A time step splitting module, which is used to uniformly set a time step splitting point globally according to the time splitting sampling strategy, divide all time steps of the denoising process into a first part and a second part. The time steps in the first part are responsible for gradually denoising the input noisy image to an intermediate state image, and the time steps in the second part are responsible for gradually denoising the intermediate state image to a clear image; A complete time step client local training module, which is used for each client participating in federated learning to locally initialize the diffusion model and randomly sample time steps from the complete time step range using local image data for training. After reaching the termination condition, each client saves the local optimal model parameters for the first part of the time step sampling in the sampling stage; A second half time step client local personalized training module, which is used for each client participating in federated learning. After re-initializing the diffusion model and setting the personalized embedding layer for privacy protection, it then randomly samples time steps from the second part again using local image data and performs local personalized training on the diffusion model. Moreover, in each sampled time step, the input image of the diffusion model needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model. When re-using local image data to perform iterative training on the diffusion model, the training method for each round is: first randomly sample a time step from the second part, then calculate the noisy image corresponding to the sampled time step through the forward diffusion process of the diffusion model, and then expand the personalized embedding layer to the same dimension as the noisy image through repeated splicing and superimpose the two to obtain a superimposed image. The obtained superimposed image is input into the denoising network to calculate the noise prediction value corresponding to the sampled time step, calculate the mean square error between the noise prediction value and the actual noise and use it as the loss function to perform gradient descent optimization on the denoising network and the personalized embedding layer in the diffusion model; A global optimization module, which is used after each client completes local personalized training of the diffusion model. The personalized embedding layer is retained on the client local, and only the remaining model parameters are uploaded to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters and sends them back to each client for further optimization. Continuously cycle the process of local personalized training and server weighted aggregation until the termination condition is reached, and each client obtains the global optimal model parameters; A stage-based image generation module, which is used in the sampling stage. Each client uses a diffusion model in combination with the locally saved local optimal model parameters to gradually execute each time step of the first part, gradually denoise the input noise image to an intermediate state image, and then use the diffusion model in combination with the global optimal model parameters to gradually execute each time step of the second part and superimpose the locally saved personalized embedding layer on the input image of each time step, so as to gradually denoise the intermediate state image to a clear image.
8. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by a processor, the method for training and sampling a diffusion model based on personalized federated learning according to any one of claims 1 to 6 is implemented.
9. A computer electronic device, characterized in that, Comprising a memory and a processor; The memory is used for storing a computer program; The processor is used for implementing the method for training and sampling a diffusion model based on personalized federated learning according to any one of claims 1 to 6 when executing the computer program.
Citation Information
Patent Citations
Low-dose CT imaging method based on context error modulation generalized diffusion model
CN116468817A
Individualized federal potential diffusion model learning method and system
CN117910601A