Diffusion model training and sampling method and system based on personalized federal learning
By introducing personalized embedding layer and time segmentation sampling strategies into the federal diffusion model, the shortcomings of diffusion models in the prior art in terms of privacy protection and generation quality are solved, and higher quality image generation and stricter category-level privacy protection are achieved.
Patent Information
- Application Number
- CN202510536167.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Existing federated learning-based proliferation model training methods have shortcomings in privacy protection and generation quality, especially in preventing category-level privacy breaches.
Using the diffusion model training and sampling method based on personalized federated learning, training and sampling is carried out under the architecture of the federated diffusion model by introducing a personalized embedding layer and time segmentation sampling strategy. The personalized embedding layer is used to protect privacy, ensuring that the generated images conform to local data distribution, while the time segmentation sampling strategy limits the denoising ability of the global model and prevents privacy leakage.
On the premise of ensuring data security and protecting category-level privacy, the quality of image generation is significantly improved, effectively preventing category-level information leakage, which is significantly better than the existing federated learning methods.
Smart Images

Figure CN120046048A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and specifically relates to a method and system for training and sampling diffusion models based on personalized federated learning. Background Art
[0002] In recent years, diffusion models have attracted much attention due to their excellent performance in tasks such as generating high-quality images and audio. However, traditional diffusion model training usually needs to be carried out on a centralized dataset, which inevitably brings the risk of privacy leakage. Especially in sensitive fields such as medical images and financial data, there are significant security risks in data sharing.
[0003] Federated learning, as a distributed training method, can collaboratively train models without sharing the original data, thus effectively protecting data privacy. Although federated learning can avoid the leakage of original data, there is still a risk of privacy leakage in the training of diffusion models. Specifically, the diffusion model may leak the class information of the client through the generated images. For example, the model of a certain client may generate images related to the data classes of other clients, resulting in privacy leakage.
[0004] To address this challenge, personalized federated learning has become a more suitable solution. Different from traditional federated learning, personalized federated learning aims to customize personalized models for each client to solve the problem of data distribution differences. However, although existing personalized federated learning methods can ensure sample-level privacy protection, for diffusion models, there may still be privacy leakage, such as the class information of the data. Therefore, how to prevent class-level privacy leakage as much as possible while improving the generation quality has become an urgent challenge to be solved. Summary of the Invention
[0005] The object of the present invention is to overcome the deficiencies in privacy protection and generation quality of the existing diffusion model training and sampling methods based on federated learning, and provide an innovative method, device and medium that can improve the generation quality of diffusion models on the premise of ensuring data security and protecting class-level privacy.
[0006] To achieve the object of the present invention, the following technical solutions are provided:
[0007] In a first aspect, the present invention provides a method for training and sampling a diffusion model based on personalized federated learning, which includes the following steps:
[0008] S1: Uniformly set a time step segmentation point globally according to the time segmentation sampling strategy, divide all time steps of the denoising process into a first part and a second part. The time steps in the first part are responsible for gradually denoising the input noisy image to an intermediate state image, and the time steps in the second part are responsible for gradually denoising the intermediate state image to a clear image;
[0009] S2: For each client participating in federated learning, initialize the diffusion model locally and randomly sample time steps from the full time step range using local image data for training. After reaching the termination condition, each client saves the local optimal model parameters for sampling in the first part of the time steps;
[0010] S3: For each client participating in federated learning, re-initialize the diffusion model and set a personalized embedding layer for privacy protection, then randomly sample time steps from the second part again using local image data and perform local personalized training on the diffusion model. In each sampled time step, the input image of the diffusion model needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model;
[0011] S4: After each client completes the local personalized training of the diffusion model, retain the personalized embedding layer locally at the client, and only upload the remaining model parameters to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters and sends them back to each client for further optimization. Continuously loop the process of local personalized training and server weighted aggregation until the termination condition is reached, and each client obtains the global optimal model parameters;
[0012] S5: In the sampling stage, each client uses the diffusion model combined with the locally saved local optimal model parameters to gradually execute each time step of the first part, gradually denoise the input noisy image to an intermediate state image, and then use the diffusion model combined with the global optimal model parameters to gradually execute each time step of the second part and superimpose the locally saved personalized embedding layer on the input image of each time step, so as to gradually denoise the intermediate state image to a clear image.
[0013] Preferably, as in the first aspect above, the time step segmentation point adopts the median value of all time steps in the denoising process.
[0014] Preferably, in the above first aspect, in step S3, when iteratively training the diffusion model by reusing local image data, the training method for each round is as follows: randomly sample a time step from the second part, then calculate the noise-added image corresponding to the sampled time step through the forward diffusion process of the diffusion model, and then expand the personalized embedding layer to the same dimension as the noise-added image by repeated splicing and superimpose the two. The superimposed image is input into the denoising network to calculate the noise prediction value corresponding to the sampled time step. Calculate the mean square error between the noise prediction value and the actual noise and use it as the loss function to perform gradient descent optimization on the denoising network and the personalized embedding layer in the diffusion model.
[0015] Preferably, in the above first aspect, in step S4, after the server receives the model parameters uploaded by each client, it weights and aggregates the model parameters uploaded by all clients with the proportion of the training data volume used by each client in the total training data volume globally to obtain the updated global model parameters.
[0016] Preferably, the termination condition is reaching the set maximum number of training rounds or the performance index of the diffusion model converges.
[0017] Preferably, the denoising network in the diffusion model uses U-Net.
[0018] Preferably, the personalized embedding layer uses a learnable embedding vector with a dimension of 512.
[0019] In a second aspect, the present invention provides a diffusion model training and sampling system based on personalized federated learning, which includes:
[0020] A time step segmentation module for uniformly setting a time step segmentation point globally according to the time segmentation sampling strategy, dividing all time steps of the denoising process into a first part and a second part. The time steps in the first part are responsible for gradually denoising the input noise image to an intermediate state image, and the time steps in the second part are responsible for gradually denoising the intermediate state image to a clear image;
[0021] A complete time step client local training module for each client participating in federated learning, initializing the diffusion model locally and randomly sampling time steps from the complete time step range using local image data for training. After reaching the termination condition, each client saves the local optimal model parameters for sampling in the first part of the time steps during the sampling stage;
[0022] In the second half of the time steps, the client local personalized training module is used to, for each client participating in federated learning, re-initialize the diffusion model and set the personalized embedding layer for privacy protection, and then randomly sample time steps from the second part and perform local personalized training on the diffusion model by re-using the local image data. In each sampled time step, the input image of the diffusion model needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model;
[0023] The global optimization module is used to, after each client completes the local personalized training of the diffusion model, retain the personalized embedding layer on the client local, and only upload the remaining model parameters to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters and sends them back to each client for further optimization. The process of local personalized training and server weighted aggregation is continuously looped until the termination condition is reached, and the global optimal model parameters are obtained on each client;
[0024] The stage-based image generation module is used to, in the sampling stage, each client uses the diffusion model combined with the local optimal model parameters saved locally, gradually execute each time step of the first part, gradually denoise the input noise image to an intermediate state image, and then use the diffusion model combined with the global optimal model parameters, gradually execute each time step of the second part and superimpose the locally saved personalized embedding layer on the input image of each time step, so as to gradually denoise the intermediate state image to a clear image.
[0025] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for training and sampling a diffusion model based on personalized federated learning as described in any one of the above first aspect solutions is implemented.
[0026] In a fourth aspect, the present invention provides a computer electronic device, which is characterized by including a memory and a processor;
[0027] The memory is used to store a computer program;
[0028] The processor is used to, when executing the computer program, implement the method for training and sampling a diffusion model based on personalized federated learning as described in any one of the above first aspect solutions.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] 1. Improve the generation quality: The present invention adopts the method of personalized federated learning, allowing multiple data owners to jointly train a personalized federated diffusion model on the premise of ensuring data security and protecting class-level privacy. By sharing parameters and updates, the model learns the common features and knowledge of other clients, ultimately improving the generation quality of images.
[0031] 2. Protect class-level data privacy: With the increasing awareness of privacy protection, many data owners are reluctant to directly share their private data with third parties for training. Traditional federated learning allows the model to be trained on local devices without sharing local data, which can protect instance-level privacy. However, different from discriminative models, the diffusion model trained by federated learning may generate images from the distributions of other clients, which still poses potential privacy leakage problems, such as revealing the class information of local data. The present invention effectively prevents other clients from generating images that match the local data classes through the global model by introducing a personalized embedding layer that is only stored locally and not uploaded to the server side in the architecture of the federated diffusion model, significantly enhancing the privacy protection ability. On this basis, the present invention also proposes a time segmentation strategy, only training a part of the time steps during the joint training process, so as to prevent the global model from obtaining the knowledge of denoising all time steps of local data, further strengthening the privacy protection. In the sampling stage, the present invention adopts a time segmentation sampling strategy, and the missing time steps of the personalized federated model are supplemented by the local model, giving full play to the advantage of the good generation quality of the personalized federated diffusion model on the premise of ensuring data security and protecting class-level privacy, and obtaining better-quality pictures. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flowchart of the steps of a method for training and sampling a diffusion model based on personalized federated learning.
[0033] Figure 2 It is a block diagram of the composition of a system for training and sampling a diffusion model based on personalized federated learning.
[0034] Figure 3 It is a schematic structural diagram of a computer electronic device. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given in conjunction with the accompanying drawings. Many specific details are set forth in the following description to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined correspondingly without conflict.
[0036] Before the specific description, several concepts mentioned in the present invention are defined as follows:
[0037] Federated learning: A distributed machine learning method that allows multiple clients to jointly train a global model without sharing the original data. Each client uses its local data to train the model locally. After several rounds of training, the model parameters are uploaded to the central server for update. Through parameter aggregation, knowledge sharing and model optimization are achieved, thus constructing a global model with better performance.
[0038] Diffusion Models: A generative model based on gradually adding and removing noise. It generates new data similar to the real data distribution by learning and simulating the denoising process. Diffusion Models define a Markov diffusion chain. The Markov diffusion chain gradually adds noise to the image data through the forward diffusion process until the data becomes approximately Gaussian distribution, and then gradually removes the noise through the reverse generation process to generate clear image data. In the present invention, the reverse generation process is called the denoising process. The denoising process consists of a series of time steps, and each time step needs to predict the noise to be removed at the current time step based on the denoising network.
[0039] Personalized embedding layer: A model component unique to the client, which is used to protect the data category-level privacy. The personalized embedding layer is randomly initialized locally by the client, ensuring that the model only generates images that conform to the local data distribution, while preventing other clients from obtaining the data distribution of the local data.
[0040] Time-sliced sampling strategy: An optimization strategy for the denoising process of the diffusion model. The denoising process is divided into two stages: the first part of the time steps are completed by the local model of the client, gradually restoring the random noise to an intermediate state; the second part of the time steps are completed by the personalized federated diffusion model, generating the final clear image in combination with the personalized embedding layer. This strategy gives full play to the advantages of the personalized model, improves the generation quality and balances privacy protection.
[0041] Privacy leakage rate: An index used to quantitatively evaluate whether the generated image contains information about the data categories of other clients. The generated image is classified by a classifier, and the proportion of the generated image classified as the category of other clients is statistically calculated to measure the category-level privacy protection effect.
[0042] Diffusion models have shown remarkable ability to generate high-quality images. However, considering data privacy, security, and accessibility, training generative models on large centralized datasets is challenging. Federated learning is a distributed training paradigm that allows models to be trained on local devices without sharing local data, effectively protecting the privacy of all parties involved. However, different from discriminative models, diffusion models trained through federated learning may generate images of other clients' distributions, which still brings potential privacy leakage problems. For example, it may leak the class information of clients' local data. The present invention proposes a training and sampling method for diffusion models based on personalized federated learning, aiming to more strictly protect the privacy of local data, that is, to achieve class-level privacy protection. This method aims to improve the generation quality while protecting data privacy and preventing class information leakage by introducing a personalized embedding layer, a time-sliced sampling strategy, and a federated learning framework. First, in the model training stage, each client uses local data and the personalized embedding layer for training to ensure that the generated images conform to the local data distribution, and effectively avoids the server and other clients from generating pictures that conform to the local data classes through the personalized embedding layer that is only stored locally. Each client only trains the latter half of the time steps of the diffusion model denoising process, thus restricting the expansion of the global denoising ability and further ensuring that the server and other clients cannot completely generate pictures that conform to the local data classes. Then, in the model aggregation and update stage, after the clients complete training locally, they upload the parameters of the denoising network (excluding the personalized embedding layer) to the central server. The server performs weighted aggregation on the received model parameters to generate a new global model and distributes it to each client. The client combines the updated global model with the local personalized embedding layer and uses local data to further optimize the model to improve the generation quality. Through the iterative federated learning training process, the clients and the server gradually optimize the model performance until the set training goal or the maximum number of training rounds is reached. In the sampling stage, the present invention proposes a time-sliced sampling strategy adapted to the federated model that only trains some time steps, dividing the denoising process into two parts. The first half of the denoising is completed by the client's local model, which is responsible for gradually restoring the random noise to an intermediate state; the second half of the denoising is completed by the personalized federated diffusion model to generate the final clear image. Through this strategy, the ability of the personalized diffusion model that only trains a part of the time steps is fully utilized, and the generation quality is improved under the premise of meeting privacy protection. The present invention can effectively improve the generation quality of the diffusion model, while ensuring data security and protecting class-level privacy by combining the personalized embedding layer and the time-sliced sampling strategy under the federated learning framework.
[0043] As Figure 1 shown, in a preferred embodiment of the present invention, a training and sampling method for a diffusion model based on personalized federated learning is provided, and the steps are as follows:
[0044] S1: Set a time-step segmentation point globally according to the time-slicing sampling strategy, dividing all time steps of the denoising process into a first part and a second part. The time steps in the first part are responsible for gradually denoising the input noisy image to an intermediate state image, and the time steps in the second part are responsible for gradually denoising the intermediate state image to a clear image.
[0045] It should be noted that the above time-step segmentation point is set for all clients globally. Based on this time-step segmentation point, all time steps of the denoising process are divided into a first part that comes first and a second part that comes later. This time-step segmentation point can be adjusted according to task requirements. For example, the midpoint or other points of the denoising time steps can be selected to balance the generation quality and privacy protection. Moreover, the time-step segmentation points of all clients need to be consistent during the training process and the sampling process. Assuming that the total number of steps in the denoising process is T, the set time-step segmentation point is , the starting time step of the denoising process is T, and the ending time step is 0.
[0046] S2: For each client participating in federated learning, initialize the diffusion model locally and randomly sample time steps from the full time-step range using local image data for training. After reaching the termination condition, each client saves the local optimal model parameters for sampling in the first part of the time steps during the sampling phase.
[0047] Each client of the present invention needs to perform non-federated training on the diffusion model using local image data in the conventional training manner of the diffusion model. During training, the aforementioned time-slicing sampling strategy does not need to be executed, the training is for the full time steps, and there is no need to introduce a personalized embedding layer. The model parameters of the diffusion model preliminarily trained by each client need to be saved as local optimal model parameters for subsequent calls during the sampling phase.
[0048] In an embodiment of the present invention, the denoising network in the diffusion model adopts U-Net.
[0049] In an embodiment of the present invention, the specific implementation method of the above step S2 is as follows:
[0050] Each client uses its local image data to perform local training on the diffusion model for the full time steps. The termination condition for training can be any of the following conditions: reaching the set maximum number of training epochs, or the Frechet Inception Distance (FID) metric of the generated images converges on the validation dataset, that is, the FID does not decrease significantly for several consecutive epochs. After training ends, save the final model parameters of the diffusion model as local optimal model parameters. The local optimal model parameters saved on the th client are denoted as .
[0051] In addition, in addition to performing the training in step S2, each client also needs to be re-initialized and perform personalized federated training according to steps S3 and S4. During this training process, a time-sliced sampling strategy needs to be considered. All time steps of the denoising process are divided into a first part and a second part, but only the time steps of the second part are trained using local data and the personalized embedding layer, and the time steps of the first part do not participate in the training.
[0052] S3: For each client participating in federated learning, after re-initializing the diffusion model and setting the personalized embedding layer for privacy protection, randomly sample time steps from the second part using the local image data again and perform local personalized training on the diffusion model. For each sampled time step, the input image of the diffusion model needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model.
[0053] In the embodiment of the present invention, the specific implementation method of the above step S3 is as follows:
[0054] S31. Model initialization: For each client participating in federated learning, each client initializes its diffusion model parameters, including the global parameters of the denoising network. At the same time, a client-specific personalized embedding layer needs to be added. Among them, the personalized embedding layer is a learnable one-dimensional embedding vector randomly initialized by the client locally, which is used for privacy protection to ensure that the model only generates images that conform to the local data distribution, and at the same time prevent other clients from obtaining the data distribution of the local data. The dimension of the one-dimensional embedding vector can be adjusted according to actual needs. In the embodiment of the present invention, the dimension of the one-dimensional embedding vector is 512.
[0055] S32. Training of the second half of the denoising process: Re-use the local image data of the client to train the diffusion model. Each client only trains the second part of the time steps of the diffusion model's denoising process (that is, the process from the intermediate state image to the clear image). During each training, the client randomly samples a time step from the time step range and randomly samples a time step , and according to the forward diffusion process of the diffusion model, calculates the noise image corresponding to the time step through the following formula:
[0056]
[0057] Among them, is the real image data, is the random Gaussian noise, is the predefined single-step noise scheduling parameter. is 's cumulative version, The above-mentioned is the input image corresponding to the time step in the denoising process of a conventional diffusion model.
[0058] S33. Superposition of personalized embedding layers: During training, the personalized embedding layer is used as part of the input and embedded into the denoising network in a superposed manner:
[0059]
[0060] Among them, is the personalized embedding layer of the th client, and is the noisy image obtained after superposition.
[0061] It should be noted that the dimension of the personalized embedding layer and may be different. Therefore, if the dimensions are different, needs to be continuously copied and concatenated until its dimension is the same as that of unrolled into a one-dimensional vector, thereby realizing the superposition between the personalized embedding layer and the input image of the time step . The final input to the denoising network is the above-mentioned noisy image , rather than the original . The way the denoising network calculates the internal noise is the same as that of the traditional diffusion model. It predicts the output noise based on the input image and the corresponding time step . Therefore, the output noise prediction value can be expressed as .
[0062] S34. Calculation of the loss function and backpropagation update of model parameters: After the denoising network calculates the noise prediction value, the standard denoising error of the diffusion model is used as the loss function :
[0063]
[0064] Among them, is the noise prediction value output by the model, represents the model parameters, is the actually added random Gaussian noise, and is the local image data distribution of the th client.
[0065] The training objective of the diffusion model is to minimize the difference between the predicted noise and the actual noise, that is, to minimize the loss function Moreover, during the training process, the personalized embedding layer and the diffusion model jointly participate in gradient descent optimization, but the personalized embedding layer is only retained locally afterwards and not uploaded to the server.
[0066] It should be noted that: during the training process, each client updates the original model parameters using the conventional gradient descent strategy. All client uploads, server aggregations, and global model distributions are based on the original model parameters without EMA processing. At the same time, to improve the stability of the training process and the generation effect, and following the established practices in the current diffusion model field, a set of model parameters generated through the Exponential Moving Average (EMA) mechanism is also maintained synchronously during the training process. The EMA mechanism is mainly used to smooth the update process of the model parameters, reduce training fluctuations, and thus effectively improve the stability and overall quality of the generated images. In all comparative experiments, to make the evaluation of the generated images more in line with the established practices in the current diffusion model field and obtain a more stable and smooth generation effect, each model uses the model parameters after EMA processing for sampling to evaluate the image generation quality; while the evaluation of the privacy leakage rate is based on the original model parameters without EMA smoothing processing, in order to more directly reflect the actual performance of each model in terms of privacy protection. This can not only ensure that the advantages of the EMA mechanism in improving model stability and effect are fully utilized in the image generation quality test, but also objectively evaluate the true performance of the model in privacy protection.
[0067] S4: Model aggregation and update: After each client completes the local personalized training of the diffusion model, the personalized embedding layer is retained locally on the client, and only the remaining model parameters are uploaded to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters, and sends them back to each client for further optimization. The process of local personalized training (note that only the local personalized training in step S2 is looped, and the training in step S1 does not need to be looped) and server weighted aggregation is continuously cycled until the termination condition is reached, and the globally optimal model parameters are obtained on each client.
[0068] In the embodiments of the present invention, the specific implementation method of the above step S4 is as follows:
[0069] S41. Model parameter upload: After each client locally personalizes the training of the diffusion model to a preset number of rounds, the personalized embedding layer is retained locally on the client, and only the remaining model parameters are uploaded to the server
[0070] S42. Parameter aggregation: After the server receives the model parameters from each client, it uses the proportion of the training data volume used by each client in the total training data volume globally as the weight to perform weighted aggregation on the model parameters uploaded by all clients to obtain the updated global model parameters. The specific weighted aggregation formula is as follows:
[0071]
[0072] Among them, is the aggregated global model parameter, is the model parameter uploaded by the i-th client, and N is the total number of clients participating in federated learning. is the th amount of training data used by the client to train the diffusion model, is the total amount of training data of all clients.
[0073] S43. Global model distribution: The server sends the aggregated global model parameter to each client. After each client receives the global model parameter , it loads it into the denoising network of the diffusion model and combines it with the locally stored personalized embedding layer to form a model for the next round of training.
[0074] The above steps S3 and S4 need to be continuously repeated and iterated. After each client completes the preset number of local iterations (such as every 10,000 iterations), the current model parameters, including the denoising network parameters and the personalized embedding layer, need to be saved for subsequent model generation and analysis. At the same time, the model parameters of the denoising network are uploaded to the server for global weighted aggregation. The termination condition of the final federated training loop can be any of the following conditions: reaching the set maximum number of training rounds, or the Frechet Inception Distance (FID) metric of the generated images converges on the validation dataset, that is, the FID does not decrease significantly for several consecutive rounds.
[0075] Finally, after the federated training ends, the model parameters with the best performance on each client (or the model parameters of the maximum number of training rounds) are saved, which are called global optimal parameters. The global optimal model parameters on the th client are denoted as . Therefore, the diffusion model based on the global optimal model parameter can be regarded as a federated diffusion model after federated learning. In addition, the personalized embedding layer of each client also retains the finally optimized result, denoted as the personalized embedding vector . This optimized personalized embedding vector needs to be combined into the diffusion model to change the input of the denoising network. Therefore, after introducing the personalized embedding vector into the federated diffusion model on each client, it is equivalent to constructing a personalized federated diffusion model for gradually performing the denoising process in the second part. In addition, each client also saves the local optimal model parameters saved after the complete time step client local training in step S2. , based on the local optimal model parameters The diffusion model can be regarded as a local diffusion model, which is used to gradually execute the denoising process in the first part. It should be noted that the model structures of the federated diffusion model and the local diffusion model are the same, and the only difference lies in the model parameters.
[0076] S5: Time-slicing sampling strategy: In the sampling stage, each client uses the diffusion model combined with the locally saved local optimal model parameters to gradually execute each time step of the first part, gradually denoising the input noisy image to an intermediate state image, and then uses the diffusion model combined with the global optimal model parameters to gradually execute each time step of the second part and superimpose the locally saved personalized embedding layer on the input image of each time step, so as to gradually denoise the intermediate state image to a clear image.
[0077] It should be noted that in each time step of the above first part, each client uses the diffusion model combined with the locally saved local optimal model parameters to gradually execute the denoising process, that is, uses the aforementioned local diffusion model to gradually execute the denoising process. And in each time step of the above second part, each client uses the diffusion model combined with the global optimal model parameters and superimposes the locally saved personalized embedding layer on the input image of each time step, so as to gradually execute the denoising process, that is, uses the aforementioned personalized federated diffusion model to gradually execute the denoising process.
[0078] In the embodiment of the present invention, for any i-th client, the specific implementation method of the above step S5 is as follows:
[0079] S51. Time-slicing setting: In the sampling stage, according to the pre-set time step segmentation point , the denoising process of the diffusion model is divided into two parts.
[0080] S52. In the first part (time step ): It is completed by the local diffusion model of the client, which is responsible for gradually denoising the input noise to an intermediate state.
[0081] In the embodiment of the present invention, in each time step of the first part, the intermediate state can be denoised according to the following formula:
[0082]
[0083] where is the image at time step , is the local optimal model parameter of client . is the noise predicted by the local diffusion model. is a preset hyperparameter for controlling the degree of added noise, is the cumulative version of. represents the added noise term, where is the standard deviation of the noise, is the random noise sampled from the standard normal distribution, which is to maintain an appropriate amount of noise added during the denoising process and ensure the randomness of the generation process.
[0084] S53. In the second part (time step ): It is completed by the personalized federated diffusion model to complete the denoising from the intermediate state to the clear image.
[0085] In the embodiment of the present invention, for each time step in the second part, the image of the intermediate state can be denoised to generate the final clear image according to the following formula: The formula is as follows:
[0086]
[0087] where, represents the global optimal model parameters of the i-th client after the training ends, represents the personalized embedding vector of the i-th client after the training ends, represents the image at the time step, is the noise predicted by the personalized federated diffusion model.
[0088] S54. Generation result output: When the denoising process is completed (time step ), the client outputs the finally generated image .
[0089] It should be particularly emphasized that in the diffusion model training and sampling method based on personalized federated learning provided by the present invention, the design and implementation of each step are of great significance. Among them, the model training step enables each client to use local data and a personalized embedding layer for training, and only optimizes the latter half of the time steps in the denoising process of the diffusion model, achieving privacy protection and personalization of the generated images. The model aggregation and update step enables the client to upload model parameters and the server to perform weighted aggregation, realizing cross-client knowledge sharing on the premise of ensuring data security and protecting category-level privacy, and improving the generation ability of the model. The iterative step ensures the continuous improvement of the model generation ability and meets the user requirements. The time-sliced sampling strategy divides the denoising process into two parts. The first half is completed by the local model, and the second half is controlled by the personalized federated diffusion model, giving play to the ability of the personalized diffusion model that has only been trained for a part of the time steps, and achieving a balance between the improvement of generation quality and privacy protection.
[0090] Based on the diffusion model training and sampling method based on personalized federated learning shown in S1-S5 in the above embodiments, it is applied to a specific example to demonstrate its effect. The specific process is as described above and will not be elaborated. The following mainly shows its specific parameter settings and implementation effects.
[0091] Embodiment
[0092] Taking the training and sampling process of the diffusion model based on personalized federated learning shown in S1-S5 in the above embodiments applied to a specific dataset as an example, the present invention will be specifically described. The specific steps are as follows:
[0093] 1) According to the aforementioned step S1, a time step split point is uniformly set globally , where T is the total number of steps in the denoising process.
[0094] 2) According to the aforementioned step S2, the data in the dataset is divided into multiple clients according to categories. The data in each client has the characteristic of non-independent and identically distributed (Non-IID). A diffusion model based on the U-Net denoising network is constructed in each client, and the local training of the diffusion model is completed using the local image data allocated to itself. After iterating until the termination condition is met, the local optimal model parameters are saved to form a local diffusion model.
[0095] 3) According to the aforementioned step S3, each client re-initializes the diffusion model, and at the same time initializes a 512-dimensional personalized embedding layer. The local data and the personalized embedding layer are used to perform local personalized training on the diffusion model, and only the latter half of the time steps in the denoising process of the diffusion model are optimized . In this process, each client only trains the local image data on its own device, and these image data will not be uploaded or shared, protecting the privacy of user data at the instance level. The personalized embedding layer ensures that the generated images conform to the local data distribution and prevents the generation of images of other client data categories, protecting the privacy of user data at the category level. Only jointly training the latter half of the time steps limits the complete denoising ability of the federated model and further protects user privacy.
[0096] 4) According to the aforementioned step S4, after a certain number of rounds of training, each client only uploads the model parameter updates of the denoising network to the central server, and the personalized embedding layer always remains on the client side locally. Subsequently, the central server receives the model parameter updates from each client. Once the parameter collection is completed, the server will initiate the model parameter aggregation process. In the aggregation stage, the central server uses a widely applied federated averaging strategy to perform weighted averaging according to the weight ratio of the data volume contribution of the clients, thereby aggregating and generating a unified global model. After the aggregation of the global model parameters is completed, they are distributed to each client. After each client receives the global model parameters, it combines them with the local personalized embedding layer to further optimize the local model. In this way, the model realizes cross-client knowledge sharing while ensuring data security and protecting class-level privacy, and improves the generation ability of the model.
[0097] After each aggregation is completed and the global model parameters are distributed to each client, the federated learning training process needs to be repeated. The client and the central server continuously optimize the model performance and gradually improve the generation quality and privacy protection effect through the repeated iterative process of training and aggregation. The loop process of local personalized training on the client side and global aggregation on the central server needs to be executed until either of the following stopping conditions A and B is met: A. The generation quality (such as the FID metric) of the model on the validation set converges; B. The training reaches the preset maximum number of iteration rounds.
[0098] 5) According to the aforementioned step S5, in the sampling stage, a time-slicing sampling strategy is adopted, and the denoising process of the diffusion model is divided into two parts: the first half (time step ): The denoising is completed by the local diffusion model of the client, and the input random noise is gradually restored to the intermediate state; the second half (time step ): The denoising is completed by the personalized federated diffusion model to generate the final clear image. The design of this time-slicing sampling strategy effectively gives play to the advantage of good generation quality of the personalized federated diffusion model on the premise of ensuring data security and protecting class-level privacy, and ensures the improvement of the generation quality.
[0099] 6) Privacy evaluation: The privacy evaluation step determines whether the generated image contains the class information of other clients by classifying the generated image with a classifier, and quantitatively evaluates the class-level privacy leakage of various federated diffusion models, so as to verify the privacy protection effect. The specific implementation method is as follows:
[0100] 6.1) Definition of privacy leakage rate: The privacy leakage rate is used to measure the proportion of the class information of other clients' data contained in the generated image. Its formula is:
[0101]
[0102] where: is the dataset of the nth client, and its distribution; is the set of images generated by the model of the client; is the joint data distribution of all clients.
[0103] 6.2) Classification model construction: Use a pre-trained classifier to identify the categories of the generated images. The classifier should have high accuracy to ensure the credibility of the evaluation results. For example: For the CIFAR-10 dataset, the VGG13 classification model is used, and the classification accuracy is 94.2%; for the CIFAR-100 dataset, the SEResNet152 classification model is used, and the classification accuracy is 79.34%; for the Food-101 dataset, the ResNet50 classification model is used, and the classification accuracy is 62.79%.
[0104] 6.3) Classification result screening: Classify the generated images, and only retain the results with a classification confidence higher than 90% for privacy leakage evaluation to reduce the impact of classification errors.
[0105] 6.4) Privacy leakage judgment: If the classification result of the generated image does not belong to the data category of its corresponding client, it is regarded as privacy leakage. Count the number of images whose classification results do not belong to the data category of their corresponding clients, and calculate the privacy leakage rate accordingly.
[0106] 6.5) Output of privacy evaluation results: Output the privacy leakage rates of each client, and compare and analyze them with other methods to evaluate the privacy protection effect.
[0107] In this embodiment, multiple comparison methods are tested, including the local update algorithm (local), the federated learning algorithm with personalized layers (fedper), the federated proximal algorithm (fedprox), the federated graph learning algorithm (pfedgraph), and the federated average algorithm (fedavg). Among them, further taking fedavg as the baseline method, two ablation experiments are set up. The first ablation experiment is the baseline method + time-sliced sampling strategy (τ = T / 2), and the second ablation experiment is the baseline method + personalized embedding layer. Since the method of the present invention is equivalent to the baseline method + time-sliced sampling strategy (τ = T / 2) + personalized embedding layer, the first ablation experiment is equivalent to removing the personalized embedding layer of the present invention, and the second ablation experiment is equivalent to removing the time-sliced sampling strategy of the present invention.
[0108] Such as Figure 1As shown, it demonstrates the improvement effects of the present invention and each comparative method on the image generation quality (FID(train) and FID(test)) and class-level privacy protection effect on the dataset. For all three metrics, the lower the better.
[0109] Table 1
[0110]
[0111] The experimental results shown in Table 1 indicate that the diffusion model training and sampling method based on personalized federated learning provided by the present invention is significantly superior to the existing methods in terms of image generation quality and class-level privacy protection. The specific performance is as follows:
[0112] In terms of image generation quality: For the two metrics of FID(train) and FID(test), the results of the method of the present invention are 12.02 and 24.49 respectively, which are better than all other baseline methods (such as FedAvg, FedPer, and pFedGraph, etc.). In contrast, the generation quality of the local training model (local) is poor (FID(train)=14.08, FID(test)=26.48) because it cannot perform cross-client knowledge sharing and the model generalization ability is limited. Therefore, the present invention effectively utilizes the advantages of federated learning and improves the image generation quality by introducing the time-sliced sampling strategy (TSS) and the personalized embedding layer (Client-Condition).
[0113] In terms of class-level privacy protection: In terms of the privacy leakage rate metric, the leakage rate of the method of the present invention is 16.43%, which is significantly lower than other federated learning methods (such as 80.00% for FedAvg and 78.74% for FedPer). This result indicates that the present invention ensures that the generated images strictly conform to the local data distribution of the client through the introduction of the personalized embedding layer, effectively preventing class-level information leakage. On this basis, the time-sliced strategy prevents the global model from obtaining the knowledge of denoising all time steps of the local data, further strengthening the privacy protection.
[0114] In summary, the present invention significantly improves the image generation quality by introducing the personalized embedding layer and the time-sliced sampling strategy, while ensuring data security and protecting class-level privacy, surpassing the existing federated learning baseline methods, and verifying its effectiveness and practical value in practical applications.
[0115] In another embodiment of the present invention, based on the same inventive concept, a diffusion model training and sampling system based on personalized federated learning is provided, as Figure 2 shown, which includes:
[0116] A time step segmentation module, which is used to uniformly set a time step segmentation point globally according to the time segmentation sampling strategy, divide all time steps of the denoising process into a first part and a second part. The time steps in the first part are responsible for gradually denoising the input noisy image to an intermediate state image, and the time steps in the second part are responsible for gradually denoising the intermediate state image to a clear image;
[0117] A complete time step client local training module, which is used for each client participating in federated learning to initialize the diffusion model locally and train it using local image data. After reaching the termination condition, each client saves the local optimal model parameters for sampling in the first part of the time steps;
[0118] A second half time step client local personalized training module, which is used for each client participating in federated learning to re-initialize the diffusion model and set a personalized embedding layer for privacy protection, and then randomly sample time steps from the second part and perform local training on the diffusion model using local image data again. Moreover, the input image of the diffusion model in each sampled time step needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model;
[0119] A global optimization module, which is used after each client completes the local training of the diffusion model. The personalized embedding layer is retained locally on the client, and only the remaining model parameters are uploaded to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters and sends them back to each client for further optimization. The process of local training and server weighted aggregation is continuously looped until the termination condition is reached, and each client obtains the global optimal model parameters;
[0120] A phased image generation module, which is used in the sampling stage. Each client uses the diffusion model combined with the locally saved local optimal model parameters to gradually execute each time step of the first part, gradually denoise the input noisy image to an intermediate state image, and then use the diffusion model combined with the global optimal model parameters to gradually execute each time step of the second part and superimpose the locally saved personalized embedding layer on the input image of each time step, so as to gradually denoise the intermediate state image to a clear image.
[0121] Each module in the above diffusion model training and sampling system based on personalized federated learning corresponds to S1~S5 of the foregoing embodiments respectively. Therefore, the specific implementation methods can also be referred to the foregoing embodiments, which will not be elaborated here.
[0122] It should be noted that according to the embodiments disclosed in the present invention, the specific implementation functions of various modules in the above diffusion model training and sampling system based on personalized federated learning can be realized by writing computer software programs, and the computer programs contain program codes for executing corresponding methods.
[0123] In another embodiment of the present invention, based on the same inventive concept, a computer-readable storage medium is provided. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the method for training and sampling a diffusion model based on personalized federated learning as described in S1~S5 above is realized.
[0124] In another embodiment of the present invention, based on the same inventive concept, a computer device is provided, as Figure 3 shown, which includes a memory and a processor;
[0125] The memory is used to store a computer program;
[0126] The processor is used to realize the method for training and sampling a diffusion model based on personalized federated learning as described in S1~S5 above when executing the computer program.
[0127] It can be understood that the above storage medium may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0128] The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0129] It should be noted that the computer device may be any physical machine with a GPU, a CPU, and an intelligent network card slot, including personal computers (PCs) and servers.
[0130] The embodiments described above are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A diffusion model training and sampling method based on personalized federated learning, characterized by The steps include: S1: According to the time segmentation sampling strategy, a time step segmentation point is uniformly set globally, and all time steps of the denoising process are divided into the first part and the second part. The time steps of the first part are responsible for gradually denoising the input noisy image to the intermediate state image, and the time steps of the second part are responsible for gradually denoising the intermediate state image to a clear image; S2: For each client participating in federated learning, the diffusion model is initialized locally and the local image data is used to randomly sample time steps from the full time step range for training. After the termination condition is reached, each client saves the local optimal model parameters for the first part of the time step sampling in the sampling phase; S3: For each client participating in federated learning, after reinitializing the diffusion model and setting a personalized embedding layer for privacy protection, the local image data is reused to randomly sample time steps from the second part and perform local personalized training on the diffusion model. The input image of the diffusion model in each sampling time step needs to be superimposed with the local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model. S4: After each client completes the local personalized training of the diffusion model, the personalized embedding layer is retained locally on the client, and only the remaining model parameters are uploaded to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters and sends them back to each client for further optimization. The process of local personalized training and server weighted aggregation is continuously repeated until the termination condition is reached, and the global optimal model parameters are obtained on each client. S5: In the sampling stage, each client uses the diffusion model combined with the locally saved local optimal model parameters to gradually execute each time step of the first part, and gradually denoises the input noisy image to an intermediate state image. Then, the diffusion model is combined with the global optimal model parameters to gradually execute each time step of the second part and superimpose the locally saved personalized embedding layer on the input image of each time step, thereby gradually denoising the intermediate state image to a clear image.
2. The diffusion model training and sampling method based on personalized federated learning according to claim 1 is characterized in that: The time step division point adopts the middle value of all time steps in the denoising process.
3. The diffusion model training and sampling method based on personalized federated learning according to claim 1 is characterized in that: In S3, when the local image data is reused to iteratively train the diffusion model, the training method for each round is: first randomly sample a time step from the second part, then calculate the noisy image corresponding to the sampled time step through the forward diffusion process of the diffusion model, and then expand the personalized embedding layer to the same dimension as the noisy image by repeated splicing and superimpose the two. The obtained superimposed image is input into the denoising network to calculate the noise prediction value corresponding to the sampling time step, and the mean square error between the noise prediction value and the actual noise is calculated and used as the loss function to perform gradient descent optimization on the denoising network and the personalized embedding layer in the diffusion model.
4. The diffusion model training and sampling method based on personalized federated learning according to claim 1 is characterized in that: In S4, after receiving the model parameters uploaded by each client, the server performs weighted aggregation on the model parameters uploaded by all clients, taking the proportion of the training data used by each client in the total training data in the global scope as the weight, to obtain updated global model parameters.
5. The diffusion model training and sampling method based on personalized federated learning according to claim 1 is characterized in that: The termination condition is that the set maximum number of training rounds is reached or the performance index of the diffusion model converges.
6. The diffusion model training and sampling method based on personalized federated learning according to claim 1 is characterized in that: The denoising network in the diffusion model adopts U-Net.
7. The diffusion model training and sampling method based on personalized federated learning according to claim 1 is characterized in that: The personalized embedding layer uses a 512-dimensional learnable embedding vector.
8. A diffusion model training and sampling system based on personalized federated learning, characterized in that: include: The time step segmentation module is used to uniformly set a time step segmentation point in the global scope according to the time segmentation sampling strategy, and divide all the time steps of the denoising process into the first part and the second part. The time steps of the first part are responsible for gradually denoising the input noisy image to the intermediate state image, and the time steps of the second part are responsible for gradually denoising the intermediate state image to a clear image. The full time step client local training module is used to initialize the diffusion model locally for each client participating in federated learning and use local image data to randomly sample time steps from the full time step range for training. After reaching the termination condition, each client saves the local optimal model parameters for the first part of the time step sampling in the sampling phase; The client local personalized training module in the second half of the time step is used to re-initialize the diffusion model and set a personalized embedding layer for privacy protection for each client participating in federated learning, and then reuse the local image data to randomly sample time steps from the second part and perform local personalized training on the diffusion model. The input image of the diffusion model in each sampling time step needs to be superimposed with a local personalized embedding layer, and the personalized embedding layer participates in the parameter optimization of the training process together with the diffusion model; The global optimization module is used to keep the personalized embedding layer on the client after each client completes the local personalized training of the diffusion model, and only upload the remaining model parameters to the server. The server performs weighted aggregation on the model parameters received from each client, updates the global model parameters and sends them back to each client for further optimization. The process of local personalized training and server weighted aggregation is continuously repeated until the termination condition is reached, and the global optimal model parameters are obtained on each client. The phased image generation module is used for, in the sampling phase, each client uses the diffusion model in combination with the locally stored local optimal model parameters to gradually execute each time step of the first part, gradually denoises the input noisy image to an intermediate state image, and then uses the diffusion model in combination with the global optimal model parameters to gradually execute each time step of the second part and superimpose the locally stored personalized embedding layer on the input image of each time step, thereby gradually denoising the intermediate state image to a clear image.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the diffusion model training and sampling method based on personalized federated learning as described in any one of claims 1 to 7 is implemented.
10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the diffusion model training and sampling method based on personalized federated learning as described in any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
Low-dose CT imaging method based on context error modulation generalized diffusion model
CN116468817A
Heterogeneity federal learning method and device
CN116502709A
Personalized differential privacy federal learning method and system
CN117556459A
Image classification method based on personalized federated learning and conditional generative adversarial network
CN117788949A
Individualized federal potential diffusion model learning method and system
CN117910601A
Cited By
Federal social recommendation method and device based on adaptive diffusion denoising
CN121278188A