Federal transfer learning medical image analysis method and system based on diffusion model, and storage medium
By employing a federated transfer learning method based on a diffusion model, each medical institution trains its model locally and transmits parameters, while the target institution generates synthetic data for diagnosis. This approach solves the problems of medical image data privacy and cross-domain discrepancies, enabling efficient and accurate medical image analysis.
Patent Information
- Application Number
- CN202510793819.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-11-07
AI Technical Summary
Medical image data involves sensitive patient information and privacy, which makes it difficult to effectively aggregate high-quality data from different institutions. This results in low efficiency and low accuracy of the medical image analysis results output by the model. Furthermore, there are significant differences in equipment, imaging conditions, and pathological features among the data from different institutions, which affects the cross-domain transfer and diagnostic performance of the model.
A federated transfer learning approach based on a diffusion model is adopted. Each medical institution trains a feature extractor, classifier, and diffusion generative model locally and transmits the model parameters to the target institution. The target institution uses the generative model to generate synthetic data from random noise and achieves cross-institutional knowledge transfer and feature alignment through pseudo-label generation and dynamic weight adjustment.
It effectively protects data privacy, reduces communication costs, improves the efficiency and accuracy of medical image analysis, prevents negative transfer, and enhances the diagnostic accuracy and robustness of the model in the target domain.
Smart Images

Figure CN120912941A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image analysis, in particular to a federated transfer learning medical image analysis method and system based on a diffusion model and a computer readable storage medium. BACKGROUND
[0002] With the advent of medical informatization and the big data era, medical image data is playing an increasingly important role. Modern medical institutions generally use computer-aided diagnosis technology to analyze various medical images through artificial intelligence means such as deep learning and image recognition to obtain medical image analysis results, thereby assisting doctors in making accurate diagnoses. In the field of medical image analysis, models such as deep convolutional neural networks have shown excellent feature extraction and classification capabilities and have achieved remarkable results in image recognition, lesion detection and other tasks.
[0003] However, medical image data usually involves patient sensitive information and privacy. In order to protect patient privacy, a large amount of high-quality data distributed in different institutions (i.e. source parties in the present application) cannot be effectively aggregated, and medical image data collected by different institutions also has great differences in equipment, imaging conditions and pathological characteristics, thereby resulting in low efficiency and low accuracy of medical image analysis results output by the model.
[0004] Therefore, the prior art still needs to be improved and developed. SUMMARY
[0005] The main purpose of the present application is to provide a federated transfer learning medical image analysis method and system based on a diffusion model and a computer readable storage medium, which aims to solve the problem in the prior art that medical image data usually involves patient sensitive information and privacy, in order to protect patient privacy, a large amount of high-quality data distributed in different institutions cannot be effectively aggregated, and medical image data collected by different institutions also has great differences in equipment, imaging conditions and pathological characteristics, thereby resulting in low efficiency and low accuracy of medical image analysis results output by the model.
[0006] To achieve the above purpose, the present application provides a federated transfer learning medical image analysis method based on a diffusion model, which comprises the following steps:
[0007] obtaining a target medical image and a preset source model, and performing model training on the preset source model according to the target medical image to obtain a target source model;
[0008] obtaining random Gaussian noise, and performing noise restoration processing and data prediction processing on the random Gaussian noise according to the model parameters of the target source model to obtain synthetic data and preliminary medical image analysis labels;
[0009] obtaining target data, calculating an overall loss according to the target data, the synthetic data and the preliminary medical image analysis label, and iteratively optimizing the preliminary medical image analysis label according to the overall loss to obtain a target medical image analysis label;
[0010] calculating a confidence result of the target data through the target source model, and determining a final medical image analysis result in the target medical image analysis label according to the confidence result.
[0011] Optionally, the federated transfer learning medical image analysis method based on a diffusion model, wherein the preset source model comprises a preset feature extractor, a preset classifier and a preset diffusion generation model; and the target source model comprises a target feature extractor, a target classifier and a target diffusion generation model.
[0012] obtaining a target medical image and a preset source model, and performing model training on the preset source model according to the target medical image to obtain a target source model, specifically comprising:
[0013] obtaining a local medical image, and performing preprocessing on the local medical image to obtain a target medical image, wherein the preprocessing comprises normalization processing, denoising processing and data enhancement processing.
[0014] determining the preset feature extractor and the preset classifier, and performing model training processing on the preset feature extractor and the preset classifier using the target medical image to obtain the target feature extractor and the target classifier.
[0015] determining the preset diffusion generation model, and performing model training processing on the preset diffusion generation model using the local medical image to obtain the target diffusion generation model.
[0016] Optionally, the federated transfer learning medical image analysis method based on a diffusion model, wherein the random Gaussian noise is obtained, and noise restoration processing and data prediction processing are performed on the random Gaussian noise according to the model parameters of the target source model to obtain synthetic data and a preliminary medical image analysis label, specifically comprising:
[0017] obtaining random Gaussian noise, and performing noise restoration processing on the random Gaussian noise through the model parameters of the target diffusion generation model to obtain synthetic data.
[0018] performing data prediction processing on the synthetic data through the model parameters of the target feature extractor and the target classifier to obtain a preliminary medical image analysis label.
[0019] Optionally, the federated transfer learning medical image analysis method based on a diffusion model, wherein acquiring target data, calculating an overall loss based on the target data, the synthetic data, and the preliminary medical image analysis labels, and iteratively optimizing the preliminary medical image analysis labels based on the overall loss to obtain target medical image analysis labels, specifically includes:
[0020] Identify target data and calculate the maximum mean difference loss between the target data and the synthetic data;
[0021] Calculate the classification loss of the preliminary medical image analysis labels, and then weight the classification loss with the maximum mean difference loss to obtain the overall loss;
[0022] The initial medical image analysis labels are iteratively optimized based on the overall loss to obtain the target medical image analysis labels.
[0023] Optionally, in the aforementioned federated transfer learning medical image analysis method based on a diffusion model, the expression for the maximum mean difference loss is:
[0024]
[0025] in, For the maximum mean difference loss, n k Let be the number of samples, and i be the number of the i-th sample. These are the model parameters for the target feature extractor. For source S′ k The i-th data in the dataset For the i-th data in the target data, For Hilbert space.
[0026] Optionally, the federated transfer learning medical image analysis method based on a diffusion model, wherein calculating the confidence result of the target data through the target source-side model and determining the final medical image analysis result in the target medical image analysis label based on the confidence result specifically includes:
[0027] The target data is input into the target feature extractor and the target classifier, and multiple prediction results are output.
[0028] Obtain the probability vector corresponding to each prediction result, and obtain the label category and corresponding confidence result for each probability vector;
[0029] Set a preset confidence threshold, and extract target prediction results from multiple prediction results whose confidence scores are greater than the preset confidence threshold;
[0030] According to the target prediction result, a voting decision process is performed on the target medical image analysis label to obtain a final medical image analysis result.
[0031] Optionally, the federated transfer learning medical image analysis method based on the diffusion model, wherein the confidence result of the target data is calculated through the target source model, and the final medical image analysis result in the target medical image analysis label is determined according to the confidence result, and then the method further comprises:
[0032] The voting weight of the source party is dynamically updated according to the final medical image analysis result.
[0033] Optionally, the federated transfer learning medical image analysis method based on the diffusion model, wherein the voting weight of the source party is dynamically updated according to the final medical image analysis result, and specifically comprises:
[0034] The consistency rate of the label category and the final medical image analysis result is calculated, wherein the expression of the consistency rate is:
[0035]
[0036] wherein r k is the consistency rate, is an indicator function, n label is the number of generated medical image analysis labels, is the label category, y i is the final medical image analysis result;
[0037] The voting weight of the source party is dynamically updated according to the consistency rate, and the expression of the dynamic update of the voting weight of the source party is:
[0038]
[0039] wherein a k is the voting weight of the source party, K is the number of source parties, j is the jth source party, and r j is the jth consistency rate.
[0040] In addition, to achieve the above-mentioned purpose, the application further provides a federated transfer learning medical image analysis system based on a diffusion model, wherein the federated transfer learning medical image analysis system based on the diffusion model comprises a source party and a target party:
[0041] The source party is configured to acquire a target medical image and a preset source model, perform model training on the preset source model according to the target medical image to obtain a target source model, and send model parameters of the target source model to the target party.
[0042] The target party is configured to obtain random Gaussian noise, and perform noise restoration processing and data prediction processing on the random Gaussian noise according to model parameters of the target source party model, to obtain synthetic data and preliminary medical image analysis labels;
[0043] The target party is further configured to obtain target data, calculate an overall loss according to the target data, the synthetic data, and the preliminary medical image analysis labels, and iteratively optimize the preliminary medical image analysis labels according to the overall loss, to obtain target medical image analysis labels.
[0044] The target party is further configured to calculate a confidence result of the target data by using the target source party model, and determine a final medical image analysis result in the target medical image analysis labels according to the confidence result.
[0045] In the present application, a target medical image and a preset source party model are obtained, and the preset source party model is trained according to the target medical image, to obtain a target source party model. Random Gaussian noise is obtained, and noise restoration processing and data prediction processing are performed on the random Gaussian noise according to model parameters of the target source party model, to obtain synthetic data and preliminary medical image analysis labels. Target data is obtained, an overall loss is calculated according to the target data, the synthetic data, and the preliminary medical image analysis labels, and the preliminary medical image analysis labels are iteratively optimized according to the overall loss, to obtain target medical image analysis labels. A confidence result of the target data is calculated by using the target source party model, and a final medical image analysis result in the target medical image analysis labels is determined according to the confidence result. In the present application, the source party training model is used, and the trained model parameters are sent to the target party for use, so that the user privacy is effectively protected. Then, the target party generates synthetic data from random noise by using the generated model trained by the source party, and generates pseudo labels in combination with the models trained by the source parties, so that the final medical image analysis result is determined through iterative optimization and confidence calculation, thereby effectively improving the output efficiency and accuracy of the medical image analysis result. BRIEF DESCRIPTION OF DRAWINGS
[0046] Figure 1 is a flowchart of a preferred embodiment of the federated transfer learning medical image analysis method based on the diffusion model of the present application;
[0047] Figure 2 is a schematic diagram of the overall execution process of a preferred embodiment of the federated transfer learning medical image analysis method based on the diffusion model of the present application;
[0048] Figure 3FIG. 1 is a schematic diagram of data distribution comparison before and after federated transfer learning of a preferred embodiment of the medical image analysis method based on federated transfer learning of a diffusion model of the present application;
[0049] Figure 4 FIG. 2 is a schematic diagram of target ROC of a preferred embodiment of the medical image analysis method based on federated transfer learning of a diffusion model of the present application;
[0050] Figure 5 FIG. 3 is a structural diagram of a preferred embodiment of the medical image analysis system based on federated transfer learning of a diffusion model of the present application. DETAILED DESCRIPTION
[0051] In order to make the objectives, technical solutions and advantages of the present application clearer and more explicit, the present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0052] With the advent of medical informatization and the big data era, medical image data plays an increasingly important role in disease diagnosis, disease assessment and treatment plan formulation. Modern medical institutions generally use computer-aided diagnosis (CAD) technology, which analyzes CT (Computed Tomography), MRI (Magnetic Resonance Imaging), X-ray, ultrasound and other medical images through deep learning, image recognition and other artificial intelligence methods, thereby assisting doctors in making accurate diagnoses.
[0053] In the field of medical image analysis, models such as deep convolutional neural networks (CNNs) have demonstrated superior feature extraction and classification capabilities and have achieved significant results in image recognition, lesion detection, and other tasks. However, there are still several bottleneck problems in practical applications, mainly reflected in the following aspects: 1. Data privacy and security issues: Medical image data often involves sensitive information and privacy of patients, which is strictly protected by laws and regulations such as GDPR (General Data Protection Regulation), HIPAA (Health Insurance Portability and Accountability Act), etc. In order to avoid the leakage of patient privacy, medical institutions face great challenges in data sharing and centralized storage. This "data island" phenomenon makes it difficult to effectively gather large amounts of high-quality data distributed in different institutions, limiting the training effect of deep learning models on large-scale data. 2. Data distribution is uneven and the cost of labeling is high: The image data collected by different medical institutions may have significant differences in imaging equipment, imaging protocols, patient groups, and pathological characteristics, resulting in large differences in data feature and sample distribution. In addition, since the labeling work of medical images usually requires professional doctors, the labeling cost is high and the period is long, making it difficult for small or primary medical institutions to obtain enough labeled data, affecting the generalization ability and robustness of the model. 3. The need for cross-domain knowledge transfer: Due to the distribution difference of data between institutions, the model trained by a single institution is often difficult to be directly applied to other institutions, resulting in performance degradation of the model in cross-institution environment. In order to solve this problem, cross-domain transfer learning (Domain Adaptation) and unsupervised multi-source domain adaptation (UMDA) technology have attracted widespread attention, which aims to reduce the difference between different domains through data distribution alignment, feature conversion and other methods, and improve the universality and adaptability of the model.
[0054] Currently, there are mainly three types of technical solutions in the field of medical image data analysis and diagnosis: 1. Centralized deep learning method: Early methods usually collect data from each institution to a central server to train deep neural networks with large-scale data. Although this method can achieve good results in some tasks, it is strictly limited in practical application due to data privacy, legal compliance, and data transmission issues. In addition, centralized data collection also faces problems such as uneven data quality, noise interference, and high data integration costs. 2. Federated Learning (FL): To solve the problem of data privacy, federated learning allows each medical institution to train the model locally and only transmits model parameters rather than raw data, thereby realizing collaborative training of cross-institution models. Although FL can protect data privacy to some extent, it requires high data distribution, and the data from each institution needs to have a large overlap in feature or sample space. At the same time, federated learning requires frequent exchange of model parameters in actual communication, which can easily lead to excessive communication overhead, and due to the difference in computing resources of each institution, it can also cause training delay and synchronization problems. 3. Cross-domain transfer and domain adaptation technology: To compensate for the differences in data distribution between different institutions, researchers have proposed various domain adaptation methods, such as Maximum Mean Discrepancy (MMD), adversarial training, feature mapping, and statistical distance minimization. These methods can achieve feature alignment between different data domains to some extent and improve the performance of the model on the target domain. However, existing centralized domain adaptation methods usually require source data and target data to be processed in the same location, which has obvious limitations in data privacy and legal compliance. 4. Generation model and data sharing technology: In recent years, generative adversarial networks (GAN), variational autoencoders (VAE), and diffusion models have made significant progress in image generation. These models can generate high-quality, realistic images from noise, providing new ideas for data augmentation and privacy protection. In particular, diffusion models model data distribution by gradually adding and removing noise, providing a stable implementation scheme for generating synthetic images similar to the real data distribution. Generation models have important application prospects in data sharing, but how to combine them with federated learning and cross-domain transfer technology to build a system that can protect privacy and realize cross-institution knowledge transfer is still a key problem to be solved.
[0055] In summary, although the existing technology has made certain progress in its respective field, there are still several key challenges in the actual medical image data diagnosis, including: 1. Data privacy and security: how to realize multi-institutional data collaborative training and cross-domain migration while protecting patient privacy is a problem that needs to be solved urgently. Traditional centralized training methods cannot meet the security compliance requirements. 2. Data distribution difference: there are great differences in equipment, imaging conditions and pathological characteristics of medical images collected by different institutions. Direct use of a single model often leads to performance degradation. How to realize cross-domain feature alignment and improve the generalization ability of the model in the target domain is the key to technological breakthrough. 3. Communication overhead and computing burden: although federated learning solves the problem of data privacy, it is easy to waste communication resources due to frequent model parameter transmission in actual application. At the same time, uneven computing capacity of each institution also affects the overall training efficiency. 4. Negative transfer risk: in multi-source data collaborative training, some low-quality or malicious source data may cause negative transfer and affect model performance. How to automatically select source data that has a positive contribution to the target domain and dynamically adjust its influence weight is one of the core technologies to ensure the stable operation of the system.
[0056] In summary, the existing technology has not yet fully solved the above problems, so there is an urgent need for an innovative solution that combines generative models, federated learning, domain adaptation and automated source selection to achieve efficient and accurate diagnosis of medical image data while ensuring data privacy and reducing communication costs.
[0057] To solve the above problems, the present application discloses a federated transfer learning framework based on diffusion model (hereinafter referred to as AutoFTL, Auto Federated Transfer Learning) for automatic diagnosis of medical image data. The core idea of this solution is that each medical institution (i.e. the source party in this application) trains a feature extractor, a classifier and a diffusion generative model locally using its labeled data, and transmits the model parameters to the target party through a secure channel; the target party generates synthetic data from random noise using the source party's generative model, and then performs domain adaptation and pseudo-label generation in combination with each source party's model, and finally determines the diagnosis result of the target image through multi-source voting and dynamic weight adjustment.
[0058] This method has the following main advantages: 1. Privacy protection: the source party does not need to transmit raw data, only the trained model parameters, thus effectively protecting data privacy. 2. Automated source selection: through the pseudo-label voting mechanism and dynamic updating of source weights, the source party that has a positive contribution to the target domain is effectively selected to prevent negative transfer. 3. Efficient domain adaptation: techniques such as Maximum Mean Discrepancy (MMD) are used to align the features of synthetic data and target data, improving the accuracy of diagnosis. 4. Low communication cost: only model parameters are transmitted instead of large amounts of raw data, significantly reducing communication overhead.
[0059] The method for medical image analysis based on a diffusion model and federated transfer learning according to the preferred embodiment of the present application comprises the following steps: Figure 1 As shown in the figure, the method for medical image analysis based on a diffusion model and federated transfer learning comprises the following steps:
[0060] In step S10, a target medical image and a preset source model are obtained, and the preset source model is trained according to the target medical image to obtain a target source model. The preset source model comprises a preset feature extractor, a preset classifier and a preset diffusion generation model; and the target source model comprises a target feature extractor, a target classifier and a target diffusion generation model.
[0061] As shown in the figure, in the present application, the system as a whole is designed as a distributed federated learning environment, and the basic idea is that each medical institution (i.e. the source party in the present application) retains its own medical image data, obtains its own model parameters through local training, and then transmits these encrypted model parameters to the target party coordinated by the center. The target party generates "synthetic data" similar to the source data distribution without obtaining the actual original data, thereby realizing cross-institution knowledge transfer, and the whole system architecture embodies the dual requirements of data privacy protection and cross-domain knowledge sharing. Figure 2 Source party module: each source party represents an independent medical institution, which has locally collected medical image data and corresponding diagnostic labels. To realize knowledge transfer, each source party needs to independently train the following three sub-modules locally: 1. Feature extractor F: this module uses a deep convolutional neural network to extract features from medical images. By learning the low, medium and high level features of the image, the original image data is mapped to a common feature space, providing high-quality feature representation for subsequent classification and domain alignment. 2. Classifier C: based on the output of the feature extractor, the classifier further processes the features and makes specific diagnostic judgments. Usually, a softmax output layer is used, and a cross-entropy loss function is used for model training to ensure that the classifier can accurately classify the image into the preset diagnostic categories. 3. Generation model U (i.e. the diffusion model in the present application): to solve the problem of data privacy, the source party does not directly transmit the original data, but trains a generation model. The diffusion model converts the original image data into a Gaussian noise sequence through a process of gradually adding noise, and then generates an image similar to the original data distribution through a reverse denoising process. This process can not only preserve the distribution information of the data, but also ensure privacy protection.
[0062]
[0063] Target module: The target is usually an institution that needs to make diagnostic decisions but has less data or no annotation. The target does not directly obtain source data, but after receiving the model parameters transmitted by each source, it generates "synthetic data" through the generation model and uses these data for domain adaptation and pseudo-label generation. The workflow of the target mainly includes: 1. Receive model parameters: Obtain the parameters of the feature extractor, classifier and generation model from each source, which have been privacy-protected and do not contain original data details. 2. Synthetic data generation: Use the received generation model to generate synthetic data from random noise to approximate the source data distribution. 3. Pseudo-label generation and domain adaptation: Obtain preliminary pseudo-labels by passing synthetic data through the source model, and use MMD and other technologies to align the target data and synthetic data in the feature space. 4. Pseudo-label voting and source weight update: Use multi-source prediction for target samples, combine the lowest confidence filtering and weighted voting to determine the final pseudo-label, and dynamically adjust the weight according to the consistency of each source to ensure that high-quality source data plays a greater role.
[0064] As shown in Figure 2 , Figure 2 The interaction between the modules and the data flow are shown, and the whole system realizes the goal of sharing knowledge across institutions and improving automatic diagnosis performance under the premise of data privacy protection.
[0065] Specifically, a local medical image is obtained, and the local medical image is preprocessed to obtain a target medical image, wherein the preprocessing includes normalization processing, denoising processing and data enhancement processing; the preset feature extractor and the preset classifier are determined, and the target medical image is used to perform model training processing on the preset feature extractor and the preset classifier to obtain the target feature extractor and the target classifier; the preset diffusion generation model is determined, and the local medical image is used to perform model training processing on the preset diffusion generation model to obtain the target diffusion generation model.
[0066] In the present application, a source end model training module is provided. At the source end, each medical institution independently completes model training based on its local labeled data. Its main tasks include two aspects: feature extraction and classifier training, and generation model (diffusion model) training.
[0067] Among them, feature extraction and classifier training: in this phase, the source party first uses the local medical image data to train the deep convolutional neural network to extract the effective features of the image. The key is to learn the deep features that can distinguish different diagnostic categories. The specific process is as follows: 1. Data preprocessing and enhancement: normalize, denoise and data enhancement on medical image data, improve image quality and model robustness. 2. Feature extractor design and training: use deep networks such as ResNet, VGG16, MobileNet to extract distinguishing features from images. (Among them, the process of feature extraction is as follows: the preprocessed image is input into the deep convolutional neural network (such as ResNet, VGG16 or MobileNet). These networks are composed of multiple convolutional layers, activation functions (such as ReLU) and pooling layers. Each convolutional layer uses a set of filters to scan local regions to extract low-level features such as edges, textures and shapes in the image; as the network deepens layer by layer, these low-level features are gradually combined into higher-level, more abstract feature representations. Finally, after processing by multiple convolutional layers, each channel of the feature map output by the network contains a certain pattern in the image, providing distinguishing feature information for subsequent classification). 3. Classifier construction and training: based on feature extraction, design fully connected layers and softmax output layers to classify images. (Among them, the process of diagnostic classification is as follows: the feature vector extracted by the feature extractor is input into one or more fully connected layers, which linearly combine the features and map them to an output vector of fixed length. Then, the output vector is converted to a probability distribution of each category by the softmax layer, where each probability value represents the confidence of the image belonging to the corresponding diagnostic category. Finally, the model selects the category with the highest probability as the diagnostic result of the image. This process continuously optimizes the classifier parameters by minimizing the cross-entropy loss during training to improve classification accuracy). 4. Loss function: for each source party, cross-entropy loss L class,k is used to guide model training, which is mathematically expressed as:
[0068]
[0069] Among them, and represent the data and labels of source S k , is the trained classifier, i.e. the target classifier in the present application, is the trained feature extractor, i.e. the target feature extractor in the present application, represents the number of data samples, and the network weights are updated through backpropagation until the model reaches a high classification accuracy on local data.
[0070] Training of the diffusion generative model: In order to achieve data sharing without leaking privacy, each source party models the medical image data using a diffusion model, and the diffusion model training is divided into two stages: 1. Forward diffusion process: starting from the original image x0 (the original image is the local medical image in the present application, and 0 represents the 0th step of the diffusion model), gradually introduce Gaussian noise to make the image gradually change into a noisy image, and the formula is:
[0071]
[0072] where t represents the number of iterations, that is, the image x t at each step is obtained by performing the above operation on the image x t-1 at the previous step, and when t is 1, it is the image generated by adding noise to the original medical data image at the first step. β t is a preset noise coefficient between 0 and 1, z t is Gaussian noise.
[0073] 2. Reverse denoising process: using the trained neural network (referred to as diffusion model) to learn to recover the image from noise, and the formula is as follows:
[0074]
[0075] where α t = 1-β t , α t represents the noise coefficient when the step number t is i, is the average value of the noise coefficients of all steps, α i is the noise coefficient at time i, U is the reverse denoising processing of the diffusion model on the noisy image x t , t is the iteration step number, z is the noise, is a normal distribution, I is a unit vector, ∏ is the symbol of multiplication, and σ t is the standard deviation.
[0076] After that, the noise is gradually restored to an image similar to the original data distribution.
[0077] Training objective: the mean square error (MSE) loss is determined, and the objective is to minimize, The expression is as follows:
[0078]
[0079] where, is the mean square error loss, is the maximum step number.
[0080] After the training is completed, each source party will transmit the trained model parameters to the target party in a secure and encrypted manner.
[0081] In step S20, random Gaussian noise is obtained, and noise restoration processing and data prediction processing are performed on the random Gaussian noise according to the model parameters of the target source model, to obtain synthetic data and preliminary medical image analysis labels.
[0082] In the present application, a target end data generation and domain adaptation module is provided: the task of the target party is to generate synthetic data similar to the source data distribution using the model parameters transmitted by the source party without directly accessing the real data, and to realize effective alignment of the target data and the generated data in the feature space through domain adaptation technology, thereby improving the diagnostic accuracy. This process mainly includes three links of data generation, pseudo-label generation and domain alignment.
[0083] Specifically, random Gaussian noise is obtained, and noise restoration processing is performed on the random Gaussian noise through the model parameters of the target diffusion generation model, to obtain synthetic data.
[0084] For the synthetic data generation process: the target party uses the generation model of each source party From the random Gaussian noise (random Gaussian noise is usually obtained by sampling from a standard normal distribution. Specifically, a noise tensor that follows N(0, 1) is generated using a normal distribution sampling function in a programming language (such as numpy.random.normal in Python), and the size of the tensor matches the input size required by the model. This noise sample is used as the starting input for the reverse denoising process of the diffusion model, thereby starting to generate synthetic images similar to the source data distribution), step-by-step reverse denoising is performed to generate synthetic images similar to the source data distribution. The specific steps are as follows: 1. Initialize a batch of noise samples following a Gaussian distribution; 2. Perform the denoising process step by step according to the formula (i.e., the formula for the reverse denoising process in the diffusion model) until the synthetic image (i.e., the synthetic data in the present application) is restored, forming a synthetic data set The generated synthetic data has a similar distribution to the real source data, but does not contain the original image details, thereby effectively protecting data privacy.
[0085] Data prediction processing is performed on the synthetic data through the model parameters of the target feature extractor and the target classifier, to obtain preliminary medical image analysis labels.
[0086] For the process of pseudo-label generation and domain alignment: after obtaining the synthetic data, the target party uses the feature extractor and classifier provided by the source party The synthetic data is predicted to generate preliminary pseudo labels (i.e. preliminary medical image analysis labels in the present application).
[0087] The process of pseudo label generation is as follows: for each generated synthetic image, the prediction probability is obtained through the network, and the class with the highest probability is selected as the preliminary pseudo label. The purpose of this step is to provide an initial class guide for the target data, thereby helping the model converge in subsequent training.
[0088] It can be understood that in the pseudo label generation process, the generated synthetic image is first input into the feature extractor trained by the source party to extract the feature representation of the image; then, the features are passed to the classifier, which is composed of a fully connected layer and a softmax layer. The softmax layer converts the output of the fully connected layer into a probability distribution of each class, and each probability value output represents the confidence of the image belonging to the class. Finally, the model selects the class with the highest prediction probability as the preliminary pseudo label of the image, thereby providing preliminary class guidance for subsequent training.
[0089] Step S30, obtaining target data, calculating the overall loss according to the target data, the synthetic data and the preliminary medical image analysis label, and iteratively optimizing the preliminary medical image analysis label according to the overall loss to obtain a target medical image analysis label.
[0090] Specifically, the target data is determined, and the maximum mean discrepancy loss between the target data and the synthetic data is calculated; the classification loss of the preliminary medical image analysis label is calculated, and the classification loss and the maximum mean discrepancy loss are weighted to obtain the overall loss; the preliminary medical image analysis label is iteratively optimized according to the overall loss to obtain the target medical image analysis label.
[0091] The expression of the maximum mean discrepancy loss is:
[0092]
[0093] wherein, is the maximum mean discrepancy loss, n k is the number of samples, i is the number of the i-th sample, is the model parameter of the target feature extractor, is the i-th data in the data set of the source party S' k is the i-th data in the target data, is the i-th data in the target data, is the Hilbert space.
[0094] The present application adopts domain adaptation technology: since the target data DT The target of the present application is to make a diagnosis automatically in the case that the target data (medical image data) has no label There may be a distribution difference between the target data and the synthetic data
[0095]
[0096] wherein, is the maximum mean difference loss between the target data D T and the synthetic data , E is the expected value, represents a piece of data in the data set of the source S' k , and x T is the data in the target data set. and P T are the data distributions of the source and the target respectively, and φ(·) represents a mapping function, which is usually selected as the output of a certain hidden layer, and specifically, the MMD loss function can be written as:
[0097]
[0098] The joint training target of the present application: the target combines the classification loss generated by the pseudo label and the MMD loss by a weight coefficient β to form an overall loss wherein, the expression of the overall loss is:
[0099]
[0100] By continuously optimizing the loss function, the target gradually aligns the feature distribution of the target data with the synthetic data (and the source data), thereby improving the classification accuracy.
[0101] Iterative optimization process: the entire domain adaptation process needs to be performed in multiple training cycles (when the loss function converges, it is ended, of course, an upper limit of iteration is also preset to avoid wasting computing resources). In each cycle, the target continuously updates the synthetic data, regenerates the pseudo label, and adjusts the model parameters using the latest pseudo label feedback. In this way, as the iteration proceeds, the target can gradually improve the diagnosis accuracy of the medical image data, and make the generated data more consistent with the target data in the feature space.
[0102] Step S40, calculating the confidence result of the target data by the target source model, and determining the final medical image analysis result in the target medical image analysis label according to the confidence result.
[0103] The pseudo label voting and source selection module is arranged in the application: due to different sources, the pseudo labels generated by the sources may have certain deviations due to data collection, model training, etc., in order to further improve the accuracy and robustness of the final diagnosis result of the target party, the target party introduces a pseudo label voting mechanism and a dynamic source weight updating strategy.
[0104] Specifically, the target data is input into the target feature extractor and the target classifier, and a plurality of prediction results are output; a probability vector corresponding to each prediction result is obtained, and a label category and a corresponding confidence result of each probability vector are obtained; a pre-confidence threshold is set, and a target prediction result in which the confidence result is greater than the pre-confidence threshold is extracted from the plurality of prediction results; the target medical image analysis label is voted and decided according to the target prediction result, and a final medical image analysis result is obtained.
[0105] Further, a consistency rate of the label category and the final medical image analysis result is calculated, wherein the expression of the consistency rate is:
[0106]
[0107] wherein r k is the consistency rate, is an indicator function, n label is the number of generated medical image analysis labels, is the label category, y i is the final medical image analysis result.
[0108] According to the consistency rate, the voting weight of the source party is dynamically updated, wherein the expression of dynamically updating the voting weight of the source party is:
[0109]
[0110] wherein a k is the voting weight of the source party, K is the number of source parties, j is the jth source party, r j is the jth consistency rate.
[0111] The pseudo label voting mechanism is arranged in the application: in this link, the target party will process each target sample (that is, the target data in the application, the target data has no label, and the purpose of the application is to automatically diagnose the target data (medical image) in the absence of labels) as follows: 1, multi-model prediction: input the target sample into the feature extractor and the classifier of each source party to obtain a plurality of prediction results. Each result is output in the form of a probability vector, and the class corresponding to the maximum probability is marked and the corresponding confidence 2. Confidence screening: To avoid the influence of low-confidence uncertain predictions on the overall effect, the target party sets a minimum confidence threshold p min Only when the prediction confidence of a certain source party is , the prediction of that source party is considered valid. 3. Voting decision: For each sample, the majority voting or weighted voting strategy is used according to the prediction results of all valid source parties to determine the final pseudo-label (i.e. the final medical image analysis result in the present application). This process fully considers the consistency between the prediction results of different source parties, providing more reliable label guidance for subsequent model training.
[0112] In the present application, dynamic source weight updating is also provided: to further improve the robustness of the system, the target party dynamically adjusts the weight of each source party in the final voting according to the consistency of each source party with the final pseudo-label, the process includes: 1. Consistency evaluation: for the target sample whose pseudo-label has been generated, calculate the consistency rate r i of each source party's prediction with the final pseudo-label y k :
[0113]
[0114] 2. Weight calculation: according to the consistency rate, different weights are given to each source party, and the update formula is:
[0115]
[0116] In this way, the source party with higher consistency will have greater influence in voting, further preventing low-quality or malicious source parties from causing negative transfer. Through the iteration of pseudo-label voting and source weight updating, the target party can continuously improve the utilization efficiency of information from each source party, ensuring that the final model achieves higher accuracy and robustness in the diagnosis task of medical image data.
[0117] As shown in Figure 3 and Figure 4 , Figure 3 is a comparison chart of data distribution before and after federated transfer learning, Figure 4 is a target ROC chart. The experimental results show that the present application not only significantly outperforms traditional methods in accuracy and area under the ROC curve (AUC, Area Under Curve), but also has a significant advantage in communication cost, and has a high application prospect.
[0118] The entire end-to-end process in the present application can be summarized as follows:
[0119] 1. Source model training phase: Feature extractor and classifier training. The source uses a deep convolutional neural network (e.g. ResNet, VGG16 or MobileNet) to build a feature extractor to extract low, medium and high-level features from medical images, and on this basis, build a classifier to diagnose and classify images. The entire process uses a cross-entropy loss function, and the network parameters are constantly updated through a backpropagation algorithm, so that the classification accuracy of the model on local data reaches the expected level.
[0120] 2. Diffusion generative model training: In order to share data distribution information without transmitting original data, the source also uses a diffusion model to model medical image data. First, the original image is gradually transformed into a noisy image through a forward diffusion process; then, a reverse denoising process is used to gradually recover an image similar to the original data distribution from Gaussian noise. The source optimizes the diffusion model by minimizing the mean square error loss, so that the model can accurately restore the main distribution characteristics of the image during the generation process. After completing the above two stages, the sources will transmit the parameters of the trained feature extractor, classifier and diffusion generative model to the target through secure encryption.
[0121] 3、Target model updating and domain adaptation stage: After receiving the model parameters from each source, the target does not directly access or store the real source data, but uses the following process to realize end-to-end model updating: a. Synthetic data generation: The target uses the diffusion generation model transmitted by each source to gradually remove noise from randomly generated Gaussian noise to generate synthetic image data similar to the distribution of each source data. These synthetic data constitute the "synthetic data set" of each source, simulating the distribution of real source data. b. Pseudo-label generation: The target inputs the generated synthetic data into the feature extractor and classifier corresponding to each source to obtain preliminary prediction results, and takes the class with the highest probability as the pseudo-label of each synthetic sample. This avoids direct access to real data and provides basic classification guidance for subsequent domain adaptation. c. Domain adaptation training: Since there may be differences in the statistical distribution between target data and synthetic data, the target uses domain alignment technology to reduce these differences. Specifically, the target compares the distribution of the two sets of data in the public feature space (for example, using the maximum mean difference MMD loss to quantify the difference), and combines the classification loss generated by the pseudo-label to form the training target of the target end. The target uses gradient descent and other optimization methods to continuously adjust the parameters of its own model in each training period to gradually align the feature expression of the target data with the synthetic data, thereby achieving cross-domain feature alignment. d. Pseudo-label voting and source weight update: To further improve the accuracy and robustness of the pseudo-label, the target uses a multi-source prediction method for each target data sample. The specific process is as follows: First, input the target sample into the model of each source to obtain the prediction probability and corresponding class; then, for each sample, the target sets a minimum confidence threshold and only retains the results of the source whose prediction confidence is higher than the threshold; finally, the target determines the final pseudo-label of the sample by majority or weighted voting mechanism based on the effective prediction results of each source. At the same time, the target calculates the consistency of the prediction results of each source and the final pseudo-label, and dynamically adjusts the weight of each source in the voting process according to the consistency. Sources with higher consistency have greater weight in subsequent voting, further suppressing the negative impact of low-quality or malicious sources on the overall result.
[0122] 4、Final output: Through multiple rounds of iterative optimization, the target continuously updates the model parameters in each training period to generate more accurate synthetic data and pseudo-labels. After the end of the entire end-to-end process, the target finally outputs the diagnostic pseudo-labels for the target data as the final result of automatic diagnosis of medical image data. The entire process not only realizes effective knowledge transfer across institutions, but also ensures the privacy and security of data from each medical institution. Through the above process, the system efficiently performs cross-institution knowledge transfer and domain adaptation without directly accessing the source data, while fully protecting the privacy of data from each institution.
[0123] The main innovations of the present application are as follows:
[0124] 1. Privacy-protected data generation mechanism: using diffusion generative model to generate synthetic data similar to the real data distribution at the source side, avoiding the transmission of original data, and realizing data privacy protection.
[0125] 2. Automatic source selection and dynamic weight update: through pseudo-label voting mechanism and consistency evaluation, automatic screening and weighting of information from each source is realized, effectively preventing negative transfer caused by low-quality or malicious source data.
[0126] 3. Joint training framework of fusion domain alignment and classification optimization: using the weighted combination of maximum mean difference (MMD) loss and pseudo-label classification loss, the target data and synthetic data are efficiently aligned in the feature space, thereby improving the final diagnostic accuracy.
[0127] 4. Distributed federated transfer learning architecture with low communication cost: only model parameters and intermediate features are transmitted, without the need to transmit original medical image data, greatly reducing communication burden while ensuring data security.
[0128] Possible design changes or variations of the present application:
[0129] 1. Generation model replacement scheme: in addition to diffusion models, generative adversarial networks (GAN), variational autoencoders (VAE), and other generative models can also be used, but the trade-off between generation quality and stability needs to be considered.
[0130] 2. Domain adaptation index extension: in addition to MMD loss, adversarial loss, Euclidean distance, or other statistical indicators can be introduced for domain alignment to further improve the generalization ability of the target model.
[0131] 3. Heterogeneous design of model architecture: allows different neural network structures for each source, with self-defined feature extractors and classifiers based on their data characteristics, and unified optimization through fusion strategies at the target side.
[0132] 4. Extension to other medical image fields: in addition to medical image data, the present application can be extended to MRI or other medical images, and even applied to federated transfer learning of non-image data.
[0133] In addition, knowledge distillation and model compression techniques can be combined to further reduce model parameter transmission and communication overhead. To address potential malicious source attacks, more comprehensive security verification and anomaly detection mechanisms can be designed in the future. In addition to image data, federated transfer learning based on graph neural networks (GNN) can also be explored to handle structured or graph data such as electroencephalogram (EEG) and other medical data.
[0134] Further, as Figure 5As shown, based on the above federal transfer learning medical image analysis method based on the diffusion model, the application also correspondingly provides a federal transfer learning medical image analysis system based on a diffusion model, wherein the federal transfer learning medical image analysis system based on the diffusion model comprises a source party and a target party:
[0135] The source party is used for acquiring a target medical image and a preset source party model, performing model training on the preset source party model according to the target medical image, obtaining a target source party model, and sending model parameters of the target source party model to the target party.
[0136] The target party is used for acquiring random Gaussian noise, and performing noise restoration processing and data prediction processing on the random Gaussian noise according to the model parameters of the target source party model, to obtain synthetic data and a preliminary medical image analysis label.
[0137] The target party is also used for acquiring target data, calculating an overall loss according to the target data, the synthetic data and the preliminary medical image analysis label, and iteratively optimizing the preliminary medical image analysis label according to the overall loss, to obtain a target medical image analysis label.
[0138] The target party is also used for calculating a confidence result of the target data through the target source party model, and determining a final medical image analysis result in the target medical image analysis label according to the confidence result.
[0139] In summary, the present application provides a federal transfer learning medical image analysis method and system based on a diffusion model and a storage medium. The method comprises the following steps: obtaining a target medical image and a preset source model, and performing model training on the preset source model according to the target medical image to obtain a target source model; obtaining random Gaussian noise, and performing noise restoration processing and data prediction processing on the random Gaussian noise according to the model parameters of the target source model to obtain synthetic data and preliminary medical image analysis labels; obtaining target data, calculating an overall loss according to the target data, the synthetic data, and the preliminary medical image analysis labels, and iteratively optimizing the preliminary medical image analysis labels according to the overall loss to obtain target medical image analysis labels; and calculating the confidence result of the target data through the target source model, and determining the final medical image analysis result in the target medical image analysis labels according to the confidence result. The present application effectively protects user privacy by training a source model and sending the trained model parameters to a target for use. Then, the target generates synthetic data from random noise using the trained source model, generates pseudo labels in combination with the trained models of each source, and finally determines the final medical image analysis result through iterative optimization and confidence calculation, thereby effectively improving the output efficiency and accuracy of the medical image analysis result.
[0140] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, such that processes, methods, articles, or terminals including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, methods, articles, or terminals. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article, or terminal including the element.
[0141] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware (such as a processor, a controller, etc.) to complete, and the program can be stored in a computer-readable computer-readable storage medium, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a magnetic disc, an optical disc, etc.
[0142] It should be understood that the application is not limited to the above examples, and those skilled in the art can make improvements or changes according to the above description, and all these improvements and changes should be within the protection scope of the appended claims of the present application.
Claims
1. A federated transfer learning medical image analysis method based on a diffusion model, characterized in that, The federal transfer learning medical image analysis method based on the diffusion model comprises: obtaining a target medical image and a preset source model, and performing model training on the preset source model according to the target medical image to obtain a target source model; obtaining random Gaussian noise, and performing noise restoration processing and data prediction processing on the random Gaussian noise according to the model parameters of the target source model to obtain synthetic data and a preliminary medical image analysis label; obtaining target data, calculating an overall loss according to the target data, the synthetic data and the preliminary medical image analysis label, and iteratively optimizing the preliminary medical image analysis label according to the overall loss to obtain a target medical image analysis label; calculating the confidence result of the target data through the target source model, and determining the final medical image analysis result in the target medical image analysis label according to the confidence result. 2.The diffusion model based federated transfer learning medical image analysis method according to claim 1, wherein, The preset source model comprises a preset feature extractor, a preset classifier and a preset diffusion generation model; the target source model comprises a target feature extractor, a target classifier and a target diffusion generation model; The obtaining of the target medical image and the preset source model, and the model training of the preset source model according to the target medical image to obtain the target source model specifically comprises: obtaining a local medical image, and performing preprocessing on the local medical image to obtain a target medical image, wherein the preprocessing comprises normalization processing, denoising processing and data enhancement processing; determining the preset feature extractor and the preset classifier, and performing model training processing on the preset feature extractor and the preset classifier using the target medical image to obtain the target feature extractor and the target classifier; determining the preset diffusion generation model, and performing model training processing on the preset diffusion generation model using the local medical image to obtain the target diffusion generation model. 3.The diffusion model based federated transfer learning medical image analysis method according to claim 2, characterized in that, The obtaining of the random Gaussian noise, and the noise restoration processing and data prediction processing of the random Gaussian noise according to the model parameters of the target source model to obtain the synthetic data and the preliminary medical image analysis label specifically comprises: obtaining random Gaussian noise, and performing noise restoration processing on the random Gaussian noise through the model parameters of the target diffusion generation model to obtain synthetic data; performing data prediction processing on the synthetic data through the model parameters of the target feature extractor and the target classifier to obtain a preliminary medical image analysis label. 4.The diffusion model based federated transfer learning medical image analysis method of claim 1, wherein, The obtaining of the target data, the calculation of an overall loss according to the target data, the synthetic data and the preliminary medical image analysis label, and the iteratively optimizing of the preliminary medical image analysis label according to the overall loss to obtain a target medical image analysis label specifically comprises: determining target data, and calculating the maximum mean difference loss between the target data and the synthetic data; calculating the classification loss of the preliminary medical image analysis label, and performing weighted processing on the classification loss and the maximum mean difference loss to obtain an overall loss; The preliminary medical image analysis label is iteratively optimized according to the overall loss, and a target medical image analysis label is obtained.
5. The diffusion model based federated transfer learning medical image analysis method according to claim 4, characterized in that, An expression of the maximum mean difference loss is: wherein, is the maximum mean discrepancy loss, n k is the number of samples, i is the i-th sample number, is the model parameter of the target feature extractor, is the i-th data in the data set of the source party S′ k is the i-th data in the target data, is the i-th data in the target data, is the Hilbert space. 6.The diffusion model based federated transfer learning medical image analysis method according to claim 2, wherein, The confidence result of the target data is calculated through the target source-side model, and a final medical image analysis result in the target medical image analysis label is determined according to the confidence result, and specifically includes: The target data is input into the target feature extractor and the target classifier, and a plurality of prediction results are output; A probability vector corresponding to each prediction result is obtained, and a label category and a corresponding confidence result of each probability vector are obtained; A pre-confidence threshold is set, and a target prediction result in which the confidence result is greater than the pre-confidence threshold is extracted from the plurality of prediction results; The target medical image analysis label is voted and decided according to the target prediction result, and a final medical image analysis result is obtained.
7. The diffusion model based federated transfer learning medical image analysis method according to claim 6, characterized in that, The confidence result of the target data is calculated through the target source-side model, and a final medical image analysis result in the target medical image analysis label is determined according to the confidence result, and specifically includes: The voting weight of the source party is dynamically updated according to the final medical image analysis result.
8. The diffusion model based federated transfer learning medical image analysis method according to claim 7, characterized in that, The voting weight of the source party is dynamically updated according to the final medical image analysis result, and specifically includes: A consistency rate of the label category and the final medical image analysis result is calculated, and an expression of the consistency rate is: wherein r k is a consistency rate, is an indicator function, n label is the number of samples for which a medical image analysis label has been generated, is a label class, y i is a final medical image analysis result; The voting weight of the source party is dynamically updated according to the consistency rate, and an expression of dynamically updating the voting weight of the source party is: Wherein, α k is the voting weight of the source party, K is the number of source parties, j is the jth source party, r j is the jth consistency rate.
9. A federated transfer learning medical image analysis system based on a diffusion model, characterized by, The federated transfer learning medical image analysis system based on the diffusion model includes a source party and a target party: The source party is configured to obtain a target medical image and a preset source-side model, perform model training on the preset source-side model according to the target medical image, obtain a target source-side model, and send model parameters of the target source-side model to the target party. The target party is configured to obtain random Gaussian noise, perform noise restoration processing and data prediction processing on the random Gaussian noise according to the model parameters of the target source-side model, and obtain synthetic data and a preliminary medical image analysis label. The target party is further configured to obtain target data, calculate an overall loss according to the target data, the synthetic data, and the preliminary medical image analysis label, and iteratively optimize the preliminary medical image analysis label according to the overall loss to obtain a target medical image analysis label. The target party is further configured to calculate a confidence result of the target data through the target source-side model, and determine a final medical image analysis result in the target medical image analysis label according to the confidence result.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a federated transfer learning medical image analysis program based on a diffusion model, and the federated transfer learning medical image analysis program based on the diffusion model is executed by the processor to implement the steps of the federated transfer learning medical image analysis method based on the diffusion model in any one of claims 1-8.