Federal medical image segmentation method based on client consistency and adaptive disturbance
By adopting client consistency and adaptive perturbation methods in federal medical image segmentation, the problems of data privacy protection, insufficient labeling data and model generalization capabilities are solved, and the medical image segmentation effect with high precision, robustness and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510052542.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-14
AI Technical Summary
Federal semi-supervised medical imaging segmentation technology faces challenges such as data privacy protection, insufficient labeling data and model generalization capabilities, especially in the case of inconsistent data distribution and noise, which affects the training effect and generalization performance of the model.
The federal medical image segmentation method based on client consistency and adaptive perturbation is adopted. By building an image segmentation network of supervised clients and unsupervised clients, the model is optimized using Gaussian noise perturbation and consistency regularization terms to enhance the robustness and generalization capabilities of the model.
Effective use of labeled data and labelless data improves the accuracy of medical image segmentation and the robustness of the model, reduces the dependence of manual labeling, and enhances the generalization ability of the model and the ability to identify difficult-to-separate lesions.
Smart Images

Figure CN119991686A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image segmentation, and in particular to a federated medical image segmentation method based on client consistency and adaptive perturbation. Background Art
[0002] Federated semi-supervised medical image segmentation is an innovative application that combines a federated learning framework, semi-supervised learning methods, and deep learning algorithms. These technologies can train high-performance segmentation models using distributed data from multiple medical institutions while protecting data privacy. At the same time, they can effectively use medical image data with limited annotations through semi-supervised learning to achieve automatic segmentation of anatomical structures and lesions in images. This technology is designed to help doctors and medical professionals locate and analyze key areas in medical images more accurately and efficiently, reduce their workload, and improve the speed and accuracy of diagnosis. In addition, while ensuring data security, this method provides strong support for solving the problem of insufficient annotated data, which is of great value for clinical decision-making and early diagnosis and treatment of patients.
[0003] Federated semi-supervised medical image segmentation technology combines federated learning, semi-supervised learning and deep learning methods, aiming to address challenges in medical image segmentation such as data privacy protection, insufficient labeled data and model generalization ability. Medical image segmentation is a key task in computer-aided diagnosis. It is mainly used to accurately separate organs, lesions or other anatomical structures in medical images from the background, providing an important basis for clinical diagnosis, surgical planning and disease detection. However, this task faces many technical challenges, including data privacy restrictions, high data annotation costs, inconsistent data distribution and other issues.
[0004] First, medical imaging data (such as CT, MRI, and ultrasound images) usually involve patient privacy and cannot be freely shared between different hospitals and medical institutions, resulting in serious data silos and affecting the training effect of the segmentation model. Traditional centralized training methods require all data to be aggregated to a unified server, but in actual applications, this method not only violates data privacy protection regulations (such as GDPR and HIPAA), but may also increase the risk of data leakage. Federated learning achieves cross-institutional collaborative training and ensures data privacy security by training models locally on each data source and only transmitting model parameters instead of raw data.
[0005] Secondly, the labeling process of medical image data relies on professional doctors, which is time-consuming and economically expensive, resulting in a scarcity of large-scale labeled data. Semi-supervised learning can use a large amount of unlabeled data to assist supervised training through pseudo-label generation, contrastive learning, and other methods, thereby improving segmentation performance. Under the federated learning framework, semi-supervised learning effectively alleviates the problem of insufficient labeled data by combining limited labeled data with a large amount of unlabeled data, thereby improving the model's performance on medical image segmentation tasks.
[0006] In addition, medical image segmentation faces the problem of inconsistent data distribution. Images acquired by different institutions may differ significantly due to equipment models, scanning parameters, and individual differences in patients, resulting in data distribution shift and affecting the generalization ability of the model. For example, the resolution, noise level, contrast, and texture features of the image may vary. To solve these problems, federated semi-supervised learning combines the collaborative training of multiple distributed data sources, integrates the feature distribution of different data sources, and optimizes the model locally, thereby enhancing the robustness and generalization performance of the segmentation model.
[0007] However, federated semi-supervised medical image segmentation also faces many challenges. On the one hand, due to the uneven quality of unlabeled data, the generated pseudo-labels may be noisy, affecting the accuracy of model training. On the other hand, the communication and synchronization of the model between different data sources may bring additional computing and bandwidth overhead, especially in high-resolution medical image segmentation tasks. In addition, the problem of data imbalance also exists. The number of samples in certain organs or lesion areas is small, which causes the model to be more inclined to identify categories with larger sample sizes during training and ignore the accurate segmentation of rare categories. To this end, common solutions include weighting the loss function, introducing data augmentation strategies, and designing more robust feature extraction modules.
[0008] In summary, federated semi-supervised medical image segmentation technology provides an efficient and accurate solution for medical image segmentation by protecting data privacy, making full use of unlabeled data and cross-institutional collaborative learning. The application of this technology is expected to play an important role in clinical diagnosis, lesion detection and surgical planning, promote the development of intelligent medical care, and further improve the diagnostic efficiency and treatment effect of patients. Summary of the invention
[0009] In view of the deficiencies of the prior art, the present invention provides a federated medical image segmentation method based on client consistency and adaptive perturbation to solve the background technical problems.
[0010] To achieve the above object, the present invention provides the following technical solution: a federated medical image segmentation method based on client consistency and adaptive perturbation, the method comprising the following contents:
[0011] Load case data: Get the heart disease segmentation database. The database names the folders with each patient's ID. Each folder stores the slice data of the visual 3D image of the case. The slice data of the visual 3D image is the 2D slices cut from the CT image along one dimension. Normalize the grayscale values of all slices of a case and record them as the original image. Then, obtain the corresponding correct segmentation image according to the case ID number as the training annotation and record it as the original image annotation.
[0012] Build and train a federated semi-supervised medical image global segmentation network based on client-side consistency and adaptive perturbation;
[0013] Construct a federated semi-supervised medical image global segmentation network based on client-to-client consistency and adaptive perturbation, including constructing a supervised client image segmentation network, an unsupervised client image segmentation network, and a server global network;
[0014] The original medical images of the loaded unsupervised client will first be sent to the unsupervised client image segmentation network, and a certain degree of Gaussian noise will be applied to the original medical images in the client as perturbation; the trained unsupervised network will be used as the regularization term of the supervised network together with the labels of the supervised client to optimize the supervised client network; after a certain number of rounds of local training, the two clients are synchronously uploaded to the server network and weightedly aggregated into a global network; the labeled data and unlabeled data are effectively used to construct an information-rich and highly generalized global segmentation network;
[0015] Construct an unsupervised client-side adaptive perturbation mechanism; apply a certain degree of Gaussian noise perturbation to the local image data of the unsupervised client; input the data before and after the perturbation into the unsupervised client-side image segmentation network at the same time to capture the relative feature information in the medical image; use a fixed step size to increase or decrease the perturbation intensity by locally verifying the model performance to adapt to the optimal perturbation intensity of each client with different learning difficulties; obtain an unsupervised client-side segmentation network with stronger performance;
[0016] Construct an interaction mechanism between clients; the supervised client uses the optimized local model parameters of the unsupervised client as the regularization term of local supervised learning to reduce the parameter distance between the supervised model and the unsupervised model; obtaining more similar local models will effectively reduce the model aggregation degradation problem caused by different optimization functions when performing global model aggregation.
[0017] Preferably, the specific implementation is as follows:
[0018] Global and local image segmentation networks; the structure uses the U-Net segmentation network, which aims to generate high-quality segmentation results, especially when dealing with complex structures and small sample data; through the symmetrical encoder-decoder structure, the contextual information in the image is effectively captured; the local network and the global network use the same network structure to achieve model aggregation;
[0019] Inter-client consistency strategy: The strategy uses the optimized unsupervised client network parameters and the supervised client local supervision data to synchronize and optimize the supervised client, effectively reducing the gap between the supervised client network parameters and the unsupervised client network parameters, reducing the model degradation problem during global model aggregation, and increasing the segmentation performance of the global model;
[0020] Adaptive perturbation strategy: Use consistency regularization to enable the unsupervised client to generalize local data. Specifically, the local data of the unsupervised client is perturbed to a certain extent and the output before and after the perturbation is made as consistent as possible. After each round of aggregation, the unsupervised client evaluates the performance of the local model and increases or decreases the perturbation intensity according to a fixed step size, so as to adapt the optimal perturbation intensity to local data images with different learning difficulties.
[0021] Preferably, the supervised client model training loss function is:
[0022] L SC =(1-γ)L SUP +γL CONS
[0023] Among them, L SC represents the total loss of the supervised client model training; γ represents the weight of the consistency loss; L SuP represents the supervision loss; L CONS represents the consistency loss between clients. The loss function for unsupervised client model training is:
[0024]
[0025] Among them, L UC represents the total loss of unsupervised client model training; N represents the total number of unsupervised client data samples; U c () indicates the output of the client; X i Indicates the original sample input in this round, X i represents the perturbed samples.
[0026] Preferably, the network structure of the federated semi-supervised medical image segmentation is: a symmetrical encoder-decoder structure; the overall network structure is divided into an encoding path (downsampling part) and a decoding path (upsampling part), and the two present a U-shaped symmetrical structure. The encoding path is mainly used to extract high-level semantic features of the image, and gradually reduces the spatial resolution through multiple [c3s1, P] structures, where [c3s1] represents two consecutive convolutional layers with a kernel size of 3×3 and a stride of 1, each of which is followed by batch normalization (BN) and ReLU activation functions, and [P] represents a maximum pooling layer with a pooling window size of 2×2 and a stride of 2, which gradually reduces the feature map resolution and doubles the number of channels. The decoding path gradually restores the image resolution through an upsampling (U) operation, and at the same time, the feature map of the corresponding layer in the encoding path is spliced with the upsampled feature map through a jump connection, thereby realizing the fusion of shallow detail features and deep semantic features. The spliced feature map is further extracted through [c3s1]. During the decoding process, the feature map resolution doubles layer by layer, while the number of channels is gradually halved until it is restored to the size of the input image. Finally, a 1×1 convolution is used to map the number of channels to the number of categories, and a Sigmoid or Softmax activation function is used to generate a pixel-level probability map. The entire network is a fully convolutional network (FCN) without a fully connected layer, and can process input images of any size.
[0027] Preferably, the medical image segmentation network structure is U-Net.
[0028] Preferably, a multimodal image data fusion strategy is introduced. If there are medical image data of multiple modalities, a fusion module is designed to fuse these data and input them into the network; the multimodal image data fusion strategy includes but is not limited to feature-level fusion, decision-level fusion or pixel-level fusion.
[0029] Preferably, an attention mechanism is introduced into the image segmentation network to enhance the network's ability to capture key features. The attention mechanism includes but is not limited to a spatial attention module, a channel attention module or a self-attention mechanism.
[0030] Preferably, a dynamic weight allocation mechanism is constructed to dynamically adjust the weight of each client model according to its performance and data quality during the global model aggregation process. The dynamic weight allocation mechanism adjusts the weight based on the verification accuracy, data quality indicators or other performance indicators of the client model.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] (1) Federated semi-supervised medical image global segmentation network based on client-to-client consistency and adaptive perturbation. The present invention proposes a new medical image segmentation method that combines supervised clients and unsupervised clients, and effectively integrates labeled data and unlabeled data through federated learning. Specifically, the supervised client and the unsupervised client use different image segmentation networks for training, and the unsupervised client applies medical images perturbed by Gaussian noise as a regularization term to optimize the segmentation network of the supervised client. This method can make up for the challenge of insufficient labeled data, especially in the field of medical imaging, where labeled data is often expensive and difficult to obtain. Through global aggregation and weighted averaging, the present invention can achieve high-quality medical image segmentation results with strong generalization ability, significantly improving the accuracy of segmentation and the robustness of the model. Compared with the prior art, the present invention can better utilize unlabeled data and reduce the reliance on manual annotation, so that medical image analysis can still achieve excellent performance in the case of insufficient data.
[0033] (2) Construction of an unsupervised client-side adaptive perturbation mechanism. In order to improve the performance of the unsupervised client-side image segmentation network, the present invention captures the relative feature information in the medical image by applying appropriate Gaussian noise perturbations. These perturbations not only increase the network's ability to recognize complex and diverse image features, but also enhance the robustness of the model so that it can adapt to the image data distribution of different clients. Through local verification performance evaluation, a fixed step size is used to adjust the perturbation intensity, so that each client can automatically adjust the perturbation intensity according to the learning difficulty of its local data. This adaptive mechanism optimizes the performance of the unsupervised client model, especially when faced with medical images with high noise or data imbalance, it can better capture potential disease characteristics. Therefore, compared with traditional methods, the present invention is more suitable for the diversity and complexity of medical imaging data, and improves the segmentation accuracy and the ability to recognize difficult-to-distinguish lesions.
[0034] (3) Innovation of the interaction mechanism between clients. In the present invention, a close cooperative relationship is established between the supervised client and the unsupervised client. By using the optimized local model parameters of the unsupervised client as the regularization term of the supervised client, the training process of the supervised model is guided. This mechanism effectively reduces the parameter gap between the supervised and unsupervised models, thereby reducing the model aggregation degradation problem caused by different optimization functions. When performing global model aggregation, the model parameters of all clients are weighted and integrated to form a more consistent and robust global model, which effectively improves the performance of the global segmentation network. Compared with traditional federated learning methods, the present invention optimizes parameter sharing and collaborative learning through model interaction between clients, ensures the quality of global model aggregation, and reduces the deviation and error caused by the distribution of different client data.
[0035] (4) Stronger generalization and accuracy of the global segmentation model. The present invention improves the performance of the medical image segmentation model in a variety of different data and tasks by combining collaborative optimization, perturbation enhancement and regularization strategies among multiple clients during the training process. Whether in terms of disease localization, morphological feature analysis or automatic identification of lesions, the present invention can provide more accurate segmentation results with high generalization capabilities. In addition, by utilizing the potential information provided by the unsupervised client, the model can explore more detailed disease characteristics, especially for some pathological features that are difficult to identify by conventional methods, and can provide more accurate diagnostic support.
[0036] (5) The interpretability of medical imaging reports is enhanced. Compared with the prior art, the present invention not only improves the segmentation accuracy, but also enhances the interpretability of model decisions. Through the combination of multi-view analysis and perturbation mechanism, the model can more accurately identify key areas in the image and provide doctors with more detailed diagnostic information. The patient's imaging features will be accurately located through the global network, and the generated medical imaging report will contain more details, allowing doctors to make more accurate diagnosis and treatment decisions based on the report.
[0037] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A block diagram of a federated medical image segmentation method based on client consistency and adaptive perturbation according to the present invention;
[0039] Figure 2 This is a flow chart of the federated medical image segmentation method based on client consistency and adaptive perturbation of the present invention. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this technical field without creative work are within the scope of protection of the present invention.
[0041] The present invention is a federated semi-supervised medical image segmentation method based on client consistency and adaptive perturbation. Figure 1 , including the following steps:
[0042] Step 1: Load case data and preprocess medical image data: Obtain a multi-medical site database for a certain disease. The case database names the folders with the ID of each different medical site, and each folder stores data sets with different feature distributions. Normalize and standardize the grayscale values of the image data of all views corresponding to a case first, and record them as original images; introduce a multimodal image data fusion strategy. If there are medical image data of multiple modalities, design a fusion module to fuse these data and input them into the network; multimodal image data fusion strategies include but are not limited to feature-level fusion, decision-level fusion, or pixel-level fusion.
[0043] Step 2: Based on the data loaded in step 1, distinguish the supervised client database of the labeled data set and the unsupervised client database without labels, and apply Gaussian perturbation to the unsupervised data to construct a perturbed database that corresponds one-to-one with the unsupervised client database. The specific perturbation method is:
[0044] G(x, y) = μ + P·Z
[0045] Among them, G(x, y) represents the disturbance value applied at the position (x, y); μ is the mean of the Gaussian distribution, which is set to 0; P is the standard deviation of the Gaussian distribution, which controls the intensity of the disturbance; Z is a random variable from a standard normal distribution with a mean of 0 and a variance of 1.
[0046] Step 3: Build a client image segmentation network. Construct a convolutional neural network and divide it into supervised client networks and unsupervised client networks. When building these networks, an attention mechanism is introduced to enhance the network's ability to capture key features. The attention mechanism can choose a spatial attention module, a channel attention module, or a self-attention mechanism. Then input the supervised dataset into the supervised client network, input the unsupervised dataset and the corresponding perturbation dataset into the unsupervised client network, calculate the image class activation map and obtain the corresponding loss. The loss of the unsupervised client is defined as:
[0047]
[0048] Among them, L UC represents the total loss of unsupervised client model training; N represents the total number of unsupervised client data samples; U C () indicates the output of the client; X i Indicates the original sample input in this round, X i represents the sample after perturbation;
[0049] Step 4: Perform feature interaction between clients. By transferring the model of the unsupervised client to the fully supervised client, we can leverage the knowledge learned by the unsupervised client through consistency regularization and organize this knowledge into the model of the fully supervised client. This helps the fully supervised client learn more robust feature representations, thereby improving the performance of the global model. By calculating the consistency loss, we can encourage the model of the fully supervised client to learn consistent representations based on the model of the unsupervised client. This helps to reduce the representation space differences between the unsupervised and fully supervised client models, making it easier to achieve model aggregation and fusion in the federated learning process. This knowledge distillation method helps to alleviate the representation space difference problem to a certain extent. In the early stages of federated learning, we use the fully supervised client to pre-train the global model, aiming to accelerate the convergence of the global model. The total loss of the supervised client is defined as:
[0050] L SC =(1-γ)L SUP +γL CONS
[0051] Among them, L SC represents the total loss of the supervised client model training; γ represents the weight of the consistency loss; L SUP represents the supervision loss; L CONS Indicates the loss of consistency between clients.
[0052] Step 5: Dynamically adaptive perturbation intensity. There are risks in directly incorporating unsupervised clients into global model aggregation, especially when using unsupervised clients trained with consistency regularization, which may easily lead to model collapse problems. Ensuring that unsupervised clients learn the correct knowledge as much as possible is a problem that must be solved. Consistency regularization is used to promote the local data of unsupervised clients. Specifically, by perturbing the local data of unsupervised clients, the output before and after the perturbation is made as consistent as possible. However, the degree of perturbation applied has a great impact on the training effect. If the perturbation is too light, it may increase the risk of model overfitting. The model may learn the details of the training data too confidently, resulting in overfitting on unseen data, reducing the robustness of the model, and making it difficult to handle the diversity in real-world scenarios. On the other hand, excessive perturbation may lead to training difficulties. Therefore, it is crucial to select appropriate perturbations for each dataset. In the federated learning paradigm, a key issue in improving the performance of global model aggregation is the feature shift problem. This means significant differences in complexity, details, and other information between different client data. In previous work, the same degree of perturbation was applied regardless of the federated learning paradigm, which improved the performance of the global model. This may be due to the selection of a good set of perturbations based on the performance of the global model. However, for each local model's dataset, the strength of the perturbation will affect the performance of the global model during the aggregation process because the generalization ability of the local model is incomplete. Therefore, it is necessary to adaptively adjust the perturbation strength for each unsupervised client. Specifically, in each federated training round, the unsupervised client evaluates whether the current perturbation strength is appropriate before and after training, and observes its performance on the validation set. Initially, the perturbation is set to a low level and tested on the validation set in each federated round. During each communication process, it is assumed that the perturbation is increased by default, and the performance of the perturbed model on the validation set is evaluated. If the performance deteriorates after the perturbation, it means that the current perturbation is too complex, and the relevant information of the increased perturbation will not be passed to the server. Instead, the original perturbation is kept and the best perturbation strength is recorded.
[0053] In federated learning, the data quality and quantity of different clients may vary significantly, which may lead to the underestimation or overestimation of the contribution of some clients in the global model aggregation process. It is necessary to dynamically adjust the weight of each unsupervised client based on its performance and data quality; specifically, after each federation round, an evaluation is performed based on the verification accuracy, data quality indicators or other performance indicators of each client model, and a weight value is calculated for each client model based on the evaluation results. The weight value can be determined based on the relative ranking of the verification accuracy, the score of the data quality indicator, or a comprehensive consideration of other performance indicators; in the global model aggregation process, a weighted average is performed based on the weight value of each client model to obtain an updated global model.
[0054] Step 6: Build a global model. Perform weighted aggregation of the supervised client network and the unsupervised client network. Specifically, map all parameters of the model one by one, calculate weights locally based on the locally verified performance, and sum the locally calculated weights multiplied by the values of the parameters themselves.
[0055] Step 7: Global network back-transmission updates the client network. After the global network is updated, the parameters are passed to all client network models for the next round of updates;
[0056] Step 8: Calculate the binary cross entropy loss to constrain the training process. During the training process, optimize all parameters of the model: adjust the weights of the supervised client network and the unsupervised client network in a direction that makes the binary cross entropy loss smaller, and thus obtain the trained supervised client network and unsupervised client network weights.
[0057] Example 1
[0058] This example performs medical image segmentation on the ACDC and Fundus datasets. The ACDC dataset contains 1902 cardiac region images and their corresponding true labels, which are divided into training, validation, and test sets in a ratio of 7:2:1. Specifically, 1331 images are assigned to the training set, 380 images to the validation set, and 191 images to the test set. The resolution of the ACDC images is fixed at 224×224 pixels. The ACDC dataset is divided into 4 parts, and each part is assigned to 4 virtual clients. The Fundus dataset uses retinal fundus images from 4 different clinical centers for optic disc and optic cup segmentation, simulating a real-world federated learning scenario. For preprocessing, an 800×800 disc region is cropped from these datasets, and the cropped region is resized to 384×384 pixels as network input. All data are preprocessed, and the ratio of training, validation, and test sets is also 7:2:1. The data are normalized separately so that the intensity values have zero mean and unit variance.
[0059] The following is a detailed description of each step and model parameter settings:
[0060] Step 1: Preprocessing of original images.
[0061] Since the data formats among different medical sites are not uniform, the grayscale values are first normalized to return the grayscale value range to [0, 1]. Then, the shape and size of each image in the ACDC database are adjusted to a single-channel medical image of [1, 224, 224], and the shape and size of each image in the Fundus database are adjusted to a single-channel medical image of [1, 384, 384].
[0062] Step 2: Dataset division.
[0063] The patient data of the corresponding sites were randomly shuffled, and the scikit-learn library was used to divide the data of the screened medical image description database, and ensure that there was no overlap between the parts: 70% of the data was divided into a training set for model training; 20% of the data was divided into a validation set for selecting the optimal model; and 10% of the data was divided into a test set for testing the effect of the model.
[0064] Step 3: Load multi-site medical imaging data.
[0065] Use the os module in Python to traverse all .png files in the folder where the data is stored, save the path, and then obtain the corresponding medical image based on the file name. Take the case stored in the folder named xray_0001 under the path ' / Fundus / data / ' as an example, first obtain the storage path of the case, that is, ' / Fundus / data / Fundus_0001', and then search for the .png file in the folder based on this path.
[0066] Step 4: Build an image segmentation network.
[0067] ① Build an unsupervised image segmentation network. Build U-Net as an unsupervised client image feature extraction network. The ACDC dataset inputs original medical image data of size [1, 224, 224], the Fundus dataset inputs original medical image data of size [1, 384, 384], the ACDC dataset inputs perturbed medical image data of size [1, 224, 224], and the Fundus dataset inputs perturbed medical image data of size [1, 384, 384]. Calculate the output difference to get the loss.
[0068] ②Build a supervised image segmentation network. Construct U-Net as a supervised client image feature extraction network. The ACDC dataset inputs original medical image data of size [1, 224, 224], and the Fundus dataset inputs original medical image data of size [1, 384, 384]. According to the corresponding labels of the data, the loss L1 is obtained, and the segmentation network model of the unsupervised client is obtained. The corresponding data is input into the network to obtain the output and the loss is calculated with the output of the local network to obtain the loss L2.
[0069] Step 5: Build an unsupervised client-side adaptive perturbation module.
[0070] After a round of learning, the unsupervised client will adjust the perturbation parameters of the Gaussian perturbation according to the performance of the local model. If the perturbation is too light, it may increase the risk of overfitting the model. The model may learn the details of the training data too confidently, resulting in overfitting on unseen data, reducing the robustness of the model, making it difficult to handle the diversity in real-world scenarios. On the other hand, excessive perturbation may lead to difficulties in training. Therefore, it is crucial to select appropriate perturbations for each data set. By dynamically adjusting this parameter, a more appropriate perturbation intensity can be adapted according to the learning difficulty of different data sets to achieve better learning results.
[0071] Step 6: Global model aggregation.
[0072] Finally, after each client has learned, the model parameters are matched according to the local performance verification effect, and then weighted summation is performed one by one to aggregate into a global model, which is passed to each client before the next round of learning begins.
[0073] This embodiment conducts experiments on the Fundus and ACDC databases respectively.
[0074] Four criteria were used to evaluate the models on the Fundus and ACDC datasets, including label-based DICE accuracy, 95% Hausdorff distance, mean symmetric surface distance (ASSD), and positive predictive value (PPV).
[0075] The method proposed in the present invention is compared with several classic methods, including LabelOnly, FedST, FedMatch, PDCFed, and IsoFed. The LabelOnly method indicates that only the data of the supervised client is used for training, which represents the lower limit of the performance in this federation process. Tables 1, 2, 3, and 4 respectively list the comparison of the accuracy of the method proposed in the present invention and other methods at the ratio of supervised client to unsupervised client of 1:3 and 2:2 on the Fundus and ACDC datasets, where ours is the method of the present invention.
[0076] Table 1 is a comparison of the performance of the Fundus dataset models with a supervised and unsupervised client ratio of 2:2.
[0077]
[0078] Table 2 is a comparison of the performance of the Fundus dataset models with a supervised and unsupervised client ratio of 1:3.
[0079]
[0080]
[0081] Table 3 is a comparison of the performance of the ACDC dataset models with a supervised and unsupervised client ratio of 2:2.
[0082]
[0083] Table 4 is a comparison of the performance of the ACDC dataset models with a supervised and unsupervised client ratio of 1:3
[0084]
[0085]
[0086] The Dice coefficient is a measure of the similarity between two sets, and is usually used to evaluate the overlap between the predicted segmentation and the true label. Its value range is 0 to 1. The higher the value, the more overlap the predicted result has with the true label, and the better the segmentation effect.
[0087] HD95 is an indicator to measure the similarity of segmentation boundaries. Specifically, it calculates the 95th percentile value of the maximum distance between the predicted segmentation boundary and the true boundary. It reflects how far the predicted segmentation is from the true boundary in space. The smaller the value of HD95, the closer the segmentation result is to the true boundary. This indicator can reveal local differences in segmentation boundaries, so it is very important for evaluating the accuracy of boundaries.
[0088] ASSD is an indicator for evaluating the average distance between the predicted segmentation surface and the true segmentation surface. It calculates the average distance between the predicted surface and the true surface, and takes into account the symmetry between the two. ASSD comprehensively considers the accuracy of the segmentation boundary. The smaller the ASSD value, the closer the predicted result is to the true surface, indicating that the model has a higher accuracy on the segmentation boundary.
[0089] PPV (Positive Predictive Value) measures the proportion of samples that are actually positive (real lesion areas) among all samples predicted as positive (for example, predicted as lesion areas) by the model. It reflects the accuracy of the model, especially the accuracy of the model in predicting positive classes. The higher the PPV value, the more accurate the model's prediction is, and the fewer samples are misclassified as positive.
[0090] It can be seen from the table that the method of the present invention achieves State-of-the-Art (SOTA) performance in the label-based metric Dice in all datasets of all scales; in HD95, ASSD and PPV, the method of the present invention also achieves State-of-the-Art (SOTA) performance on the ACDC dataset and the Fundus dataset, which shows that the method of the present invention has good adaptability to different types of datasets.
[0091] The method of the present invention uses consistency between clients to improve the aggregation efficiency and model performance of the global model aggregation. Secondly, the method of the present invention optimizes the degree of disturbance in the learning process for the unsupervised client learning process, and improves the generalization ability and performance stability of the model. The combination of these technologies enables the method of the present invention to achieve good results on different data sets, showing its versatility and scalability.
[0092] Among them, FedST, FedMatch, PDCFed, and IsoFed are all well known in the art, and the LabelOnly algorithm also presents the lower limit of the algorithms in the art by using only the label client, thereby enhancing the feasibility of the algorithm.
[0093] This application provides a new solution for the field of federated semi-supervised medical image segmentation. By comprehensively considering model aggregation and optimizing the design of the unsupervised client learning process, multiple modules are organically combined to form an end-to-end federated semi-supervised medical image segmentation framework, which successfully addresses the challenges faced by existing methods.
Claims
1. A federated medical image segmentation method based on client consistency and adaptive perturbation, characterized in that: The method includes the following: Load case data: Get the heart disease segmentation database. The database names the folders with each patient's ID. Each folder stores the slice data of the visual 3D image of the case. The slice data of the visual 3D image is the 2D slices cut from the CT image along one dimension. Normalize the grayscale values of all slices of a case and record them as the original image. Then, obtain the corresponding correct segmentation image according to the case ID number as the training annotation and record it as the original image annotation. Build and train a federated semi-supervised medical image global segmentation network based on client-side consistency and adaptive perturbation; Construct a federated semi-supervised medical image global segmentation network based on client-to-client consistency and adaptive perturbation, including constructing a supervised client image segmentation network, an unsupervised client image segmentation network, and a server global network; The original medical images of the loaded unsupervised client will first be sent to the unsupervised client image segmentation network, and a certain degree of Gaussian noise will be applied to the original medical images in the client as perturbation; the trained unsupervised network will be used as the regularization term of the supervised network together with the labels of the supervised client to optimize the supervised client network; after a certain number of rounds of local training, the two clients are synchronously uploaded to the server network and weightedly aggregated into a global network; the labeled data and unlabeled data are effectively used to construct an information-rich and highly generalized global segmentation network; Construct an unsupervised client-side adaptive perturbation mechanism; apply a certain degree of Gaussian noise perturbation to the local image data of the unsupervised client; input the data before and after the perturbation into the unsupervised client-side image segmentation network at the same time to capture the relative feature information in the medical image; use a fixed step size to increase or decrease the perturbation intensity by locally verifying the model performance to adapt to the optimal perturbation intensity of each client with different learning difficulties; obtain an unsupervised client-side segmentation network with stronger performance; Construct an interaction mechanism between clients; the supervised client uses the optimized local model parameters of the unsupervised client as the regularization term of local supervised learning to reduce the parameter distance between the supervised model and the unsupervised model; obtaining more similar local models will effectively reduce the model aggregation degradation problem caused by different optimization functions when performing global model aggregation.
2. The method for federated medical image segmentation based on client consistency and adaptive perturbation according to claim 1, characterized in that: The specific implementation method is as follows: Global and local image segmentation networks; the structure uses the U-Net segmentation network, which aims to generate high-quality segmentation results, especially when dealing with complex structures and small sample data; through the symmetrical encoder-decoder structure, the contextual information in the image is effectively captured; the local network and the global network use the same network structure to achieve model aggregation; Inter-client consistency strategy; The strategy uses the optimized unsupervised client network parameters and the supervised client local supervision data to optimize the supervised client synchronously, effectively reducing the gap between the supervised client network and the unsupervised client network parameters, reducing the model degradation problem when the global model is aggregated, and increasing the segmentation performance of the global model; Adaptive perturbation strategy: Use consistency regularization to enable the unsupervised client to generalize local data. Specifically, the local data of the unsupervised client is perturbed to a certain extent and the output before and after the perturbation is made as consistent as possible. After each round of aggregation, the unsupervised client evaluates the performance of the local model and increases or decreases the perturbation intensity according to a fixed step size, so as to adapt the optimal perturbation intensity to local data images with different learning difficulties.
3. The method for federated medical image segmentation based on client consistency and adaptive perturbation according to claim 1, characterized in that: The supervised client model training loss function is: L SC =(1-γ)L SUP +γL CONS Among them, L SC represents the total loss of the supervised client model training; γ represents the weight of the consistency loss; L SUP represents the supervision loss; L CONS represents the consistency loss between clients. The loss function for unsupervised client model training is: Among them, L UC represents the total loss of unsupervised client model training; N represents the total number of unsupervised client data samples; U C () indicates the output of the client; X i Indicates the original sample input in this round, X i ′ represents the sample after perturbation.
4. The method for federated medical image segmentation based on client consistency and adaptive perturbation according to claim 2, characterized in that: The network structure of the federated semi-supervised medical image segmentation is: a symmetrical encoder-decoder structure; the overall network structure is divided into an encoding path (downsampling part) and a decoding path (upsampling part), and the two present a U-shaped symmetrical structure. The encoding path is mainly used to extract high-level semantic features of the image, and gradually reduces the spatial resolution through multiple [c3s1, P] structures, where [c3s1] represents two consecutive convolutional layers with a kernel size of 3×3 and a stride of 1, each followed by batch normalization (BN) and ReLU activation functions, and [P] represents a maximum pooling layer with a pooling window size of 2×2 and a stride of 2, which gradually reduces the feature map resolution and doubles the number of channels. The decoding path gradually restores the image resolution through upsampling (U) operations, and at the same time, the feature map of the corresponding layer in the encoding path is spliced with the upsampled feature map through jump connections, thereby realizing the fusion of shallow detail features and deep semantic features. The spliced feature map is further extracted through [c3s1]. During the decoding process, the feature map resolution doubles layer by layer, while the number of channels is gradually halved until it is restored to the size of the input image. Finally, a 1×1 convolution is used to map the number of channels to the number of categories, and a sigmoid activation function is used to generate a pixel-level probability map. The entire network is a fully convolutional network (FCN) without a fully connected layer, and can process input images of any size.
5. The method for federated medical image segmentation based on client consistency and adaptive perturbation according to claim 2, characterized in that: Its medical image segmentation network structure is U-Net.
6. The method for federated medical image segmentation based on client consistency and adaptive perturbation according to claim 1, characterized in that: A multimodal image data fusion strategy is introduced. If there are medical image data of multiple modalities, a fusion module is designed to fuse these data and input them into the network; the multimodal image data fusion strategy includes but is not limited to feature-level fusion, decision-level fusion or pixel-level fusion.
7. The method for federated medical image segmentation based on client consistency and adaptive perturbation according to claim 1, characterized in that: An attention mechanism is introduced into the image segmentation network to enhance the network's ability to capture key features. The attention mechanism includes but is not limited to a spatial attention module, a channel attention module or a self-attention mechanism.
8. The method for federated medical image segmentation based on client consistency and adaptive perturbation according to claim 1, characterized in that: Build a dynamic weight allocation mechanism to dynamically adjust the weight of each client model according to its performance and data quality during the global model aggregation process. The dynamic weight allocation mechanism adjusts the weight based on the verification accuracy, data quality indicators or other performance indicators of the client model.
Citation Information
Patent Citations
Federal learning-based multi-task mixed supervision medical image segmentation method and system
CN115187783A
Medical image segmentation method based on multi-mode self-supervision
CN115601352A
Novel federal semi-supervised learning lung tumor segmentation method and device
CN116310330A
MR medical image colorectal cancer segmentation method and system based on semi-supervised learning
CN117409201A
Semi-supervised medical image segmentation method based on cascade parallel CNN-Transformer and full disturbance
CN117456177A
Cited By
Privacy protection medical image semantic segmentation method for dynamic data flow
CN122090060A
A privacy protection medical image semantic segmentation method for dynamic data flow
CN122090060B