A medical image segmentation method based on federated learning and attention mechanism
Through the combination of federated learning and attention mechanism, the joint training of multi-party medical image data is achieved, the problem of insufficient data in a single institution is solved, the generalization ability and privacy protection of the model are improved, and the medical image segmentation effect is significantly improved.
Patent Information
- Application Number
- CN202310116822.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-02-15
AI Technical Summary
Due to its privacy and scarcity, medical image data is difficult for a single institution to obtain sufficient data for effective training of deep learning models, which in turn affects the prediction ability of the model.
The medical image segmentation method based on federated learning and attention mechanism is adopted to realize multi-party joint model training through the collaborative work of the central server and the local server. The specific steps include the central server sending the initial global model to the local server, each local server trains the local model based on its own data and uploads parameters, the central server performs model aggregation and optimization to generate the global model, and feeds it back to the local server to update the local model.
It effectively improves the generalization ability of deep learning models, ensures the security of patient privacy data, and the segmentation effect of the model is significantly better than that of unilateral data training models.
Smart Images

Figure CN116245886B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and medical image processing technology, and in particular, to a medical image segmentation method based on federated learning and attention mechanism. Background Art
[0002] Medical images are private, small in number, and difficult to obtain, so it is very difficult to train models related to medical images. On the one hand, because medical images involve patient privacy, the security of the medical image data source must be guaranteed when using medical images for learning. It is precisely because of the privacy of medical images that a single institution can only obtain limited medical image data for deep learning, which cannot meet the training requirements of existing image processing models, resulting in poor prediction capabilities of the model and an inability to provide strong data support in actual diagnostic scenarios. Summary of the invention
[0003] Based on the above-mentioned technical problem of limited training data for medical image processing models, a medical image segmentation method based on federated learning and attention mechanism is provided. The present invention combines multiple parties to train models. After each party optimizes the model based on its own data, the information of each party's model is integrated, so that each party can obtain a global model that has learned all user data sets, effectively improving the generalization ability of the deep learning model.
[0004] The technical means adopted by the present invention are as follows:
[0005] A medical image segmentation method based on federated learning and attention mechanism is applied to a distributed image processing system, wherein the distributed image processing system includes a central server and a plurality of local servers holding different local data. The method includes the following steps:
[0006] S1. The central server sends an initial global model to each of the local servers; the initial global model is an Attention U-Net model, including an encoder with a five-layer path and a decoder with a five-layer path;
[0007] S2. The local server trains the received initial model based on its own medical image data, and uploads the generated local model parameters after training to the central server; each local model places the federated fusion information layer feature map into the decoder structure corresponding to the local model for iterative learning at the decoder stage, and the federated fusion information layer feature map is obtained according to the following method: in the decoder with five-layer paths of the global model, each layer path performs a concat operation with the feature map transmitted from the encoder of the same layer while performing an upsampling operation, thereby generating a federated fusion information layer feature map;
[0008] S3. The central server generates a global model by performing model aggregation optimization based on the parameters of each local model, wherein the model aggregation optimization includes weighted averaging the model parameters of each network level of all local models;
[0009] S4, repeatedly perform steps S2-S3 until the loss function value and the number of training rounds of the global model reach the preset conditions, then terminate the training;
[0010] S5. The central server feeds back the global model parameters to each local server. Each local server replaces the local model parameters with the global model parameters to update the local model, and performs medical image segmentation based on the updated local model.
[0011] Furthermore, in the encoder structure of the five-layer path in the initial global model, each layer of the path undergoes two sets of feature extraction operations, each set of feature extraction operations undergoes 3*3 convolution, BatchNormlization, and ReLu activation function, and a maximum pooling layer with a convolution kernel of 2*2 is used between each layer to downsample the upper feature image;
[0012] The decoder structure of the five-layer path needs to undergo an upsampling operation on each layer, and then receive the feature map transmitted by the encoder at the same level for information fusion. The convolution operation in the decoder is the same as the convolution operation in the encoder, so that the feature map is sequentially subjected to 3*3 convolution, BatchNormlization, and ReLu activation function. The decoder needs to undergo a 1*1 convolution operation before outputting the model prediction result.
[0013] Furthermore, each layer structure of the decoder in the initial global model includes an attention module, and the input of the attention module is the feature map extracted by the encoding path of the same layer and the feature map output by the decoding path of the previous layer.
[0014] Furthermore, the output of the attention module is obtained according to the following method:
[0015] The feature map extracted by the encoding path of the same layer and the feature map output by the decoding path of the previous layer are used as the input of the attention module. The attention coefficient is obtained by 1*1 convolution, ReLu function, Sigmoid function and resampling operation.
[0016] The attention coefficient is multiplied pixel by pixel with the feature map output by the previous layer decoding path to obtain the output feature map of the attention module.
[0017] Furthermore, the central server sends the initial global model to each of the local servers, including sending training parameters to each of the local servers, wherein the training parameters include local iteration rounds, global aggregation rounds, initialization learning rate, optimizer, optimizer parameters and batch size.
[0018] Furthermore, the local server trains the received initial model based on its own medical image data, including: preprocessing the local data, wherein the preprocessing includes image flipping, image rotation, and image cropping; and training the local model based on the preprocessed image according to the received training parameters.
[0019] Compared with the prior art, the present invention has the following advantages:
[0020] The present invention can protect the privacy of medical image data of patients held by all parties from being leaked while performing joint computations by multiple parties. It solves the problem of multi-party joint model training in a data island state. The effect of the established federated learning model is significantly better than the effect of model training based on unilateral data. The U-Net medical image segmentation algorithm based on federated learning and attention mechanism proposed in the present invention has better medical image segmentation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0022] Figure 1 This is a schematic diagram of the model of a medical image segmentation method based on federated learning and attention mechanism of the present invention.
[0023] Figure 2 This is a model training architecture diagram of a medical image segmentation method based on federated learning and attention mechanism in the present invention.
[0024] Figure 3 This is the U-Net network structure diagram.
[0025] Figure 4 This is a schematic diagram of the parallel dual-path attention mechanism module added to the U-Net network in the present invention. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0027] The present invention discloses a medical image segmentation method based on federated learning and attention mechanism, which is applied to a distributed image processing system, wherein the distributed image processing system includes a central server and a plurality of local servers holding different local data. The central server is used to issue an initial global model and perform model fusion according to local model parameters to generate a global model. The local server is used to perform model training based on its own data to generate a local model.
[0028] For the global model located in the central server, it adopts an encoder-decoder structure, in which the encoder and decoder are symmetrical structures, such as Figure 1 As shown in the figure, the encoder layer and decoder layer of the same level transfer information through information fusion operation. During the transfer process, the attention module is used, which can make the network model pay more attention to the valuable parts of the input image information and ignore the parts with low relevance to the results. In the encoder, the network of the same level needs to undergo convolution operation, batch normativation operation and ReLU function, and then pass through the pooling layer to compress the feature image information. The pooling layer can also reduce the size of the feature image. The reduced feature image reduces the number of parameters required for the subsequent network. In the encoder stage, the input feature map is restored to the size of the upper layer feature map through the upsampling operation. This operation can prepare for the subsequent feature fusion. Another input of the attention mechanism is the feature map output by the decoder. In the final stage of the decoder, a 1x1 convolution operation is used to perform binary classification for each pixel, so that the output of the entire network is a binary image.
[0029] For the local model located on the local client, its main learning framework is similar to the global model, and both have an encoder structure and a decoder structure. The difference between the local model and the global model is that the local model needs the assistance of the global model during the training process. Specifically, when the decoder stage of the local model performs feature image recovery, it not only samples the feature map from its own network, but also accepts the feature map inferred by the global model for the current medical image at the same level of network structure, that is, Figure 1 After the local model receives the feature map of the federated fusion information layer, it adds it to the feature map output by the attention module at the current level. This operation enables the local model to acquire knowledge from the global model during training.
[0030] Based on the above model structure, the present invention discloses a medical image segmentation method based on federated learning and attention mechanism, such as Figure 2 As shown, the following steps are included:
[0031] S1. The central server sends an initial global model to each of the local servers; the initial global model is an improved Attention U-Net model, including an encoder with a five-layer path and a decoder with a five-layer path.
[0032] Specifically, it is assumed that there are n participants (n>=2). In this embodiment, it is assumed that n=2, that is, there are two participants A and B and a trusted third-party server C. The central server C sends the initialized model to participants A and B, and sends the parameters of model training to participants A and B, including the local iteration round h, the global aggregation round t, the initialized learning rate r, the optimizer Adam, the optimizer parameters β1 and β2, and the batch size.
[0033] This application mainly adopts the improved U-Net network as the deep learning network, which has the advantages of simple network structure, small network model size and fast training speed. Based on this, the present invention uses the attention mechanism module to enhance the U-Net network's ability to perceive local information, effectively improving the segmentation performance of the federated U-Net model. Figure 3 As shown, it includes a contraction path and an expansion path, which can also be referred to as an encoder and a decoder. The present invention refers to an encoder and a decoder structure. In the encoder structure, it is divided into 5 layers of paths, each layer of paths passes through two sets of feature extraction operations respectively, and each set of feature extraction operations passes through 3*3 convolution, BatchNormlization, and ReLu activation functions respectively. The upper feature image is downsampled using a maximum pooling layer with a convolution kernel of 2*2 between each layer. This operation can shrink the image size to half of the original size and increase the number of channels to twice the original size. In the decoder, it is also divided into 5 layers of paths, each layer of paths needs to undergo an upsampling operation, and then the feature map transmitted by the encoder at the same level is subjected to information fusion, and the fusion method is to perform a concat operation. The convolution operation in the decoder is the same as the convolution operation in the encoder, and they all let the feature map pass through 3*3 convolution, BatchNormlization, and ReLu activation functions in sequence. The decoder needs to undergo a 1*1 convolution operation before outputting the model prediction result, so that the feature vector learned by the entire network is converted into a possible classification category.
[0034] The present invention adds a parallel dual-path attention mechanism module to the traditional U-Net network. Its working principle is as follows: Figure 4As shown in the figure, assume that the two inputs of the parallel dual-path attention mechanism module are x and g, where g is the deep network output of the U-Net network structure, and x is the network output of the previous layer of g. They all need to go through 1*1 convolution, ReLu function, Sigmoid function and resampling operation to finally get the attention coefficient α. Then, the input x is passed through the pooling layer and the fully connected layer in turn to get the attention coefficient μ. Finally, the attention coefficient α is multiplied element by element with μ and the input feature map x, and the formula is as follows:
[0035]
[0036] α=σ 2 [m(x,g)] (2)
[0037]
[0038]
[0039] where σ 1 Represents the ReLU activation function, σ 2 Represents the Sigmoid activation function, b x 、b g With b β Is the bias term of the convolution operation. Input x, g pass through W x With W g The convolution operation represented by is added, and then the convolution operation represented by β is performed after the RelU activation function to obtain the output m, as shown in formula (1). The attention coefficient α is obtained by passing m through the Sigmoid activation function, as shown in formula (2). The input x is pooled at the channel level, where C x Represents the characteristic pixels of each channel of input x, average(C x ) indicates channel-level average pooling. Then the output of the pooling operation is passed through the fully connected layer to obtain the attention coefficient μ, where represents a fully connected operation, b f is the bias term of the fully connected layer, as shown in formula (3). Finally, the input feature map x is multiplied element by element with the attention coefficients α and μ to obtain the output feature map x' as shown in formula (4).
[0040] S2. The local server trains the received initial model based on its own medical image data, and uploads the generated local model parameters after training to the central server; each local model places the federated fusion information layer feature map into the decoder structure corresponding to the local model for iterative learning in the decoder stage, and the federated fusion information layer feature map is obtained according to the following method: in the decoder with five-layer paths of the global model, each layer of the path performs a concat operation with the feature map transmitted from the encoder of the same layer while performing an upsampling operation, thereby generating a federated fusion information layer feature map.
[0041] Specifically, after each participant obtains the model sent by the central server, it performs local training data preprocessing operations, including but not limited to image flipping, image rotation, and image cropping. Participant A and Participant B perform local model training on their respective local data sets. During training, the size of each batch of data is the same as the Batch size sent by the central server. Each participant's model is iterated locally for h rounds, and then the iterated local model client model is uploaded to the central server for aggregation. In the decoder stage, the local client model uses the center server model to learn the feature map of the federated fusion information layer, and places the feature map into the decoder structure corresponding to the client model for iterative learning.
[0042] S3. The central server generates a global model by performing model aggregation optimization based on the parameters of each local model. The model aggregation optimization includes weighted averaging the model parameters of each network level of all local models.
[0043] Specifically, the central server receives the client models of the current round iteration of each participant, and then performs a model aggregation operation, which performs a weighted average of the model parameters of all client models at each network level. Finally, the center server model after aggregation in the current round is obtained.
[0044] S4. Repeat steps S2-S3 until the loss function value of the global model and the number of training rounds reach the preset conditions, then terminate the training.
[0045] Specifically, after updating the center sever model, the center model determines whether to stop training based on the loss function value and the number of training rounds.
[0046] S5. The central server feeds back the global model parameters to each local server. Each local server replaces the local model parameters with the global model parameters to update the local model, and performs medical image segmentation based on the updated local model.
[0047] Specifically, after completing global training, the central server sends the center server model to each participant. Each participant uses the model sent by the center server to perform subsequent segmentation tasks.
[0048] The present invention will be further described below with reference to examples.
[0049] Suppose that there are multiple medical institutions holding medical image data, and the data they have has certain specificity. For example, medical institution A is located in northern China and medical institution B is only in southern China. Therefore, the patients received by these two medical institutions have typical north-south differences. If medical institution A wants to better serve patients from the southern region, then it needs more samples from patients in the south. The traditional solution is that medical institution A directly purchases data from medical institution B. However, at this stage, all countries are proposing privacy protection bills, so traditional solutions are no longer feasible. The present invention can perform joint modeling work among multiple medical institutions under the premise of ensuring that user data is not leaked. By using the algorithm proposed by the present invention, a medical image segmentation algorithm can be trained using data set samples from multiple users.
[0050] First, medical institution A and medical institution B (or more participants) prepare their own available medical data. After the trusted third-party server confirms the model to be trained, medical institution A and medical institution B use their local medical data sets to train the local segmentation model of the asset after receiving the model. After the specified local model iteration rounds are completed, a global model aggregation operation is performed, that is, medical institution A and medical institution B simultaneously upload their trained local models to the trusted third-party server. Since the communication between the local medical institution and the central server is only the transmission model, there is no problem of user data leakage in this process. After the central server receives the local models from all users, it performs model aggregation and model distribution, and a global iteration ends. Repeat the above steps until the model converges.
[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A medical image segmentation method based on federated learning and attention mechanism, applied to a distributed image processing system, wherein the distributed image processing system includes a central server and a plurality of local servers holding different local data. It is characterized in that The method comprises the following steps: S1. The central server sends an initial global model to each of the local servers; the initial global model is an Attention U-Net model, including an encoder with a five-layer path and a decoder with a five-layer path; S2. The local server trains the received initial model based on its own medical image data, and uploads the generated local model parameters after training to the central server; each local model places the federated fusion information layer feature map into the decoder structure corresponding to the local model for iterative learning at the decoder stage, and the federated fusion information layer feature map is obtained according to the following method: in the decoder with five-layer paths of the global model, each layer path performs a concat operation with the feature map transmitted from the encoder of the same layer while performing an upsampling operation, thereby generating a federated fusion information layer feature map; S3. The central server generates a global model by performing model aggregation optimization based on the parameters of each local model, wherein the model aggregation optimization includes weighted averaging the model parameters of each network level of all local models; S4, repeatedly perform steps S2-S3 until the loss function value and the number of training rounds of the global model reach the preset conditions, then terminate the training; S5. The central server feeds back the global model parameters to each local server. Each local server replaces the local model parameters with the global model parameters to update the local model, and performs medical image segmentation based on the updated local model.
2. According to claim 1, a medical image segmentation method based on federated learning and attention mechanism, It is characterized in that The encoder structure of the five-layer path in the initial global model, each layer of the path undergoes two sets of feature extraction operations, each set of feature extraction operations undergoes 3*3 convolution, BatchNormlization, ReLu activation function, and a maximum pooling layer with a convolution kernel of 2*2 is used between each layer to downsample the upper feature image; The decoder structure of the five-layer path needs to undergo an upsampling operation on each layer, and then receive the feature map transmitted by the encoder at the same level for information fusion. The convolution operation in the decoder is the same as the convolution operation in the encoder, so that the feature map is sequentially subjected to 3*3 convolution, BatchNormlization, and ReLu activation function. The decoder needs to undergo a 1*1 convolution operation before outputting the model prediction result.
3. According to claim 2, a medical image segmentation method based on federated learning and attention mechanism, It is characterized in that Each layer structure of the decoder in the initial global model includes an attention module, and the input of the attention module is the feature map extracted by the encoding path of the same layer and the feature map output by the decoding path of the previous layer.
4. According to claim 3, a medical image segmentation method based on federated learning and attention mechanism, It is characterized in that The output of the attention module is obtained as follows: The feature map extracted by the encoding path of the same layer and the feature map output by the decoding path of the previous layer are used as the input of the attention module. The attention coefficient is obtained by 1*1 convolution, ReLu function, Sigmoid function and resampling operation. The attention coefficient is multiplied pixel by pixel with the feature map output by the previous layer decoding path to obtain the output feature map of the attention module.
5. The medical image segmentation method based on federated learning and attention mechanism according to claim 1, It is characterized in that The central server sends the initial global model to each of the local servers. Including sending training parameters to each local server, the training parameters including local iteration rounds, global aggregation rounds, initialization learning rate, optimizer, optimizer parameters and batch size.
6. The medical image segmentation method based on federated learning and attention mechanism according to claim 4, It is characterized in that The local server trains the received initial model based on its own medical image data, including: preprocessing the local data, wherein the preprocessing includes image flipping, image rotation and image cropping; and training the local model based on the preprocessed image according to the received training parameters.
Citation Information
Patent Citations
CT image detection method based on federated learning and related device
CN112508907A
Multi-center brain tumor prognosis lifetime prediction method and system based on federated learning
CN113571203A