A Medical Image Segmentation Method Based on Federated Learning and Improved Pooling Pyramid
Through federated learning and improving the U-Net network structure of the pooled pyramid, the problem of insufficient user privacy leakage and segmentation accuracy in medical image segmentation is solved, and privacy protection and high-precision segmentation are achieved under joint modeling of multiple participants.
Patent Information
- Application Number
- CN202310421906.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-19
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-04-19
AI Technical Summary
In medical image segmentation jointly modeled by multiple participants, the existing technology has the problem of user privacy data leakage, and the model training effect is poor, especially in the case of uneven data distribution and regional differences, the segmentation accuracy is insufficient.
The federated learning framework is adopted to combine the improved pooled pyramid U-Net network structure, and the model parameters are initialized through the central server, and each participant performs local training and aggregation, and the improved pooled pyramid module is used to perform feature map connection and attention mechanism to protect privacy data from leakage and improve segmentation accuracy.
It realizes the accuracy of medical image segmentation and model generalization capabilities while protecting user privacy, and solves the problem of multi-party joint modeling under the data island state.
Smart Images

Figure CN116452806B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a medical image segmentation method based on federated learning and improved pooling pyramid. Background Art
[0002] Medical image data is difficult to obtain and is strongly interfered by devices. Therefore, the model training related to medical images is relatively difficult. Due to the increasing attention to data privacy in various countries in recent years and considering the privacy characteristics of medical images, special attention needs to be paid to the problem of user privacy leakage in medical image-related work. Therefore, the multi-party joint training method with privacy protection capabilities is becoming more and more worthy of research.
[0003] The multi-party joint training medical image segmentation method is a popular research direction. Under the constraints of multi-party joint modeling and user privacy protection, there are still limitations in this field. If the data sets of each training participant are directly aggregated for training, it will lead to the interaction and flow of data between each participant, and then lead to the problem of data leakage. Summary of the Invention
[0004] In view of the technical problem of user privacy data leakage caused by aggregating medical images from all parties for model training, the present invention provides a medical image segmentation method based on federated learning and improved pooling pyramid, which overcomes the problem of privacy leakage through the method of federated learning, and at the same time improves the segmentation accuracy of medical images based on the improved pooling pyramid method.
[0005] The technical means adopted by the present invention are as follows:
[0006] A medical image segmentation method based on federated learning and improved pooling pyramid, comprising the following steps:
[0007] S1. The central server issues the initialized neural network model parameters, and the neural network model is used for medical image segmentation;
[0008] S2. Each participant receives the initialized neural network model parameters and uses the initialized neural network model parameters as the local neural network model parameters;
[0009] S3. Train the local neural network model based on the local data and generate new local neural network model parameters;
[0010] S4. Each participant uploads their respective new local neural network model parameters to the central server, and the central server performs aggregation processing on the neural network model parameters to generate aggregated neural network model parameters;
[0011] S5. The central server distributes the aggregated network parameters to each participant, and each participant updates the local neural network model parameters using the aggregated neural network model parameters;
[0012] S6. Each participant determines whether the local model has converged based on the current local neural network model parameters. If it has converged, the participant saves the current local neural network model parameters; otherwise, it returns to S3;
[0013] S7. Each participant performs medical image segmentation based on the converged local neural network model.
[0014] Furthermore, the neural network model adopts an improved pooling pyramid U-Net network structure. The improved pooling pyramid U-Net network structure adopts a symmetric encoder-decoder structure. For the feature map transfer between encoder layers, an improved pooling pyramid is used for connection, and the encoder layer and decoder layer of the same level are connected through skip connections.
[0015] Furthermore, the improved pooling pyramid applies max pooling and adaptive average pooling operations. The adaptive average pooling operation outputs the feature map to four scale sizes, and the max pooling takes the maximum value of the feature map within the kernel range.
[0016] Furthermore, the improved pooling pyramid U-Net network structure adds an attention mechanism module to the bottom and top structures.
[0017] Furthermore, the adaptive average pooling operation outputs the feature map to four scale sizes, including 1*1, 2*2, 3*3, and 4*4 scale sizes.
[0018] Furthermore, the feature map output after max pooling, the 1*1 scale feature map, the 2*2 scale feature map, the 3*3 scale feature map, and the 4*4 scale feature map are successively subjected to bilinear interpolation operations and then superimposed on the original feature map to obtain the output feature map.
[0019] Compared with the prior art, the present invention has the following advantages:
[0020] 1. The present invention can solve the problem of multi-party joint modeling in the state of data islands. The model established using the federated learning framework has better performance than the single-party training model, and at the same time protects the user's private data from being leaked. The model update method used in the present invention can ensure the model accuracy in the case of multi-party joint training, so that the medical image segmentation algorithm based on federated learning and improved pooling pyramid proposed by the present invention has better medical image segmentation effect.
[0021] 2. To enable the model obtained from federated learning training to have stronger generalization ability, it is necessary to make the local model have good prediction ability. Therefore, the present invention proposes a medical image segmentation model based on U-Net, designs an improved version of the pooling pyramid module, and then federates the U-Net network with the improved pooling pyramid model. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a model training flowchart of the medical image segmentation method based on federated learning and improved pooling pyramid of the present invention.
[0024] Figure 2 It is a structure diagram of the U-Net network.
[0025] Figure 3 It is a schematic diagram of the improved pooling pyramid module structure in the embodiment of the present invention.
[0026] Figure 4 It is a structure diagram of the U-Net network with an improved pooling pyramid in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] As Figure 1 shown, the present invention provides a medical image segmentation method based on federated learning and improved pooling pyramid, including the following steps:
[0029] S1. The central server distributes the initialized neural network model parameters, and the neural network model is used for medical image segmentation.
[0030] Suppose there are n parties for joint training. In this example, there are two parties, A and B, and the trusted third-party server is C. First, the central server C initializes the model parameters to be trained, and then sends the initialized model to parties A and B. The central server also needs to specify the following parameters, including the local iteration rounds h, the global aggregation rounds t, the initialized learning rate r, the optimizer Adam, and the Batch size.
[0031] S2. Each party receives the initialized neural network model parameters and uses the initialized neural network model parameters as the local neural network model parameters.
[0032] Parties A and B obtain the initialized model and start training. During local training, each party performs preprocessing operations on the local training data, including but not limited to image flipping and image cropping.
[0033] S3. Train the local neural network model based on the local data and generate new local neural network model parameters.
[0034] Parties A and B first perform local training on their own datasets, using the Batch size sent by the central server during training. Each party's model iterates h times locally, and then uploads its local model to the central server.
[0035] S4. Each party uploads its new local neural network model parameters to the central server, and the central server performs aggregation processing on the neural network model parameters of each party to generate aggregated neural network model parameters.
[0036] S5. The central server distributes the aggregated network parameters to each party, and each party uses the aggregated neural network model parameters to update the local neural network model parameters.
[0037] The central server receives each party's client model and then performs weighted average aggregation operation on the model parameters. Finally, the central server obtains the global model aggregated in the current round. After the global model is updated, it is judged whether to stop training according to constraints such as the loss function value and the number of training rounds. After the central server completes the global training, it distributes the global model to parties A and B.
[0038] S6. Each party judges whether the local model converges based on the current local neural network model parameters. If it converges, the party saves the current local neural network model parameters; otherwise, it returns to S3.
[0039] S7. Each party performs medical image segmentation based on the converged local neural network model.
[0040] Each participant updates the local model using the optimal model during the training process, and the local participant uses the optimal model for subsequent image segmentation tasks.
[0041] In the present invention, the central server and each participant both adopt an improved pooling pyramid U-Net network structure as the basic model. Specifically:
[0042] As Figure 2 shown, the U-Net network has a contracting path and an expanding path, corresponding to the encoder and decoder respectively. In the encoder structure of U-Net, a 5-layer convolutional path is used for feature extraction, and each set of feature extraction operations passes through a 3*3 convolution, Batch Normalization, and ReLu activation function respectively. A convolutional kernel is used for max-pooling operation between each layer, and this operation can downsample the upper-layer feature image. The U-Net decoder also has 5 layers of paths. Each layer of the path needs to go through an upsampling operation, and then a concat operation is performed on the feature map transmitted from the encoder of the same level. The convolutional operations in the decoder are the same as those in the encoder, and they both make the feature map pass through a 3*3 convolution, Batch Normalization, and ReLu activation function in sequence. Then it goes through a 1*1 convolution operation.
[0043] Figure 3 is a schematic diagram of the improved pooling pyramid module. This module applies two pooling operations, namely max-pooling and adaptive average pooling. Among them, the adaptive average pooling will output the feature map to 4 scales, and the max-pooling will take the maximum value of the feature map within the kernel range. After obtaining 5 kinds of feature maps, a bilinear interpolation algorithm is used to restore the feature map, and then it is stitched with the original input image, and then subsequent convolutional operations are performed.
[0044] Figure 4It is an improved pooling pyramid U-Net network structure, which adopts a symmetric encoder-decoder structure. For the transfer of feature maps between encoder layers, an improved pooling pyramid is used for connection. There is a skip connection structure between the encoder layer and the decoder layer of the same level, and an attention mechanism module is included in the bottom and top structures. The attention mechanism module enables the network model to ignore the parts with low relevance to the result and makes the network model pay more attention to the valuable parts in the feature map. Each layer of the encoder network needs to go through convolutional operations, Batch Normalization operations, and ReLU functions. After that, it passes through an improved pooling pyramid module for compression and extraction of feature image information. The role of the improved pooling pyramid is to reduce the size of the feature image and reduce the number of network training parameters. The feature maps in the decoding stage are restored to the original size of the feature maps through upsampling operations. In the final stage of the decoder, binary classification needs to be performed on each pixel point of the output image. This operation uses 1*1 convolutional operations, and the final output of the network is a black-and-white image, that is, the segmentation result.
[0045] The following further illustrates the solution and effect of the present invention through a specific application example.
[0046] Suppose there are two medical institutions A and B that respectively hold medical image data related to skin diseases. They both want to train a medical image segmentation model for skin diseases, but currently they will encounter at least two problems. First, the data owned by medical institutions A and B is specific. For example, medical institution A is located in a plateau area and medical institution B is located in a plain area; or medical institution A is located in an inland area and medical institution B is located in a coastal area. Therefore, the patients faced by these two medical institutions have regional differences. Second, the distribution of patient samples in medical institutions is unbalanced. For example, medical institution A is a hospital in the provincial capital city of a first-tier city in China, while medical institution B is a township hospital in a non-first-tier city in China. Therefore, medical institution A has more sample quantities, and medical institution B has fewer sample quantities. This leads to a large difference in the number of patient samples in each medical institution. If medical institution B only uses its own sample set for training, the model trained by institution B cannot have good generalization performance. The present invention can assist medical institutions in joint modeling work to solve the above problems while protecting the privacy data of users from being leaked.
[0047] First, each institution participating in federated learning needs to prepare its own local dataset. Taking two medical institutions as an example, after medical institution A and medical institution B have prepared their datasets, they confirm a feasible third-party server and then receive the initial model. Medical institution A and medical institution B perform model training on their local datasets with the aim of enabling the model in this round to learn knowledge. The federated learning algorithm imposes constraints on the number of local training rounds for each participant, requiring them to update the model within the specified number of rounds and then upload and aggregate the locally trained models. The central server is responsible for receiving the models uploaded by each participant, performing model aggregation on the central server, and distributing the aggregated model. Thus, one global iteration ends. Repeat the above steps until the model converges.
[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A medical image segmentation method based on federated learning and improved pooling pyramid, characterized in that The steps include: S1. The central server issues the initialized neural network model parameters, and the neural network model is used for medical image segmentation; the neural network model adopts an improved pooling pyramid U-Net network structure, and the improved pooling pyramid U-Net network structure adopts a symmetric encoder-decoder structure. For the transfer of feature maps between encoder layers, an improved pooling pyramid is used for connection. The encoder layer and the decoder layer of the same level are connected by skip connections. The improved pooling pyramid applies max pooling and adaptive average pooling operations. The adaptive average pooling operation outputs the feature map to four scale sizes. The max pooling takes the maximum value of the feature map within the kernel range. The five scale features output after max pooling and the output after adaptive average pooling are successively superimposed on the original feature map through bilinear interpolation operations to obtain the output feature map; S2. Each participant receives the initialized neural network model parameters and uses the initialized neural network model parameters as the local neural network model parameters; S3. Train the local neural network model based on local data and generate new local neural network model parameters; S4. Each participant uploads their respective new local neural network model parameters to the central server, and the central server performs aggregation processing on the neural network model parameters to generate aggregated neural network model parameters; S5. The central server issues the aggregated network parameters to each participant, and each participant uses the aggregated neural network model parameters to update the local neural network model parameters; S6. Each participant determines whether the local model converges based on the current local neural network model parameters. If it converges, the participant saves the current local neural network model parameters; otherwise, return to S3; S7. Each participant performs medical image segmentation based on the converged local neural network model.
2. The medical image segmentation method based on federated learning and improved pooling pyramid according to claim 1, wherein The improved pooling pyramid U-Net network structure adds an attention mechanism module to the bottom and top structures.
3. A medical image segmentation method based on federated learning and improved pooling pyramid according to claim 1, characterized in that, The adaptive average pooling operation outputs the feature map to four scale sizes, including four scale sizes of 1*1, 2*2, 3*3, and 4*4.
Citation Information
Patent Citations
Semantic segmentation method for up-sampling decoding of convolution attention mechanism
CN113486897A
Brain tumor classification detection method and system based on feature pyramid network structure and channel attention mechanism
CN115170885A