A Lung Image Segmentation Model and Method for a Lightweight Network

Through the lightweight network architecture, combined with the improved reverse residual depth separation module, hollow space pyramid pooling layer and attention gate mechanism, the problem of large amount of calculation and insufficient accuracy in lung image segmentation is solved, and efficient and accurate lung medical image segmentation is achieved.

CN116188783BActive Publication Date: 2025-07-29ANHUI POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310191850.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-02
Publication Date
2025-07-29
Estimated Expiration
2043-03-02

AI Technical Summary

Technical Problem

The existing U-Net network model has a large amount of calculation and low calculation efficiency in lung image segmentation, which is difficult to adapt to the segmentation needs of lung medical images, and the segmentation accuracy is insufficient.

Method used

The lightweight network architecture is adopted, including improved reverse residual depth separation module, hollow space pyramid pooling layer and attention gate mechanism, simplifying the network structure, reducing the amount of parameters and calculations, and improving feature extraction and recovery capabilities.

Benefits of technology

It greatly reduces the calculation amount and parameters of the network model, improves the segmentation accuracy, especially the segmentation effect along the object boundary, and improves the generalization ability and computing efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188783B_ABST
    Figure CN116188783B_ABST
Patent Text Reader

Abstract

The present invention discloses a lung image segmentation model for a lightweight network, which includes a decoding part and an encoding part; the encoding part includes: two convolutional layers and two Maxpooling layers, which are used to extract features from the input image; the feature map extracted by the encoding part is sent to the decoding part after passing through an atrous spatial pyramid pooling layer; the decoding part includes two convolutional layers and two Up-sampling layers; it is used to output after up-sampling and restoring the feature map extracted by the encoding part and then passing through a 1×1 convolutional layer. The segmentation model further includes an attention gate AG layer, and the attention gate AG layer compensates the high-level feature map of the decoding part based on the high-level feature map of the encoding part. The advantages of the present invention are: 1. Greatly reduce the computational amount and the number of parameters of the network model, and reduce unnecessary computing power consumption; 2. Improve the accuracy of the network model and ensure the segmentation effect of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and particularly to a segmentation method of a lightweight network model with a new architecture for lung images. Background Art

[0002] The segmentation of lung images plays an important role in the diagnosis of lung diseases. In the field of image processing, traditional segmentation methods require manual segmentation by professional doctors, which is time-consuming and laborious. With the development of computer vision, intelligent segmentation using deep learning methods has become a hot topic in the field of image processing research. The FCN fully convolutional network is the earliest and classic encoder-decoder structure. Although the operation of pooling first and then upsampling can restore some spatial information, there is still a small amount of information that is difficult to restore. The subsequent U-Net network is also developed from the fully convolutional network. At that time, its unique encoder-decoder and skip connection structure was one of the best models for medical image segmentation with small samples.

[0003] However, in the real world, it is very difficult to apply the U-Net network model. First of all, the U-Net model itself requires a large amount of computing resources. And computing power can be directly converted into power consumption. From the perspectives of environment and economy, minimizing power consumption should be our pursuit goal. This leads to a question: Why should we waste computing resources and apply an unnecessarily computationally expensive model to simple images? Ideally, when the images are simple, we should use small networks, and when the images are complex, we should use large networks. And lung medical images are a kind of simple images, and the content and information they contain are far less than natural pictures, so a segmentation model with fewer parameters and less computational complexity should be used. Moreover, in lung image segmentation, the accuracy of model segmentation is the most important evaluation index. Although U-Net has good performance in the field of image segmentation, it cannot highly adapt to the segmentation of lung medical images. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art, and provide a lung image segmentation model and method of a lightweight network, to solve the defects of large computational amount and low computational efficiency existing in the U-Net network model for lung image segmentation in the prior art, and provide a new lightweight network that can simplify the network as much as possible, reduce the complexity of the network model, reduce the number of parameters and computational amount required by the model, and improve the computational efficiency of the network.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is: A lung image segmentation model of a lightweight network, including a decoding part and an encoding part;

[0006] The encoding part includes: two convolutional layers and two Maxpooling layers, which are used to extract features from the input image;

[0007] The feature map extracted by the encoding part is sent to the decoding part after passing through the atrous spatial pyramid pooling layer;

[0008] The decoding part includes two convolutional layers and two upsampling layers; it is used to perform upsampling restoration on the feature map extracted by the encoding part and then output after passing through a 1×1 convolutional layer.

[0009] The segmentation model further includes an attention gate AG layer, and the attention gate AG layer compensates the high-level feature map of the decoding part based on the low-level feature map of the encoding part.

[0010] The two convolutional layers and two Maxpooling layers in the decoding part are respectively: the first convolutional module, the first downsampling module, the second convolutional module, and the second downsampling module;

[0011] The input image is sent from the input end Input to the first convolutional module. After being processed by the first convolutional module, it is sent to the downsampling module for downsampling and then sent to the second convolutional module. After being processed by the second convolutional module, it is sent to the second downsampling module for downsampling.

[0012] The decoding part includes two convolutional layers and two upsampling layers, which are respectively: the third convolutional module, the fourth convolutional module, the first upsampling module, and the second upsampling module. The feature map after passing through the atrous spatial pyramid pooling layer is sent to the first upsampling module for upsampling operation and then sent to the third convolutional module for processing, and then sent to the second upsampling module for upsampling processing, and then sent to the fourth convolutional module for processing to complete the decoding process.

[0013] The first convolutional module, the second convolutional module, the third convolutional module, and the fourth convolutional module corresponding to the convolutional layers are all improved inverted residual depthwise separable modules. The improved inverted residual depthwise separable module includes: depthwise convolution, pointwise convolution, and residual connection. Essentially, it is the combination of two depthwise separable convolutions plus skip residuals, and the structure is similar to Figure 4 the left inverted residual depthwise separable convolution in , but there are also differences. First, various complex 1×1 convolutional operations for dimension increase and decrease are not added inside the improved inverted residual depthwise separable convolution. Second, the improved convolution is different in the skip manner of the residual connection.

[0014] The Atrous Spatial Pyramid Pooling layer consists of four convolutional layers and one pooling layer, which process the input image five times respectively: the first processing is that the input image undergoes a convolutional operation of 1×1×64 and then a BN operation to obtain the first feature map; the second to fourth processings are that the input image undergoes convolutional operations of 3×3×64 to obtain the second to fourth feature maps; the fifth processing is to perform global average pooling on the input image to obtain the fifth feature map. The feature maps obtained from the five processings are stacked together, and after dimensionality reduction through a 1×1 convolution, the feature map processed by the Atrous Spatial Pyramid Pooling layer is sent to the decoding part;

[0015] Among the second to fourth ones, three convolutional layers all use depthwise separable convolution for convolutional operations.

[0016] In the decoding part, the high-level feature map output by the Atrous Spatial Pyramid Pooling layer module and the low-level feature map output by the second convolutional module are compensated by the Attention Gate (AG) layer and then sent to the third convolutional module for processing;

[0017] The high-level feature map output by the first upsampling module and the low-level feature map output by the first convolutional module are compensated by the Attention Gate (AG) layer, and the obtained feature map is sent to the fourth convolutional module, and then output after passing through a 1×1 convolutional layer.

[0018] The compensation of the Attention Gate (AG) layer includes: the high-level feature map g in the decoding stage and the corresponding low-level feature map x in the encoding stage l are processed through a 1×1 convolution so that the two feature maps have the same size and number of channels, and then the obtained feature map after adding them together is sent to the ReLu activation function layer for processing and then sent to a 1×1 convolution of Ψ with the number of channels becoming 1; the operation formula includes:

[0019] After that, the sigmoid function and bilinear interpolation resampling are used to restore the same feature map size as x l to generate an attention coefficient table, and the formula is as follows:

[0020] Finally, the input feature (x l ) is multiplied by the attention coefficient (α) calculated in AG. At this time, the values in the target area will become larger, thereby suppressing irrelevant areas and obtaining a new shallow feature map

[0021] A method for segmenting lung images of a lightweight network, the segmentation method includes:

[0022] 1. Production of the dataset

[0023] This semantic segmentation mainly uses 2D images from the LUNA 16 (Lung Nodule Analysis 16) competition lung imaging dataset, including CT images and label maps. Both are single-channel images with a resolution of 512x512, and there are 266 of each. Among them, 234 are used as the training set and 32 as the validation set. To test the generalization ability of this model on other datasets, another liver dataset with 398 images and their corresponding label images is used. Among them, 320 are for the training set and 78 for the validation set. And the images of the two datasets are adjusted to 256×256 pixels.

[0024] 2. Establish a lung imaging segmentation model; the model is the lung imaging segmentation model of the lightweight network described above;

[0025] 3. Train the established lung imaging segmentation model;

[0026] 4. Use the trained lung imaging segmentation model for lung imaging segmentation. Input the lung medical image to be segmented into the lung imaging segmentation model, and the output of the model is the lung imaging segmentation result;

[0027] 5. Explanation of generalization ability: Usually, it is expected that the network trained with training samples has strong generalization ability. The PM-UNet lung segmentation model takes this ability into account. Without using data augmentation to expand the dataset, it improves the generalization ability of the model. At the time of designing the model, the atrous spatial pyramid pooling fuses features of multiple dimensions to obtain more effective information and improve the generalization ability, which has also been verified on the liver dataset.

[0028] The advantages of the present invention are as follows: 1. Greatly reduce the computational amount and the number of parameters of the network model, and reduce unnecessary computing power consumption; 2. Improve the accuracy of the network model, especially the segmentation result along the object boundary, and ensure the segmentation effect of the model. 3. It is applicable to the image segmentation of lung medical images, simplifies the network as much as possible, reduces the complexity of the network model, reduces the number of parameters and the computational amount required by the model, and improves the computational efficiency of the network. 4. It has strong generalization ability and also has a good segmentation effect on other medical images. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The following briefly describes the content expressed in each drawing of the specification of the present invention and the marks in the drawings:

[0030] Figure 1 It is the convolutional receptive field diagram provided by the present invention;

[0031] Figure 2 It is the standard convolution diagram provided by the present invention;

[0032] Figure 3The depthwise separable convolution diagram provided by the present invention;

[0033] Figure 4 The improved inverted residual depthwise separable module diagram provided by the present invention;

[0034] Figure 5 The atrous spatial pyramid pooling diagram provided by the present invention;

[0035] Figure 6 The attention gate diagram provided by the present invention;

[0036] Figure 7 The flowchart of the PM-UNet overall model provided by the present invention;

[0037] Figure 8 The comparison diagram of the number of parameters and computational complexity between the PM-UNet and the UNET models provided by the present invention;

[0038] Figure 9 The segmentation IoU result diagram of the lung dataset provided by the present invention;

[0039] Figure 10 The segmentation IoU result diagram of the liver dataset provided by the present invention;

[0040] Figure 11 The comparison diagram of the lung segmentation effects between the PM-Unet and the M-UNET models provided by the present invention;

[0041] Figure 12 The comparison effect diagram of segmenting lung and liver images using the present model and the prior art model. Detailed implementation manners

[0042] Next, with reference to the accompanying drawings, through the description of the optimal embodiments, the specific implementation manners of the present invention will be further described in detail.

[0043] The purpose of the present invention is to provide a lightweight neural network PM-UNet model with a new architecture for segmenting lung images. Compared with the traditional U-Net, there are two improvements: 1. Greatly reduce the computational complexity and the number of parameters of the network model, reducing unnecessary computing power consumption; 2. Improve the accuracy of the network model and ensure the segmentation effect of the model. These two points achieve the goal of simplifying the network as much as possible without affecting the segmentation accuracy of the model, reducing the complexity of the network model, reducing the number of parameters and computational complexity required by the model, and improving the computational efficiency of the network.

[0044] To achieve the above object, the specific steps are as follows:

[0045] (1) Change the original four-layer upsampling and downsampling into two layers. Under the new architecture, the same features repeatedly extracted on simple medical images can be reduced, simplifying the complex network;

[0046] (2) Combine the depthwise separable convolution in MobileNet V2 in the compression model to upgrade the decoder network from ordinary convolution to an improved inverse depthwise separable convolution module to optimize the network computing cost;

[0047] (3) After simplifying the network structure twice, there will inevitably be a decline in network performance. To solve this problem, we added Atrous Spatial Pyramid Pooling (ASPP) and Attention Gate (AG) as compensation, which can increase the network accuracy instead of decreasing it.

[0048] Figure 1 The following gives two constraint conditions for the architecture design. First, define some required parameter symbols: the input lung image size is z, the convolution kernel size is k, the number of convolution layers is l, and the minimum c value is t.

[0049] The first constraint mathematics is the size of the learning ability. To quantitatively measure the learning ability of the convolution layer, we define the c-value (c-value) of the convolution layer as follows. The formula is as follows:

[0050]

[0051]

[0052] The architecture of the deep model can be determined by the total number of each layer n and the number of structures in each stage a i For example, n = 3 and a1, a2, a3 = 4, 3, 2 indicate that the model has 3 stages, and the number of convolution layers in the first, second, and third stages is 4, 3, and 2 respectively. Among them, 2 l k is the actual filter size (Real Filter Size) of the l-th layer, is the receptive field size (Receptive Field Size) of the last layer, and t is the minimum value of the c-value. According to general design, t = 1 / 6 is a good c-value lower bound for all convolution layers of different tasks.

[0053] The second constraint is that the receptive field of the top layer should not be larger than the image area. The mathematical formula is as follows:

[0054]

[0055] The left side of the formula is the receptive field (Receptive Field) of the topmost convolution layer, 2 i-1 (k - 1) is the increment of the Receptive Field of the i-th convolution layer, 2i-1 (k - 1)a i is the total increment of the Receptive Field of the i-th convolutional layer, and the size of the image on the right side of the equation is z. The above two constraints lay the theoretical foundation for the new architecture. In summary, the receptive field should have excellent learning ability, but the receptive field of the top layer should not be larger than the image area.

[0056] For example, in the traditional 5-layer U-Net model structure, each layer has two convolutions, with a total of 23 convolution operations. In this design, the input lung image size is 256×256×3, and the size of the convolution kernel k in the original U-Net is 3×3×3. Through the above theoretical derivation, the minimum value that meets the c-value can be obtained, and the receptive field of the top convolutional layer is not larger than the 16×16 lung image at the bottom layer. However, in the new convolutional neural network PM-UNet, on the premise of not affecting the model segmentation accuracy, the network is simplified as much as possible, reducing the number of parameters and computational complexity required by the model, reducing the network model complexity, and improving the computational efficiency of the network. Cut half of the four upsampling and downsampling networks of U-Net, making U-Net change from a five-layer network to a three-layer network. At this time, the receptive field cannot meet the first constraint condition, that is, the lower bound of the learning ability c-value. At this time, t is much less than 1 / 6, and the receptive field is insufficient. Therefore, the Bottleneck at the bottom of the "U" - shaped network is replaced by the Atrous Spatial Pyramid Pooling (ASPP) to expand the receptive field.

[0057] Figure 2 For standard convolution: The number of channels of the convolution kernel is the same as the number of channels of the input. Each channel of the convolution kernel corresponds to an input channel, and after performing convolution operations separately on each channel and then adding them together. Let the input feature be F: D F ×D F ×M, and the output feature be G: D F ×D F ×N, and the convolution kernel be: D K ×D K , then the formula for the computing power consumption of standard convolution is:

[0058] D K ×D K ×M×N×D F ×D F (4)

[0059] Figure 3 For depthwise convolution: The convolution kernel has only one channel and is responsible for one channel of the input. Pointwise convolution: After depthwise convolution, the N feature maps are stacked together, and then weighted and combined into a new feature map using a 1×1×N convolution kernel. Then, under the same assumption, the formula for the computing power consumption of depthwise separable convolution at this time is:

[0060] D K ×D K ×M×D F ×D F +M×N×D F ×D F (5)

[0061] As can be seen from the above formula, the ratio of the two convolution operation amounts is as follows:

[0062]

[0063] Through the formula, if we use the well-known 3×3 convolution, the computational cost of depthwise separable convolution is between one-eighth and one-ninth of that of standard convolution. It can be seen that using depthwise separable convolution reduces the required parameters compared to ordinary convolution. Therefore, in this application, we introduce and adopt depthwise convolution.

[0064] Figure 4 Inspired by the convolution in MobileNet V2 on the left, it is improved into the reverse residual depthwise separable convolution module on the right, replacing the two ordinary 3×3 convolutions in one layer with the improved convolution module. The core of the MobileNet V2 network is the inverted residual structure and linear bottlenecks. The improved reverse residual depthwise separable convolution module in this paper uses the depthwise convolution, pointwise convolution, and residual connection in the inverted residual structure. The most basic is the combination of two reverse residual depthwise separable convolutions, without adding various complex dimensionality increase and decrease operations in the internal inverted residual structure. First, because the increase in the dimension of the tensor (Tensor) is accompanied by a decrease in the calculation speed, the purpose of this paper is to use depthwise separable convolution to reduce the number of parameters and computational cost of the model. However, if the feature maps of the convolutional layer are all using low-dimensional tensors (Tensors) to extract features, then it is impossible to extract enough overall information. Therefore, in this paper, according to the traditional way of increasing layer by layer in U-Net, the dimension of the feature maps (feature maps) is increased. The working principle of this module is to change the previous ordinary convolution operation that considers both channels and regions into a convolution that first only considers regions and then considers channels. The separation of channels and regions is realized. A complete convolution operation is split into two operations like the factorization in mathematics. This can greatly reduce the number of parameters and computational cost required by the model. And the internal identity mapping well improves the way of forward and backward information transmission, thus greatly promoting the optimization of the network.

[0065] Workflow of the improved inverted residual depthwise separable convolution module: The input is the feature map of the lung image. In the first step, depthwise convolution is performed for lightweight filtering. Each input channel corresponds to a filter. Then, the BN (Batch Normalization) operation is added. Finally, the Relu6 activation function is used to accelerate the convergence speed of the network and control the problem of gradient vanishing. In the second step, pointwise convolution, that is, 1×1 convolution, linearly combines the feature maps after depthwise convolution. The above two steps are similar to depthwise separable convolution plus BN and Relu6 operations. The third and fourth steps are the same as the above two steps, which is to replace the two 3×3 convolutions in each layer of the original U-Net. In the fifth step, the input feature map is connected residually with the output feature map through 1×1 convolution, and the stacked feature map is input to the next layer.

[0066] As Figure 5 shown, its workflow is that the input image is processed five times and then stacked together. In the first processing, the input image undergoes a 1×1×64 convolution operation, and after adding the BN operation, the first feature map is obtained. In the second to fourth processings, the input image undergoes 3×3×64 convolution and the BN operation. However, these three convolutions here are depthwise separable convolutions, which can reduce the number of parameters and computational complexity while expanding the receptive field. For the ordinary convolution process (dilation rate = 1), at this time, the receptive field after 3×3 convolution is 3, and the formula for the receptive field of dilated convolution is:

[0067] k'=(d - 1)×(k - 1)+k (7)

[0068] where k’ is the receptive field size after dilated convolution, and k is the convolution kernel size. For 3×3 dilated convolutions with dilation coefficients of 6, 12, and 18 respectively, when the dilation coefficient is the maximum of 18, the receptive field of the convolution at this time is 37, while the input image after two downsamplings is only 64×64, and the receptive field at this time is close to but does not exceed the entire relevant image area. It should be noted that in order to make the sizes of the output feature maps consistent, the dilation coefficient rate is equal to the padding coefficient. Due to the influence of the edge effect (padding increases with a large dilation rate), so our fifth image processing is global average pooling to reduce the influence of the edge effect and obtain an image with global features. At this time, the feature maps processed five times are stacked together and sent to the decoding process through 1×1 convolution for dimensionality reduction.

[0069] As Figure 6As shown, in order to alleviate the problem that deconvolution can only restore partial features and cannot restore the original data 100%, we used a new attention gate (AG) model for medical imaging. This model is different from the traditional U-Net cascade operation in which the feature information of high and low layers is simply superimposed together, resulting in excessive extraction of shallow redundant information. In order to better restore the image, we add an attention gate (AG) model to the jump connection that fuses high-level and low-level information. It can automatically learn to focus on the key information of the target. This allows the attention gate (AG) model to highlight the useful and significant features of the image while suppressing interference from irrelevant areas. In the figure, g is the high-level feature map of the decoding stage, and x l is the low-level feature map in the encoding stage, W g is the gating signal, W x is the attention gate. First, g, x l Both of them go through a 1×1 convolution to ensure that the two feature maps have the same size and number of channels. Then, when they are added together, the value of the target area of the high-level information will become larger. Then, using the ReLu activation function, the number of channels after the 1×1 convolution of Ψ becomes 1. The calculation formula is as follows:

[0070]

[0071] Then use the sigmoid function and bilinear interpolation to restore the resampling of x l For the same feature map size, generate the attention coefficient table with the following formula:

[0072]

[0073] Finally, input features (x l ) By multiplying the attention coefficient (α) calculated in AG, the value of the target area will become larger, thereby suppressing irrelevant areas and obtaining a new shallow feature map

[0074] Module 1 is the first convolution ( Figure 7 Named InvertedResidual Block1) output, module 2 is the second convolution ( Figure 7 Named InvertedResidual Block2) output, module 3 is the void space pyramid pooling layer module ( Figure 7 Named as ASPP) output, module 4 is the first upsampling module ( Figure 7 UP2) output.

[0075] Note that the working process of the gate (AG) in this model is: Module 1 compensates Module 4, and Module 1 is x lThe low-level feature map in the encoding stage is used as the input, and Module 4 uses the high-level feature map in the decoding stage of g as the input. Both Module 1 and Module 4 will go through 1×1 convolution, and then the two are added together. Then, the ReLu activation function is used, and the number of channels becomes 1 after the 1×1 convolution of Ψ. After that, the sigmoid function and bilinear interpolation resampling are used to restore the same feature map size as x l to generate the attention coefficient table. Finally, Module 1 is used as the input feature map and multiplied by the generated attention coefficient table (α) to output a new shallow feature map The skip connection at this time is to use the new shallow feature map as the input and superimpose it with the high-level feature map in the decoding stage of Module 4

[0076] Similarly, for another attention gate, Module 2 compensates for Module 3. Module 2 uses the low-level feature map in the encoding stage of x l as the input, and Module 3 uses the high-level feature map in the decoding stage of g as the input. Both Module 2 and Module 3 will go through 1×1 convolution, and then the two are added together. Then, the ReLu activation function is used, and the number of channels becomes 1 after the 1×1 convolution of Ψ. After that, the sigmoid function and bilinear interpolation resampling are used to restore the same feature map size as x l to generate the attention coefficient table. Finally, Module 2 is used as the input feature map and multiplied by the generated attention coefficient table (α) to output a new shallow feature map The skip connection at this time is to use the new shallow feature map as the input and superimpose it with the high-level feature map in the decoding stage of Module 3

[0077] As Figure 7 shown, the algorithm framework of this model retains the unique encoding-decoding and skip connection operation structures of U-Net, and performs model pruning and improvement on its structure, reducing the original four-layer upsampling and downsampling to two times. In PM-UNet, the left side is the encoding path, which contains two improved inverted residual depthwise separable modules( Figure 7It is named InvertedResidualBlock and has two downsamplings, namely Maxpooling layers, to learn more semantic segmentation information by gradually reducing the size of the feature map. On the right is the decoding path, which also has two improved inverted residual depthwise separable modules and two upsamplings, namely Transposed Convolution layers, to restore the gradually reduced feature map caused by downsampling. And an attention decoding mode (AG) is used to suppress the influence of background pixels in the image on the segmentation effect and highlight the target areas containing key information. In addition, at the bottom of the "U" - shaped network, Atrous Spatial Pyramid Pooling (ASPP) is added to solve the problem that the similarity between the lung parenchyma and other tissues and organs in lung images easily leads to an unsatisfactory actual segmentation effect, especially the segmentation results along the object boundaries.

[0078] To illustrate the advantages of the proposed PM - Unet network model in lung image segmentation, the present application is verified through the following experiments:

[0079] The framework of this model is built with Pytorch. The relevant model code is written in Jupyter notebook. The computer in the laboratory is NVIDIA GTX3050TI. The Batchsize for training and validation is 2. The Epoch for Dataset 1 and Dataset 2 is set to 80. The loss function of this model is Cross Entropy Loss. The parameters are updated with the Adam optimizer. The learning rate is set to 0.001, and learning rate decay is used. The learning rate is reduced by 90% every 7 rounds to accelerate the model convergence speed. In the image segmentation task, the accuracy (ACC) of segmented pixels and the Intersection over Union (IoU) between the segmented image and the actual label image are important indicators to evaluate the segmentation effect. These two evaluation indicators are also used in the medical lung image segmentation in this paper to evaluate the segmentation performance of the MAU - Net network model. The calculation formulas are as follows:

[0080]

[0081]

[0082] In the above formulas (10) and (11), GT represents the label annotated by experts, and SR represents the prediction result obtained by the model in this paper. TP is the number of lung pixels correctly classified as lung parenchyma; TN is the number of background pixels correctly classified as background. FP is the number of background pixels misclassified as lung parenchyma; FN is the number of lung pixels misclassified as background.

[0083] Figure 8As shown, the complexity of a model generally refers to the amount of computation and the number of parameters required for the model to run. When the input image tensor is 1×3×256×256 and each parameter is of the floating-point data type (float), that is, one parameter is 4 bytes, M-Net upgrades the traditional U-Net network from ordinary convolution to an improved depthwise separable convolution lightweight network. Moreover, M-Net+ASPP is a model with atrous spatial pyramid pooling added, and M-Net+atten is a model with an attention gate network added to M-Net. PM-UNet is a network model that simultaneously uses atrous spatial pyramid pooling and the attention gate. Table 1 shows the number of parameters and the amount of computation required for each model under four different upsampling and downsampling layers. The values on the left in the table are the number of parameters in the network, and the values on the right are the floating-point computation amounts.

[0084] Table 1 Comparison Results of Models

[0085]

[0086] From the above Figure 1 table, it can be seen that the number of parameters and the amount of computation of the PM-UNet network model with two-layer downsampling studied in this paper have decreased by approximately 97.8% and 85.6% compared to the traditional U-Net network, and the model complexity has been greatly reduced. And using the method of time.time for the running time, under cuda acceleration, it takes 1489.9 seconds for the PM-UNet network model to run 80 rounds. In other words, this model can process approximately 14 groups of lung images within one second. Moreover, the added compensation modules, atrous spatial pyramid pooling and the attention gate, each of these two modules accounts for about 56% and 3% of the number of parameters of the improved model PM-UNet, and about 21% and 7% of the amount of computation. The next section shows that while greatly saving computational power, the model performance has not been lost and even has been improved.

[0087] Table 2 shows the segmentation results of lung images under each model network. To illustrate the segmentation effect of the improved model, several comparative ablation experiments were conducted, and two datasets were used. One is the LUNA lung nodule detection dataset, and the other is the liver dataset used to verify the generalization ability of the model. The experimental results are as follows:

[0088] Table 2 Results of the Validation Set of Dataset 1

[0089]

[0090] As can be seen from Table 2, the increase in the number of layers of the lightweight network M-Net is accompanied by an improvement in performance. However, it can be seen that the M-Net after three layers does not have a significant jump in segmentation performance due to the increase in network complexity. Therefore, we build the model based on the three-layer M-Net. The addition of ASPP is to compensate for the insufficient receptive field caused by the reduction in the number of layers and to improve the problem of boundary segmentation. However, new problems have emerged. As shown in Table 2, the LOSS of the three-layer M-Net+aspp network for segmenting images is too large, significantly larger than that of other segmentation models. The main reason is related to the grid effect of the dilated convolution used in the atrous spatial pyramid, resulting in insufficient ability to extract detailed features of the image, that is, there is some loss of image information during the segmentation process, and an underfitting phenomenon occurs, which is fatal for pixel-level image segmentation tasks. The added attention mechanism is to improve the attention of the model network to useful information and can solve the underfitting problem that is difficult to converge. As shown in Table 1 and Table 2, although from the perspective of the tasks of the lightweight network model, the number of parameters of the network is increased by 3% and the computational amount is increased by 7%, but compared with the 93% decrease in LOSS and the accelerated convergence speed, it is still worthwhile. The lightweight network plus the two compensation modules of the atrous spatial pyramid and the attention gate significantly improve the segmentation effect. Moreover, the addition of each layer is accompanied by a certain improvement in performance. However, considering the segmentation computing power and segmentation effect comprehensively, the three-layer PM-UNet should be selected as the segmentation network.

[0091] As Figure 9 shown, the IoU results of the traditional U-Net and the improved model running for 80 rounds on the validation set of Dataset 1. In the segmentation results of the lung images in Dataset 1, the solid line is the improved model PM-UNet, and the dashed line is the traditional model U-Net. After 80 rounds of training, there is a significant gap in the evaluation index IoU. PM-UNet is stable at about 0.958, while U-Net is stable at about 0.932. The evaluation index IoU of PM-UNet is greater than that of U-Net, and PM-UNet is more stable and can achieve fast convergence. It can be concluded that the performance of PM-UNet is superior to that of U-Net.

[0092] Table 3 shows the liver image segmentation. Since it is only used for verifying the generalization ability, only two model experiments are conducted considering the computing power. The IoU results of the traditional U-Net and the improved three-layer PM-UNet running for 80 rounds on the validation set of Dataset 2.

[0093] Table 3 Results of the Validation Set of Dataset 2

[0094]

[0095] As Figure 10As shown in the figure, in the liver image segmentation results of dataset 2 in the generalization experiment, after 80 rounds of training, there is an obvious gap in the evaluation index IoU. PM-UNet is stable at about 0.926, while U-Net is stable at about 0.838.

[0096] From the experimental results, we can see that in both datasets, whether it is the convergence speed or stability, the improved three-layer PM-UNet model has a certain improvement in lung segmentation and a good performance in liver segmentation on the basis that the factorization convolution and topological connection are introduced, and the amount of calculation and the number of parameters are sharply reduced.

[0097] Above Figure 11 As shown in the figure, the original segmentation image is on the left, the label image is in the middle, and the segmentation effect images of each model are on the right. The model on the left is M-UNet, and the one on the right is PM-Unet. From top to bottom in each image are the second, third, and fourth layers of M-UNet and PM-UNet. From the above segmentation models, it can be seen that the compensation module added in this paper has an obvious improvement effect. Moreover, it can be seen that there is not much difference between the segmentation images of the third and fourth layers of PM-UNet. This also proves from the segmentation effect that when considering the computational efficiency, the three-layer PM-UNet is the best choice in this paper. Below Figure 12 This is the comparative segmentation image of the three-layer PM-UNet and the traditional U-Net.

[0098] As Figure 12 shown in the figure, these are the segmentation images generated using two datasets. A and B are the LUNA pulmonary nodule detection datasets, and C and D are the liver datasets. A and C are the traditional U-Net. In A, it can be clearly seen that there are many defects in the box, especially in the pulmonary end, which are not well segmented. In the box of the C segmentation image, there are redundant segmentation areas. Therefore, the segmentation effect of the traditional network U-Net is not excellent. While B and D are the improved three-layer PM-UNet network. In B, not only the areas on the label image can be segmented, but also the local areas not involved in the label image in the box can be segmented. On the fuzzy boundary, it can be seen that the pulmonary parenchyma contour is completely segmented without disconnection or deformation, and the segmentation effect is better. In short, there are no major defects in the overall segmentation of the lungs and liver parenchyma.

[0099] Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made by adopting the method concept and technical solution of the present invention, they are all within the protection scope of the present invention.

Claims

1. A lung image segmentation model for a lightweight network, characterized in that: It includes a decoding part and an encoding part; The encoding part includes: two improved inverted residual depthwise separable modules and two Maxpooling layers, which are used to extract features from the input image; The feature map extracted by the encoding part is sent to the decoding part after passing through the atrous spatial pyramid pooling layer; The decoding part includes two improved inverted residual depthwise separable modules and two Up sampling layers; it is used to perform upsampling restoration on the feature map extracted by the encoding part and then output after passing through a 1×1 convolutional layer; The segmentation model also includes an attention gate AG layer, and the attention gate AG layer compensates the high-level feature map of the decoding part based on the high-level feature map of the encoding part; The two convolutional layers and two Maxpooling layers in the decoding part are respectively: the first convolutional module, the first downsampling module, the second convolutional module, and the second downsampling module; The input image of the input end Input is sent to the first convolutional module, and after being processed by the first convolutional module, it is sent to the downsampling module for downsampling and then sent to the second convolutional module. After being processed by the second convolutional module, it is sent to the second downsampling module for downsampling processing; The decoding part includes two convolutional layers and two Up sampling layers, which are respectively: the third convolutional module, the fourth convolutional module, the first upsampling module, and the second upsampling module. The feature map after passing through the atrous spatial pyramid pooling layer is sent to the first upsampling module for upsampling operation and then sent to the third convolutional module for processing, and then sent to the second upsampling module for upsampling processing, and then sent to the fourth convolutional module for processing to complete the decoding process; The first convolutional module, the second convolutional module, the third convolutional module, and the fourth convolutional module corresponding to the convolutional layer are all improved inverted residual depthwise separable modules. The improved inverted residual depthwise separable module includes: two pairs of combinations of depthwise convolution and pointwise convolution plus 1×1 convolutional residual connection; The atrous spatial pyramid pooling layer includes four convolutional layers and one pooling layer, and the input image is processed five times respectively: the first processing is that the input image passes through a 1×1×64 convolutional operation and then adds a BN operation to obtain the first feature map. The second to fourth processes are that the input image passes through a 3×3×64 convolution and adds a BN operation to obtain the second to fourth feature maps; the fifth processing is to perform global average pooling on the input image to obtain the fifth feature map. The feature maps of the five processes are stacked together and then sent to the decoding part after passing through a 1×1 convolution for dimensionality reduction processing to obtain the feature map processed by the atrous spatial pyramid pooling layer; Among the second to fourth, three convolutional layers all use depthwise separable convolution for convolution operations; In the decoding part, the high-level feature map output by the atrous spatial pyramid pooling layer module and the low-level feature map output by the second convolutional module are compensated by the attention gate AG layer and then sent to the third convolutional module for processing; The high-level feature map output by the second upsampling module and the feature map obtained by compensating the low-level feature map output by the first convolutional module by the attention gate AG layer are sent to the fourth convolutional module, and then output after passing through a 1×1 convolutional layer.

2. The lung image segmentation model of a lightweight network according to claim 1, characterized in that: Note that the AG layer compensation of the door includes: the high-level feature map g in the decoding stage and the corresponding low-level feature map x in the encoding stage l are processed through 1×1 convolution so that the two feature maps have the same size and number of channels. Then, the obtained feature map after addition is sent to the ReLu activation function layer for processing and then to the 1×1 convolution of Ψ, and the number of channels becomes 1. The operation formula includes: After that, the resampling of the sigmoid function and bilinear interpolation is used to restore the same feature map size as x l to generate an attention coefficient table. The formula is as follows: Finally, input the feature x l Multiply it by the attention coefficient α calculated in the AG. At this time, the value of the target region will become larger, thereby suppressing the irrelevant regions and obtaining a new shallow feature map Among them refers to the low-level feature map in the encoding stage; g i is the high-level feature map in the decoding stage; σ1 refers to the ReLu activation function; σ2 refers to the sigmoid function; W g is the gating signal, W x is the attention gate; b h and b Ψ are the bias terms of the convolution in the attention gate model; Ψ is the convolution operation.

3. A method for segmenting lung images of a lightweight network, characterized in that: The segmentation method includes: Establishing a lung image segmentation model; the model is a lung image segmentation model of the lightweight network as described in claim 1 or 2; Training the established lung image segmentation model; Performing lung image segmentation using the trained lung image segmentation model, inputting the lung medical image to be segmented into the lung image segmentation model, and the output of the model is the lung image segmentation result.