Lung Segmentation Method for X-ray Images Based on a 4-Convolution-Layer Stereo Pyramid Network
Through the 4-convolution layer three-dimensional pyramid network structure and attention mechanism, the problems of large amount of parameters, overfitting and information loss in lung image segmentation are solved, and feature extraction of high resolution and rich context information is achieved, improving the segmentation effect.
Patent Information
- Application Number
- CN202210861243.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-20
AI Technical Summary
The prior art has problems such as large amount of parameters, overfitting and information loss in lung image segmentation, resulting in poor segmentation effect.
The 4-convolution layer three-dimensional pyramid network structure is adopted to enhance feature extraction through layer jump connection and attention mechanism, maintain image resolution and suppress irrelevant features.
Maintain high resolution and rich context information on limited data sets, reduce the amount of network parameters, avoid overfitting, and improve the effect of lung image segmentation.
Smart Images

Figure CN115294154B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing. In particular, it relates to a method for segmenting the lungs in X-ray images based on a four-convolution-layer stereo pyramid network. Background Art
[0002] CXR (Chest X Ray) imaging technology has low cost and low-dose radiation, and is one of the most effective methods for screening pneumonia diseases at present, and is an important reference basis for doctors in the diagnosis process of pneumonia diseases.
[0003] The Chinese patent application with the application number CN202110589697.0 discloses a method for segmenting lung X-ray images based on UNet. First, the lung X-ray image and the label image are preprocessed. By adding pixel points with a pixel value of 0, the image is made square, and then Gaussian filtering is used for denoising and the Contrast Limited Adaptive Histogram Equalization (CLAHE) algorithm is used to enhance the image to obtain a training set. Then, the encoding part of the network consists of a migrated MobileNet network model with a 5-layer network structure. The decoding part improves UNet by changing the convolution method and adding residual blocks. Finally, the segmented lung X-ray image test data is introduced into the trained improved UNet model to obtain the final segmentation result.
[0004] Although the above method can solve the problem of lung image segmentation to a certain extent, the neural network structure adopted has a large number of parameters, resulting in an excessive operating burden on the system. At the same time, due to the limitation of training samples, using too deep a network will cause overfitting. And context information and detail information are lost during the convolution process, and the image resolution cannot be well maintained. Summary of the Invention
[0005] In view of the technical problem in the above-mentioned prior art that the image resolution is affected due to the loss of context information and detail information, a method for segmenting the lungs in X-ray images based on a four-convolution-layer stereo pyramid network is provided. The present invention segments the bilateral lungs of the CXR image through a four-convolution-layer stereo pyramid network structure to obtain a clear segmentation image, and uses a relatively low number of parameters to overcome the problem of imperfect segmentation of the lung boundary due to pathological conditions or poor imaging quality in the gray-scale CXR image.
[0006] The technical means adopted by the present invention are as follows:
[0007] A method for segmenting the lungs in X-ray images based on a four-convolution-layer stereo pyramid network, comprising:
[0008] Obtaining a CXR image dataset, and performing data augmentation processing on the CXR image dataset to generate a training dataset;
[0009] Construct a 4-convolution-layer stereo pyramid network structure, and train the 4-convolution-layer stereo pyramid network structure based on the training dataset to obtain the optimal network structure parameters. The 4-convolution-layer stereo pyramid network structure includes four encoder blocks, four decoder blocks, and a fusion module that are symmetrically arranged. Each encoder block and the decoder block at the corresponding position are connected by a skip connection. The fusion module is used to fuse the images with the same size of the output of each layer;
[0010] Obtain the CXR image to be processed and input it into the 4-convolution-layer stereo pyramid network structure applying the optimal network structure parameters for processing, and finally output the segmentation result.
[0011] Further, each encoder block in the 4-convolution-layer stereo pyramid network structure includes two consecutive 3×3 convolutional layers and a max pooling layer, and a BN layer and a ReLU layer are connected behind each convolutional layer.
[0012] Further, before the encoder blocks and the decoder blocks at the corresponding positions are connected by skip connections, it also includes feeding the output features of each encoder block and the output features of the decoder block at the corresponding position into the attention mechanism module.
[0013] Further, the attention mechanism module includes a channel attention module and a spatial attention module that are concatenated in sequence.
[0014] Further, perform data augmentation processing on the CXR image dataset, including: performing contrast enhancement processing on the CXR image data based on the CLAHE algorithm to generate an augmented dataset.
[0015] Further, in the 4-convolution-layer stereo pyramid network structure, bilinear interpolation is used for the output of each layer of the fusion module to restore the output feature map to the same size as the previous layer.
[0016] Further, in the 4-convolution-layer stereo pyramid network structure, the output feature maps with the same scale of each layer are concatenated in the channel dimension and pooled to reduce the dimension, and then feature fusion is performed to generate the model output.
[0017] Compared with the prior art, the present invention has the following advantages:
[0018] In the present invention, on a limited dataset, while maintaining the high resolution of the image through the improved encoder-decoder-skip structure, it has rich context information. At the same time, an attention mechanism is introduced into the skip structure to enhance the weight of key features and suppress irrelevant features. Finally, using 4 convolutional layers can achieve the same effect as 5 convolutional layers and reduce the number of network parameters.
[0019] For the above reasons, the present invention will also promote related image processing and applications, such as the identification and classification of pneumonia diseases and other applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of a method for segmenting the lungs in X-ray images based on a 4-convolution-layer spatial pyramid network of the present invention.
[0022] Figure 2 It is a structural diagram of a 4-convolution-layer spatial pyramid network for segmenting the lungs in CXR images of the present invention.
[0023] Figure 3 It is a structural diagram of the attention mechanism module of the present invention.
[0024] Figure 4 It is a structural diagram of the feature map of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0027] Such as Figures 1-4As shown in the figure, the present invention provides a four-convolution-layer three-dimensional pyramid network structure for segmenting the double lungs of CXR images, providing a strong basis for the confirmation of lung lesions. The specific flowchart is as follows Figure 1 shown, mainly including:
[0028] S1. Obtain a CXR image dataset, and perform data augmentation processing on the CXR image dataset to generate a training dataset.
[0029] Specifically, the present invention first enhances the image contrast of the CXR image through the CLAHE algorithm and expands the dataset to obtain the training data for training the network structure.
[0030] S2. Construct a four-convolution-layer three-dimensional pyramid network structure, and train the four-convolution-layer three-dimensional pyramid network structure based on the training dataset to obtain optimal network structure parameters, where the four-convolution-layer three-dimensional pyramid network structure includes four encoder blocks, four decoder blocks, and a fusion module that are symmetrically arranged. Each encoder block is connected to the decoder block at the corresponding position by a skip connection, and the fusion module is formed by fusing images with the same size of each layer output; the present invention finds the model with the highest test accuracy as the optimal model through a finite number of iterative learning, and also continuously backpropagates to optimize the parameters during the learning process.
[0031] Specifically, the four-convolution-layer three-dimensional pyramid network structure in the present invention is as follows Figure 2 shown
[0032] The network includes four encoder blocks ( Figure 2 the four convs after input in the figure) and four decoder blocks (four deconvolution layers that are symmetrically structured with respect to the encoding network), and are connected by skip connections. Through this structure, the image resolution can be well maintained. Each encoder block consists of two consecutive 3×3 convolutional layers and a max pooling layer. A BN layer and a ReLU layer are connected after each convolutional layer. The specific representation is shown in formula (1).
[0033]
[0034] In this application, x l is used to represent that the image processes local information layer by layer through the convolutional layer at layer l, and gradually extracts the feature map of the high-dimensional image. In formula (1), i represents the spatial dimension, c represents the number of channels, c′ represents the number of channels in the (l-1)th layer, the decoder uses bilinear interpolation to restore the image resolution, and then fuses the image with the corresponding image in the encoding network structure to endow the image with richer context information.
[0035] Before the skip connection, the image needs to be sent into the attention mechanism module. The structure of the attention mechanism module is as followsFigure 3 As shown, in this mechanism, x l is the feature map of the l-th layer in the encoding part, and g is the feature map of the decoding part at the corresponding encoding position (g and x l are feature maps with the same input size and number of channels). By calculating the attention weights of the feature map in both the channel and spatial dimensions, the feature map can be adaptively adjusted. Among them, g i,c contains more accurate information. Adding the information in g to x l can enable the attention coefficient to be better trained and updated. The formula of the channel attention module is shown in (2) and (3).
[0036]
[0037]
[0038] Among them represents the i-th feature map in the l-th layer. σ is the sigmoid function, and is a 1×1 convolution operation. F g 、F l 、F int represent the number of channels. F int is smaller than F g 、F x . H represents the height and W represents the width. Before entering the spatial attention, let be multiplied by to update . The formula of the spatial attention module is shown in (4) and (5).
[0039]
[0040]
[0041] Among them, Ψ represents a 1×1 convolution, σ 1 represents the ReLU function, σ 2 represents the sigmoid function, represents the bias term.
[0042] Through training, the attention coefficient α ∈ [0, 1] is obtained, making the value in the target area approach 1 and the value in the irrelevant area approach 0. Finally, is multiplied by α, and the result of the multiplication will focus the attention on the target area. Its comprehensive formula can be expressed as:
[0043]
[0044] Finally, using the feature pyramid structure (as shown in Figure 4 ), The resolution of the image is maintained and activated through the skip connection structure, which is represented by formula (7).
[0045]
[0046] Where C represents the concatenation operation, and σ represents the sigmoid activation.
[0047] Under the above conditions, an improved multi-scale fusion pyramid structure is added. The output of each layer of the fusion module uses bilinear interpolation to restore the feature map to the same size as the previous layer, which is represented by formula (8).
[0048]
[0049] Where U represents bilinear interpolation to obtain richer texture and semantic information.
[0050] In the overall network structure, the feature maps with the same scale are connected in the channel dimension and pooled to reduce the dimension for feature fusion, as shown in formula (9).
[0051]
[0052] Where f is a 1×1 convolution, represents the feature map after feature mapping of the l-1 layer encoding network. Using this structure, high-resolution low-level features and high semantic information of high-level features are maintained through layer-by-layer fusion.
[0053] S3. Obtain the CXR image to be processed and input it into the 4-convolution-layer stereo pyramid network structure with the optimal network structure parameters for processing, and finally output the segmentation result. In this embodiment, it is preferably to input the test set image into the trained network to output the bilateral lung segmentation map.
[0054] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An X-ray image lung segmentation method based on a 4-convolution-layer stereo pyramid network, characterized in that, it includes: Obtain a CXR image dataset, and perform data augmentation on the CXR image dataset to generate a training dataset; Construct a 4-convolution-layer stereo pyramid network structure, and train the 4-convolution-layer stereo pyramid network structure based on the training dataset to obtain optimal network structure parameters. The 4-convolution-layer stereo pyramid network structure includes four encoder blocks, four decoder blocks and a fusion module that are symmetrically arranged. Each encoder block and the decoder block at the corresponding position are connected by a skip connection. The fusion module is used to fuse the images with the same size of the output of each layer. Before each encoder block and the decoder block at the corresponding position are connected by a skip connection, it also includes sending the output features of each encoder block and the output features of the decoder block at the corresponding position into an attention mechanism module; in the 4-convolution-layer stereo pyramid network structure, the output of each layer of the fusion module uses bilinear interpolation to restore the output feature map to the same size as the previous layer, connect the output feature maps with the same scale of each layer in the channel dimension and use a pooling operation to reduce the dimension, and then perform feature fusion to generate the model output; Obtain the CXR image to be processed and input it into the 4-convolution-layer stereo pyramid network structure with optimal network structure parameters for processing, and finally output the segmentation result.
2. The X-ray image lung segmentation method based on a 4-convolution-layer stereo pyramid network according to claim 1, characterized in that, Each encoder block in the 4-convolution-layer stereo pyramid network structure includes two consecutive 3×3 convolutional layers and a max pooling layer, and a BN layer and a ReLU layer are connected behind each convolutional layer.
3. The X-ray image lung segmentation method based on a 4-convolution-layer stereo pyramid network according to claim 1, characterized in that, The attention mechanism module includes a channel attention module and a spatial attention module that are concatenated in sequence.
4. The X-ray image lung segmentation method based on a 4-convolution-layer stereo pyramid network according to claim 1, characterized in that, Performing data augmentation on the CXR image dataset includes: performing contrast enhancement processing on the CXR image data based on the CLAHE algorithm to generate an augmented dataset.
Citation Information
Patent Citations
A lung X-ray image segmentation method based on UNet
CN113223021B
Lung image segmentation method and device, storage medium and computer equipment
CN113643308A
Multi-spectral image fusion method based on Y-shaped pyramid network
CN114283104A