Lung X-ray image segmentation system based on deep learning

By introducing multi-scale feature extraction and global context information perception modules into the lung X-ray image segmentation system, combining non-local and spatial attention mechanisms, the segmentation difficulties caused by morphological and grayscale similarity in lung X-ray image segmentation are solved, and a higher precision lung area segmentation is achieved.

CN120339624APending Publication Date: 2025-07-18ZHONGBEI UNIV
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510495710.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the morphology and size differences of lung regions in lung X-ray image segmentation, and the lung regions are highly similar in grayscale and texture characteristics to other regions, resulting in poor segmentation and prone to over-segmentation or under-segmentation.

Method used

The lung X-ray image segmentation system based on deep learning is adopted, combined with the multi-scale feature extraction module, the global context information perception module, the non-local attention mechanism and the spatial attention mechanism, and through the expanded convolution and feature fusion method, the adaptability to the lung region and the perception ability of the global structure is enhanced.

Benefits of technology

It improves the ability to identify complex and irregular areas, enhances the segmentation accuracy of target areas, reduces the risk of over-segment or under-segment, and improves the segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339624A_ABST
    Figure CN120339624A_ABST
Patent Text Reader

Abstract

The invention discloses a lung X-ray image segmentation system based on deep learning, which comprises a data preprocessing module, a lung X-ray image segmentation network module, a training module and a segmentation result output module, and is characterized in that the lung X-ray image segmentation network module constructs a lung X-ray image segmentation network based on deep learning, and the lung X-ray image segmentation network is recorded as MSDA-Net; and the training module is used for training the constructed MSDA-Net. A multi-scale feature extraction module is introduced into the network, and expansion convolution and feature fusion methods with different expansion rates are combined, so that the adaptability of the model to a lung region in an X-ray image is enhanced, a key target region in the image is highlighted, irrelevant background noise is suppressed, the significant difference of the lung region in form and size can be effectively handled, and the accuracy of the model is improved. The method improves the recognition capability of a complex and irregular region, and improves the segmentation effect of a target region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The field of the present invention relates to the field of medical image processing, and particularly relates to a lung X-ray image segmentation system based on deep learning. Background Art

[0002] As a non-invasive and low-cost medical imaging method, X-ray images are widely used in the screening and diagnosis of various lung diseases such as pulmonary tuberculosis and pneumonia. Traditionally, clinicians need to manually identify the lesion area, which is very time-consuming and laborious. Automated methods can reduce the workload of doctors, avoid misjudgment caused by long-term fatigue, greatly improve the efficiency of doctors' clinical diagnosis and treatment, and help the early diagnosis and treatment of related diseases. However, traditional segmentation methods, such as threshold-based methods, edge detection-based methods, and graph-based methods, still rely on manual means and require doctors to manually divide the region of interest first. Moreover, in the face of some complex backgrounds and features, the segmentation effects of these methods are often not ideal. With the development of deep learning technology, convolutional neural networks have gradually become one of the core technologies in computer vision and other fields. However, due to the inherent limitations of convolution, convolution operations often only involve adjacent pixels, making it impossible for the network to obtain long-range dependencies of features. Compared with convolutional neural networks, Transformer assigns different weights to each element of the input, calculates the correlation between each position in the sequence and other positions, thereby establishing long-range dependencies of features, capturing the context information of the feature map, and effectively overcoming the limitations of convolutional neural networks in long-range dependence modeling.

[0003] However, achieving accurate segmentation of lung X-ray images still faces significant challenges. On the one hand, the lung region has differences in morphology and size distribution, and traditional segmentation methods show limitations in capturing features of different scales, resulting in poor segmentation effects for target regions of different scales and irregular morphologies. On the other hand, the lung region is highly similar to other regions in terms of gray level and texture characteristics. It is difficult to effectively distinguish the boundary only relying on local features, and there is a lack of sufficient extraction of the global structure, easily resulting in over-segmentation or under-segmentation phenomena, leading to a decrease in segmentation accuracy. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a lung X-ray image segmentation system based on deep learning in view of the deficiencies of the prior art.

[0005] The technical solution of the present invention is as follows:

[0006] A lung X-ray image segmentation system based on deep learning, including a data preprocessing module, a lung X-ray image segmentation network module, a training module, and a segmentation result output module. The data preprocessing module collects the original medical image data and the corresponding manually segmented result data of the target area to form the original image dataset I and the label dataset L respectively; and preprocesses the original image dataset and the label dataset, and constructs a training dataset and a test dataset. The lung X-ray image segmentation network module constructs a lung X-ray image segmentation network based on deep learning, denoted as MSDA-Net. The training module trains the constructed MSDA-Net.

[0007] In the described lung X-ray image segmentation system, the lung X-ray image segmentation network module constructs a lung X-ray image segmentation network based on deep learning. The network structure uses UNet as the backbone network, including an encoder, four multi-scale feature extraction modules, a global context information perception module, and a decoder. The encoder includes an encoding module and a max pooling layer. One encoding module contains two convolutional layers with a convolutional kernel of 3×3 and a padding value of 1, a batch normalization layer, and a ReLU activation layer; five encoding module calculations and four max pooling layers are performed on the original input image for downsampling operations. The multi-scale feature extraction module contains dilated convolutions with a convolutional kernel of 3×3 with different dilation rates and a feature fusion module (FFM); the global context information perception module includes a non-local attention module and a spatial attention module. The features from the encoder are respectively input into the non-local attention module and the spatial attention module, and the results of these two branches are added element by element, and finally the output result G of the global context information perception module is obtained.

[0008] In the described lung X-ray image segmentation system, the four multi-scale feature extraction modules are used to extract and fuse the multi-scale information in the encoder layer by layer to enhance the feature expression ability; among them, the input of each module is constructed as follows:

[0009] For the 1st, 2nd, and 3rd multi-scale feature extraction modules, the input is the feature map X extracted by the current layer encoder l and the output M of the previous layer multi-scale feature extraction module l+1 :

[0010] M l = MSFE(X l , M l+1 ) l = 1, 2, 3

[0011] For the 4th multi-scale feature extraction module, in addition to inputting the feature map X4 extracted by the current layer encoder, the output result G of the global context information perception module is also input to supplement the global receptive field information, thereby improving the model's understanding ability of high-level semantics;

[0012] M4 = MSFE(X4, G).

[0013] For the lung X-ray image segmentation system described above, for the four multi-scale feature extraction modules, the specific process is as follows: First, concatenate the feature map X of the l-th layer l and the L-channel of the feature map to obtain X, where l = 1, 2, 3, 4. Then, input the feature map X into the first two branches. Input the results of the first branch and the second branch into the feature fusion module to obtain F1. Then, input the feature map X into the third branch to obtain x3. Input x3 and F1 into the feature fusion module to obtain F2. Finally, add F2 and X element-wise to obtain the final output result F, as shown in the following formula:

[0014] x1 = Conv 1×1 (X)

[0015] x2 = Conv 2×2 (X)

[0016] F1 = FFM(x1, x2)

[0017] x3 = Conv 4×4 (X)

[0018] F2 = FFM(x3, F1)

[0019] F = F2 + X

[0020] Where Conv k×k represents a dilated convolution with a dilation rate of k and a 3×3 convolution kernel, and FFM(·) represents the feature fusion module.

[0021] For the lung X-ray image segmentation system described above, the specific process of the feature fusion module includes: First, concatenate two feature maps f1 and f2 to obtain f 12 , pass through a 1×1 convolutional layer and a 3×3 convolutional layer, obtain the weight att through the Softmax function, then obtain the attention weights att1 and att2 respectively through the Split(·) operation. Multiply the feature maps f1 and f2 with the attention weights att1 and att2 element-wise to obtain f1' and f2', and finally add f1' and f2' element-wise to obtain f1'2, as shown in the following formula:

[0022] f 12 = Concat[f1, f2]

[0023] att = Softmax[Conv3(Conv1(f 12 ))]

[0024] att1, att2 = Split(att)

[0025] f1' = att1 × f1

[0026] f2' = att2 × f2

[0027] f1'2 = f1' + f2'

[0028] Where Concat[·] represents the channel concatenation operation, and Conv k represents the convolution operation with a convolution kernel of k×k.

[0029] The feature map f 12 is then respectively passed through the global average pooling layer and the global max pooling layer, through the 1×1 convolution operation, for channel concatenation and the Sigmoid activation function to obtain the attention weight att3, and f 12 is element-wise multiplied by att3 to obtain f3, and f3 is element-wise added to f1'2 to obtain the finally output result F final , as shown in the following formula:

[0030] att3 = Sigmoid[Concat(AvgPool(f 12 ), Maxpool(f 12 ))]

[0031] f3 = f 12 × att3

[0032] F final = f3 + f1'2.

[0033] In the described lung X-ray image segmentation system, the non-local attention module is used to calculate the similarity between input elements. In the non-local attention module, for each element x i , its similarity f(x i , x j ) with all other elements xj is calculated, and then all elements x j are weighted and summed according to these similarities to represent g(x j ), and the weighted output y i is obtained. Then normalization processing C(x) is performed so that the sum of all weights is 1;

[0034]

[0035] In the described lung X-ray image segmentation system, in the spatial attention module, the feature map is input into the convolutional layer to reduce the channel dimension to C' = C / 8, reducing the overall computational amount, and generating through reshaping and transposing operations to generate and F is input into the convolutional layer to generate Formed by the reshape operation Next, perform matrix multiplication and Softmax normalization operations between M and N to obtain the position attention map As shown in the following formula:

[0036]

[0037] S ij Represents the influence of the i-th position on the j-th position. Then perform matrix multiplication between W and the transpose of S ij And reshape the result into Finally, multiply it by the parameter α and perform an element-wise summation operation on the feature F to obtain the final output

[0038]

[0039] Among them, α is a learnable parameter, initialized to 0, and gradually assigns more weights through learning. The model that combines the non-local attention module and the spatial attention module can more effectively aggregate long-distance features and key spatial position relationships, improve the segmentation accuracy of the target region boundary, and enhance the model's understanding ability of global structure information.

[0040] For the described lung X-ray image segmentation system, the decoder includes a decoding module and an output layer; the decoding module includes two convolutional layers with a convolutional kernel of 3×3 and a padding value of 1, a batch normalization layer, and a ReLU activation layer, and also includes a bilinear interpolation layer and a channel splicing layer; the output layer is a 1×1 convolution, a batch normalization layer, and a ReLU activation layer; a total of four decoding module operations and one output layer operation are performed.

[0041] The present invention has the following beneficial effects:

[0042] 1. By introducing a multi-scale feature extraction module into the network and combining dilated convolutions with different dilation rates and feature fusion methods, the adaptability of the model to the lung region in the X-ray image is enhanced, the key target regions in the image are highlighted, irrelevant background noise is suppressed, and it can effectively cope with the significant differences in the shape and size of the lung region, improving the recognition ability for complex and irregularly shaped regions, thereby improving the segmentation effect of the target region.

[0043] 2. To solve the high similarity in gray scale and texture characteristics between the lung region and other regions, a global context information perception module is introduced into the network. By integrating the non-local attention mechanism and the spatial attention mechanism, the long-distance dependence relationship between features is captured, and the model's perception ability of the global structure is enhanced. Description of the Drawings

[0044] Figure 1: Flowchart of the lung X-ray image segmentation method based on deep learning;

[0045] Figure 2 : Flowchart of data preprocessing;

[0046] Figure 3 : Structure diagram of the MSDA-Net network;

[0047] Figure 4 (a): Diagram of the multi-scale feature extraction module;

[0048] Figure 4 (b): Feature fusion module (FFM);

[0049] Figure 5 : Diagram of the global context information perception module;

[0050] Figure 6 : Flowchart of the lung X-ray image segmentation system based on deep learning; Specific implementation manners

[0051] The present invention will be described in detail below in conjunction with specific embodiments.

[0052] A lung X-ray image segmentation method and system based on deep learning.

[0053] In a first aspect, the present invention provides a lung X-ray image segmentation method based on deep learning, and the overall process is as Figure 1 shown. It includes four steps: data preprocessing, constructing a lung X-ray image segmentation network, training the network, and outputting the segmentation result. The specific operations are as follows:

[0054] S1: Collect the original medical image data and the corresponding manually segmented result data of the target area to form the original image dataset I and the label dataset L respectively; and perform data preprocessing on the original image dataset and the label dataset, and construct the training dataset and the test dataset;

[0055] S2: Construct a lung X-ray image segmentation network based on deep learning, denoted as MSDA-Net.

[0056] S3: Train the constructed MSDA-Net.

[0057] S4: Use the trained network to segment the target area.

[0058] The data preprocessing method process in step S1 includes the steps: three-dimensional medical image slicing, two-dimensional image normalization, and two-dimensional image scaling. Each part will be described separately below.

[0059] If the medical image is a 3D image, each original medical image data and the corresponding target region label data are sliced into 2D along the cross-section. To accelerate the model convergence and improve the model performance, first, all 2D images are normalized. Then, the normalized medical image data and the corresponding labels are respectively used to construct the original image training dataset and the test dataset according to the ratio of 80% and 20%. All image data are scaled using the resize() function, and the image size is scaled to 224×224. The specific process is as Figure 2 shown.

[0060] In step S2, a deep learning-based lung X-ray image segmentation network is constructed. The network structure is as Figure 3 shown. The UNet is used as the backbone network, including an encoder, four multi-scale feature extraction modules, a global context information perception module, and a decoder, specifically as follows:

[0061] The encoder includes an encoding module and a max pooling layer. One encoding module contains two convolutional layers with a convolutional kernel of 3×3 and a padding value of 1, a batch normalization layer, and a ReLU activation layer. The original input image is subjected to five encoding module calculations and four max pooling layer downsampling operations.

[0062] The multi-scale feature extraction module is as Figure 4 (a) shown, and contains dilated convolutions with a convolutional kernel of 3×3 with different dilation rates and a feature fusion module (FFM).

[0063] In this network structure, a total of four multi-scale feature extraction modules (Multi-Scale Feature Extraction Module, MSFE) are designed to layer-by-layer extract and fuse the multi-scale information in the encoder, enhancing the feature expression ability. Among them, the inputs of each module are as follows:

[0064] For the 1st, 2nd, and 3rd multi-scale feature extraction modules, the input is the feature map X l extracted by the current layer encoder and the output M l+1 of the previous layer multi-scale feature extraction module:

[0065] M l = MSFE(X l , M l+1 ) l = 1, 2, 3

[0066] For the 4th multi-scale feature extraction module, in addition to the input of the feature map X4 extracted by the current layer encoder, the output result G of the global context information perception module is also input to supplement the global receptive field information, thereby improving the model's understanding ability of high-level semantics.

[0067] M4 = MSFE(X4, G)

[0068] To simplify the description process, in the following process description, the feature map passed in from the previous stage is uniformly denoted as L:

[0069] When l = 1, 2, 3, L = M l+1

[0070] When l = 4, L = G

[0071] The specific operations are as follows: First, concatenate the feature map X of the l-th layer l and the feature map L in the channel dimension to obtain X, where l = 1, 2, 3, 4. Then, input the feature map X into the first two branches. Input the results of the first branch and the second branch into the feature fusion module to obtain F1. Next, input the feature map X into the third branch to obtain x3. Input x3 and F1 into the feature fusion module to obtain F2. Finally, add F2 and X element-wise to obtain the final output result F, as shown in the following formula:

[0072] x1 = Conv 1×1 (X)

[0073] x2 = Conv 2×2 (X)

[0074] F1 = FFM(x1, x2)

[0075] x3 = Conv 4×4 (X)

[0076] F2 = FFM(x3, F1)

[0077] F = F2 + X

[0078] where Conv k×k represents the dilated convolution with a dilation rate of k and a 3×3 convolutional kernel, and FFM(·) represents the feature fusion module.

[0079] The feature fusion module (FFM) is as shown in Figure 4 (b). The specific operations of the feature fusion module include: First, concatenate two feature maps f1 and f2 in the channel dimension to obtain f 12 , pass through a 1×1 convolutional layer and a 3×3 convolutional layer, obtain the weight att through the Softmax function, then obtain the attention weights att1 and att2 respectively through the Split(·) operation on att. Multiply the feature map f1 and the feature map f2 with the attention weights att1 and att2 element-wise to obtain f1' and f2', and finally add f1' and f2' element-wise to obtain f1'2, as shown in the following formula:

[0080] f 12 = Concat[f1, f2]

[0081] att = Softmax[Conv3(Conv1(f 12 ))]

[0082] att1, att2 = Split(att)

[0083] f1' = att1 × f1

[0084] f2' = att2 × f2

[0085] f1'2 = f1' + f2'

[0086] Among them, Concat[·] represents the channel concatenation operation, and Conv k represents the convolution operation with a convolution kernel of k×k.

[0087] The feature map f 12 is then respectively passed through the global average pooling layer and the global maximum pooling layer, through the 1×1 convolution operation, for channel concatenation and the Sigmoid activation function to obtain the attention weight att3. Multiply f 12 element-wise with att3 to obtain f3, and add f3 and f1'2 element-wise to obtain the finally output result F final , as shown in the following formula:

[0088] att3 = Sigmoid[Concat(AvgPool(f 12 ), Maxpool(f 12 ))]

[0089] f3 = f 12 × att3

[0090] F final = f3 + f1'2

[0091] The global context information perception module is as Figure 5 shown, mainly including the non-local attention module and the spatial attention module. The specific process is as follows: The features from the encoder are respectively input into the non-local attention module and the spatial attention module, and then the results of these two branches are added element-wise, and finally the output result G of the global context information perception module is obtained.

[0092] The non-local attention module is used to calculate the similarity between input elements. In the non-local attention module, for each element x i , calculate its similarity f(x j , x i , x j ) with all other elements x j , and then weighted sum the representations g(x j) to obtain the weighted output y i . Then perform normalization processing C(x) so that the sum of all weights is 1.

[0093]

[0094] In the spatial attention module, the feature map is input into the convolutional layer to reduce the channel dimension to C' = C / 8, reducing the overall computational amount, and generating generating through shaping and transposing operations and F is input into the convolutional layer to generate forming through the reshape operation Next, perform matrix multiplication and Softmax normalization operations between M and N to obtain the position attention map as shown in the following formula:

[0095]

[0096] S ij represents the influence of the i-th position on the j-th position. Then perform matrix multiplication between W and the transpose of S ij and reshape the result into Finally, multiply it by the parameter α and perform an element-wise summation operation on the feature F to obtain the final output

[0097]

[0098] where α is a learnable parameter, initialized to 0, and gradually assigns more weights through learning. The model that combines the non-local attention module and the spatial attention module can more effectively aggregate long-distance features and key spatial position relationships, improve the segmentation accuracy of the target region boundary, and enhance the model's understanding ability of global structural information.

[0099] The decoder includes a decoding module and an output layer. One of the decoding modules includes two convolutional layers with a convolutional kernel of 3×3 and a padding value of 1, a batch normalization layer, and a ReLU activation layer, and also includes a bilinear interpolation layer and a channel concatenation layer. The output layer is a 1×1 convolution, a batch normalization layer, and a ReLU activation layer. A total of four decoding module operations and one output layer operation are required.

[0100] In step S3, the model is trained using the training dataset in step S1. The specific steps are as follows: First, the images in the training dataset are input into the network for calculation to obtain the calculation results of this round of iteration. The results are compared with the corresponding labels, and the loss value is calculated using the loss function. Second, the gradients are calculated by the Adam optimizer and the parameters in the neural network are updated through backpropagation. Then, the above process is iterated until the model converges to obtain the network training model. Finally, the model is verified using the test dataset. During the training process, a combined strategy of weighted binary cross-entropy loss function and weighted intersection over union loss function is adopted. Compared with the traditional binary cross-entropy loss function, the weighted binary cross-entropy loss function assigns more weights to the pixels at the boundary, while the weighted intersection over union loss function assigns higher weights to the pixels with higher prediction uncertainty, prompting the model to focus on optimizing these fuzzy regions. The loss function is shown as follows:

[0101]

[0102] where P is the prediction result, G is the ground truth result, w and h are the width and height of the input image, (x, y) represents the coordinates of each pixel in the image, and ε is the weight hyperparameter.

[0103] In a second aspect, the present invention proposes a lung X-ray image segmentation system based on deep learning. The specific process is as Figure 6 shown, mainly including:

[0104] Data input module 101: Collect a lung X-ray image dataset, including original CT image files and annotated image files.

[0105] Data preprocessing module 102: Perform preprocessing operations on the collected data. If the data is a three-dimensional image, slice the data along the Z-axis to form two-dimensional images, normalize all images, and then construct a training dataset and a test dataset according to the ratios of 80% and 20% respectively. Scale all image data, and scale the image size to 224×224.

[0106] Segmentation model construction module 103: Establish an image segmentation model according to the above method, including establishing an encoder, a multi-scale feature extraction module, a global context information perception module, and a decoder.

[0107] Model training module 104: Input the original CT images and corresponding labels in the training set into the model for training. Calculate the error between the prediction result and the ground truth label using the loss function, update the weights in the network according to the gradient through backpropagation and using the Adam optimizer. This process is iterated continuously until the model converges. Finally, predict the CT images in the test set and output the prediction results.

[0108] Evaluation module 105: Evaluate the performance of the segmentation results, and use metrics such as the Dice coefficient and sensitivity to evaluate the model's effectiveness.

[0109] Example 1: Construct a lung X-ray image segmentation method based on the JSRT and Montgomery lung X-ray datasets

[0110] Combine the publicly available datasets from the Japan Society of Radiological Technology (JSRT) and Montgomery to form an experimental dataset. The original data and the corresponding target region label data in the dataset are respectively used to form dataset I and label dataset L, and further preprocessing operations are performed on the data. Manually segment the results of each data and the corresponding target region along the Z-axis to form two-dimensional images, and normalize all the images; construct a training dataset and a test dataset according to the ratio of 80% and 20% for the normalized data and the corresponding target region label data, and scale all the image data so that the image size is scaled to 224×224.

[0111] Taking the input of an X-ray data as an example, construct and train a deep learning-based lung X-ray image segmentation method (MSDA-Net). The specific steps are as follows:

[0112] First, construct an encoder. The encoder includes an encoding module and a max pooling layer. One encoding module contains two convolutional layers with a convolutional kernel of 3×3 and a padding value of 1, a batch normalization layer, and a ReLU activation layer. Perform five convolutional encoding module calculations and four max pooling layer operations on the original input image for downsampling.

[0113] Secondly, construct a multi-scale feature extraction module. The multi-scale feature extraction module contains dilated convolutions with a convolutional kernel of 3×3 with different dilation rates and a feature fusion module (FFM). Specifically, it can be divided into three branches. The first branch is a dilated convolution with a dilation rate of 1; the second branch contains a dilated convolution with a dilation rate of 2; the third branch contains a dilated convolution with a dilation rate of 4. First, input the input feature map X into the first two branches respectively. Input the results of the first branch and the second branch into the feature fusion module to get F1. Then input the feature map X into the third branch to get x3. Input x3 and F1 into the feature fusion module to get F2. Finally, add F2 and X element-wise to get the final output result F.

[0114] Thirdly, construct a global context information perception module. The global context information perception module mainly includes non-local attention and spatial attention. The specific process is as follows: Input the features from the encoder into the non-local attention module and the spatial attention module respectively, and then add the results of these two branches element-wise to get the final output result.

[0115] Then, a decoder is constructed, which includes a decoding module and an output layer. One of the decoding modules includes two 3×3 convolutional kernels with a padding value of 1, a batch normalization layer, a ReLU activation layer, and also includes a bilinear interpolation layer and a channel concatenation layer. The output layer is a 1×1 convolution, a batch normalization layer, and a ReLU activation layer. The input image has to go through four decoding modules and one output layer in total.

[0116] Finally, the model is trained. The specific steps are as follows: First, the images in the training dataset are input into the network for calculation to obtain the calculation results of this round of iteration, and the results are compared with the corresponding labels, and the loss value is calculated using the loss function; Second, the gradients are calculated by the Adam optimizer and the parameters in the neural network are updated through backpropagation; Then, the above process is iterated until the model converges, and the test dataset is used to verify the model. During the training process, a combined strategy of weighted binary cross-entropy loss function and weighted intersection over union loss function is adopted.

[0117]

[0118] The above is the solution of the lung X-ray image segmentation method based on deep learning in this embodiment.

[0119] Table 1 Experimental results of MSDA-Net on the Chest X-ray dataset

[0120]

[0121] As shown in Table 1, compared with the classic UNet model on the Chest X-ray dataset, MSDA-Net has obvious improvements in various performance indicators. Among them, the Dice coefficient has increased from 96.97% to 97.35%, the Jaccard index has increased from 95.64% to 95.93%, and the F1-score has increased from 97.77% to 97.92%, indicating that MSDA-Net has more advantages in terms of segmentation accuracy. Especially in the Sensitivity index, MSDA-Net has reached 97.94%, which is 1.07% higher than the UNet model, indicating that MSDA-Net can identify the lesion area more comprehensively and effectively reduce the risk of missed detection. Overall, MSDA-Net can enhance the sensitivity of the model to the lesion area while ensuring high accuracy.

[0122] Example 2: Construct a lung X-ray image segmentation system based on deep learning according to the segmentation results

[0123] The solution of the lung X-ray image segmentation system based on deep learning is basically the same as the solution of the above-mentioned lung X-ray image segmentation method based on deep learning. The lung X-ray image segmentation system based on deep learning in this embodiment mainly includes the following steps:

[0124] Data input module 101: Collect a dataset of lung X-ray images, including original CT image files and annotated image files.

[0125] Data preprocessing module 102: Perform preprocessing operations on the collected data. If the data is a three-dimensional image, slice the data along the Z-axis to form two-dimensional images, and normalize all images. Then, construct a training dataset and a test dataset according to the ratios of 80% and 20% respectively, and scale all image data so that the image size is scaled to 224×224.

[0126] Segmentation model module 103: Establish an image segmentation model according to the above method, including establishing an encoder, a multi-scale feature extraction module, a global context information perception module, and a decoder.

[0127] Training model module 104: Input the original CT images and corresponding labels of the training set into the model for training. Calculate the error between the predicted result and the true label using a loss function, update the weights in the network according to the gradient through backpropagation and using the Adam optimizer. This process is iterated until the model converges. Finally, predict the CT images of the test set and output the prediction results.

[0128] Evaluation module 105: Evaluate the performance of the segmentation results, and use metrics such as the Dice coefficient and sensitivity to evaluate the model effect.

[0129] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations shall fall within the protection scope of the appended claims of the present invention.

Claims

1. A lung X-ray image segmentation system based on deep learning, characterized in that, It includes a data preprocessing module, a lung X-ray image segmentation network module, a training module, and a segmentation result output module. The data preprocessing module collects the original medical image data and the corresponding manually segmented result data of the target region to form the original image dataset I and the label dataset L respectively; and preprocesses the original image dataset and the label dataset, and constructs a training dataset and a test dataset. The lung X-ray image segmentation network module constructs a deep learning-based lung X-ray image segmentation network, denoted as MSDA-Net. The training module trains the constructed MSDA-Net.

2. The lung X-ray image segmentation system according to claim 1, wherein The lung X-ray image segmentation network module constructs a deep learning-based lung X-ray image segmentation network. The network structure uses UNet as the backbone network, including an encoder, four multi-scale feature extraction modules, a global context information perception module, and a decoder. The encoder includes an encoding module and a max pooling layer. One encoding module contains two convolutional layers with a convolutional kernel of 3×3 and a padding value of 1, a batch normalization layer, and a ReLU activation layer. Five encoding module calculations and four max pooling layer operations are performed on the original input image for downsampling. The multi-scale feature extraction module contains dilated convolutions with a convolutional kernel of 3×3 with different dilation rates and a feature fusion module (FFM). The global context information perception module includes a non-local attention module and a spatial attention module. The features from the encoder are respectively input into the non-local attention module and the spatial attention module, and the results of these two branches are added element by element. Finally, the output result G of the global context information perception module is obtained.

3. The lung X-ray image segmentation system according to claim 2, wherein The four multi-scale feature extraction modules are used to extract and fuse the multi-scale information in the encoder layer by layer to enhance the feature expression ability. Among them, the input of each module is composed as follows: For the 1st, 2nd, and 3rd multi-scale feature extraction modules, the input is the feature map X extracted by the current layer encoder l and the output M of the previous layer multi-scale feature extraction module l+1 : M l = MSFE(X l , M l+1 ), l = 1, 2, 3 For the 4th multi-scale feature extraction module, in addition to inputting the feature map X4 extracted by the current layer encoder, the output result G of the global context information perception module is also input to supplement the global receptive field information, thereby improving the model's understanding ability of high-level semantics. M4 = MSFE(X4, G).

4. The lung X-ray image segmentation system according to claim 3, characterized in that The specific process of the four multi-scale feature extraction modules is as follows: First, the feature map X of the l-th layer is concatenated with the feature map L in the channel dimension to obtain X, where l = 1, 2, 3, 4. Then, the feature map X is input into the first two branches. The results of the first branch and the second branch are input into the feature fusion module to obtain F1. Next, the feature map X is input into the third branch to obtain x3. x3 and F1 are input into the feature fusion module to obtain F2. Finally, F2 and X are added element-wise to obtain the final output result F. l ​ 5. The lung X-ray image segmentation system according to claim 4, characterized in that, The specific calculation formulas of the four multi-scale feature extraction modules are as follows: x1 = Conv 1×1 (X) x2 = Conv 2×2 (X) F1 = FFM(x1, x2) x3 = Conv 4×4 (X) F2 = FFM(x3, F1) F = F2 + X Among them, Conv k×k represents the dilated convolution with a dilation rate of k and a 3×3 convolution kernel, and FFM(·) represents the feature fusion module.

6. The lung X-ray image segmentation system according to claim 2, wherein The specific process of the feature fusion module is as follows: First, the two feature maps f1 and f2 are concatenated along the channel dimension to obtain f 12 , which passes through a 1×1 convolutional layer and a 3×3 convolutional layer. The weight att is obtained through the Softmax function. Then, the att is used to obtain the attention weights att1 and att2 respectively through the Split(·) operation. The feature maps f1 and f2 are multiplied element-wise with the attention weights att1 and att2 respectively to obtain f1' and f2'. Finally, f1' and f2' are added element-wise to obtain f’ 12 , as shown in the following formula: f 12 = Concat[f1, f2] att = Softmax[Conv3(Conv1(f 12 ))] att1, att2 = Split(att) f1' = att1 × f1 f2' = att2 × f2 f’ 12 = f1' + f2' Among them, Concat[·] represents the channel concatenation operation, and Conv k represents the convolution operation with a convolution kernel of k×k; Take the feature map f 12 Then, respectively pass through the global average pooling layer and the global max pooling layer, perform 1×1 convolution operation, channel concatenation and Sigmoid activation function to obtain the attention weight att3. Multiply f 12 element-wise with att3 to obtain f3. Add f3 and f’ 12 element-wise to obtain the final output result F final , as shown in the following formula: att3 = Sigmoid[Concat(AvgPool(f 12 ), Maxpool(f 12 ))] f3 = f 12 × att3 F final = f3 + f' 12 .

7. The lung X-ray image segmentation system according to claim 2, wherein The non-local attention module is used to calculate the similarity between input elements. In the non-local attention module, for each element x i , calculate its difference with all other elements x j The similarity f(x i ,x j ), and then weighted sum all elements x according to these similarities j The representation of g(x j ), and get the weighted output y i ; Then normalize C(x) so that the sum of all weights is 1; 8. The lung X-ray image segmentation system according to claim 2, characterized in that, In the spatial attention module, the feature map is input into the convolutional layer to reduce the channel dimension to C' = C / 8, reducing the overall computational load, and generating generating through shaping and transposing operations and F is input into the convolutional layer to generate forming through the reshape operation Next, matrix multiplication and Softmax normalization operations are performed between M and N to obtain the position attention map 9. The lung X-ray image segmentation system according to claim 8, wherein The specific calculation formulas are as follows: S ij represents the influence of the \(i\)-th position on the \(j\)-th position; then perform matrix multiplication between \(W\) and the transpose of \(S\) ij and reshape the result into Finally, multiply it by the parameter \(\alpha\) and perform an element-wise summation operation on the feature \(F\) to obtain the final output Among them, α is a learnable parameter, initialized to 0, and gradually assigns more weights through learning. By combining the non-local attention module and the spatial attention module, the model can more effectively aggregate long-distance features and key spatial position relationships, improve the segmentation accuracy of the target region boundary, and enhance the model's understanding ability of global structural information.

10. The lung X-ray image segmentation system according to claim 1, wherein The decoder includes a decoding module and an output layer; the decoding module includes two convolutional layers with a convolution kernel of 3×3 and a padding value of 1, a batch normalization layer, and a ReLU activation layer, and also includes a bilinear interpolation layer and a channel splicing layer; the output layer is a 1×1 convolution, a batch normalization layer, and a ReLU activation layer; a total of four decoding module operations and one output layer operation are performed.

Citation Information

Cited By

  • Medical image segmentation model training method, segmentation method, equipment and storage medium

    CN120580251A

  • Generation method and system based on lung DR image to CT breathing motion image

    CN120976356A

  • Multi-branch collaborative network-based uncertainty perception pulmonary nodule segmentation method

    CN121095273A

  • Lung airway segmentation and repair method and system

    CN121121096A

  • A method and system for lung airway segmentation and repair

    CN121121096B