Weak supervision lung X-ray image segmentation method based on improved cyclic generative adversarial network

By using an improved LS-CycleGAN network, leveraging asymmetric data annotation and many-to-many mapping, and combining SE-Blocks and MOA-Blocks modules, the problems of poor dataset annotation quality and insufficient sample size are solved, achieving efficient and accurate segmentation of lung X-ray images and supporting medical diagnosis.

CN120953295APending Publication Date: 2025-11-14HUNAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410589160.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-13
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing deep learning networks face problems such as poor dataset labeling quality and insufficient sample size in lung X-ray image segmentation tasks, leading to unstable model training and overfitting, making it difficult to generate accurate segmentation masks.

Method used

An improved LS-CycleGAN network is adopted, which uses asymmetric data annotation for many-to-many mapping. Combined with SE-Blocks and MOA-Blocks modules, the generator and discriminator perform fully convolutional discrimination. The LPIPS2 loss function is used to improve the training efficiency of the generator and the segmentation accuracy.

Benefits of technology

Under weak supervision, the generator can better focus on the edge details of lung X-ray images, the accuracy of the discriminator is improved, and the segmentation results have significantly improved IoU, Dice, VOE and RVD indices, supporting fast and accurate segmentation for medical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953295A_ABST
    Figure CN120953295A_ABST
Patent Text Reader

Abstract

The invention provides a weak supervision lung X-ray image segmentation method based on an improved cyclic generative adversarial network, is suitable for a network model-LS-CycleGAN for lung X-ray image segmentation, and aims to solve the problem of too few accurate segmentation labels in actual clinical treatment. In the LS-CycleGAN, the improved generator is more focused on the edge details of the image, the introduction of SEBlock improves the defect that the original network is easy to over-fit, and the introduction of MOABlock helps to improve the quality of the generated image. The output result of the improved discriminator is not a Boolean value any more, the discriminating accuracy is improved by outputting a discriminating score matrix through the full convolutional network, in addition, the improved and designed discriminator network can be independently and preferentially trained, the training time is saved, the loss of the generator is helped to be converged more quickly when the discriminator network and the generator are jointly trained, and the reliability of the discriminator network is improved. And the overall training time is shortened. Most importantly, many-to-many mapping can be generated by relying on a loop structure of the network model, a data original image in a training set and a segmentation label, the limitation that segmentation images and labels need one-to-one mapping is overcome, the model can be suitable for more clinical conditions, and the defects that in clinical medical examination, the data size is small, and the efficiency is high are overcome. And the label segmentation quality is poor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a weakly supervised lung X-ray image segmentation method based on deep learning. Specifically, this invention proposes a technique for segmenting lung X-ray images with asymmetric data labels based on an improved recurrent generative adversarial network (RGAN) deep learning network. This technique can serve as an auxiliary tool for physicians to diagnose abnormalities in lung imaging and has promising application prospects. Background Technology

[0002] Lung disease is the third leading cause of death worldwide. According to data provided by the World Health Organization, nearly 4 million people die annually from lung tissue lesions between 2019 and 2022, with 454.6 million cases of lung disease. Lung X-ray is a medical imaging technique used to observe lung structure and detect lung-related diseases. Many common lung diseases require lung X-ray as an adjunct to treatment. For long-term smokers or high-risk individuals with a family history of lung cancer, doctors can use lung X-ray for health screening to detect lung cancer early (Smith RA, Andrews KS, Brooks D, et al. Cancer screening in the United States, 2019: A review of current American Cancer Society guidelines and current issues in cancer screening[J]. CA: acancerjournal for clinicians, 2019, 69(3):184-210.). If lung nodules are found during other examinations, a lung X-ray can provide more detailed information to help doctors determine the nature of the nodules (benign or malignant), assess the lung nodules, or evaluate airway narrowing and lung tissue damage (Bankier AA, MacMahon H, Colby T, et al. Fleischner Society: glossary of terms for thoracic imaging[J]. Radiology, 2024, 310(2):e232558.), which helps determine the extent of the condition and monitor disease progression.

[0003] Traditional methods for segmenting lung parenchyma from X-ray images of the lungs typically involve manual labeling and segmentation by medical professionals. However, thanks to the rapid development of deep learning, there are now many ways to use computers for image segmentation (Hofmanninger J, Prayer F, Pan J, et al. Automatic lung segmentation in routine imaging is primarily a data diversity problem, not a methodology problem[J]. European Radiology Experimental, 2020, 4: 1-13.). In the field of medical image processing, there are already a variety of common network structures capable of handling most segmentation tasks. For example, there is the U-net (Ronneberger O, Fischer P, Brox TU-net: Convolutional networks for biomedical image segmentation[C] / / Medical image computing and computer-assisted intervention–MICCAI 2015:18th international conference,Munich,Germany,October 5-9,2015,proceedings,part III 18.Springer InternationalPublishing,2015:234-241.), a region-based convolutional neural network. It extracts and abstracts features from the input image through an encoder, and then performs layer-by-layer upsampling and feature fusion through a decoder to finally obtain the segmentation result. Another example is the Mask R-CNN (He K, Gkioxari G, Dollár P, et al.Mask r-cnn[C] / / Proceedings of the IEEE international conference on computer Vision. 2017: 2961-2969. This network combines the ideas of object detection and semantic segmentation. It extracts candidate regions through an object detection network, and then further segments each candidate region to generate a binary mask for each pixel, thereby obtaining a higher quality segmentation result; or it integrates a backbone network (usually ResNet (He K, Zhang X, Ren S, et al.).Deepresidual learning for image recognition[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition.2016:770-778.), Xception(CholletF. andpatternrecognition.2017:1251-1258.)), dilated convolution (Szegedy C, Vanhoucke V, Ioffe S, etal. Rethinking the inception architecture for computer vision [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition.2016:2818-2826.), multi-scale pyramid pooling (He K, Zhang X, Ren S, et al. al.Spatial pyramidpooling indeep convolutional networks for visual recognition[J].IEEE transactions onpattern analysis and machine The DeepLab model (Chen LC, Papandreou G, Kokkinos I, et al.) is a network model that incorporates key technologies such as fully connected conditional random fields (Sainath TN, Vinyals O, Senior A, et al. Convolutional, long short-term memory, fully connected deep neural networks [C] / / 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2015: 4580-4584.) and other technologies.Deeplab: Semantic image segmentation with deep convolutional nets, atrousconvolution, and fully connected crfs[J]. IEEE transactions on pattern analysis and machine intelligence, 2017, 40(4):834-848.). .

[0004] In addition, many segmentation networks are applied in the practice of medical image segmentation. However, when most network models face medical image segmentation tasks, two major factors affect the segmentation effect of the model. One is the accuracy of the data labeling in the training set, which will directly affect the loss function of the model during the training process. If the data labeling is not accurate, the model will have difficulty learning useful information during the training process, resulting in underfitting of the model training results. The other is that the training dataset has a small sample size, the model has poor generalization ability, and is prone to overfitting. When adjusting complex model parameters, the small amount of data will make the process of estimating parameters difficult and the results will be inaccurate, thus making the model training process unstable and unable to converge

[22] . Therefore, reducing the influence of the above two factors on the training process has become a major focus of improving the accuracy of deep learning network models in lung X-ray image segmentation.

[0005] Therefore, when segmenting lung X-ray images, problems such as small dataset size and poor annotation quality arise. This invention proposes a novel deep learning network architecture: LS-CycleGAN, a weakly supervised segmentation network based on an improved CycleGAN suitable for lung X-ray segmentation. This model aims to segment lung X-ray images using asymmetric annotation training and generate segmentation masks, thereby helping medical personnel quickly and accurately segment lung X-ray images, thus contributing to the diagnosis and treatment of lung diseases. This invention validates the method using the publicly available dataset LIDC-IDRI.

[0006] Some terms:

[0007] Machine learning: Machine learning is a branch of artificial intelligence that aims to enable computers to learn and extract patterns from data by building and training models, thereby enabling prediction and decision-making on unknown data.

[0008] Deep learning: Deep learning is a branch of machine learning that simulates the working principle of human neural networks, and achieves learning and analysis of data by building deep neural network models.

[0009] Convolutional Neural Networks (CNNs): A CNN is a neural network model specifically designed for processing data with a grid structure, such as images. It has important applications in image processing and computer vision.

[0010] The computational cost of a convolutional neural network (CNN) refers to the total number of computational operations between its layers. A common metric for computational cost is the number of floating-point multiplication operations (FLOPs).

[0011] The number of parameters in a convolutional neural network is an important indicator of the network's size and complexity. A larger number of parameters usually means a more complex network structure and a more powerful expressive ability.

[0012] Depth of Convolutional Neural Networks: The depth of a convolutional neural network refers to the number of layers in the network. Depth is an important attribute of convolutional neural networks, and it has a significant impact on the network's expressive and learning capabilities.

[0013] Width of a Convolutional Neural Network: The width of a convolutional neural network refers to the number of channels or nodes in each layer of the network. Width is an important attribute of convolutional neural networks, affecting the network's capacity and complexity.

[0014] Squeeze and Excitation (SE) is an attention mechanism used in computer vision tasks. The SE module learns the weights between channels, enabling the network to adaptively adjust the importance of channels, thereby improving the model's expressive power.

[0015] MoA-Transformer: A multi-resolution overlapping attention (MOA) module that can be inserted after each stage of the LocalTransformer to facilitate information communication with nearby windows and all non-local windows by providing global information exchange across all windows of the LocalTransformer with minimal computational cost and parameters. Summary of the Invention

[0016] The purpose of this invention is to provide a new weakly supervised segmentation model, called LS-Cyclegan, which can use the original lung X-ray image and partial annotations to learn a many-to-many mapping between two domains and generate a segmentation mask for the lung X-ray image. The advantages of this invention are: (1) Compared with the mainstream fully supervised segmentation network model, a new segmentation approach is proposed, breaking through the limitation of a one-to-one mapping between the image to be segmented and the segmentation mask. An asymmetric training image domain is adopted, and the model characteristics are used to generate a many-to-many mapping to achieve the purpose of weakly supervised training. (2) The improved generator can pay more attention to the edge details of the lung parenchyma in the lung X-ray. Under the action of the improved SE-Blocks and MOA-Blocks, the network can simultaneously pay attention to the edge details and global features of the image. (3) The improved discriminator network can perform fully convolutional discrimination on the generated image. Instead of outputting a single Boolean value, it outputs a discrimination score, which greatly improves the discrimination accuracy. In order to achieve the above objectives, this invention adopts the following technical solution:

[0017] (1) Divide the samples that need to be input into the model for training into a training set and a validation set in a ratio of 5:1;

[0018] (2) For the training set, the pydicom library was used to read the original DICOM lung X-ray image file information and batch-processed it into png format to facilitate the LS-cycleGAN network to read the input. Multiple training set structures were set up. TrainA contained the original lung CT images in the image domain, and Train B contained the segmentation mask labels in the image domain. A total of five training structures (TrainA-Train B) were set up: 20-20, 50-20, 100-20, 150-20, and 200-20. The number of segmentation masks was kept at a low level, and the number of original lung CT images was increased in turn to increase the number of many-to-many mappings, thereby stimulating the loss functions of the generator network and the discriminator network to converge quickly.

[0019] (3) After multiple training sessions, common medical image segmentation metrics such as IoU, Dice, VOE, and RVD on the validation set are monitored and saved. The weights that perform best on these metrics are selected as the final weights to obtain the optimized model.

[0020] (4) Select a test sample and adjust the image size to 256×256. Then perform center cropping on the image, cropping the image from the center to 256×256 so as to retain the central part of the image. This part of the image is processed as domain A (TrainA), and random segmentation mask labels are selected as domain B (TrainB).

[0021] (5) Use the processed datasets of domain A and domain B as input to the optimization model, and output the mask label of the image to be segmented (domain A) according to the model output. Attached Figure Description

[0022] Figure 1 This is a structural diagram of the improved LS-CycleGAN model in this invention;

[0023] Figure 2 This is a diagram of the improved generator structure in this invention;

[0024] Figure 3 This is a structural diagram of the SE-Blocks module in this invention;

[0025] Figure 4 This is a structural diagram of the MOA-Blocks module in this invention;

[0026] Figure 5 This is a structural diagram of the improved discriminator in this invention;

[0027] Figure 6 This is a visual quantization diagram of the training loss in this invention;

[0028] Figure 7 This is a diagram illustrating the evaluation metrics for training results in an asymmetric dataset, as presented in this invention.

[0029] Figure 8 This is a diagram illustrating the evaluation metrics for training results in a symmetric dataset, as presented in this invention.

[0030] Figure 9 This is a comparison diagram of the segmentation results and labels in the test set of this invention;

[0031] Specific implementation methods and model performance evaluation

[0032] This invention was implemented on a PC provided by the Intelligent Information Processing Laboratory of Hunan University of Technology. The hardware environment for training the network model consisted of an AMD Ryzen 95900HX, an Nvidia RTX 3060, 32GB of RAM, and a Linux operating system. The deep learning framework used was PyTorch. To objectively compare the performance of each network model and avoid the influence of the training mechanism on the experiment, the parameters designed in the experiment were processed in the same way. Each network model was trained for 200 epochs on the same dataset, with a batch size of 16 samples.

[0033] The present invention is generally divided into the following contents:

[0034] 1. The improved LS-CycleGAN generator structure pays more attention to edge details in grayscale images, making it suitable for lung X-ray image segmentation tasks. The design of the SE-Blocks block in the middle layer of the ResNet structure can prevent overfitting caused by a small number of training samples. The addition of MOA-Blocks blocks in the skip links helps generate more detailed segmentation results.

[0035] 2. The discriminator network was reconstructed and designed as a fully convolutional network to help output the discrimination matrix graph and improve the discrimination accuracy. Furthermore, the use of an Auto-Encoder allows for independent training before the generator, stimulating the generator network's training during actual training, reducing training time, and resolving the model training imbalance problem.

[0036] A novel recurrent adversarial loss, LPIPS2, was designed to adapt to weakly supervised tasks. In lung X-ray image segmentation, LPIPS2 significantly outperformed the original network's L2 loss, helping to stabilize the connections between multiple generators.

[0037] Model Performance Evaluation: In this invention, to accurately verify the model's weakly supervised segmentation performance on lung X-rays, performance comparisons were conducted on five asymmetric datasets and three symmetric datasets. Four common performance evaluation metrics were selected: IoU (i.e., Jaccard), Dice, VOE, and RVD. IoU is the overlap region between the predicted segment and the label divided by the joint region between the predicted segment and the label, measuring the overall accuracy of the classification model. The Dice coefficient is defined as twice the intersection divided by the sum of pixels, measuring the accuracy in image edge details. VOE can be called volume overlap error, representing the edge error rate. RVD represents the difference in mathematical volume between the two, measuring the overall error rate of the segmentation results. Specific experimental results are as follows: Figure 7 Figure 8 As shown, experiments were conducted on eight different asymmetric dataset structures. The training results exhibited significant differences depending on the number of elements in the two image domains of the training set with various combinations. While maintaining the number of elements in the TrainB labeled image domain, increasing the number of elements in the TrainA dataset resulted in an upward trend in various evaluation metrics. Finally, in the 200-20 dataset combination, the four performance metrics—IoU, Dice, VOE, and RVD—reached 0.963, 0.925, 0.043, and 0.005, respectively. This demonstrates that the improved LS-CycleGAN can generate many-to-many mappings as expected, compensating for the negative impact of insufficient segmentation labels during training. Figure 9As shown, the segmented images on the test set were visualized and compared, yielding good results. Therefore, LS-CycleGAN is determined to have excellent performance in weakly supervised lung X-ray segmentation tasks, effectively assisting medical professionals in clinical diagnosis.

[0038] Training Efficiency Evaluation: In this invention, to evaluate the model training efficiency, a training visualization analysis is performed on the improved generator, discriminator, and recurrent adversarial loss. To facilitate the adjustment of hyperparameters during training and the observation of training process details, we quantify multiple types of loss during training. By analyzing the three main losses during training, we can see that the convergence speed is relatively fast before 50 epochs due to the stimulation of the pre-trained discriminator. The loss gradient decreases rapidly before 50-100 epochs, and G_loss also converges relatively quickly when D_loss converges. The generated image quality is relatively unstable during 50-100 epochs, so setting the training epoch to 200 yields the best results. However, when the training epoch is around 130, all three losses can converge. The overall gradient curve tends to smooth out after 100 epochs, indicating a relatively fast training efficiency.

Claims

1. A weakly supervised lung X-ray image segmentation method based on an improved recurrent generative adversarial network, used for segmenting asymmetric data-labeled lung X-ray images, characterized in that: Based on the original network, the LS-CycleGAN network model undergoes several optimizations and improvements. This model consists of two parts: a generator network and a discriminator network. First, the generator network is improved, adopting an Auto-Encoder + Skip-connection network structure. The residual network module abandons the original residual network structure and uses the proposed SE_Blocks structure. Its unique global connectivity compensates for the original residual network's limitation of only extracting features locally, allowing the model to extract features globally and possess multi-scale invariance, thus improving the quality of the generated images. The discriminator network uses an Auto-Encoder network structure instead of the original CycleGAN generator structure. Its core advantage is that the discriminator's training is no longer constrained by the generator; the discriminator can be trained first, and its optimization can stimulate the generator's training. This solves the training imbalance problem that occurs in the original model. These modifications enable the network to learn a many-to-many mapping between two domains using only the original lung X-ray image and partial annotations, generating a segmentation mask for the lung X-ray image.

2. The method for segmenting asymmetric data-labeled lung X-ray images based on an improved recurrent generative adversarial network (RGAN) deep learning network according to claim 1, wherein the RGAN model is improved, characterized in that: ① The Squeeze and Excitation (SE) function in the SE_Blocks module is a channel attention mechanism that can improve network performance by adaptively adjusting the channel weights of feature maps. However, it only processes information globally within a single channel, ignoring information interaction in the spatial domain and failing to fully utilize its information. Residual connections, on the other hand, directly add the input to the output, allowing the network to learn residual information more easily, thus mitigating the vanishing and exploding gradient problems. This cross-layer connection design makes the network easier to optimize. Residual connections and training. SE_Blocks can effectively pass information and gradients, helping the network learn deeper features. For large images, this means the network can better capture details, textures, and structural information in the image. Compared to traditional convolutional neural networks, SE_Blocks has stronger expressive power and better feature extraction capabilities, enabling it to capture the essential features of images more accurately. ② A modified MOA (Multi-Resolution Overlapping Attention) module is added to the skip links. A global attention module, MOA, is introduced between the skip links at each stage, which fully utilizes the global information of the Local Transformer to generate global features. The original module only contains multiplication and addition operations, so it does not significantly increase the computational cost or the number of parameters. A 1×1 convolution kernel is added before the output of this module to convolve the synthesized features and the input features, thus incorporating local information. ③ The original Generative Adversarial Network (GAN) generator is designed to output only a single evaluation value (True or False), which evaluates the entire image generated by the generator. For image similarity discrimination, this design loses a lot of judgment on image details. Unlike the original GAN, the LS-CycleGAN discriminator network is designed as a fully convolutional network. After the image passes through various convolutional layers, it is not input into fully connected layers or activation functions. Instead, convolution maps the input into an N*N matrix. This matrix is ​​equivalent to the final evaluation value in the original GAN, used to evaluate the generator's generated image. Instead of using a single value to measure the entire image, an N*N matrix is ​​used to evaluate the entire image, achieving better evaluation results.

3. The technical solution for segmenting asymmetric data-labeled lung X-ray images using a deep learning network based on an improved recurrent generative adversarial network, as described in claim 1, is as follows: ① Divide the samples that need to be input into the model for training into a training set and a validation set in a 5:1 ratio; ② For the training set, the pydicom library is used to read the original DICOM lung X-ray image files and batch-process them into PNG format for easy input reading by the LS-cycleGAN network. Multiple training set structures are set up: Train A contains the original lung X-ray images in the image domain, and Train B contains segmentation mask labels in the image domain. Five training structures (Train A-Train B) are set up: 20-20, 50-20, 100-20, 150-20, and 200-20. The number of segmentation masks is kept relatively low, and the number of original lung X-ray images is increased sequentially to increase the number of many-to-many mappings, stimulating the generator and discriminator network loss functions to converge quickly. ③ After multiple training sessions, common medical image segmentation metrics such as IoU, Dice, VOE, and RVD on the validation set are monitored and saved. The weights that perform best on these metrics are selected as the final weights to obtain the optimized model. ④ Select a test sample and resize the image to 256×256. Then, crop the image to 256×256 from the center to preserve the central portion of the image. This portion of the image is treated as domain A (TrainA), and a random segmentation mask label is selected as domain B (TrainB). ⑤ Use the processed datasets of domain A and domain B as input to the optimization model, and output the mask label of the image to be segmented (domain A) according to the model output.