Suppression Method for Rib Images in X-Ray Chest Radiographs Based on Attention Generative Adversarial Network
Through the attention-based generative adversarial network, X-chest rib image suppression is solved, and the problems of incomplete rib suppression and loss of texture details in the prior art are generated, clear bone inhibition images are generated, which improves the accuracy of lung disease detection.
Patent Information
- Application Number
- CN202310290621.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-23
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-03-23
AI Technical Summary
The existing methods have problems of loss of texture details, blurred image generation or incomplete rib suppression during rib suppression in X-chest X-ray, which affects the accuracy of lung disease detection.
Attention-based generative adversarial network is adopted, and backbone networks of generators and discriminators are constructed, data augmentation is performed in combination with rotation and translation operations, and feature correction is used for SKnet attention mechanism and spatial attention blocks to generate clear bone suppression images.
While effectively suppressing rib images, it retains detailed information in the non-rib area, improving the accuracy and image quality of lung disease detection.
Smart Images

Figure CN116542868B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and relates to a method for suppressing rib images of X-ray chest films based on an attention-based generative adversarial network. Background Art
[0002] In the examination of pulmonary nodules, although CT technology is becoming more and more popular, as a most common medical imaging technology, X-ray is widely used in the screening of lung diseases due to its advantages such as low radiation dose and low cost. However, an X-ray chest film is a two-dimensional radiographic image of X-rays passing through a three-dimensional human body. Anatomical tissue structures of the human body such as ribs and other soft tissue structures of the human body will overlap in the image and block some soft tissues of the human body. If the lesion location happens to fall in these overlapping areas, the lesion image will be blocked by the ribs in the X-ray chest film, seriously affecting the detection results of lung diseases. By using bone suppression technology to suppress the rib structure in the X-ray chest film, the soft tissue image blocked by the rib structure can be obtained, which can improve the detection accuracy of doctors and computer-aided systems for lung diseases to a certain extent, and has important clinical significance.
[0003] Dual Energy Subtraction (DES) is a relatively common X-ray chest film bone suppression technology in clinical practice. It obtains X-ray images of two different energy levels, and based on the subtraction of the two images, highlights the soft tissues of the X-ray chest film and suppresses the imaging of other anatomical tissues such as ribs; however, due to relying on special equipment and the imaging process being affected by factors such as the patient's heartbeat, breathing, and movement, the imaging quality is unstable. At the same time, obtaining X-ray images of two different energy levels requires a large radiation dose, which is harmful to the patient's body, and there are also many limitations in clinical practice.
[0004] The method based on deep learning has greatly improved the bone suppression effect. However, due to the complex overlapping tissue structures in the X-ray chest film and the similar textures of different tissue structures, there may be problems of inaccurate positioning and recognition of the bone structure to be suppressed when the existing methods capture the rib structure in the X-ray chest film. Therefore, rib suppression in X-ray chest films is still an important problem faced in the screening of lung diseases in X-ray chest films. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for suppressing rib images of X-ray chest films based on an attention-based generative adversarial network, which overcomes the problems of texture detail loss, blurred generated images, or incomplete rib suppression caused by directly generating bone suppression images by the existing methods, and further obtains a clear bone suppression image without changing the detail information of the non-rib area.
[0006] The technical solution adopted by the present invention is a method for suppressing rib images in X-ray chest films based on an attention-based generative adversarial network, which specifically includes the following steps:
[0007] Step 1, preprocess the training set and test set images in the dataset;
[0008] Step 2, construct a rib image suppression network model for X-ray chest films based on an attention-based generative adversarial network;
[0009] Step 3, use the training set data preprocessed in Step 1 to train the model constructed in Step 2 to obtain a trained rib image suppression network model for X-ray chest films based on an attention-based generative adversarial network;
[0010] Step 4, put the test set images preprocessed in Step 1 into the model trained in Step 3 to finally obtain the soft tissue images after bone suppression.
[0011] The features of the present invention also lie in:
[0012] The preprocessing process of Step 1 is: rotate and translate the images in the dataset to achieve data augmentation.
[0013] In Step 2, for the rib image suppression network model for X-ray chest films based on an attention-based generative adversarial network, a generative adversarial network is used as the backbone network, and the backbone network includes a generator and a discriminator.
[0014] The process of suppressing rib images in X-ray chest films by the rib image suppression network model for X-ray chest films based on an attention-based generative adversarial network is as follows:
[0015] Step A), take the image X in the dataset as the input of the generator, and the encoder Encoder extracts multiple image features {C1, C2, C3, C4} at different scales in turn as shown in the following formula (1):
[0016] {C1, C2, C3, C4} = Encoder(X) (1);
[0017] Step B), fuse and correct the features at the C4 scale through the internal and external attention modules of the fusion attention mechanism. The internal correction adopts the principle of the SKnet attention mechanism to complete the feature correction in the channel dimension as shown in formula (2); add a spatial attention block Gate outside the SKnet to calculate the spatial coefficient g, as shown in formula (3); pass the input features through a convolution with dilation rate r = 1 and then through a sigmoid operation to calculate the spatial change threshold value g of the feature C4, and obtain the final aggregated feature F through the weighted sum of this value, C5, and the input feature C4 as shown in formula (4):
[0018] C5 = Sknet(C4) (2);
[0019] g = Gate(C4) (3);
[0020] F = C4 * g + C5 * (1 - g) (4);
[0021] Step C), according to the aggregated feature F obtained in Step B), through the decoder Decoder of the generative adversarial network, the rib tissue map is generated through network adaptive learning. The formula is as shown in (5), and the rib attention map and the residual rib attention map obtained from the generator are fused with the input X-ray chest image to obtain the bone-suppressed soft tissue image S o , and the formula is as shown in (6):
[0022] R = Decoder(F) (5);
[0023] S o = R + X(1 - R) (6).
[0024] The specific process of Step 3 is as follows: Using the adversarial loss function, reconstruction loss function, and attention loss function to constrain the network model obtained in Step 2, and then performing backpropagation to update the parameters to obtain a trained X-ray chest image rib image suppression network model based on attention.
[0025] In Step 3, the adversarial loss function is represented by the following formula (7): where D and G represent the discriminator and the generator respectively, G is a function of the standard X-ray chest image, generating a virtual bone-suppressed image for the discriminator to distinguish, and E represents the error; P data represents the data distribution, x ∼ P data(x) and s ∼ P data(s) represent selecting s and x from the data distributions of chest X-rays and rib-suppressed images respectively:
[0026]
[0027] The reconstruction loss function L rec (G) is represented by the following formula (8): where S represents the ground truth of the bone-suppressed image, X is the input conventional X-ray chest image, and the L1 loss ||.||1 is used to calculate the pixel gap. The generator G will learn the mapping between the bone-suppressed image and the conventional X-ray chest image:
[0028]
[0029] The attention loss function L R is represented by the following formula (9): where W, H, and C are the width, height, and number of channels of the attention map R respectively:
[0030]
[0031]
[0032] The total loss function L is as follows:
[0033]
[0034] Among them, λ1, λ2, and λ3 are hyperparameters that control the relative importance of each part of the loss term.
[0035] The beneficial effect of the present invention is that the present invention proposes a new rib fusion bone suppression algorithm based on the generative adversarial network model. Aiming at the problem of the loss of detailed information of lung texture in the generation of virtual soft tissues by previous bone suppression methods, the present invention achieves the purpose of bone suppression by changing the contrast between the ribs and lung tissues in the picture and reducing the presence of the ribs, thereby avoiding the loss of information in areas other than the ribs to obtain high-quality soft tissue images. The method of the present invention has been verified on the public dataset to have high metrics and excellent performance. Brief Description of the Drawings
[0036] Figure 1 It is a schematic flowchart of the method for suppressing rib images in X-ray chest films based on the attention-based generative adversarial network of the present invention;
[0037] Figure 2 It is a schematic overall structure diagram of the method for suppressing rib images in X-ray chest films based on the attention-based generative adversarial network of the present invention;
[0038] Figure 3 It is a fusion adversarial attention module of the method for suppressing rib images in X-ray chest films based on the attention-based generative adversarial network of the present invention;
[0039] Figure 4 It is a process diagram of the fusion and generation of bone suppression images in the method for suppressing rib images in X-ray chest films based on the attention-based generative adversarial network of the present invention;
[0040] Figure 5 It is a result comparison diagram of the method for suppressing rib images in X-ray chest films based on the attention-based generative adversarial network of the present invention. Detailed Embodiments
[0041] The method for suppressing rib images in X-ray chest films based on the attention-based generative adversarial network of the present invention, as Figure 1 shown, specifically includes the following steps:
[0042] Step 1, preprocess the training set and test set images in the dataset; specifically:
[0043] The data is enhanced by operations such as rotation and translation. This paper uses the Japanese Society of Radiological Technology (JSRT) dataset. After data enhancement, there are 4080 paired data sets, each with a resolution of 1024*1024 and a bit depth of 8 PNG files. The divided training data set is used for training, and the test set is used for testing.
[0044] Step 2, Figure 2 The figure shows the overall structure of the X-ray rib image suppression method based on the attention-based generative adversarial network. The model uses the generative adversarial network as the backbone network, mainly consisting of a generator part that generates images and a discriminator that determines the authenticity of the generated images. The generator consists of an encoder, a fusion attention module, and a decoder. The X-ray rib image suppression network model based on the attention-based generative adversarial network sequentially extracts features from the data set, fusion correction, and feature information decoupling, and finally completes the X-ray rib image suppression based on the attention-based generative adversarial network. The specific steps are as follows:
[0045] In step 2.1, the image X in the dataset is used as the input of the generator. The encoder extracts multiple image features of different scales {C1, C2, C3, C4} (as shown in formula 1), whose sizes are 1 / 2, 1 / 4, 1 / 8 and 1 / 16 of the input image size, respectively. This is shown in formula (1):
[0046] {C1,C2,C3,C4}=Encoder(X) (1);
[0047] Step 2.2, such as Figure 3 As shown in the figure, the C4 scale features are fused and corrected by the internal and external attention modules of the fusion attention mechanism (Converging Attention Blocks) to improve the network performance.
[0048] Internal correction adopts the principle of SKnet attention mechanism to complete the feature correction of channel dimension as shown in formula (2). In order to maintain the consistency of the global structural distribution of features after the correction of pixel detail information, the present invention adds a spatial attention block (Gate) outside SKnet to calculate the spatial coefficient g, as shown in formula (3). Combining the feature aggregation of spatial and channel changes while ensuring the overall structural layout of the feature attention map, the convergence speed of the network to generate the target image is accelerated. The stacked fusion attention block greatly enriches the network depth of the generator, allowing the generator to capture as much rib area information as possible, forming a more complete rib attention map.
[0049] Specifically, feature C4 is respectively passed through convolution operations with a kernel size of 3x3 and dilation convolutions with r = 1, 2, 4, 8 in sequence. Different dilation convolutions enable the sub-kernels to focus on different receptive fields, so as to fully consider the overall and detailed information of the image features. The feature information of different scales is fused in an element-wise addition manner. Then, global average pooling (GAP) is used to obtain the global information on each channel. Two fully connected operations (FC) complete the information interaction in the channel dimension. Softmax normalization is used to obtain the weight values corresponding to different scales, which are respectively multiplied by the corresponding features to be corrected for feature correction. Then, the corrected feature information is fused in an element-wise addition manner to obtain the final fused feature C5, thus completing the internal channel correction. To maintain the consistency of the global structure distribution of the pixel detail information after correction, the present invention adds a spatial attention block outside the SKnet to calculate the spatial coefficient g. The input feature is passed through a convolution with dilation r = 1 and then through a sigmoid operation to calculate the spatial change threshold value g of feature C4. This value is weighted and summed with C5 and the input feature C4 through a formula to obtain the finally aggregated feature F as shown in the following formula (4):
[0050] C5 = Sknet(C4) (2);
[0051] g = Gate(C4) (3);
[0052] F = C4 * g + C5 * (1 - g) (4);
[0053] Step 2.3, according to the aggregated feature F obtained in step 2.2, through the decoder Decoder of the generative adversarial network, a rib tissue map is adaptively learned and generated through the network. The formula is as shown in (5). The rib attention map (rib attention R) and the residual rib attention map (residual rib attention 1 - R) obtained from the generator are fused with the input X - ray chest image (input) to obtain the bone - suppressed soft tissue image S o (output), the formula is as shown in (6), and the image fusion visualization process is as Figure 4 shown.
[0054] R = Decoder(F) (5);
[0055] S o = R + X(1 - R) (6).
[0056] Step 3: Train the model using the preprocessed training set data in Step 1: Use the adversarial loss function, reconstruction loss function, and attention loss function to constrain the network model obtained in Step 2, and then perform backpropagation for parameter update. After 200 rounds of training, where 1 round means training all the preprocessed images once, finally obtain the trained X-ray rib image suppression network model of the attention-based generative adversarial network. The specific loss functions used are as follows:
[0057] Adversarial loss function: D and G represent the discriminator and generator respectively. G is a function of the standard X-ray chest film, which generates virtual bone-suppressed images for the discriminator to distinguish. E represents the error; P data represents the data distribution; and x ∼ P data(x) and s ∼ P data(s) represent selecting s and x from the data distributions of chest X-rays and rib-suppressed images respectively. Use the adversarial discriminator D to distinguish whether the generated image is a real soft tissue image:
[0058]
[0059] Reconstruction loss function: where S represents the ground truth of the bone-suppressed image, X is the input conventional X-ray chest film, use L1 loss ||.||1 to calculate the pixel gap, and the generator G will learn the mapping between the bone-suppressed image and the conventional X-ray chest film:
[0060]
[0061] Attention loss function: The mask automatically learned by the network is easily saturated to 1, which will cause the generator to not achieve the ideal effect. To prevent this situation, total variation regularization is performed on the mask R for constraint. W, H, and C are the width, height, and number of channels of the attention map R respectively. To generate the target soft tissue image So:
[0062]
[0063] Therefore, the total loss of the present invention is:
[0064]
[0065] Among them, λ1, λ2, and λ3 are hyperparameters that control the relative importance of each part of the loss term.
[0066] Step 4: Put the test set images processed in Step 1 into the model trained in Step 3, and finally obtain the soft tissue images after bone suppression. The comparison results are as Figure 5 shown. Figure 5 The first column is the input image, Figure 5 The second to fifth columns are the bone-suppressed images obtained by the existing benchmark methods,Figure 5 Column VI is the bone suppression image obtained by this method, and Column VII is the labeled bone suppression image. From Figure 5 the comparison results, in the case of effective suppression of rib images, the present invention still retains clear detail information, and even the nodular structure of the trachea is still clear.
Claims
1. An X-ray chest rib image suppression method based on an attention-based generative adversarial network, characterized in that: Specifically, the following steps are included: Step 1, preprocess the training set and test set images in the dataset; Step 2, construct an X-ray chest rib image suppression network model based on an attention-based generative adversarial network; In the said Step 2, for the X-ray chest rib image suppression network model based on an attention-based generative adversarial network, a generative adversarial network is used as the backbone network, and this backbone network includes a generator and a discriminator; The process of suppressing the X-ray chest rib image by the X-ray chest rib image suppression network model based on an attention-based generative adversarial network is as follows: Step A), taking the image X in the dataset as the input of the generator, the encoder Encoder sequentially extracts image features at multiple different scales As shown in the following formula (1): (1) Step B), The features of the scale are fused and corrected through the internal and external attention modules of the fusion attention mechanism. The internal correction adopts the principle of the SKnet attention mechanism to complete the feature correction of the channel dimension as shown in formula (2); A spatial attention block Gate is added outside the SKnet to calculate the spatial coefficient g, as shown in formula (3); The input features are convolved through a convolution with dilation rate r = 1 and then passed through a sigmoid operation to calculate the feature spatial variation threshold value g, and this value is combined with and the input features The final aggregated feature F is obtained through weighted sum by the formula as shown in formula (4) below: Step C), according to the aggregated feature F obtained in Step B), through the decoder Decoder of the generative adversarial network, the rib tissue map is generated by adaptive learning of the network. The formula is as shown in (5). The rib attention map and the residual rib attention map obtained from the generator are fused with the input X-ray chest film to obtain the bone-suppressed soft tissue image , and the formula is as shown in (6): Wherein, R represents the rib attention map; Step 3, use the preprocessed training set data in Step 1 to train the model constructed in Step 2 to obtain the trained X-ray chest rib image suppression network model based on an attention-based generative adversarial network; Step 4, put the preprocessed test set images in Step 1 into the model trained in Step 3 to finally obtain the soft tissue images after bone suppression.
2. The suppression method of X-ray chest rib images based on the attention-based generative adversarial network according to claim 1, characterized in that: The preprocessing process of the said Step 1 is: perform rotation and translation operations on the images in the dataset to achieve data augmentation.
3. The method for suppressing X-ray chest rib images based on an attention-based generative adversarial network according to claim 1, characterized in that: The specific process of the said Step 3 is: use an adversarial loss function, a reconstruction loss function, and an attention loss function to constrain the network model obtained in Step 2, and then perform backpropagation to update the parameters to obtain the trained X-ray chest rib image suppression network model based on an attention-based generative adversarial network.
4. The suppression method of X-ray chest film rib images based on the attention-based generative adversarial network according to claim 3, characterized in that: In the said step 3, the adversarial loss function is represented by the following formula (7): where D and G respectively represent the discriminator and the generator, G is a function of the standard X-ray chest film, generating a virtual bone suppression image for the discriminator to distinguish, and E represents the error; represents the data distribution, and represent selecting x and s from the data distributions of the chest X-ray and the rib suppression image respectively: (7); Reconstruction loss function It is represented by the following formula (8): where S represents the ground truth of the bone suppression image, X is the input conventional X-ray chest radiograph, the L1 loss ||.||1 is used to calculate the pixel gap, and the generator G will learn the mapping between the bone suppression image and the conventional X-ray chest radiograph: (8) Attention loss function It is represented by the following formula (9): where W, H, and C are the width, height, and number of channels of the attention map R, respectively: (9); The total loss function L is: (10) Among them, and are hyperparameters that control the relative importance of each part of the loss term.
Citation Information
Patent Citations
Chest DR dual-energy digital subtraction image generation method
CN113052930A