A bone suppression method for high-resolution chest X-ray images based on BS-LDM model
The BS-LDM model is used to automatically generate high-resolution soft tissue images, which solves the problems of resolution and radiation risk in the existing bone suppression method, achieves high-quality soft tissue image generation, and improves diagnostic accuracy and support.
Patent Information
- Application Number
- CN202410337624.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-23
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-03-23
AI Technical Summary
Existing bone suppression methods for chest X-ray images cannot effectively suppress bone structures while maintaining high resolution, high clarity, and detailed textures. In addition, the high-cost dual-energy subtraction technology carries a high radiation risk.
A high-resolution chest X-ray image bone suppression method based on the BS-LDM model is adopted. Through vector quantization, generative adversarial networks and conditional diffusion models, high-resolution and high-definition soft tissue images are automatically generated. The implicit diffusion model is used to perform image processing in the latent space, and the image space is mapped through the encoder and decoder. The conditional diffusion model is combined to control the generation results and reduce the influence of bone structure.
It achieves the generation of clearly visible soft tissue structures without increasing radiation dose, reduces misdiagnosis and missed diagnosis, improves the diagnostic accuracy of lung lesions, and provides support for auxiliary diagnosis and treatment.
Smart Images

Figure CN118212156B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image restoration and enhancement, and in particular to a bone suppression method for high-resolution chest X-ray images based on a BS-LDM model. Background Art
[0002] Chest X-ray (CXR) is a widely used imaging technique for lung screening. However, the diagnostic accuracy of CXR is limited due to the overlap of bone structures with lung tissue. To improve this, a technique called bone suppression has been employed. Dual-energy subtraction (DES) is a well-known bone suppression method, but it is expensive and requires high radiation exposure. Therefore, research is underway to develop cost-effective alternatives.
[0003] Initial CXR bone suppression methods were often based on statistical methods, such as those proposed by Simko and Juhasz et al. These methods require precise segmentation and boundary annotation but fail to incorporate high-level semantic information about the skeleton. Recently, many bone suppression methods have leveraged deep learning algorithms to generate soft tissue. Some of these methods approach bone suppression as an end-to-end image denoising task, while others focus on delineating bone gradients in CXR images to distinguish soft tissue from bony elements.
[0004] Yang et al. pioneered a cascaded multi-scale convolutional neural network (CNN) trained in the gradient domain of CXR images for bone suppression. While this model achieved impressive performance, it did not maintain high levels of perceptual or structural integrity. In another study, Gusarev et al. conceptualized bone as noise by utilizing a combination of autoencoders and deep CNN features to eliminate bone structure, but unfortunately, this resulted in blurry images. Taking cues from generative adversarial networks (GANs), Zhou et al. introduced a multi-scale conditional adversarial network (MCA-Net) designed to generate soft tissue images that preserve essential anatomical structure. Furthermore, Rajaraman et al. developed the ResNet-BS model for bone suppression in CXR images and validated its utility through subsequent analysis tasks. Chen et al. proposed BS-Diff, which develops a conditional diffusion model and adds a basic enhancement module for bone suppression. Despite this, this model still has limitations in terms of resolution, detail, and computational requirements.
[0005] In summary, the current clinical and scientific challenges are as follows:
[0006] From a clinical perspective, the analysis mainly includes: the bones in the soft tissue images obtained by DES can be basically completely suppressed and an appropriate data inclusion standard is formulated.
[0007] From a scientific perspective, the main issues include: the trained model needs to be able to effectively suppress bones; the trained model cannot introduce or reduce other substances while suppressing bones (i.e., it just suppresses bones); the trained model can maintain good texture details, such as vascular structure and clarity; the trained model must be robust and able to remove motion artifacts caused by the patient's heartbeat and breathing when shooting DES. Summary of the Invention
[0008] In response to the shortcomings of the existing technology, the present invention proposes a high-resolution chest X-ray image bone suppression method based on the BS-LDM (Bone Suppression-Latent Diffusion Model). Based on the input chest X-ray image, this method automatically generates a high-resolution, high-definition soft tissue image that includes spatial features and texture details.
[0009] In order to solve the above technical problems, the technical solution of the present invention is:
[0010] A high-resolution chest X-ray image bone suppression method based on the BS-LDM model comprises the following steps:
[0011] S1, collect image data and preprocess;
[0012] S1-1. Use dual-energy silhouette equipment to acquire chest X-ray images and matching soft tissue images of the same patient;
[0013] S1-2. The collected paired images were screened according to the inclusion criteria and the automatic registration operation of discrete Fourier transform was used to maximize the image similarity to achieve the best alignment state. Including data other than annotations would interfere with the prediction effect of the model. The inclusion criteria included: age > 18 years, no history of chest surgery or trauma, chest X-rays were taken using dual-energy imaging conditions, the imaging position met the standard requirements for chest anteroposterior position, the patient had a normal chest, the chest cavity diagnosis was normal, and the emphysema diagnosis was normal.
[0014] S2. Build a high-resolution chest X-ray image bone suppression network model based on the BS-LDM model. The high-resolution chest X-ray image bone suppression network model based on the BS-LDM model includes a vector quantization generative adversarial network and a conditional diffusion model. The core of the implicit diffusion model is to perform the forward and reverse processes of the diffusion model in the latent space. The vector quantization generative adversarial network maps the image in the pixel space to the latent space through the encoder, and maps the latent variables in the latent space back to the pixel space through the decoder. The conditional diffusion model allows the generation result to be controlled according to the intention. Its core idea is to learn a conditional reverse process. Without changing the forward process, the sampled x0 has high fidelity to the data distribution. During training, first sample From a completely paired data distribution That is, soft tissue x0 and chest X-ray image Learn a conditional diffusion model, providing As input to the reverse process, the formula is as follows:
[0015]
[0016] in, is a conditional reverse process with mean μ θ and variance ∑ θ U-Net based networks can be used (input is x t and t) are estimated.
[0017] S3. Use the preprocessed chest X-ray image as both input and label for the self-reconstruction task, repeatedly train the vector quantized generative adversarial network, optimize the network parameters, and continuously perform iterative optimization to minimize the difference between the true value image (label) and the model output image;
[0018] S4. Obtain a chest X-ray image and a matching soft tissue image through DES. Use the preprocessed chest X-ray image as input and the matching soft tissue image obtained by DES as a label for image generation. Repeatedly train the conditional diffusion model, optimize the network parameters, and continuously perform iterative optimization to minimize the difference between the true value image (label) and the model output image.
[0019] S5. Use a dual-energy silhouette device to collect the patient's chest X-ray image. After preprocessing, input it into the trained high-resolution chest X-ray image bone suppression network model based on the BS-LDM model. The model passes through the encoder of the vector quantization generative adversarial network, the conditional diffusion model, and the decoder of the vector quantization generative adversarial network to finally generate a soft tissue image. Among them, the vector quantization generative adversarial network maps the image in the pixel space to the latent space through the encoder, and maps the latent variables in the latent space back to the pixel space through the decoder. In the conditional diffusion model, the model accepts the splicing of Gaussian noise and chest X-ray images as input. After multiple sampling and denoising, the latent variables of the predicted soft tissue image are obtained;
[0020] Preferably, in step S1, the pre-processed images are uniformly adjusted to 1024×1024.
[0021] Preferably, in step S3, the loss function of the model training is:
[0022]
[0023] in, is the mean absolute error loss; To quantify the losses; Perceptual loss for pre-trained visual geometry neural networks; An adversarial loss on the patch discriminator for a pixel-by-pixel super-resolution model. is the weight of the mean absolute error loss, λ Qua =1 is the weight of quantization loss, λ Per = 0.001 is the weight of the perceptual loss of the pre-trained visual geometry neural network, λ Adv = 0.01 is the weight of the adversarial loss on the patch discriminator based on the pixel-to-pixel super-resolution model.
[0024] Preferably, in step S4, the loss function of the model training is mean square error.
[0025] The present invention has the following characteristics and beneficial effects:
[0026] By adopting the above technical solution, the present invention truly realizes the medical demand of automatically removing bones to generate soft tissue based on chest X-ray images, specifically including: 1. The boneless X-ray chest photo obtained by the relevant deep learning algorithm can make the soft tissue structures (such as lungs, heart, blood vessels, etc.) in the chest image more clearly visible, which can help diagnose lung lesions overlapping with the rib area, reduce misdiagnosis or missed diagnosis, and effectively enable doctors to more easily observe and evaluate the nature, size and location of some lung diseases (such as: intrapulmonary nodules, pneumonia, tumors), and provide assistance and intervention for further diagnosis and treatment. 2. Patients do not need to undergo high-dose radiation examination equipment such as DES, which is currently commonly used in clinical practice, and can also avoid image artifacts caused by heartbeat and respiratory movements brought by such equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 It is a schematic diagram of the process of the present invention;
[0029] Figure 2 This is the overall architecture of the high-resolution chest X-ray image bone suppression network model based on the BS-LDM model in an embodiment of the present invention;
[0030] Figure 3 Schematic diagram comparing soft tissue images generated by the bone suppression network model of high-resolution chest X-ray images based on the BS-LDM model in an embodiment of the present invention and soft tissue images generated by other methods. DETAILED DESCRIPTION
[0031] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0032] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, features defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "multiple" means two or more.
[0033] The present invention provides a high-resolution chest X-ray image bone suppression method based on the BS-LDM model, such as Figure 1 The specific operations are as follows:
[0034] S1. A dual-energy silhouette device was used to collect chest X-ray images and matching soft tissue images of the same patient, both with a size of 2021 × 2021. Including data other than the annotations would interfere with the model's prediction results, so according to the inclusion criteria (as follows):
[0035] (1) Aged > 18 years, with no history of chest surgery or trauma;
[0036] (2) Perform chest X-ray;
[0037] (3) chest radiography using dual-energy imaging conditions;
[0038] (4) There is no obvious chest deformity such as the radiographic position not meeting the standard requirements for chest alignment or S-shaped scoliosis of the spine;
[0039] (5) There was no diagnosis of pneumothorax, pleural effusion, or hydropneumothorax in either thoracic cavity;
[0040] (6) There is no diagnosis of emphysema on either side.
[0041] The collected paired images were screened according to the inclusion criteria, and preprocessing operations such as automatic registration and local adaptive image enhancement were performed.
[0042] The dataset used in this example consists of 167 paired anterior and posterior DES chest X-ray images collected from a partner hospital. These images were captured by a digital radiography (DR) machine equipped with a double-exposure DES device (Discovery XR656, GE Healthcare). The images were initially stored in DICOM format with a 14-bit depth but were later converted to PNG files for convenience. The pixel dimensions of all chest X-ray images were 2021×2021, with a pixel size range of 0 to 0.1943 mm. Eight paired X-rays were excluded due to operator error, obvious motion artifacts, and visible pleural effusion and pneumothorax. The final experimental dataset consisted of 159 paired images. The entire dataset was divided into a training set and a test set with a ratio of 8:2. To save memory, all images were resized to 1024×1024 pixels. The present invention also employed image registration to optimize information fusion between paired images and contrast-limited adaptive histogram equalization to enhance local contrast. Subsequently, all image pixel values were normalized to [-1, 1].
[0043] S2. Build a high-resolution chest X-ray image bone suppression network model based on the BS-LDM model. The high-resolution chest X-ray image bone suppression network model based on the BS-LDM model includes a vector quantization generative adversarial network and a conditional diffusion model. The core of the implicit diffusion model is to perform the forward and reverse processes of the diffusion model in the latent space. The vector quantization generative adversarial network maps the image in the pixel space to the latent space through the encoder, and maps the latent variables in the latent space back to the pixel space through the decoder. The conditional diffusion model allows the generation result to be controlled according to the intention. Its core idea is to learn a conditional reverse process. Without changing the forward process, the sampled x0 has high fidelity to the data distribution. During training, first sample From a completely paired data distribution That is, soft tissue x0 and chest X-ray image Learn a conditional diffusion model, providing As input to the reverse process, the formula is as follows:
[0044]
[0045] in, is a conditional reverse process with mean μ θ and variance ∑ θ U-Net based networks can be used (input is x t and t) are estimated.
[0046] S3. In the vector quantization generative adversarial network training phase, the preprocessed chest X-ray image is used as both input and label for the self-reconstruction task. The vector quantization generative adversarial network is repeatedly trained, the network parameters are optimized, and iterative optimization is continuously performed to minimize the difference between the true value image (label) and the model output image. The loss function is:
[0047]
[0048] in, is the mean absolute error loss; To quantify the losses; Perceptual loss for pre-trained visual geometry neural networks; An adversarial loss on the patch discriminator for a pixel-by-pixel super-resolution model. is the weight of the mean absolute error loss, λ Qua =1 is the weight of quantization loss, λ Per = 0.001 is the weight of the perceptual loss of the pre-trained visual geometry neural network, λ Adv = 0.01 is the weight of the adversarial loss on the patch discriminator based on the pixel-to-pixel super-resolution model.
[0049] S4. In the training phase of the conditional diffusion model, the preprocessed chest X-ray image is used as input, and the matching soft tissue image obtained by DES is used as the label for the image generation task. The conditional diffusion model is repeatedly trained, the network parameters are optimized, and iterative optimization is continuously performed to minimize the difference between the true value image (label) and the model output image. The loss function is the mean square error. The noise added in the forward process of the conditional diffusion model is offset noise, that is, additional bias noise sampled from a Gaussian distribution is superimposed on the standard noise addition process to improve the color generation effect.
[0050] S5. In the model inference stage, a dual-energy silhouette device is used to collect the patient's chest X-ray image, which is pre-processed and then input into the trained high-resolution chest X-ray image bone suppression network model based on the BS-LDM model. The model is successively processed through the encoder of the vector quantization generative adversarial network, the conditional diffusion model, and the decoder of the vector quantization generative adversarial network to finally generate a soft tissue image. Among them, the vector quantization generative adversarial network maps the image in the pixel space to the latent space through the encoder, and maps the latent variables in the latent space back to the pixel space through the decoder. In the conditional diffusion model, the model accepts the splicing of Gaussian noise and chest X-ray images as input, and obtains the predicted latent variables of the soft tissue image after multiple sampling and denoising. In the reverse process of the conditional diffusion model, the present invention adopts a dynamic clipping strategy, that is, the data value limit [-s, s] is clipped after each sampling, and the clipping interval size s decreases as the current sampling step number t increases, as shown below:
[0051] s=ω·t+b
[0052] Wherein, ω=0.0021 represents the slope of the function, and b=1.5 represents the intercept of the function.
[0053] Based on the above embodiments, Figure 3 As shown, a schematic diagram comparing the soft tissue image generated by the high-resolution chest X-ray image bone inhibition network model based on the BS-LDM model and the soft tissue images generated by other methods is shown. It can be seen that the soft tissue image generated by the high-resolution chest X-ray image bone inhibition network model based on the BS-LDM model in the embodiment of the present invention is extremely similar to the soft tissue image obtained by the high-quality dual-energy silhouette, so that it is impossible to determine whether the image is generated by the model or taken by the dual-energy silhouette device. At the same time, the soft tissue image generated by the model clearly and accurately captures and synthesizes the relevant tiny lesions, which shows that the present invention can effectively provide assistance and intervention for the diagnosis and treatment of lung diseases.
[0054] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. It will be apparent to those skilled in the art that various changes, modifications, substitutions, and variations of these embodiments, including components, without departing from the principles and spirit of the present invention are still within the scope of protection of the present invention.
Claims
1. A high-resolution chest X-ray image bone suppression method based on the BS-LDM model, characterized in that: The steps include: S1, collect image data and preprocess; S1-1. Use dual-energy silhouette equipment to acquire chest X-ray images and matching soft tissue images of the same patient; S1-2. The collected paired images are screened according to the inclusion criteria and the images are aligned by maximizing the image similarity using the automatic registration operation of discrete Fourier transform. S2. Building a high-resolution chest X-ray image bone suppression network model based on the BS-LDM model, wherein the high-resolution chest X-ray image bone suppression network model based on the BS-LDM model includes a vector quantization generative adversarial network and a conditional diffusion model; The vector quantization generative adversarial network maps the image in the pixel space to the latent space through the encoder, and maps the latent variables in the latent space back to the pixel space through the decoder. The conditional diffusion model allows the generation result to be controlled according to the intention. Its core idea is to learn a conditional reverse process. Without changing the forward process, the sampled x0 has high fidelity to the data distribution. During training, first sample From a completely paired data distribution That is, soft tissue x0 and chest X-ray image Learn a conditional diffusion model, providing As input to the reverse process, the formula is as follows: in, is a conditional reverse process with mean μ θ and variance ∑ θ Both use the U-Net based network for estimation, and the input of the U-Net network is x t and t; S3. Use the preprocessed chest X-ray image as both input and label for the self-reconstruction task, repeatedly train the vector quantized generative adversarial network, optimize the network parameters, and continuously perform iterative optimization to minimize the difference between the true value image and the vector quantized generative adversarial network output image; S4. Obtain a chest X-ray image and a matching soft tissue image through DES. Use the preprocessed chest X-ray image as input and the matching soft tissue image obtained by DES as a label for image generation. Repeatedly train the conditional diffusion model, optimize the network parameters, and continuously perform iterative optimization to minimize the difference between the true value image and the model output image. S5. Use a dual-energy silhouette device to collect the patient's chest X-ray image, which is preprocessed and then input into the trained high-resolution chest X-ray image bone suppression network model based on the BS-LDM model. First, the vector quantization generative adversarial network maps the image in the pixel space to the latent space through the encoder, and maps the latent variables in the latent space back to the pixel space through the decoder. In the conditional diffusion model, the model accepts the splicing of Gaussian noise and chest X-ray images as input, and obtains the predicted latent variables of the soft tissue image after multiple sampling and denoising.
2. The high-resolution chest X-ray image bone suppression method based on the BS-LDM model according to claim 1, characterized in that: The inclusion criteria included: age > 18 years, no history of chest surgery or trauma; chest anteroposterior X-ray using dual-energy radiography; radiographic setup meeting the standard requirements for anteroposterior chest position; normal chest cage; normal chest cavity diagnosis; and normal emphysema diagnosis.
3. The high-resolution chest X-ray image bone suppression method based on the BS-LDM model according to claim 1, characterized in that: In step S1, the pre-processed images are uniformly adjusted to 1024×1024.
4. The high-resolution chest X-ray image bone suppression method based on the BS-LDM model according to claim 1, characterized in that: In step S3, the loss function of the vector quantization generative adversarial network training is: in, is the mean absolute error loss; To quantify the losses; Perceptual loss for pre-trained visual geometry neural networks; An adversarial loss on the patch discriminator for a pixel-by-pixel super-resolution model. is the weight of the mean absolute error loss, λ Qua =1 is the weight of quantization loss, λ Per = 0.001 is the weight of the perceptual loss of the pre-trained visual geometry neural network, λ Adv = 0.01 is the weight of the adversarial loss on the patch discriminator based on the pixel-to-pixel super-resolution model.
5. The high-resolution chest X-ray image bone suppression method based on the BS-LDM model according to claim 1, characterized in that: In step S4, the loss function of the conditional diffusion model training is the mean square error.
6. The high-resolution chest X-ray image bone suppression method based on the BS-LDM model according to claim 1, characterized in that: In step S5, in the reverse process of the conditional diffusion model, a dynamic clipping strategy is adopted, that is, the data value limit is clipped after each sampling, and the clipping interval size decreases as the current sampling step number increases.
7. The high-resolution chest X-ray image bone suppression method based on the BS-LDM model according to claim 1, characterized in that: In step S5, the soft tissue image generated by the conditional diffusion model is fed into an enhancement module based on an autoencoder, and the enhanced soft tissue image is output.
Citation Information
Patent Citations
X-ray chest radiograph bone suppression processing method based on wavelet decomposition and convolutional neural network
CN107038692A
X-ray chest radiograph rib image suppression method based on joint prediction filtering generation network
CN116612026A