A flexible image noise reduction method and system using a feature representation field of view that has been untangled.
The system addresses the challenge of varying noise levels in medical images by using untangled feature representations and domain-independent learning to enhance noise removal across different imaging parameters, achieving improved image quality and structural fidelity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2026-04-01
AI Technical Summary
Existing denoising models for medical images, particularly in CT scans, are ineffective when applied to images acquired with different imaging parameters, such as reduced radiation doses, due to varying noise levels and artifact profiles, limiting their generalizability and applicability.
A system and method that utilizes a standard and low-quality image module to generate and reconstruct anatomical structural and noise features, with a loss calculation module to train the model using domain-independent learning, untangling feature representations to adapt to varying noise levels and imaging parameters.
The method enables a single trained model to effectively remove noise from images with varying noise levels, improving image quality and structural fidelity across different imaging conditions.
Smart Images

Figure 0007839163000001 
Figure 0007839163000002 
Figure 0007839163000003
Abstract
Description
[Technical Field]
[0001]
[0001] The present disclosure relates to a system and method for training and tuning a neural network model that provides a flexible solution for denoising low-dose images using domain-independent learning with untangled feature representations. [Background technology]
[0002]
[0002] Conventionally, most imaging modalities have an impact on the physical process of acquisition or reconstruction, resulting in artifacts such as noise in the final image. To train a denoising algorithm such as a neural network model, pairs of noisy image samples and noise-free image samples are presented to the neural network model, and the network attempts to minimize a cost function by removing the noise from the noisy image and recovering the corresponding noise-free ground truth image.
[0003]
[0003] However, any change in the parameters used to acquire the image will change the shape or amount of the corresponding image artifact. Therefore, the denoising model used for denoising standard images will be less effective when applied to images acquired using various acquisition parameters, such as reduced radiation doses in computed tomography (CT) scans.
[0004]
[0004] The increasing use of CT scans in modern medical practice has raised concerns about the associated radiation dose, making dose reduction a clinical goal. However, reducing the radiation dose tends to significantly increase noise and other artifacts in reconstructed images, impairing diagnostic information. A great deal of effort has been made to reduce noise in low-dose CT scans and thereby convert them into high-quality images.
[0005]
[0005] Regarding image noise reduction, machine learning techniques, including the use of convolutional neural networks (CNNs), have been studied. However, existing methods are usually tailored to specific noise levels and cannot be well generalized to noise levels that are not covered by the training sets used to train the corresponding CNNs.
[0006]
[0006] In CT imaging, multiple factors, including kilovolt peak (kVp), milliampere-seconds (mA-seconds), slice thickness, and patient size, all affect the noise level of the reconstructed image. Changing any of these imaging parameters results in different noise level or different artifact profiles, and therefore conventionally, different models have been required to denoise images acquired with such different imaging parameters. This limits the applicability of CNN-based methods in actual denoising. [Overview of the project] [Problems that the invention aims to solve]
[0007]
[0007] Therefore, there is a need for a method that can remove noise from images acquired with imaging parameters different from those used in the training set of the corresponding method. Furthermore, there is a need for a single trained model that can be used to remove noise from images with varying noise levels, including CT images acquired with lower radiation doses than the images in the training set. [Means for solving the problem]
[0008]
[0008] A system and method for removing noise from medical images are provided. In one embodiment, a standard image module is configured to generate standard anatomical structural features and standard noise features from a standard image, and to reconstruct a standard image from the standard anatomical structural features and standard noise features. A low-quality image module is similarly configured to generate low-quality anatomical structural features and low-quality noise features from a low-quality image, and to reconstruct a low-quality image from the low-quality anatomical structural features and low-quality noise features.
[0009]
[0009] A loss calculation module is provided that allows the system and method to be trained. The loss calculation module typically calculates a loss criterion based at least in part on 1) a comparison of reconstructed standard images with standard images, and 2) a comparison of reconstructed low-quality images with low-quality images.
[0010]
[0010] The loss evaluation criteria calculated by the loss calculation module are incorporated into a loss function for tuning the standard image module and the low-quality image module using machine learning. When low-quality anatomical structural features are supplied to the standard image module, the standard image module outputs a reconstructed standard transfer image that includes a noise level lower than the noise level represented by the low-quality anatomical structural features and the low-quality noise features.
[0011]
[0011] In some embodiments, the standard image module includes a standard anatomical structure encoder, a standard noise encoder, and a standard generator. Upon receiving a standard image, the standard anatomical structure encoder outputs standard anatomical structure features, the standard noise encoder outputs standard noise features, and the standard generator reconstructs the standard image from the standard anatomical structure features and standard noise features.
[0012]
[0012] In some such embodiments, the low-quality image module also has a low-quality anatomical structure encoder, a low-quality noise encoder, and a low-quality generator. When receiving a low-quality image, the low-quality anatomical structure encoder outputs low-quality anatomical structure features, the low-quality noise encoder outputs low-quality noise features, and the low-quality generator reconstructs the low-quality image from the low-quality anatomical structure features and the low-quality noise features.
[0013]
[0013] In some such embodiments, the loss calculation module calculates a loss evaluation criterion for the standard generator based at least in part on a comparison between the reconstructed standard image and the standard image. The loss calculation module also calculates a loss evaluation criterion for the low-quality generator based at least in part on a comparison between the reconstructed low-quality image and the low-quality image.
[0014]
[0014] The loss calculation module further calculates a loss evaluation criterion for the standard anatomical structure encoder based on a comparison with a segmentation label for the standard image, and the loss calculation module further calculates a loss evaluation criterion for the low-quality anatomical structure encoder based on a comparison with the output of the standard anatomical structure encoder.
[0015]
[0015] In some embodiments, the loss evaluation criterion for the low-quality anatomical structure encoder is an adversarial loss evaluation criterion.
[0016]
[0016] In some embodiments, a system implementing the described method further includes a segmentation network, and a segmentation mask of the reconstructed standard image is evaluated based on a comparison with a segmentation label for the standard image.
[0017]
[0017] In some such embodiments, standard anatomical structural features are supplied to a low-quality generator, which outputs a reconstructed low-quality transition image, the low-quality transition image containing a noise level higher than that represented by the standard anatomical structural features and standard noise features. The segmentation mask of the low-quality transition image is then evaluated based on a comparison with segmentation labels relating to the standard image.
[0018]
[0018] In some embodiments, the loss evaluation criterion for the standard transition image is evaluated based on comparison with the standard image reconstruction, and the loss evaluation criterion for the standard transition image is an adversarial loss evaluation criterion.
[0019]
[0019] In some embodiments, a standard image module and a low-quality image module are trained simultaneously.
[0020]
[0020] In other embodiments, a standard image module is trained before the low-quality image module is trained, and the values of the variables created during the training of the standard image module are kept constant during the training of the low-quality image module. In some such embodiments, after the training of the standard image module and the low-quality image module, the system further trains the standard generator while keeping the values of the standard anatomical structure encoder, the standard noise encoder, and the low-quality anatomical structure encoder constant.
[0021]
[0021] In some embodiments, standard anatomical structural features and low-quality anatomical structural features each correspond to a single anatomical structure. [Brief explanation of the drawing]
[0022] [Figure 1]
[0022] This is a schematic diagram of a system according to one embodiment of the present disclosure. [Figure 2]
[0023] This figure shows an imaging device according to one embodiment of the present disclosure. [Figure 3]
[0024] This is a diagram of a training pipeline used in one embodiment of the present disclosure. [Figure 4A]
[0025] This figure shows an exemplary training method used in the training pipeline shown in Figure 3. [Figure 4B] This figure shows an exemplary training method used in the training pipeline shown in Figure 3. [Figure 4C] This figure shows an exemplary training method used in the training pipeline shown in Figure 3. [Figure 5]
[0026] This figure shows how a model trained using the training pipeline in Figure 3 can be used to remove noise from images. [Modes for carrying out the invention]
[0023]
[0027] The description of exemplary embodiments in accordance with the principles of this disclosure is intended to be read together with the accompanying drawings, which should be considered part of the entire description. Any reference to direction or orientation in the description of embodiments disclosed herein is for illustrative purposes only and is not intended to limit the scope of this disclosure. Relative terms such as “lower,” “upper,” “horizontal,” “vertical,” “higher,” “lower,” “up,” “down,” “top,” and “bottom,” as well as their derivatives (e.g., “horizontally,” “downward,” “upward,” etc.), should be interpreted as referring to the orientation described or shown in the drawings discussed at that time. These relative terms are for illustrative purposes only and do not require the device to be configured or operated in a particular orientation unless expressly indicated so. Terms such as “attached,” “bonded,” “connected,” “joined,” and “interconnected” refer, unless otherwise specified, to relationships in which structures are fixed or attached to one another, directly or indirectly through intervening structures, as well as both movable and rigid attachments or relationships. Furthermore, the features and advantages of this disclosure are illustrated with reference to the exemplary embodiments. This disclosure should not be particularly limited to such exemplary embodiments, which illustrate several possible non-limiting combinations of features, either individually or in other combinations of features. The scope of this disclosure is defined by the claims appended herein.
[0024]
[0028] This disclosure describes one or more of the best possible modes of carrying out the disclosure as currently conceivable. This description is not intended to be understood in an exclusive sense, but rather, by reference to the accompanying drawings, provides an example of the disclosure presented solely for illustrative purposes, informing those skilled in the art of the advantages and features of the disclosure. In various drawings, similar reference numerals represent similar or similar parts.
[0025]
[0029] It is important to note that the embodiments disclosed are merely examples of many advantageous uses of the innovative teachings herein. The descriptions made in this application do not necessarily limit any of the various claimed disclosures. Furthermore, some descriptions may apply to some features of the invention but not to others. Generally, unless otherwise indicated, singular elements may be plural and vice versa without loss of generality.
[0026]
[0030] Generally, image processors that use image noise reduction algorithms or models to remove noise from medical images base their approach on the level and form of noise expected to be present in the corresponding image. This expected level and form of noise is typically based on various parameters used to acquire the image.
[0027]
[0031] For computed tomography (CT)-based medical imaging, images are processed using various image processors, such as machine learning algorithms in the form of convolutional neural networks (CNNs). These image processors, in the case of machine learning algorithms, are trained on a variety of corresponding anatomical regions and structures with specific noise levels. The noise level of an image is, in this case, a function of several factors, including kilovolt peaks (kVp), milliampere-seconds (mA-seconds), slice thickness, and patient size.
[0028]
[0032] While denoising image processors such as CNNs are formed in this case based on expected noise levels and standardized parameters, including standardized radiation doses, the systems and methods disclosed herein effectively fit such models to images acquired using various acquisition parameters, such as low radiation doses.
[0029]
[0033] The following considerations are specific to embodiments of CT-based medical imaging, but similar systems and methods can be used for other imaging modalities, such as magnetic resonance imaging (MRI) or positron emission tomography (PET).
[0030]
[0034] Figure 1 is a schematic diagram of a system 100 according to one embodiment of the present disclosure. The system 100 typically comprises a processing device 110 and an imaging device 120, as shown in the figure.
[0031]
[0035] The processing device 110 applies a processing routine to the received image. The processing device 110 comprises a memory 113 and a processor circuit 111. The memory 113 stores multiple instructions. The processor circuit 111 is coupled to the memory 113 and configured to execute instructions. The instructions stored in the memory 113 include not only processing routines but also data related to multiple machine learning algorithms, such as various convolutional neural networks for image processing.
[0032]
[0036] The processing device 110 further comprises an input unit 115 and an output unit 117. The input unit 115 receives information such as images from the imaging device 120. The output unit 117 outputs information to a user or a user interface device. The output unit 117 includes a monitor or display.
[0033]
[0037] In some embodiments, the processing device 110 is directly connected to the imaging device 120. In alternative embodiments, the processing device 110 is separate from the imaging device 120 and receives the image to be processed via a network or other interface located in the input unit 115.
[0034]
[0038] In some embodiments, the imaging device 120 comprises an image data processing device and a spectral or conventional CT scan unit that generates CT projection data when scanning an object (e.g., a patient).
[0035]
[0039] Figure 2 shows an exemplary imaging device according to one embodiment of the present disclosure. A CT imaging device is shown, and the following discussion concerns CT images, but it will be understood that similar methods can be applied to other imaging devices, and that images to which such methods are applied can be acquired in a wide variety of ways.
[0036]
[0040] In the imaging device according to embodiments of the present disclosure, the CT scan unit is adapted to perform multiple axial and / or helical scans of an object to generate CT projection data. In the imaging device according to embodiments of the present disclosure, the CT scan unit comprises an energy-resolving photon counting image detector. The CT scan unit comprises a radiation source that emits radiation across the object when acquiring projection data.
[0037]
[0041] In the imaging device according to the embodiments of the present disclosure, the CT scan unit further performs a scout scan that is different from the primary scan, thereby generating separate images associated with the scout scan and the primary scan, although the images are different but contain the same object.
[0038]
[0042] The CT scan unit 200, for example, a computed tomography (CT) scanner, comprises, in the example shown in Figure 2, a fixed gantry 202 and a rotating gantry 204 rotatably supported by the fixed gantry 202. The rotating gantry 204 rotates around the longitudinal axis of the examination area 206 of the object when acquiring projection data. The CT scan unit 200 includes a support 207 for supporting a patient within the examination area 206 and is configured to allow the patient to pass through the examination area during the imaging process.
[0039]
[0043] The CT scan unit 200 includes a radiation source 208, such as an X-ray tube, which is supported by a rotating gantry 204 and configured to rotate with the rotating gantry 204. The radiation source 208 includes an anode and a cathode. A voltage applied to the source between the anode and cathode accelerates electrons from the cathode to the anode. The flow of electrons allows for a flow of current from the cathode to the anode, which generates radiation across the examination area 206.
[0040]
[0044] The CT scan unit 200 includes a detector 210. The detector 210 demarcates an arc at a certain angle on the opposite side of the examination area 206 from the radiation source 208. The detector 210 comprises a one-dimensional or two-dimensional array of pixels, such as pixels of a direct conversion detector. The detector 210 is adapted to detect radiation crossing the examination area and generate a signal indicating the energy of the radiation.
[0041]
[0045] The CT scan unit 200 further comprises generators 211 and 213. Generator 211 generates tomography projection data 209 based on the signal from the detector 210. Generator 213 receives the tomography projection data 209 and generates a raw image of the object based on the tomography projection data 209.
[0042]
[0046] Figure 3 is a schematic diagram of a training pipeline used in one embodiment of the present disclosure. The training pipeline 300 is typically performed by a processing device 110, and the corresponding method is performed by a processor circuit 111 based on instructions stored in memory 113. Memory 113 further has a database structure that stores data necessary to perform a denoising method, such as the training pipeline 300 or an embodiment of a model that utilizes the data generated by the training pipeline. In some embodiments, the database is stored externally and accessed by the processing device 110 via a network interface.
[0043]
[0047] The method implemented in training pipeline 300 includes a step of training a learning algorithm, such as a CNN, for removing noise from low-dose CT images, using domain-independent learning with untangled feature representations.
[0044]
[0048] Domain adaptation is based on source domain D s From a given dataset, such as a source-labeled dataset belonging to a specific known target domain D t It can be defined as transferring knowledge to a target unlabeled dataset belonging to a specific domain. Domain-independent learning does not involve annotating each sample with domain labels, and both the target unlabeled dataset and the source dataset can consist of data from multiple domains (for example, in the target domain {D t1 ,D t2 ,...,D tn}, and in the source domain {D s1 ,D s2 ,...,D sn It is defined in a similar manner, except that}). Domain-independent learning is achieved by using detangled feature representation learning, which detangles style from content.
[0045]
[0049] The training pipeline 300 can thus untangle the entanglement of the anatomical structure features a from the noise features n for the CT images processed by the pipeline and reconstruct the underlying image using the anatomical structure features and the noise features. When evaluating the noise removal performance of the model, a radiologist scores the images for two types of quality, namely structural fidelity and image noise suppression. Structural fidelity is the ability of an image to accurately depict the anatomical structures within the field of view, and image noise appears as irregular patterns on the image, degrading the image quality. By extracting the anatomical structure features from low-quality images and combining the anatomical structure features with the low noise levels characteristic of higher-quality images, the model generated by the described training pipeline provides images that score highly on both metrics.
[0046]
[0050] As shown, the training pipeline 300 has a standard image module 310 and a low-quality image module 320. The standard image module 310 has a standard anatomical structure encoder E a CT 330, a standard noise encoder E n CT 340, and a standard generator G CT 350. A standard image source 360, which is either the CT scan unit 200 of FIG. 2 or an image database, then supplies a standard image 370 to the standard image module 310. When the standard image 370 is received, it is supplied to the standard anatomical structure encoder 330 and the standard noise encoder 340.
[0047]
[0051] The standard anatomical structure encoder 330 then outputs a standard anatomical structure feature a CT 380, and the standard noise encoder 340 outputs a standard noise feature n CT 390. Both the standard noise feature 390 and the standard anatomical structure feature 380 are then Supplied to standard generator 350 used to reconstruct a standard image X CT CT 370’ from the supplied standard anatomical structure feature 380 and standard noise feature 390 ru.The standard image module 310 can therefore decompose the standard image 370 into constituent features 380, 390, and then reconstruct the standard image 370' from these constituent features.
[0048]
[0052] The low-quality image module 320 has components in parallel with the components discussed for the standard image module 310. The low-quality image module 320 therefore has a low-quality anatomical structure encoder E a LDCT 430, Low-quality noise encoder E n LDCT 440, and low-quality generator G LDCT It has 450. The low-quality image source 460 then supplies the low-quality image 470 to the low-quality image module 320. The low-quality image source 460 is the CT scan unit 200 in Figure 2, where the parameters discussed above are reset to degrade image quality. The low-quality image source is based on a reduced radiation dose compared to the acquisition of the standard image 370, for example, to produce a low-dose CT scan (LDCT) image. Alternatively, the low-quality image source 460 may be an image database. Once the low-quality image 470 is received, it is supplied to the low-quality anatomical structure encoder 430 and the low-quality noise encoder 440.
[0049]
[0053] The low-quality anatomical structure encoder then processes the low-quality anatomical structure feature a LDCT Outputting 480, the low-quality noise encoder 440 outputs low-quality noise feature n LDCT Output 490. Both the low-quality noise feature 490 and the low-quality anatomical structure feature 440 are Low-quality generator 450 is supplied Next, from the supplied low-quality anatomical structural features 480 and low-quality noise features 490, a low-quality image X LDCT LDCT 470' can be reconfigured ru. The low-quality image module 320 can therefore decompose the low-quality image 470 into constituent features 480, 490, and then reconstruct the low-quality image 470' from these constituent features.
[0050]
[0054] A loss calculation module 500 is provided to calculate various loss functions when training the standard image module 310 and the low-quality image module 320. Therefore, the loss evaluation criterion of the standard generator 350 is at least partially based on a comparison between the standard image 370 and the reconstructed standard image 370'. Such a loss evaluation criterion is the reconstruction loss 510 of the standard generator 350, which is used during training to ensure that the reconstructed image 370' produced by the standard generator is exactly like the original supplied corresponding standard image 370.
[0051]
[0055] The loss evaluation criterion for the low-quality generator 450 is similarly based, at least in part, on a comparison between the low-quality image 470 and the reconstructed low-quality image 470' generated by the low-quality generator 450. The loss evaluation criterion is again the reconstruction loss 520 of the low-quality generator 450, and is used during training to ensure that the reconstructed image 470' from the low-quality generator is exactly as originally supplied as the corresponding low-quality image 470.
[0052]
[0056] The anatomical structure encoders 330 and 430 are also evaluated by the loss calculation module 500. The loss evaluation criterion for the standard anatomical structure encoder 330 is evaluated based on a comparison of standard anatomical structure features 380 with segmentation labels 540 generated for the corresponding standard image 370. Such segmentation labels are generated manually when the standard image 370 is acquired or read from a database. The loss evaluation criterion for such a segmentation label is the segmentation loss 530.
[0053]
[0057] The loss evaluation criterion for the low-quality anatomical structure encoder 430 is evaluated based on a comparison between the low-quality anatomical structure features 480 and the standard anatomical structure features 380 generated by the standard anatomical structure encoder 330. This comparison will be discussed in more detail below, but for the training methods shown in Figures 4A-4C, such a loss evaluation criterion would typically be an adversarial loss of 550.
[0054]
[0058] To evaluate the anatomical structure encoders 330 and 340, the loss calculation module 500 generates segmentation masks M for the corresponding anatomical structure features 380 and 480. a CT 570a and M a LDCT A segmentation network 560 is provided to create each of the 570b segments. The segmentation mask 570a for the standard anatomical structural features 480 is then compared with the segmentation label 540, and the segmentation mask 570b for the low-quality anatomical structural features 480 is then compared with the segmentation mask 570a for the standard anatomical structural features 380.
[0055]
[0059] Loss evaluation criteria are incorporated into the loss function to tune each generator 350, 450 and anatomical structure encoder 330, 430 using machine learning. Thus, the performance of each generator pipeline can be improved by adjusting the variables that determine the output of the corresponding module using the reconstruction losses 510, 520 for generators 350, 450, respectively. This is done by implementing standard machine learning techniques or by implementing exemplary training methods discussed below with reference to Figures 4A-4C.
[0056]
[0060] In some embodiments, the loss calculation module further generates additional loss evaluation criteria. Such loss evaluation criteria include a segmentation loss 580 for reconstructing the standard image 370' generated by the standard generator 350. Thus, the segmentation network 590 is applied to the reconstructed image 370' to form a segmentation mask M CT CT Generate 600, then segmentation mask M CT CT However, it is evaluated based on the segmentation label 540 for the corresponding image 370. The segmentation loss 580 is used in the training pipeline 500 and is taken into account at every training process.
[0057]
[0061] As shown in the figure, the anatomical structure features 380 output from the standard anatomical structure encoder 330 are connected to the low-quality generator 450, and similarly, the anatomical structure features 480 output from the low-quality anatomical structure encoder 430 are connected to the standard generator 350. When the generator 350 is supplied with the standard anatomical structure features 380 from the standard anatomical structure encoder 330, it outputs a reconstructed standard image 370' based on the standard anatomical structure features and the standard noise features 390 generated by the standard noise encoder 340. On the other hand, when the standard generator 350 is supplied with the low-quality anatomical structure features 480 generated by the low-quality anatomical structure encoder 430, it outputs a reconstructed standard transition image X LDCT CT Output 470''.
[0058]
[0062] The reconstructed standard transcription image 470'' contains a noise level lower than the noise level represented by the low-quality anatomical structure features 480 and the low-quality noise features 490. The reconstructed standard transcription image is constructed by the generator 350 based on the low-quality anatomical structure features 480 and the noise features 390 generated by the standard noise feature encoder 340. Such noise features 390 are, for example, the average of standard noise features generated from the corresponding standard image 370 during training, in the case of a transfer image.
[0059]
[0063] Similarly, a transition image is generated using the low-quality generator 450. The generator 350 is therefore supplied with standard anatomical structural features 380 from the standard anatomical structural encoder 330, and the reconstructed low-quality transition image X CT LDCT Output 370''. The reconstructed low-quality transition image 370'' includes noise levels based on standard anatomical structural features 380 and low-quality noise features 490.
[0060]
[0064] To further enhance the quality of any model trained using the training pipeline 500, additional loss evaluation criteria are generated using the transition images 370'' and 470''. The reconstructed standard-quality transition image 470'' is then evaluated using an adversarial loss 610, comparing the transition image to the reconstructed image 370'' output from the standard generator 350.
[0061]
[0065] The reconstructed low-quality transfer image 370'' is similarly evaluated using a loss criterion. Since the low-quality transfer image 370'' contains standard anatomical structural features 380, the training pipeline 300 will normally have access to the corresponding segmentation labels 540. The loss criterion for the low-quality transfer image 370'' is therefore the segmentation loss 620, and the segmentation network 630 is appropriate for the segmentation mask M CT LDCT Generate 640.
[0062]
[0066] It should be further noted that in some embodiments, the segmentation networks 560, 590, and 630 themselves are evaluated based on the segmentation losses 530, 580, and 620, which are the result of comparing the resulting segmentation masks 570a, 580, and 620 with the segmentation labels 540.
[0063]
[0067] In some embodiments, the entire training pipeline 300 is trained simultaneously. Thus, both the standard image module 310 and the low-quality image module 320 are supplied with a variety of images for the purpose of training each module. As the network improves, the training pipeline is instructed to generate transition images 470'' and 370'' based on loss evaluation criteria, anatomical structural features, and reconstructed images, which are then evaluated in parallel.
[0064]
[0068] In some embodiments, the training pipeline 300 is trained sequentially, as will be discussed in detail below with reference to Figures 4A to 4C.
[0065]
[0069] The described training pipeline 300 is used to generate a model that produces a transition image 470'', which in this case can be used to remove noise from the previous low-quality image 470. The training pipeline 300 is used to create modules for transitioning anatomical structural features from a wide variety of source images acquired using various imaging parameters. For example, different encoders are trained to transition images acquired using different amounts of radiation.
[0066]
[0070] Furthermore, the same model created using the training pipeline 300 is used across different anatomical features, although in some embodiments, the use of the model trained using the pipeline is limited to specific anatomical structures. The low-quality image 470 is not an image of the same patient or organ as the standard image 370, but in this case, it would relate to the same anatomical structure in a different image. The model is trained with respect to images of, for example, the head, abdomen, or a specific organ.
[0067]
[0071] Figures 4A to 4C show exemplary training methods used in the training pipeline 300 in Figure 3.
[0068]
[0072] In some embodiments, as shown in Figure 4A, the standard image module 310 is trained before the low-quality image module 320 is trained. In such embodiments, as shown in Figure 4B, when the low-quality image module 320 is trained, the variables associated with the standard image module, including the variables incorporated into the standard image encoders 330 and 340, the variables of the standard generator 350, and the variables of the segmentation network 560 associated with the standard anatomical structural features 380, are kept constant.
[0069]
[0073] In this embodiment, the training pipeline 300 trains the standard image module 310 and the low-quality image module 320 separately, and then further trains the standard generator 350 while retaining values for the standard anatomical structure encoder 330, the standard noise encoder 340, and the low-quality anatomical structure encoder 430.
[0070]
[0074] Therefore, when training a model using the method described in the claims, the method for training the pipeline 300 according to this disclosure first prepares a standard image module 310. The standard image module 310 includes a standard anatomical structure encoder 330 that extracts standard anatomical structure features 380 from a standard image 370, and a standard noise encoder 340 that extracts standard noise features 390 from a standard image.
[0071]
[0075] The standard image module 310 also includes a standard generator 350 that generates a reconstructed standard image 370' from standard anatomical structural features 380 and standard noise features 390.
[0072]
[0076] This method then prepares a low-quality image module 320. The low-quality image module 320 includes a low-quality anatomical structure encoder 430 that extracts low-quality anatomical structure features 480 from a low-quality image 470, and a low-quality noise encoder 440 that extracts low-quality noise features 490 from a low-quality image.
[0073]
[0077] The low-quality image module 410 also includes a low-quality generator 450 that generates a reconstructed low-quality image 470' from low-quality anatomical structural features 480 and low-quality quasi-noise features 490.
[0074]
[0078] The standard image module 310 is then trained by receiving multiple standard images 370. The standard images will be received from an image source 360, which is either a CT scan unit 200 or, alternatively, an image database. For each received standard image 370, the system then compares the reconstructed standard image 370' output from the standard image generator 350 with the corresponding standard image and extracts a standard reconstruction loss criterion. The standard encoders 330, 340 and the standard generator 350 are then adjusted by updating the variables of the standard image module 310 based on this loss criterion.
[0075]
[0079] This method generates a segmentation mask 570a in the segmentation network 560 that corresponds to the standard anatomical structure features 380 extracted from each standard image 370 by the standard anatomical structure encoder 330, while the standard image module 310 is still being trained. The segmentation mask 570a for each standard image 370 is then compared with the segmentation label 540 associated with the corresponding standard image 370 to generate a standard anatomical loss criterion 530. At least one variable of the standard anatomical structure encoder 330 is then updated based on the standard anatomical loss criterion.
[0076]
[0080] As shown in Figure 4A, these initial training steps are performed as part of the training pipeline 300 related to the standard image 370, independently of the low-quality image module 320. Additional training elements are performed during this initial part of training to improve the models used for the standard encoders 330, 340 and the standard generator 350. For example, the reconstructed standard image 370' is submitted to an additional segmentation network 590 to generate a segmentation mask 600, which is then compared to known segmentation labels 540 for the corresponding image 370 to generate additional segmentation loss evaluation criteria.
[0077]
[0081] After the standard image module 510 is trained and yields acceptable results, the low-quality image module 320 is trained, as shown in Figure 4B. Therefore, multiple low-quality images 470 are received by the low-quality image module 320 from the image source 460. As discussed above, the image source is either a CT scan unit 200, an image database, or any combination thereof.
[0078]
[0082] The system then compares the reconstructed low-quality image 470' output from the low-quality image generator 450 for each received low-quality image 470 with the corresponding low-quality image and extracts the low-quality reconstruction loss 520. The low-quality encoders 430, 440 and the low-quality generator 450 are then adjusted by keeping the variables of the standard image module 310 constant while updating the variables of the low-quality image module 320 based on the loss evaluation criteria.
[0079]
[0083] The segmentation network 560, which is initially trained when training the standard image module 310, is similarly kept constant when training the low-quality image module 320. The segmentation network 560 is then used to generate a segmentation mask 570b corresponding to the low-quality anatomical structure features 480 extracted from each low-quality image 470 by the low-quality anatomical structure encoder 430. The segmentation mask 570b is then compared to at least one of the segmentation masks 570a previously generated by the standard image module 310, usually the average of the segmentation masks 570a, in order to generate an adversarial loss metric.
[0080]
[0084] The adversarial loss metric is then used to update at least one variable of the low-quality anatomical structure encoder 430.
[0081]
[0085] Once the low-quality image module 320 is trained in this manner, the low-quality anatomical structure features 480 generated by the low-quality anatomical structure encoder 430 are then fed to the standard generator 350 of the standard image module 310. The standard generator 350 then outputs a reconstructed standard transition image 470'' based on the low-quality anatomical structure features 480 and at least one of the standard-quality noise features 390, and possibly the average of the standard-quality noise features 390.
[0082]
[0086] The reconstructed standard transfer image 470'' is then evaluated by comparing the reconstructed standard transfer image 470'' with the standard image reconstructor 370'' or the average of several such reconstructors in order to generate an adversarial loss 610. The standard generator 350 is kept constant and therefore not adjusted based on the adversarial loss evaluation criterion, but the low-quality anatomical structure encoder 430 is further adjusted based on the adversarial loss evaluation criterion.
[0083]
[0087] Furthermore, as discussed above, a reconstructed low-quality transition image 370'' is generated using the low-quality generator 450. Such a transition image 370'' is analyzed by the segmentation network 630 based on standard anatomical structural features 380 generated by the standard image module 320 to generate a corresponding segmentation map 640, which is then evaluated against segmentation labels 540 to generate a segmentation loss 620. The segmentation loss evaluation criterion may then be used to further refine the low-quality module 320.
[0084]
[0088] In some embodiments, the model created in the described training pipeline 300 is then used to create a denoising standard transition image 470''. In other embodiments, such as shown in Figure 4C, the training pipeline is then further trained. As shown, the associated encoders 330, 340, and 430 are kept constant while the standard generator 350 is further tuned. Training is performed by creating an additional reconstructed standard transition image 470'', which can then be compared to one or more standard reconstructs 370'' to generate additional adversarial loss data 610, which can then be used to further tune the standard generator 350.
[0085]
[0089] Thus, training the described training pipeline 300 presents several design checks that facilitate domain-independent incorporation of anatomical structures and preservation of anatomical structures.
[0086]
[0090] Figure 5 illustrates the use of a model trained using the training pipeline 330 in Figure 3 to denoise images. When denoising images, the model first uses a low-quality anatomical structure encoder 430 to encode low-dose anatomical structure data from low-quality images 470 into low-quality anatomical structure features 480. The low-quality anatomical structure features 480 are then fed to a standard generator 350 along with standard noise features 390 or an average of such noise features to produce a reconstructed standard transition image 470''.
[0087]
[0091] This training pipeline 300 can be used to train the conversion of low-dose CT or other low-quality images to multiple normal-dose CT images or standard images with different noise levels. In the case of CT images, this noise level can represent the overall average of normal-dose CT data, or the average of a group of CT data with specific characteristics reconstructed using a specific algorithm such as filtered back projection or iterative reconstruction. Trained CT noise encoder E n CT 330 is used on all or a subset of normal dose CT images 370 in the training set to acquire these noise features, encoding the normal dose CT images into CT noise features 390. Then, its own representative CT noise features, n CT1 , n CT2 To obtain these features, the average of the CT noise features within each group is taken, and each of these CT noise features can then be used as the average noise feature 390 discussed above. Users can select from predefined CT noise features 390 based on their specific needs to remove noise from low-dose CT data. The training pipeline 300 also performs interpolation of the CT noise features extracted from various groups so that users can continuously adjust the CT noise.
[0088]
[0092] Although the methods described herein are described in relation to CT scan images, a wide range of imaging techniques, including various medical imaging techniques, are intended, and it will be understood that images generated using a wide variety of imaging techniques can be effectively denoised using the methods described herein.
[0089]
[0093] The methods described herein are implemented on a computer, on dedicated hardware, or a combination of both, as a method implemented on a computer. The executable code of the methods described herein is stored in a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, and online software. Preferably, the computer program product includes non-temporary program code stored on a computer-readable medium for performing the methods described herein when the program product is executed on a computer. In one embodiment, the computer program includes computer program code adapted to perform all steps of the methods described herein when the computer program is executed on a computer. The computer program is embodied on a computer-readable medium.
[0090]
[0094] While this disclosure has described some described embodiments in some detail and with some specificity, it is not intended to limit itself to any such detail, embodiment, or particular embodiment. Rather, in consideration of the prior art, the claims should be interpreted with reference to the attached claims to give the broadest possible interpretation and thereby effectively encompass the scope intended by this disclosure.
[0091]
[0095] All examples and conditional statements listed herein are intended for educational purposes to help readers understand the principles of the Disclosure and the concepts to which the inventors contribute to further advancing the art, and should be construed as not limiting such specifically listed examples and conditions. Furthermore, all descriptions herein listing the principles, aspects, and embodiments of the Disclosure, as well as specific examples thereof, are intended to encompass both structural and functional equivalents of the Disclosure. Moreover, such equivalents are intended to include both currently known equivalents and equivalents to be developed in the future, i.e., any elements to be developed that perform the same function regardless of their structure.
Claims
1. A system for removing noise from medical images, wherein the medical images include standard images, which are normal dose CT images, and low-quality images, which are low-dose CT images with low image quality, and the system removes noise from medical images. A standard image module generates standard anatomical structural features and standard noise features from the standard image, and reconstructs the standard image from the standard anatomical structural features and standard noise features to generate a reconstructed standard image. A low-quality image module generates low-quality anatomical structural features and low-quality noise features from the low-quality image, and reconstructs the low-quality image from the low-quality anatomical structural features and low-quality noise features to generate a reconstructed low-quality image. A loss calculation module that calculates a loss evaluation criterion based at least partially on 1) a comparison of the reconstructed standard image with the standard image, and 2) a comparison of the reconstructed low-quality image with the low-quality image. A system comprising: a loss evaluation criterion incorporated into a loss function for adjusting the standard image module and the low-quality image module using machine learning; and when the low-quality anatomical structural features are supplied to the standard image module, the standard image module outputs a reconstructed standard transition image containing a noise level lower than the noise level of the reconstructed low-quality image generated by the low-quality anatomical structural features and the low-quality noise features.
2. The standard image module comprises a standard anatomical structure encoder, a standard noise encoder, and a standard generator. Upon receiving the standard image, the standard anatomical structure encoder outputs the standard anatomical structure features, the standard noise encoder outputs the standard noise features, and the standard generator reconstructs the standard image from the standard anatomical structure features and the standard noise features. The low-quality image module comprises a low-quality anatomical structure encoder, a low-quality noise encoder, and a low-quality generator. Upon receiving the low-quality image, the low-quality anatomical structure encoder outputs the low-quality anatomical structure features, the low-quality noise encoder outputs the low-quality noise features, and the low-quality generator reconstructs the low-quality image from the low-quality anatomical structure features and the low-quality noise features. The loss calculation module calculates the loss evaluation criterion of the standard generator, at least in part, based on a comparison between the reconstructed standard image and the standard image. The loss calculation module calculates the loss evaluation criterion for the low-quality generator, at least partially based on a comparison of the reconstructed low-quality image with the low-quality image. The loss calculation module calculates the loss evaluation criterion for the standard anatomical structure encoder based on a comparison with the segmentation labels for the standard image, The loss calculation module calculates the loss evaluation criterion for the low-quality anatomical structure encoder based on a comparison with the output of the standard anatomical structure encoder. The system according to claim 1.
3. The system according to claim 2, wherein the loss evaluation criterion for the low-quality anatomical structure encoder is an adversarial loss evaluation criterion.
4. The system according to claim 1, further comprising a segmentation network, wherein the segmentation mask of the reconstructed standard image is evaluated based on a comparison with segmentation labels relating to the standard image.
5. The system according to claim 4, wherein when the standard anatomical structural features are supplied to the low-quality generator, the low-quality generator outputs a reconstructed low-quality transition image, the reconstructed low-quality transition image contains a noise level higher than the noise level represented by the standard anatomical structural features and the standard noise features, and the segmentation mask of the reconstructed low-quality transition image is evaluated based on a comparison with the segmentation labels relating to the standard image.
6. The system according to claim 1, wherein the loss evaluation criterion for the reconstructed standard transition image is evaluated based on comparison with the reconstructed standard image, and the loss evaluation criterion for the reconstructed standard transition image is an adversarial loss evaluation criterion.
7. The system according to claim 1, wherein the standard image module and the low-quality image module are trained simultaneously.
8. The system according to claim 1, wherein the standard image module is trained before the low-quality image module is trained, and the values of the variables created during the training of the standard image module are kept constant during the training of the low-quality image module.
9. The system according to claim 8, wherein, after training the standard image module and the low-quality image module, the system further trains a standard generator while keeping the standard anatomical structure encoder, the standard noise encoder, and the low-quality anatomical structure encoder constant values.
10. The system according to claim 1, wherein the standard anatomical structural features and the low-quality anatomical structural features each correspond to a single anatomical structure.
11. A method for removing noise from medical images, wherein the medical images include a standard image which is a normal dose CT image and a low-quality image which is a low-dose CT image with low image quality, and the method is The standard image module generates standard anatomical structural features and standard noise features from the standard image, The steps include: Reconstructing the standard image from the standard anatomical structural features and the standard noise features using the standard image module to generate a reconstructed standard image; The low-quality image module generates low-quality anatomical structural features and low-quality noise features from the low-quality image, The low-quality image module is used to reconstruct the low-quality image from the low-quality anatomical structural features and the low-quality noise features to generate a reconstructed low-quality image. The loss calculation module performs the steps of: 1) calculating a loss evaluation criterion based at least in part on a comparison between the reconstructed standard image and the standard image, and 2) comparing the reconstructed low-quality image and the low-quality image. A method comprising: a loss evaluation criterion incorporated into a loss function for adjusting the standard image module and the low-quality image module using machine learning; and when the low-quality anatomical structural features are supplied to the standard image module, the standard image module outputs a reconstructed standard transition image having a noise level lower than the noise level of the reconstructed low-quality image generated by the low-quality anatomical structural features and the low-quality noise features.
12. The standard image module comprises a standard anatomical structure encoder, a standard noise encoder, and a standard generator. Upon receiving the standard image, the standard anatomical structure encoder outputs the standard anatomical structure features, the standard noise encoder outputs the standard noise features, and the standard generator reconstructs the standard image from the standard anatomical structure features and the standard noise features. The low-quality image module comprises a low-quality anatomical structure encoder, a low-quality noise encoder, and a low-quality generator. Upon receiving the low-quality image, the low-quality anatomical structure encoder outputs the low-quality anatomical structure features, the low-quality noise encoder outputs the low-quality noise features, and the low-quality generator reconstructs the low-quality image from the low-quality anatomical structure features and the low-quality noise features. The loss calculation module calculates the loss evaluation criterion of the standard generator, at least in part, based on a comparison between the reconstructed standard image and the standard image. The loss calculation module calculates the loss evaluation criterion for the low-quality generator, at least partially based on a comparison of the reconstructed low-quality image with the low-quality image. The loss calculation module calculates the loss evaluation criterion for the standard anatomical structure encoder based on a comparison with the segmentation labels for the standard image, The loss calculation module calculates the loss evaluation criterion for the low-quality anatomical structure encoder based on a comparison with the output of the standard anatomical structure encoder. The method according to claim 11.
13. The method according to claim 12, wherein the loss evaluation criterion for the low-quality anatomical structure encoder is an adversarial loss evaluation criterion.
14. The method according to claim 11, further comprising the steps of generating a segmentation mask for the reconstructed standard image in the segmentation network, and evaluating the segmentation mask based on a comparison with segmentation labels relating to the standard image.
15. The method according to claim 14, wherein when the standard anatomical structural features are supplied to the low-quality generator, the low-quality generator outputs a reconstructed low-quality transition image, the reconstructed low-quality transition image includes a noise level higher than the noise level represented by the standard anatomical structural features and the standard noise features, and the segmentation mask of the reconstructed low-quality transition image is evaluated based on a comparison with the segmentation labels relating to the standard image.
16. The method according to claim 11, wherein the loss evaluation criterion for the reconstructed standard transition image is evaluated based on comparison with the reconstructed standard image, and the loss evaluation criterion for the reconstructed standard transition image is an adversarial loss evaluation criterion.
17. The method according to claim 11, wherein the standard image module and the low-quality image module are trained simultaneously.
18. The method according to claim 11, wherein the standard image module is trained before the low-quality image module is trained, and the values of the variables created during the training of the standard image module are kept constant during the training of the low-quality image module.
19. The method according to claim 18, wherein, after training the standard image module and the low-quality image module, the method further trains a standard generator while keeping the standard anatomical structure encoder, the standard noise encoder, and the low-quality anatomical structure encoder constant.
20. The method according to claim 11, wherein the standard anatomical structural features and the low-quality anatomical structural features each correspond to a single anatomical structure.
Citation Information
Patent Citations
A Video Super-Resolution Reconstruction Method Based on Residual Convolutional Networks
CN112348745B
Patient-specific deep learning image denoising methods and systems
JP2020064609A
Medical image processing device, medical image processing method, and program
JP2020166814A
Image Rescaling
JP2023517486A
Cone-beam CT image enhancement using generative adversarial networks
US20190333219A1