Controllable Standardless Noise Removal of Medical Images

A neural network trained on sequences of medical images without clean references adjusts noise removal based on acquisition parameters, addressing limitations of existing methods and enhancing image quality across varying imaging conditions.

JP2025520779APending Publication Date: 2025-07-03KONINKLIJKE PHILIPS NV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024576400
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-06
Filing Date
2023-07-04
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing noise removal methods for medical images, particularly in CT imaging, require clean reference images for training, which are difficult to obtain and result in limited generalization across varying imaging parameters, leading to suboptimal noise removal and excessive smoothing or incomplete removal.

Method used

A method that trains a neural network using a sequence of connected images without a clean reference, modeling the distribution of clean and noisy data, and adjusts noise removal based on acquisition parameters, enabling flexible noise removal across different noise levels and imaging conditions.

Benefits of technology

Enables effective noise removal in medical images acquired with various parameters, improving image quality and diagnostic accuracy by using adjacent frames to predict and adjust noise levels, reducing reliance on clean reference images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520779000001_ABST
    Figure 2025520779000001_ABST
Patent Text Reader

Abstract

A method for training a machine learning model for noise removal is provided. The method includes the step of capturing a target image data frame. The target image data frame is one image data frame of a series including imaging data of a subject. The method further includes the step of capturing at least one preceding image data frame and at least one subsequent image data frame in the series. The content of the preceding and subsequent image data frames each partially overlaps with the content of the target image data frame. The method further includes the step of capturing acquisition parameters associated with the series of image data frames, and the step of generating a prediction of the noise-removed target image data frame based on the preceding and subsequent image data frames. The method trains a machine learning algorithm based on a noise model based on the prediction and the acquisition parameters. A system and a noise removal method are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] The present disclosure broadly relates to a system and method for training and using a neural network model for removing noise from an image without using a reference image during training. In particular, the present disclosure relates to a system and method for training and using such a neural network model in the context of computer tomography (CT) images.

Background Art

[0002]

[0002] Conventionally, in imaging modalities such as computer tomography, there are phenomena that lead to artifacts (such as noise) in the final image during acquisition physics or reconstruction. In radiation medicine, low-dose computed tomography (LDCT) is widely used, but the reduction of X-ray dose increases the noise level, which affects the diagnostic performance.

[0003]

[0003] The technology that functions best for noise reduction is based on a deep learning model trained on a pair of clean / noisy images of the same anatomical structure. Therefore, in order to train a noise removal algorithm using machine learning such as a neural network model, usually, a pair of noisy image samples and noise-free image samples are supplied to the neural network model, and the network tries to minimize the cost function by removing noise from the noisy image and restoring the corresponding noise-free ground truth image.

[0004]

[0004] It is difficult to obtain a noise-free image, i.e., a clean image. This is because these images typically require a high radiation dose to generate high-quality images. Obtaining such data in a clinical environment typically takes at least twice as long as a normal examination and can increase the patient's exposure to radiation. Further, even if such an approach is taken, the paired images may not be ideally aligned due to patient movement and can generate individual artifacts in the noise-removed images. Thus, obtaining a pair of images that can be used for training is particularly difficult in a clinical environment.

[0005]

[0005] Currently, the paired data sets are created by adding synthetic noise to actual high-quality images. However, such noise models may fail if the underlying mathematical assumptions do not match the acquisition settings. Thus, there is a need for noise removal techniques that enable training without making clean high-dose images available.

[0006]

[0006] Further, the training of neural network models for noise removal is typically specialized for a particular image quality, where the image quality depends on the acquisition parameters of the image. Thus, an image acquired at half of a typical radiation dose will have a different amount of noise than an image acquired at one-fourth of a typical radiation dose.

[0007]

[0007] Therefore, the neural network model used for noise removal is usually specific to a group of acquired parameters, and a change in the parameters used to acquire an image can, as a result, cause a change in the form or amount of artifacts in the corresponding image. Therefore, a noise removal model used to remove noise from an image acquired using a first group of parameters becomes less effective when applied to an image acquired using different acquisition parameters, such as a reduced radiation dose in the context of a CT scan. Therefore, the noise level of the images used during training limits the generalization ability of the trained noise removal model. For example, if the training set includes images with a medium noise level, the algorithm will not be able to accurately remove noise from images with a high noise level. Usually, the application of the algorithm will result in excessive smoothing or incomplete noise removal.

[0008]

[0008] In CT imaging, multiple factors including the peak tube voltage measured in (kVp), the tube current measured in milliampere-seconds (mAs), the slice thickness, the column position, and the patient size can all affect the noise level of the reconstructed image. The result is that changing any of these imaging parameters can, as a result, cause different noise levels or different artifact profiles, which, as before, hinders the generalization of noise removal capabilities and requires different models to remove noise from images acquired with such different imaging parameters. This limits the applicability of CNN-based methods in practical noise removal.

[0009]

[0009] The supervised noise removal method depends on the characteristics of the data used for training. Therefore, the degree of noise removal can only be adjusted by selecting appropriate data for training or by other post-processing techniques such as weighted combination of the initial noisy image and its noise residual. This technique is known as overcorrection.

[0010]

[0010] Current self-supervised noise removal methods do not require a low-noise reference image but have several weaknesses. These methods do not use the information contained in a sequence of connected images, such as CT projections. Instead, all images are denoised individually, which leads to suboptimal image quality and a longer execution time. Some of these methods require additional data, complicating the operation of the noise removal method. Furthermore, like supervised methods, they cannot adjust the noise removal level.

Summary of the Invention

Problems to be Solved by the Invention

[0011]

[0011] Therefore, there is a need for a controllable noise removal method that can be trained without a clean reference image and can be used to denoise images acquired with various imaging parameters. Furthermore, there is a need for a single trained model that can be used to denoise images with various noise levels, including CT images acquired at a lower radiation dose than the images used for training.

Means for Solving the Problems

[0012]

[0012] A system and method for removing noise from medical images are provided in which a reference image that should be used as a clean image is not used for training. The proposed method instead relies on the similarity present in a sequence of connected images and models the distribution of clean and noisy data. A special noise module may enable the noise removal algorithm to generalize for different noise levels and adjust the degree of noise removal by adjusting a parameter interpretable manually or automatically. The ability to adjust the noise module and train with only noisy data makes the proposed reference-free noise removal method very flexible for clinical applications.

[0013]

[0013] In some embodiments, a method for training a machine learning model for noise removal is provided. The method includes the step of capturing a target image data frame. The target image data frame is one image data frame in a series of image data frames including imaging data of a subject.

[0014]

[0014] The method further includes the step of capturing at least one previous (preceding) image data frame in the series of the image data frames that is before the target image data frame in the series. The content of the at least one previous image data frame at least partially overlaps with the content of the target image data frame. The method further includes the step of capturing at least one subsequent (following) image data frame in the series of the image data frames that is after the target image data frame in the series. The content of the at least one subsequent image data frame at least partially overlaps with the content of the target image data frame.

[0015]

[0015] The method further includes the step of capturing acquisition parameters related to the acquisition of the image data frames in the series of the image data frames, and the step of generating a prediction of the denoised target image data frame based on the at least one previous image data frame and the at least one subsequent image data frame.

[0016]

[0016] The method then trains a machine learning algorithm for denoising the target image data frame based on a noise model based on the prediction of the denoised target image data frame and the acquisition parameters.

[0017]

[0017] In some embodiments, the prediction of the noise-removed target image data frame is an estimation of the mean and standard deviation of the noise-removed target image data frame based on the representation of at least one anatomical feature structure extracted from each of at least one previous image data frame and at least one subsequent image data frame. The representations from the at least one previous image data frame and the at least one subsequent image data frame are then fused to form a prediction of the noise-removed target image data frame.

[0018]

[0018] In some such embodiments, the representation of at least one anatomical feature structure is transmitted between frames using a convolutional memory unit. Such a convolutional memory unit can be a convolutional long short-term memory unit that communicates information between frames in a sequence of image data frames.

[0019]

[0019] In some embodiments, the at least one previous image data frame is a plurality of image data frames prior to the target image data frame in the sequence of image data frames. The at least one subsequent image data frame can similarly be a plurality of image data frames after the target image data frame in the sequence of image data frames.

[0020]

[0020] In some embodiments, the prediction of the noise-removed target image data frame is output by a trained convolutional neural network to which at least one previous image data frame and at least one subsequent image data frame are supplied.

[0021]

[0021] In some embodiments, the loss function for training a machine learning algorithm is based on the distribution of noise predicted based on the acquisition parameters. Such a noise distribution can be of a Poisson-Gaussian distribution.

[0022]

[0022] In some such embodiments, the training method is repeated for a series of projection frames obtained using different acquisition parameters, and the adjustment variable is extracted or generated for the machine learning algorithm based on results related to different acquisition parameters. In some such embodiments, the adjustment variable is trained based on the variance of the tube current associated with the acquisition of a related series of image data frames.

[0023]

[0023] In some other such embodiments, the adjustment variable is a scaling factor that determines to what extent noise should be identified to be removed by the machine learning algorithm.

[0024]

[0024] In some embodiments, the acquisition parameters are extracted from DICOM files related to a series of projection frames. In some embodiments, the machine learning algorithm is a convolutional neural network.

[0025]

[0025] In some embodiments, the imaging data is CT imaging data, and each image data frame of the series of image data frames is a projection frame. Each projection frame includes imaging data of the same subject acquired from different angles.

[0026]

[0026] There is also provided a noise removal method in which a machine learning algorithm is trained in the manner described above. The method includes taking in a series of noisy image data frames including a noisy target image data frame to be noise-removed. The method includes the step of selecting a value of the adjustment variable based on the acquisition parameters of the series of noisy image data frames.

[0027]

[0027] The method includes applying a trained machine learning algorithm to a series of noisy image data frames using a selected value of an adjustment variable. The method then generates a first denoised image data frame based on an estimation of an average and a standard deviation based on the series of noisy image data frames, a distribution of noise based on acquisition parameters of the series of noisy image data frames, and a noisy target image data frame.

[0028]

[0028] In some such embodiments, the adjustment variable is trained to correspond to an acquisition parameter in training data, and the selected value of the adjustment variable is different from the actual value of the corresponding acquisition parameter of the noisy image data frame. The machine learning algorithm, in this case, identifies more noise in the noisy image data frame when using the selected value than when using the actual value.

[0029]

[0029] In some embodiments, the noise removal method further includes reconstructing an image based on a plurality of denoised image data frames including the first denoised image data frame, and outputting the reconstructed image to a user.

[0030]

[0030] Also provided is a machine learning training system including a memory storing a plurality of instructions and a processor circuit coupled to the memory and configured to execute the instructions to implement the training method described above.

Brief Description of the Drawings

[0031]

Figure 1

[0031] FIG. 1 is a schematic diagram of a system according to an embodiment of the present disclosure.

Figure 2

[0032] FIG. 2 shows an exemplary imaging device according to an embodiment of the present disclosure.

Figure 3

[0033] Figure 3 shows a pipeline for training a model used to remove noise from an image according to the present invention.

Figure 4

[0034] Figure 4 shows the use of a model for removing noise from an image according to the present invention.

Figure 5A

[0035] Figure 5A is a flowchart showing a method for training a model for removing noise from an image according to the present disclosure.

Figure 5B

[0036] Figure 5B is a flowchart showing a method for removing noise from an image according to the present disclosure.

Figure 6

[0037] Figure 6 is a schematic diagram of the use of adjustment variables in a model for removing noise from an image according to the present invention.

Mode for Carrying Out the Invention

[0032]

[0038] The description of exemplary embodiments in accordance with the principles of the present disclosure is intended to be read in conjunction with the accompanying drawings, which are to be regarded as an integral part of the overall description. In the description of the embodiments disclosed herein, any reference to direction or orientation is for convenience of explanation only and is not intended to limit the scope of the present disclosure in any way. Relative terms such as "lower," "upper," "horizontal," "vertical," "above," "below," "on," "under," "top," "bottom," etc., and their derivatives (such as "horizontally," "downwardly," "upwardly," etc.) should be construed to refer to the orientation that is being described at that time or shown in the drawings being described. These relative terms are for convenience of explanation only and do not require the device to be configured or operated in a particular orientation unless explicitly so stated. Terms such as "attached," "fixed," "connected," "coupled," "interconnected," etc., unless otherwise specified, indicate a relationship in which structures are directly or indirectly fixed or attached to each other through intervening structures, and both movable or fixed attachments or relationships. Further, the features and advantages of the present disclosure are described with reference to the illustrated embodiments. Accordingly, the present disclosure should not be explicitly limited to exemplary embodiments showing combinations of features that may exist alone or in other combinations, which are not limiting. That is, the scope of the present disclosure is defined by the claims appended hereto.

[0033]

[0039] The present disclosure describes the presently contemplated best mode of implementing the present disclosure. This description is not intended to be understood in a limiting sense and provides an example of the present disclosure presented for illustrative purposes only with reference to the accompanying drawings to show the advantages and configurations of the present disclosure to those skilled in the art. In the various figures of the drawings, like reference characters indicate like or similar parts.

[0034]

[0040] It is important to note that the disclosed embodiments are only some advantageous use cases of the innovative teachings herein. Generally, the descriptions made in the specification of this application do not necessarily limit any of the disclosures described in the various claims. Further, some descriptions may apply to some inventive features but may not apply to others. Generally, unless otherwise indicated, elements in the singular may be plural and vice versa without loss of generality.

[0035]

[0041] Generally, images acquired for use in a medical environment require some processing to remove noise from these images. Such noise removal is essential in a medical environment where the images are likely to be used for diagnosis and treatment. This is because the accuracy and precision in such images can improve their usefulness. Such noise removal is typically implemented using a machine learning-based algorithm such as a convolutional neural network (CNN).

[0036]

[0042] The CNN used for noise removal requires training to appropriately recognize noise in the context of medical imaging. Conventionally, such a CNN is trained using pairs of images, where each pair includes a first "noisy" (noisy) image and a second clean image, and the clean image is used as the ground truth. In this case, the CNN is trained to compare the noisy image with the clean image and process the noisy image to approximate the output image to the clean image. To train the CNN in this way, a training set containing a large number of image pairs is required. Further, to achieve consistent results, the training set typically includes images acquired using a consistent set of parameters, and the resulting CNN is typically limited to new images acquired using the same or similar acquisition parameters.

[0037]

[0043] As described above, the generation of such a training set is difficult and time-consuming. Therefore, the systems and methods described herein do not require pairs of images and instead rely on adjacent frames within a sequence of image data frames. Such an approach enables the reproduction of the image data of a target frame from the image data from adjacent frames within the sequence of image data frames. Such adjacent frames may provide different representations of the content of the target frame with different noise distributions. By combining a plurality of adjacent frames, a prediction regarding the target frame can be generated. Such a prediction may be an estimation of the mean and standard deviation of the noise-removed version of the target frame. Then, the CNN can be trained using the noisy target frame as ground truth data, in which case the training loss function is based on the prediction generated based on a plurality of adjacent frames.

[0038]

[0044] When training a CNN based on a sequence of images, the noise in the CT projection depends on technical acquisition parameters that can be extracted, for example, from the acquisition description in a DICOM file. In this case, one approach can consider a set of acquisition parameters for the group from which the sequence of the image data frames was acquired. These acquisition parameters can then be used to model the predicted noise distribution, which is then used in the generation of the prediction regarding the target frame. Alternatively, or in combination with such an approach, a model of the predicted noise distribution can be used in combination with the prediction regarding the target frame to further inform the loss function.

[0039]

[0045] The use of acquisition parameters makes the model more adaptable and, as a result, less dependent on the characteristics of the dataset used during training. In this way, the systems and methods described herein can be used to generate an adjustable CNN for removing noise from image data. In such an embodiment, the noise removal model can generate an estimated value of the noise variance based on the acquisition parameters. In this case, the noise estimation module can be trained with a CNN such that the acquisition parameters are taken into account in the noise removal process. By training the method using a sequence of frames obtained using different acquisition parameters, the CNN is trained to take the acquisition parameters into account during the noise removal process, thereby enabling the CNN to function over a range of acquisition parameters.

[0040]

[0046] Further, adjustable variables can be extracted from the model, thereby enabling the manipulation of noise estimation based on the acquisition parameters. During CNN-based noise removal, the above adjustable variables can be used to increase or decrease the amount of noise to be removed from the image.

[0041]

[0047] FIG. 1 is a schematic diagram of a system 100 according to an embodiment of the present disclosure. As shown, system 100 typically includes a processing device 110 and an imaging device 120.

[0042]

[0048] The processing device 110 can apply processing routines to images or measurement data (such as projection data) received from the imaging device 120. The processing device 110 may include a memory 113 and a processor circuit 111. The memory 113 can store a plurality of instructions. The processor circuit 111 is coupled to the memory 113 and configured to execute the instructions. The instructions stored in the memory 113 include processing routines, as well as data related to the processing routines such as machine learning algorithms and various filters for processing images.

[0043]

[0049] The processing device 110 may further include an input unit 115 and an output unit 117. The input unit 115 can receive information such as images or measurement data from the imaging device 120. The output unit 117 can output information such as the filtered image to the user or a user interface device. The output unit may include a monitor or a display.

[0044]

[0050] In some embodiments, the processing device 110 may be directly related to the imaging device 120. In an alternative embodiment, the processing device 110 is separate from the imaging device 120, and the processing device 11 receives the image or measurement data for processing via a network or other interface in the input unit 115.

[0045]

[0051] In some embodiments, the imaging device 120 may include an image data processing device and a spectral CT scan unit or a conventional CT scan unit that generates CT projection data when scanning a subject (e.g., a patient).

[0046]

[0052] FIG. 2 shows an exemplary imaging device 200 according to an embodiment of the present disclosure. A CT imaging device is shown, and the following description is generally in the context of CT images, but the same methods are applicable in the context of other imaging devices, and it will be understood that images to which these methods can be applied can be obtained in a variety of ways.

[0047]

[0053] In the imaging device according to an embodiment of the present disclosure, the CT scan unit may be configured to perform a multi-axis scan and / or a helical scan of the subject to generate CT projection data. Accordingly, a plurality of scans may be recorded as a series of scans (a series of scans), each including image data and recorded as an image data frame. In the imaging device according to an embodiment of the present disclosure, the CT scan unit may include an energy-resolved photon-counting image detector. The CT scan unit may include a radiation source that emits radiation across the subject when acquiring projection data.

[0048]

[0054] In the example shown in FIG. 2, a CT scan unit 200, for example a computed tomography (CT) scanner, includes a stationary gantry 202 and a rotating gantry 204 rotatably supported by the stationary gantry 202. The rotating gantry 204 can rotate about a vertical axis around an examination region 206 of a subject when acquiring projection data. The CT scan unit 200 may include a support part 207 configured to support a patient within the examination region 206 and to pass the patient through the examination region during an imaging process.

[0049]

[0055] The CT scan unit 200 can include a radiation source 208 such as an X-ray tube, and the radiation source is supported by the rotating gantry 204 and configured to rotate with the rotating gantry 204. The radiation source 208 may include an anode and a cathode. A source voltage applied between the anode and the cathode accelerates electrons from the cathode to the anode. The flow of electrons results in a current from the cathode to the anode and generates radiation that traverses the examination region 206.

[0050]

[0056] The CT scan unit 200 may include a detector 210. The detector 210 forms an arc at an angle on the opposite side of the examination region 206 with respect to the radiation source 208. The detector 210 may include a one-dimensional or two-dimensional pixel array such as a direct conversion detector pixel. The detector 210 is configured to detect radiation that traverses the examination region 206 and generate a signal indicating its energy.

[0051]

[0057] Generally, a CT scan unit acquires a series of projection frames (a series of projection frames) as the rotating gantry 204 rotates around a patient. Thus, depending on the amount of movement of the gantry between frames, the projection data of each acquired frame somewhat overlaps with adjacent frames and consists of imaging data of the same object, i.e., the patient, acquired at different angles. In the context of this application, a frame can include different types of data such as raw signal data, tomographic projection data, or image data at various stages of processing. All of these types of data include imaging data, and thus each frame in the series of frames includes some type of imaging data.

[0052]

[0058] Thus, in some embodiments, the imaging data within a frame can be processed to different extents before transmitting the data to the processing device 110. Thus, the CT scan unit 200 includes generators 211 and 213. The generator 211 generates tomographic projection data 209 based on signals from the detector 210. The generator 213 receives the tomographic projection data 209 and, in some embodiments, generates a series of raw image data frames 311 of the subject based on the tomographic projection data 209. In some embodiments, the tomographic projection data 209 can be supplied to the input portion 115 of the processing device 110, while in other embodiments, the series of raw image data frames 311 is supplied to the input portion of the processing device.

[0053]

[0059] As will be described in more detail below, in addition to a series of image data frames being provided to the input portion 115 of the processing device 110, the imaging device 120 can also supply data that defines acquisition parameters associated with the series of image frames. Such acquisition parameters can include the incident photon beam distribution spanning the position of the detector columns, as well as the tube current and / or voltage used. Such acquisition parameters are supplied in the context of the Digital Imaging and Communications in Medicine (DICOM) standard and can be transmitted, for example, together with the underlying series of image data frames. Thus, the acquisition parameters can be extracted from a DICOM file.

[0054]

[0060] FIG. 3 shows a pipeline 300 for training a model used to remove noise from an image according to the present disclosure. FIG. 4 shows an example 400 of using a model for removing noise from an image according to the present disclosure. As shown and described in further detail below with reference to FIG. 5, a method for training the model first takes in or is supplied with a series 310 of image data frames, which are then used to remove noise from any of these image data frames. In this case, a target image data frame 315 is directly supplied to a model training module 320, while another prediction module 330 predicts the content of the target image data frame.

[0055]

[0061] The prediction module 330 receives image data from at least one image data frame before the target image data frame 315 and from at least one image data frame after the target image data frame 315, and generates a prediction of the content of the target image data frame from these frames. In doing so, the prediction module 330 extracts the content of the image data frame, such as the anatomical feature structure of the subject in the image data, from the image data frame supplied by a feature structure extraction module 340. This feature structure extraction module 340 can be a separately and individually trained learning algorithm such as a CNN.

[0056]

[0062] The content extracted from the image data frame is then considered over a plurality of frames using a convolutional memory unit such as a convolutional long short-term memory unit (ConvLSTM) 350. The extracted feature structures are then combined along the time axis (at 360) and fused by a feature structure combiner module 370.

[0057]

[0063] The fused feature structure can be the anatomical feature structure of the subject in the image data and can be used to predict the content of the target image data frame already supplied to the model training module 320. In this case, the output of the prediction module 330 is the prediction of the target image data frame. For example, the prediction can be the estimated values of the mean and standard deviation of the denoised target frame. Next, the model training module 320 trains a machine learning algorithm such as a CNN using the noisy target image data frame 315 as the ground truth. In the illustrated embodiment, the CNN trained by the model training module 320 is the denoising module 330.

[0058]

[0064] In some embodiments as illustrated, a noise estimation module 380 can be provided to predict the noise distribution in the prediction. In this case, the noise estimation module 380, which will be described in more detail below, can provide a prediction in the form of a noise model based on the acquisition parameters 390 associated with the series of image data frames 310, and the acquisition parameters can take the form of DICOM files or metadata associated with the underlying image data. As will be described in more detail below, the noise estimation module 380 itself can be a machine learning algorithm such as a CNN, and the model training module 320 can also train the noise estimation module 380.

[0059]

[0065] In such an embodiment, the prediction provided by the noise estimation module 380 is supplied to the model training module 320 together with the prediction regarding the target image data frame, and more appropriately defines a loss function for use in the target image data frame 315 as the ground truth data.

[0060]

[0066] Once trained, the model 400 is deployed as shown in FIG. 4. Thus, optionally, the target image data frame 410 and at least one previous image data frame and one subsequent image data frame extracted from the series 415 of image data frames are supplied to the noise removal model 420, and a noise estimate is generated by the noise estimation module 430 based on the acquisition parameters 440, and then used by the prediction engine 470 to output a denoised image 450 derived from the target image data frame 410.

[0061]

[0067] As described in detail below, the adjustment variable 460 derived during training 300 can be supplied to the noise estimation module 430. Such adjustment variables are similar to the physical characteristics of image acquisition recorded in the acquisition parameters 440 and can be selectable by the user. Thus, the user can select a value for an adjustment variable different from the actual characteristics stored in the corresponding acquisition parameters 440 to scale the output of the noise estimation module. As another example, the adjustment variable 460 is a general value associated with the noise level in the model and can be scaled regardless of the corresponding physical characteristics.

[0062]

[0068] FIG. 5A is a flowchart showing a method of training a model for removing noise from an image according to the present disclosure.

[0063]

[0069] As illustrated and described in the context of FIG. 3, the method captures (500) the target image data frame to be considered. The target image data frame is one of the series of image data frames 310 containing the imaging data of the subject. The imaging data can be extracted from a CT scan system such as those described above with respect to FIG. 2. In such an embodiment, the image data frame is a projection frame and is part of a series of CT projection frames. Thus, each projection frame typically has imaging data of the same subject acquired from different angles.

[0064]

[0070] The method then obtains (510) at least one previous image data frame in the series of image data frames, prior to the target image data frame (obtained at 500), within the series. The content of the at least one previous image data frame at least partially overlaps the content of the target image data frame. The method further obtains (520) at least one subsequent image data frame in the series of image data frames, subsequent to the target image data frame, within the series. The content of the at least one subsequent image data frame at least partially overlaps the content of the target image data frame.

[0065]

[0071] In this way, the method sequentially obtains the image data frames before and after the target frame. The previous and subsequent frames are typically adjacent frames and, in addition to at least partially overlapping the target frame, also at least partially overlap each other. In this way, the previous and subsequent frames have at least some of the same content. When frames are described as overlapping, it is understood that the overlap relates to an image associated with or generated from the image data frame, or to the actual subject of the image data frame. For example, when an image is acquired in a linear process, the image associated with the image data frame may include some overlap and thus adjacent images with the same image content. However, when an image is acquired axially or helically, as in the case of CT projection frames, the image content does not overlap, but the subject of the images will overlap. This is because these images typically include the same subject acquired from different angles. When the angle between the acquired projections is small, it can be said that the content of the image data frames overlaps. For example, in some embodiments, it can be said that the image data frames overlap as long as the same feature structure side is visible, and in such embodiments, the image data frames will overlap as long as the angle between the projections is less than 90 degrees. In other embodiments, the degree of overlap can be measured by evaluating the difference between the images. This can be, for example, by means of a structural similarity index measure (SSIM) or mean squared error (MSE).

[0066]

[0072] In some embodiments, the method obtains a plurality of previous image data frames and obtains a plurality of image data frames preceding the target image data frame within the series of image data frames. Similarly, the method obtains a plurality of subsequent image data frames, thereby obtaining a plurality of image data frames subsequent to the target image data frame in the series of image data frames. Each of the frames thus obtained has at least some overlapping content, and the number of adjacent frames obtained may depend on how much the content of the frames overlaps and may depend on the computing capacity and memory capacity of the device implementing the method. Such values can also be adjusted by the user.

[0067]

[0073] The flowchart shows obtaining each identified frame individually, but it is understood that the method can instead obtain the series 310 of image data frames as a file from the imaging system or access a database containing such a series and then identify the target image data frame 315 within that data along with the preceding and subsequent image data frames.

[0068]

[0074] The method then obtains acquisition parameters (530) related to the acquisition of the image data frames in the series 310 of image data frames. As described above, such acquisition parameters may be provided as a DICOM file 390 related to or extracted from the series 310 of image data frames or may be included as metadata in the image data frames themselves.

[0069]

[0075] The method then generates a prediction (540) for the denoised target image data frame based on at least one previous image data frame (obtained at 510) and at least one subsequent image data frame (obtained at 520). Such a prediction may take the form of an estimate of the mean and standard deviation of the clean version of the target data frame 315.

[0070]

[0076] In some embodiments, the method generates the prediction of the noise-removed target image data frame (540) based on the representation of at least one anatomical feature structure extracted from each of at least one previous image data frame and at least one subsequent image data frame. The representation can be extracted from the previous and subsequent image data frames by a feature structure extraction module (feature structure extractor) 340 (550), and the representations extracted from each image data frame are fused by a feature structure combiner module 370 (560) to form a prediction of the noise-removed target image data frame.

[0071]

[0077] In some such embodiments, the representation of at least one anatomical feature structure is transferred between frames using a convolutional memory unit, which can be a convolutional long short-term memory unit 350 for transmitting information between frames of a series of image data frames.

[0072]

[0078] In some embodiments, multiple machine learning algorithms such as CNNs can be implemented in different steps of the method. For example, the feature structure extractor module 340 can implement a CNN that extracts feature structures such as anatomical feature structures from various image data frames (at 550). Similarly, the feature structure combiner module 370 can implement a CNN that fuses the representations of the feature structures (at 560) to form a prediction. Similarly, the prediction of the noise-removed target image data frame can be generated and output by a single trained convolutional neural network to which at least one previous image data frame and at least one subsequent image data frame are supplied. The prediction of the noise-removed target image data frame can take the form of the mean and standard deviation of the noise-removed version of the target frame.

[0073]

[0079] The self - supervised model training module 320 is supplied with a prediction regarding the denoised target image data frame, the target image data frame, and a noise model based on the acquisition parameters (captured at 530), and trains the prediction module 330 to denoise the target image data frame 315 based on the supplied data (570). The prediction module 330 can be a CNN. The training process can utilize the noisy target data frame 315 as ground truth and use the prediction (generated at 540) as the basis for a loss function.

[0074]

[0080] In some embodiments, the loss function used to train a machine learning algorithm is based on the distribution of noise predicted based on the acquisition parameters. In such embodiments, after the acquisition parameters are captured (at 530), a noise model including a prediction of the noise distribution is generated in the noise estimation module 380 based on these parameters (580). Such a prediction of the noise distribution can be utilized for generating the loss function used in training (at 570) and can be simultaneously trained or fine - tuned by the training process.

[0075]

[0081] Further details regarding the generation of the noise model by the noise estimation module will be described in more detail later with reference to FIG. 6.

[0076]

[0082] The training method described herein is typically repeated for a series of a large number of image data frames. Such a repeated series can include series acquired using different acquisition parameters. Thus, after training an algorithm using a series of image data frames 310, the method determines (580) whether an additional series is available for training. If available, the method captures a target image data frame from the additional series of image data frames 310 (500), captures the previous frame (510), and captures the subsequent frame (520).

[0077]

[0083] Each series 310 of the image data frame usually comprises an individual DICOM file 390 that defines the associated acquisition parameters. Thus, the method can take in such acquisition parameters for each iteration (530). In this way, the training of the algorithm (570) can take into account the acquisition parameters detailed in the DICOM file 390.

[0078]

[0084] As part of the generation of the noise prediction (at 580), the model generated is usually based on details extracted from the DICOM file as described later with reference to FIG. 5B. Thus, each detail extracted from the DICOM file effectively adjusts the model output as the noise prediction (at 580). During the training of the algorithm, the method can associate at least one of the details extracted from the DICOM file with a specified adjustment variable or scaling variable, which can be manually adjusted during model inference once the model has been trained.

[0079]

[0085] Thus, the adjustment variable can be extracted or generated based on the results associated with different acquisition parameters for the machine learning algorithm. As described above, the adjustment variable can be trained based on a specific variable identified in the DICOM file and thus associated with that specific variable. For example, the adjustment variable can be based on the variation of the tube current associated with the acquisition of the relevant series of the image or image data frame.

[0080]

[0086] The adjustment variable can be a scaling factor that determines to what extent the noise identified by the machine learning algorithm should be removed. Thus, the model output as the noise prediction (580) can generate a prediction of the form the noise can take, and the adjustment variable can determine to what extent the noise taking that form should be removed from the target image data frame.

[0081]

[0087] The expected form of the noise predicted by such an adjustment and noise estimation module 380 will be described in more detail later with reference to FIG. 6.

[0082]

[0088] FIG. 5B is a flowchart showing a method for removing noise from an image according to the present disclosure.

[0083]

[0089] When a machine learning algorithm, i.e., a CNN, is trained and no further additional sequences are supplied for training (at 580), the machine learning algorithm can be used for removing noise from an image.

[0084]

[0090] Accordingly, the method takes in a noisy target image data frame 410 to be denoised (590), and also takes in at least one previous image data frame and at least one subsequent image data frame together (595). The method separately takes in acquisition parameters related to the noisy image data frame (600). Such acquisition parameters can be extracted from a DICOM file 440 associated with or extracted from the image data frame 410.

[0085]

[0091] As described above, the adjustment variables can be extracted or generated during the training process described with respect to FIG. 5A. Accordingly, the adjustment variables can be applied to or associated with the acquisition parameters during model inference (610). The values of the adjustment variables can be selected and applied based on the acquisition parameters of the noisy image data frame (610). For example, the adjustment variables can be associated with and set according to details extracted from the DICOM file (610).

[0086]

[0092] When an adjustment variable is set (610), the method generates a noise prediction based thereon (620) and applies a trained machine learning noise removal algorithm, optionally, to a series of noisy image data frames including a target noisy image data frame 410 and at least one previous image data frame and at least one subsequent image data frame (630). The method then generates an estimate of the mean and standard deviation of the noise-removed target image data frame and operates on the selected value of the adjustment variable with respect to the noise prediction and the estimate combined with the target image data frame.

[0087]

[0093] The method then predicts a clean image data frame of the target image data frame and generates a first noise-removed image data frame (640).

[0088]

[0094] In some embodiments, the first noise-removed image data frame (generated at 640) is first considered before being presented to the user to determine whether the quality of the noise removal method is satisfactory (650). If it is satisfactory, the first noise-removed image data frame (generated at 640) can be output to the user. If it is not satisfactory, the examiner modifies the adjustment variable (610) to improve the quality of the first noise-removed image data frame. In some embodiments, the user himself considers the first noise-removed image data frame and determines whether the applied adjustment variable should be modified (at 610) and processes the image data frame again.

[0089]

[0095] As described above, the adjustment variable can be trained to correspond to the acquisition parameters in the training data. Thus, after the correction (at 610), the selected value of the adjustment variable may differ from the actual value of the corresponding acquisition parameter associated with the noisy image data frame. When the adjustment variable functions to apply a scaling factor, using the selected value can lead the machine learning algorithm to identify and remove more noise within the noisy image data frame than using the actual value of the corresponding acquisition parameter.

[0090]

[0096] For example, when the adjustment variable is proxy information for the acquired tube current, a low tube current typically corresponds to a high noise level in the image. Thus, when a large value is used for the adjustment variable (at 610), less noise will be removed from the image data frame and the resulting data will be more accurate. However, when a small value is used for the adjustment variable, noise removal will be more severe and the resulting data will be cleaner, but at the expense of some anatomical details that may be smoothed out. Thus, a user, such as a clinician, can determine whether a cleaner output or a more accurate output is preferred depending on the particular diagnostic task at hand.

[0091]

[0097] In some embodiments, the user can first apply the selected adjustment variable to the method (at 610) without first examining the output. In this case, this can remove more noise from the noisy image data frame during processing than the default value.

[0092]

[0098] If it is determined that the quality of the noise removal method is satisfactory (650), the method can reconstruct an image from the image data file (660). Such reconstruction may be necessary if the image data frame is, for example, a projection frame from a CT scan that is processed prior to reconstruction. In some embodiments, the reconstruction of the image may be based on a plurality of noise removal image data frames including the first noise removal image data frame. Further, the reconstruction may be an image of a volume or other configuration rather than a single image. The method then ends by outputting the reconstructed image or volume to the user (670).

[0093]

[0099] FIG. 6 shows the estimation of noise in an image in the noise estimation module 380 of FIG. 3.

[0094]

[0100] The following description relates to the noise of CT projections, but a similar noise estimation method can be implemented in other imaging modalities. The noise in CT projections can be approximated by parameters obtained from the description of the acquisition process using a mixed Poisson-Gaussian distribution. Such a distribution:

Equation

[0095]

[0101] Here, p is the normalized projection data without noise, and T is the transmission data without noise. T follows a Gaussian distribution:

Equation

Equation

[0096]

[0102] Electronic noise variance can be obtained from the measured value of the dark current, while the number of incident photons can be estimated from an air scan. However, such values are not usually readily available. Therefore, in the method described herein, the provided noise estimation module 380 eliminates the need to accurately know these parameters. Instead, these parameters can be estimated by an optimization algorithm while training the main noise removal model. Nevertheless, the theory is useful for performing pre-estimation. When the CT scanner uses bowtie filtering and automatic exposure control, the noise parameter λ corresponding to the incident photon flux should depend on the detector column position 700 and the tube current 710. In this case, the trainable noise estimation module 380 takes in the CT acquisition characteristics from the aforementioned DICOM file 390, etc. as inputs, and the noise variance:

Number

[0097]

[0103] Then, the tube current 710 can be used to extract the slope and bias coefficient, and the column position 700 can be used to extract the prediction of the photon number distribution. The extracted slope and bias coefficient can then be used to scale the predicted photon number distribution to derive a prediction of the number of incident photons λ. The method can separately estimate the electronic noise variance to derive a prediction of σ e 2 This can then be used to estimate the noise variance σ n 2

[0098]

[0104] The main noise removal model 320 and the noise estimation module 380 can be trained together, or the parameters of the noise estimation module can be pre-computed. The training is performed using the available noisy CT projections 310 and their acquisition parameters. If the test data has other noise characteristics (e.g., bowtie filtering is not used), the noise estimation module 380 can be adjusted according to such characteristics, while the main noise removal model 320 can remain unchanged. In such scenarios where the acquisition parameters are changed, both the noise removal model 320 and the noise estimation module 380, or the noise estimation module 380 alone, can be retrained directly with the new data.

[0099]

[0105] The methods described herein have been described in relation to CT scan images, but various imaging techniques including various medical imaging techniques are also envisioned, and it is understood that images generated using a variety of imaging techniques can be effectively denoised using the methods described herein. Since the described methods include individual noise models, changes to the noise models allow the methods to be used to process other medical image sequences. Accordingly, related modalities can include, for example, dynamic positron emission tomography (PET), dynamic magnetic resonance (MR), microscopic fluorescence images, ultrasound images, and others.

[0100]

[0106] The method according to the present disclosure can be implemented on a computer as a computer-implemented method, within dedicated hardware, or in a combination of both. The executable code for the method according to the present disclosure can be stored in a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Preferably, the computer program product may include non-transitory program code stored in a computer-readable medium for executing the method according to the present disclosure when the program product is executed on a computer. In one embodiment, the computer program may include computer program code adapted to execute all steps of the method according to the present disclosure when the computer program is executed on a computer. The computer program can be embodied on a computer-readable medium.

[0101]

[0107] Although the present disclosure has been described to some length and in some detail with respect to several described embodiments, it is not intended to be limited to any such detail or embodiment or to any combination thereof, and with respect to the appended claims, to provide the broadest possible interpretation of such claims in view of the prior art and, accordingly, to be construed as effectively encompassing the intended scope of the present disclosure.

[0102]

[0108] All examples and conditional statements recited herein are intended for educational purposes to assist the reader in understanding the principles of the present disclosure and the concepts contributed by the inventors to the development of the art, and are to be construed as not being limited to such specifically recited examples and conditions. Further, all descriptions herein that refer to the principles, aspects, embodiments, and specific examples disclosed herein are intended to encompass both their structural and functional equivalents. Further, such equivalents are intended to include not only currently known equivalents but also equivalents developed in the future, i.e., any elements developed to perform the same function regardless of structure.

Claims

1. A method for training a machine learning model for noise removal, comprising: capturing a target image data frame that is one of the image data frames in a series of image data frames including imaging data of a subject; capturing at least one previous image data frame in the series of image data frames that is before the target image data frame in the series, wherein the content of the at least one previous image data frame at least partially overlaps with the content of the target image data frame; capturing at least one subsequent image data frame in the series of image data frames that is after the target image data frame in the series, wherein the content of the at least one subsequent image data frame at least partially overlaps with the content of the target image data frame; capturing acquisition parameters related to the acquisition of the image data frames in the series of image data frames; generating a prediction of the noise-removed target image data frame based on the at least one previous image data frame and the at least one subsequent image data frame; training a machine learning algorithm for removing noise from the target image data frame based on the prediction of the noise-removed target image data frame and a noise model based on the acquisition parameters. A method comprising the above steps.

2. The prediction of the noise-removed target image data frame is an estimation of the mean and standard deviation of the noise-removed target image data frame based on the representation of at least one anatomical feature structure extracted from each of the at least one previous image data frame and the at least one subsequent image data frame, and the representations from the at least one previous image data frame and the at least one subsequent image data frame are fused to form the prediction of the noise-removed target image data frame. The method according to Claim 1.

3. The representation of the at least one anatomical feature structure is transmitted between frames using a convolutional memory unit. The method according to Claim 2.

4. The method according to claim 3, wherein the convolutional memory unit is a convolutional long short-term memory unit that transmits information between frames in the series of the image data frames.

5. The method according to claim 2, wherein the at least one previous image data frame is a plurality of image data frames before the target image data frame in the series of the image data frames, and the at least one subsequent image data frame is a plurality of image data frames after the target image data frame in the series of the image data frames.

6. The method according to claim 2, wherein the prediction of the noise-removed target image data frame is output by a trained convolutional neural network to which the at least one previous image data frame and the at least one subsequent image data frame are supplied.

7. The method according to claim 1, wherein a loss function for training the machine learning algorithm is based on a distribution of predicted noise based on the acquisition parameters.

8. The method for training a machine learning model for the noise removal is repeated for a series of projection frames obtained using different acquisition parameters, and an adjustment variable is extracted or generated based on results related to different acquisition parameters for the machine learning algorithm.

9. The method according to claim 8, wherein the adjustment variable is trained based on a variance of tube current related to acquisition of a related series of image data frames.

10. The method according to claim 8, wherein the adjustment variable is a scaling factor that determines how much noise identified by the machine learning algorithm should be removed.

11. The method according to claim 7, wherein the distribution of the predicted noise is based on a Poisson-Gaussian distribution.

12. The method according to claim 1, wherein the acquisition parameters are extracted from a DICOM file related to the series of the projection frames.

13. The method according to claim 1, wherein the machine learning algorithm is a convolutional neural network.

14. The imaging data is CT imaging data, each image data frame in the series of the image data frames is a projection frame, and each projection frame includes imaging data of the same subject acquired from different angles. The method according to claim 1.

15. Executing the method according to claim 8; Capturing a series of noisy image data frames including a noisy target image data frame to be denoised; Selecting a value of the adjustment variable based on acquisition parameters of the series of the noisy image data frames; Applying the trained machine learning algorithm to the series of the noisy image data frames using the selected value of the adjustment variable; Generating a first denoised image data frame based on an estimation of an average and a standard deviation based on the series of the noisy image data frames, a distribution of noise based on the acquisition parameters of the series of the noisy image data frames, and the noisy target image data frame. A noise removal method having the above.

16. The adjustment variable is trained to correspond to acquisition parameters in training data, the selected value of the adjustment variable is different from an actual value of the corresponding acquisition parameter of the noisy image data frame, and the machine learning algorithm identifies more noise in the noisy image data frame when using the selected value than when using the actual value. The noise removal method according to claim 15.

17. A machine learning training system having a memory storing a plurality of instructions and a processor circuit coupled to the memory, wherein the processor circuit executes the instructions to Capture a plurality of image data frames having a series of image data frames including imaging data of a subject; Identify a target image data frame in the series of the image data frames; Generate a prediction of the denoised target image data frame based on at least one previous image data frame of the series of image data frames before the target image data frame in the series and at least one subsequent image data frame after the target image data frame in the series, where each of the at least one previous image data frame and the at least one subsequent image data frame at least partially overlaps with the target image data frame. Capture acquisition parameters related to the acquisition of the image data frames in the series of image data frames, and Train a machine learning algorithm for denoising the target image data frame based on the prediction of the denoised target image data frame and a noise model based on the acquisition parameters. Machine learning training system.

18. The machine learning training system according to claim 17, wherein the loss function for training the machine learning algorithm is based on the distribution of noise predicted based on the acquisition parameters.

19. The machine learning training system according to claim 18, wherein the training is repeated for a series of image data frames acquired using different acquisition parameters, and the adjustment variable is extracted or generated based on the results related to different acquisition parameters for the machine learning algorithm.

20. The machine learning training system according to claim 18, wherein the imaging data is CT imaging data, each image data frame of the series of image data frames is a projection frame, and each projection frame includes imaging data of the same subject acquired from different angles.

Citation Information

Patent Citations

  • Noise reduction method in digital x ray frame series

    JP2013127773A

  • Medical image processing apparatus and medical image processing system

    JP2019069145A

  • Image processing device, image processing method, and program

    JP2020103880A

  • Image processing device, image processing method, and x-ray ct device

    JP2021019714A

  • X-ray control method, x-ray imaging device, and non-temporary computer readable medium

    JP2023133250A