Method, medical image processor and program

JP2023108605A5Pending Publication Date: 2026-01-09UNIVERSTIY OF CALIFORNIA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023004927
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-06
Filing Date
2023-01-17
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing methods for image denoising in medical imaging, such as Deep Image Prior (DIP), face challenges with overfitting and inefficiencies in removing noise from images like PET images, particularly when using anatomical prior information.

Method used

A method involving a double-overparameterized (DOP) neural network trained with mismatch learning rates, utilizing CT or MRI images as anatomical priors, to generate denoised PET images by combining convergence noise with the network's output, effectively addressing overfitting and improving noise removal.

Benefits of technology

The DOP method achieves significant noise reduction in PET images without overfitting, preserving image details and offering a better contrast-to-noise trade-off compared to traditional Gaussian filters, enabling accurate and efficient denoising.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To accurately remove noise from an image.SOLUTION: A method according to an embodiment for training a deep image prior (DIP) neural network to remove image noise includes receiving a first medical image comprising a first image of an anatomical structure, receiving a second medical image comprising a second image of the anatomical structure, and training the DIP neural network to produce a denoised image in which noise is removed from an input image so that convergent noise combined with output of the DIP neural network at the end of the training approximates the first medical image by inputting the second medical image into the DIP neural network during training and combining the convergent noise with the output of the DIP neural network. The output of the DIP neural network represents the denoised image.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed in this specification and the drawings relate to a method, a medical image processing apparatus, and a program.

Background Art

[0002] Deep Image Prior (DIP) is a teacherless method for image restoration and has been successfully applied to image noise removal of positron emission tomography (PET). However, the DIP-based method relies on early stopping of the training process to avoid overlearning on noisy images. Recently, You et al. proposed a double over-parameterized (DOP) approach in their paper "Robust recovery via implicit bias of discrepant learning rates for double over-parameterization" available at <https: / / arxiv.org / abs / 2006.08857> on the arxiv.org website. The DOP method by You et al. can be extended for use in PET image restoration by utilizing CT images from the same patient as anatomical prior information.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

[0004] One of the problems that the embodiments disclosed herein and in the drawings aim to solve is to remove noise from images with high accuracy. However, the problems that the embodiments disclosed herein and in the drawings aim to solve are not limited to the above problem. Problems corresponding to the effects of each configuration shown in the embodiments described later can also be positioned as other problems. [Means for solving the problem]

[0005] The method of the embodiment is a method for training a Deep Image Prior (DIP) neural network for denoising an image. The method includes receiving a first medical image including a first image of an anatomical structure, receiving a second medical image including a second image of the anatomical structure, and training the DIP neural network to generate a denoised image from the input image by inputting the second medical image to the DIP neural network during training and combining convergence noise with the output of the DIP neural network, such that at the end of training, the convergence noise combined with the output of the DIP neural network approximates the first medical image. The output of the DIP neural network represents the denoised image. [Brief explanation of the drawing]

[0006] [Figure 1A] Figure 1A is a flowchart illustrating a method for performing image noise reduction in a medical imaging environment. [Figure 1B] Figure 1B is an example of a noisy PET image that has been denoised through the denoising process shown in Figure 1A. [Figure 1C] Figure 1C is an exemplary input image that acts as anatomical prior information. [Figure 2A]Figure 2A is the original PET image generated using OSEM with PSF. [Figure 2B] Figure 2B is a denoised version of the original PET image in Figure 2A, after denoising using a Gaussian filter and DIP. [Figure 2C] Figure 2C is a denoised version of the original PET image in Figure 2A, after denoising using a Gaussian filter and DIP. [Figure 2D] Figure 2D is a PET image generated by denoising the original PET image in Figure 2A using the techniques described herein. [Figure 3A] Figure 3A is the original PET image generated using TOF OSEM with PSF, including a simulated lesion inserted for quantitative analysis. [Figure 3B] Figure 3B is a denoised version of the original PET image in Figure 3A, after denoising using a Gaussian filter and DIP, where simulated lesions are visible in each. [Figure 3C] Figure 3C is a denoised version of the original PET image in Figure 3A, after denoising using a Gaussian filter and DIP, where simulated lesions are visible in each. [Figure 3D] Figure 3D is one of a series of PET images generated by denoising the original PET image of Figure 3A using the techniques described herein, where a simulated lesion is visible. [Figure 3E] Figure 3E is one of a series of PET images generated by denoising the original PET image of Figure 3A using the techniques described herein, where a simulated lesion is visible. [Figure 3F] Figure 3F is one of a series of PET images generated by denoising the original PET image of Figure 3A using the techniques described herein, where a simulated lesion is visible. [Figure 3G] Figure 3G is one of a series of PET images generated by denoising the original PET image of Figure 3A using the techniques described herein, where a simulated lesion is visible. [Figure 3H] Figure 3H is one of a series of PET images generated by denoising the original PET image of Figure 3A using the technique described herein, where a simulated lesion is visible. [Figure 3I] Figure 3I is one of a series of PET images generated by denoising the original PET image of Figure 3A using the technique described herein, where a simulated lesion is visible. [Figure 4A] Figure 4A is a flowchart illustrating an alternative configuration using the technology described herein. [Figure 4B] Figure 4B is an explanatory diagram of the contents of the g and h vectors before training. [Figure 4C] Figure 4C is an explanatory diagram of the contents of the g and h vectors after training. [Figure 5] Figure 5 is a perspective view of a positron emission tomography (PET) scanner according to one aspect of this disclosure. [Figure 6] Figure 6 is a schematic diagram of the PET scanner shown in Figure 5, according to one aspect of this disclosure. [Figure 7] Figure 7 is a schematic diagram of a computed tomography (CT) imaging system according to one aspect of the present disclosure. [Figure 8] Figure 8 is a graph showing the training loss curves of the method described herein for various values ​​of α, compared with the DIP method. [Figure 9] Figure 9 shows a comparison of DOP and DIP (plotted every 100 epochs), and also compares TOF OSEM with PSF reconstruction using a Gaussian post-filter and without (plotted per iteration), showing the contrast vs. noise curves of the inserted lesions marked in Figure 3A. [Modes for carrying out the invention]

[0007] When considered in connection with the accompanying drawings, a thorough understanding of the present disclosure and many of its attendant advantages will be readily obtained, as the same will become better understood by reference to the following detailed description.

[0008] As used herein, the term "plurality" is defined as two or more. Also, as used herein, references to "one embodiment", "a particular embodiment", "an embodiment", "an implementation", "an example", or the like mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, appearances of such phrases or various places throughout this specification are not necessarily all referring to the same embodiment. Further, the particular features, structures, or characteristics may be combined in any suitable manner without limitation in one or more embodiments.

[0009] This disclosure describes a method for denoising medical images using a neural network trained during a denoising process. In one embodiment, computed tomography (CT) images or magnetic resonance imaging (MRI) images from the same patient (subject) acquired by a positron emission tomography (PET) scanner (PET device) or single-photon emission computed tomography (SPECT) scanner (SPECT device) can be used as anatomical prior information. As shown in Figures 1A and 1C, CT images can be used as input to the neural network to be trained, but MRI images can be used similarly. As shown in Figures 1A and 1B, the denoised image is a noisy PET image of the same region previously acquired from the patient using a CT scanner (see, for example, Figure 7), but SPECT images can be used similarly. In this specification, various noisy images refer to various images containing noise.

[0010]

number

[0011] In this specification, a clean image refers to, for example, an image with little noise or no noise. In Figure 1A, a U-shaped encoder-decoder network is illustrated with skip connections to represent a clean image. The number of characteristic channels is listed below each layer. The number of trainable parameters for DOP is 9,161,401 (θ: 3,155,641; g,h: 6,005,760). The DIP network uses the same network structure as θ. Each layer of the network includes 3D convolution, ReLU activation, and batch normalization. The noise layer is modeled by two g and h vectors of the same size as the training label image (illustrated as a noisy PET image, sometimes called the target image) and added to the decoder output. The resulting denoising effect was evaluated using a patient dataset acquired with Canon Inc.'s "Cartesion TOF (Time of Flight) PET / CT scanner". 18 The FDG radioactivity was 100 MBq, and PET / CT imaging was initiated 60 minutes after injection. Radiation scans were performed on 6 patient positions for 1.5 minutes each. PET images were reconstructed using OSEM (ordered subset expectation maximization method) with 3 replicates and 12 subsets (matrix size 337 × 337 × 129, voxel size 2.11 × 2.11 × 2.11 mm). 3 ). A noisy PET image was cropped to 136 × 184 × 120 pixels and used as a training label image, while a aligned CT image of the same size was used as the network input. (As a result, the g and h vectors were also 136 × 184 × 120 pixels, respectively). The input image does not need to be the same size as the training label image, and especially when random noise is used as input, the input image may be downsampled or upsampled to match the size of the training label image.

[0012] Figure 1A illustrates an architecture using the ADAM optimizer, but other optimizers (e.g., stochastic gradient descent optimizer) can be used instead. Note that an optimizer is also referred to as an optimization algorithm. The initial learning rate is 5e. -4 The selected vectors g and h were also initially small random values ​​(e.g., 5e with a Gaussian distribution). -4 It is filled with values ​​of the order of . By initializing g and h with small random values, the difference of the convolution of g and h (i.e., g°gh°h) converges to the noise, and even if the noise is not sparse, at least in the context of the PET image.

[0013] Different learning rates were used for θ and g,h, and their ratio was controlled by the parameter α. The network was trained for 2000 epochs for each bed position. The effect of the mismatched learning rate ratio α on the quality of image reconstruction is shown in Figures 3D to 3I and can be compared with that from the DIP method and conventional Gaussian post-filters, as shown in Figures 3B and 3C. In general, the method in Figure 2D showed a significant reduction in noise compared to the original noisy PET image in Figure 2A and the Gaussian post-filtered image in Figure 2B. Although the learning curves were close to zero for both DOP and DIP, the denoised image in Figure 2D did not show the problem of overfitting compared to the DIP method in Figure 2C. The mismatched learning rate ratio α controlled the convergence point. As the value of α increases, the resulting image becomes smoother. As shown in Figures 3E and 3F, α=3 or 4 produces good image denoising while preserving most of the details. Quantitatively, the DOP method yielded a better contrast-to-background noise trade-off than the DIP method (Figure 3C) and the Gaussian filter method (Figure 3B). Furthermore, the DOP method using a mismatched learning rate ratio can improve PET image quality without adjusting the network width or premature termination. The mismatched learning rate ratio α can be adjusted to control image smoothness based on domain-specific applications.

[0014] To address domain-specific imaging problems, two or more neural networks can be trained (continuously or in parallel) using different parameters α, and each of the resulting images can be displayed to the medical professional so that they can select an image with an appropriate level of noise and smoothness. The system can track the imaging conditions of an initial imaging study and the parameters α previously used by the medical professional, and initially perform imaging denoising using the same parameters α that were used in a number of previous image denoising processes (e.g., by tracking the most frequently used parameters α in general or specific to the imaging protocol). Subsequent denoising can be performed with other close values ​​of parameter α. For example, if the medical professional most frequently uses parameter α=3 for imaging process X1, denoising for subsequent imaging process X1 will use parameter α, and then denoising will be performed using parameters α=2.5 and α=3.5. The resulting denoised images can be displayed to the medical professional as soon as they become available. Alternatively, the system can use parallel processing techniques to perform n denoising steps in parallel (e.g., parameter α=3, parameter α=3.5, and parameter α=2.5). Other denoising steps that are further removed from the most commonly used ones can be performed later. In X2, a different imaging protocol, the most commonly used parameter α is 4, and the system can start with that parameter value instead when processing X2-style images.

[0015] As described herein, a convolution-based function of vectors g and h is used to represent the noise of the target image, which is modeled separately, so that the neural network is not trained to generate noisy images. In some embodiments, the image noise should have a distribution such as Poisson or Gaussian. In alternative embodiments, at least one of the vectors g and h is trained to constrain the learned noise pattern during training to follow a Gaussian or Poisson distribution.

[0016] The above describes noise reduction of PET images in Figures 1A-1C, 2D, and 3D-3I, but this technique can be applied to other types of images as well. For example, when denoising gated cardiac CT images, ungated cardiac CT images can be used as anatomical prior information.

[0017] Furthermore, the described technique can be used even when low-noise images are not available to be used as anatomical prior information. It is also possible to use random noise as the original input image, as shown in Figure 4A. Figures 4B and 4C show the contents of the g and h vectors before and after training in both Figure 1A and Figure 4A.

[0018] In yet another alternative embodiment, the noise in the denoising image can be modeled using two or more noise vectors (e.g., g and h) so that the system can learn two or more types of noise. For example, three noise vectors can be used.

[0019] Figure 8 shows the training loss curves of the method described herein for various α values ​​and compares them with the DIP method. Figure 9 is a graph of the contrast vs. noise curves of the inserted lesions marked in Figure 3A, comparing DOP and DIP (plotted every 100 epochs), and also comparing TOF OSEM with PSF reconstruction using a Gaussian post-filter and without (plotted per iteration).

[0020] Here, the device for training the neural network (DIP neural network) described above may be a medical diagnostic device (medical image processing device) such as the PET scanner 800 or CT scanner described later, or it may be an external server (training device) other than such a medical diagnostic device. That is, the training device or medical image processing device performs a method (process) for training the DIP neural network. For example, a method for training the DIP neural network includes receiving a first medical image (e.g., a PET image or SPECT image) containing a first image of an anatomical structure, and receiving a second medical image (e.g., a CT image or MRI image) containing a second image of an anatomical structure. The method for training the DIP neural network further includes inputting the second medical image to the DIP neural network during training and training the DIP neural network to generate an image from which noise has been removed from the input image, by combining the convergence noise with the output of the DIP network, so that at the end of training, the convergence noise combined with the output of the DIP network approximates the first medical image. Here, the output of the DIP network represents the image with the noise removed. A trained DIP neural network can remove noise from an image with high accuracy. Therefore, by using such a trained DIP neural network to remove noise from an image, high-precision noise removal can be achieved.

[0021] Furthermore, in the method for training a DIP neural network, training the aforementioned DIP neural network involves performing a double-over parameterization training process against convergence noise.

[0022] Furthermore, in a method for training a DIP neural network, training the DIP neural network as described above includes: initializing a first noise vector (e.g., a g vector) and a second noise vector (e.g., a h vector); training the DIP neural network to generate a denoised image by training the first noise vector and the second noise vector such that a convolution-based function based on the first noise vector and the second noise vector converges to a value equal to the noise of the first medical image; and training the DIP neural network to approximate the first medical image obtained by subtracting the convolution-based function.

[0023] Here, if the first medical image is a PET image of the subject, the second medical image may be a CT image of the subject aligned to the PET image, or an MRI image of the subject aligned to the PET image.

[0024] Furthermore, if the first medical image is a SPECT image of the subject, the second medical image may be a CT image of the subject aligned to the SPECT image, or an MRI image of the subject aligned to the SPECT image.

[0025] Furthermore, if the first medical image is an ungated cardiac computed tomography (CT) image of the subject, the second medical image may be a gated cardiac CT image of the subject aligned to the ungated cardiac CT image.

[0026] Furthermore, the device that removes noise from an image using a trained DIP neural network may be a medical diagnostic device (medical image processing device) such as the PET scanner 800 or the CT scanner described later. For example, the medical image processing device inputs an image to the trained DIP neural network. The DIP neural network then generates and outputs an image from which the noise has been removed from the input image. The medical image processing device acquires the noise-removed image output from the DIP neural network and displays the acquired image on a display.

[0027] Figures 5 and 6 show a PET scanner 800 including a number of GRDs (e.g., GRD1, GRD2, ~GRDN), each configured as a rectangular detector module. According to one embodiment, the detector ring includes 40 GRDs. In another embodiment, there are 48 GRDs, and a larger number of GRDs are used to produce a larger bore size for the PET scanner 800.

[0028] Each GRD may include a two-dimensional array of individual detector crystals that absorb gamma rays and emit scintillation photons. These scintillation photons can be detected by a two-dimensional array of photomultiplier tubes (PMTs), also located within the GRD. Optical conductors can be placed between the array of detector crystals and the PMTs. Furthermore, each GRD may contain numerous PMTs of varying sizes, each arranged to receive scintillation photons from multiple detector crystals. Each PMT can generate an analog signal indicating when a scintillation event occurs, and the energy of the gamma ray generating the detection event. Additionally, a photon emitted from one detector crystal can be detected by two or more PMTs, and based on the analog signals generated by each PMT, the detector crystal corresponding to the detection event can be determined, for example, using Anger logic and crystal decoding.

[0029] Figure 6 shows a schematic diagram of a PET scanner system having a Gamma-Ray Photon Counting Detector (GRD) positioned to detect gamma rays emitted from the subject's obj. The GRD can measure the timing, position, and energy corresponding to the detection of each gamma ray. In one embodiment, the gamma-ray detector is arranged in a ring shape, as shown in Figures 5 and 6. The detector crystal may be a scintillator crystal having individual scintillator elements arranged in a two-dimensional array. The scintillator elements may be any known scintillator material. The PMTs may be arranged so that light from each scintillator element is detected by multiple PMTs, enabling Anger arithmetic and crystal decoding of the scintillation event.

[0030] Figure 6 shows an example of the PET scanner 800 configuration, where the subject OBJ to be imaged is placed on the bed 816, and the GRD modules GRD1 through GRDN are arranged circumferentially around the subject OBJ and the bed 816. The GRDs can be fixedly connected to annular components 820 which are fixedly connected to the gantry 840. The gantry 840 houses many components of the PET imaging apparatus. The gantry 840 of the PET imaging apparatus also includes an opening through which the subject OBJ and the bed 816 can pass, allowing the GRDs to detect gamma rays emitted in the opposite direction from the subject OBJ due to annihilation events, and to determine the coincidence of gamma ray pairs using timing and energy information.

[0031] Figure 6 also shows the circuitry and hardware for acquiring, storing, processing, and distributing gamma-ray detection data. This circuitry and hardware includes a processor 870, a network controller 874, a memory 878, and a Data Acquisition System (DAS) 876. The PET imaging apparatus also includes data channels for routing detection measurement results from the GRD to the DAS 876, processor 870, memory 878, and network controller 874. The Data Acquisition System 876 can control the acquisition, digitization, and routing of detection data from the detector. In one embodiment, the DAS 876 controls the movement of the patient bed 816. The processor 870 performs functions including reconstructing images from the detection data, pre-reconstruction processing of the detection data, and post-reconstruction processing of the image data, as described herein, according to the methods described herein.

[0032] The processor 870 may be configured to perform the methods described herein. For example, a trained DIP neural network is stored in memory 878, and the processor 870 retrieves the trained DIP neural network from memory 878 and uses the retrieved trained DIP neural network to remove noise from an image. The processor 870 is an example of a processing circuit. Memory 878 is an example of a storage unit. The processor 870 may include a CPU that can be implemented as individual logic gates, such as an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other Complex Programmable Logic Device (CPLD). The FPGA or CPLD implementation may be coded in VHDL, Verilog, or other hardware description language, and the code may be stored directly in the electronic memory within the FPGA or CPLD, or in separate electronic memory. Furthermore, the memory may be non-volatile, such as ROM, EPROM, EEPROM, or flash memory. The memory may also be volatile, such as static or dynamic RAM, and a processor such as a microcontroller or microprocessor may be provided for managing the electronic memory as well as for the interaction between the FPGA or CPLD and the memory.

[0033] Alternatively, the CPU within the 870 processor may execute a computer program comprising a set of computer-readable instructions that perform the methods described herein, the program being stored in any of the aforementioned non-temporary electronic memory and / or a hard disk drive, CD, DVD, flash drive, or any other known storage medium. Furthermore, the computer-readable instructions may be provided in the form of utility applications, background daemons, or components of an operating system, or a combination thereof, and may be executed in cooperation with processors such as Intel's Celeron, Xenon, i3, i5, i7, or i9, or AMD's Ryzen or Opteron processors, and with Microsoft Vista, UNIX®, Solaris, LINUX®, Apple, MAC-OS, and other operating systems known to those skilled in the art. Furthermore, the CPU may be implemented as multiple processors that operate concurrently and cooperatively to execute instructions.

[0034] In one embodiment, the reconstructed image can be displayed on a display. The display may be an LCD display, a CRT display, a plasma display, an OLED, an LED, or any other display known in the art.

[0035] Memory 878 may be a hard disk drive, CD-ROM drive, DVD drive, flash drive, RAM, ROM, or other electronic storage device known in the art.

[0036] Network controller 874, such as an Intel Ethernet® PRO network interface card manufactured by Intel Corporation in the United States, can interface between various parts of the PET imaging system. Furthermore, network controller 874 can also interface with an external network.

[0037] To make it clear, this external network could be a public network such as the Internet, or a private network such as a LAN or WAN network, or any combination thereof, and may also include a PSTN or ISDN subnetwork. The external network could also be wired, such as an Ethernet® network, or wireless, such as a cellular network including EDGE, 3G, and 4G wireless cellular systems. The wireless network could also be WiFi, Bluetooth®, or any other known wireless communication format.

[0038] According to one embodiment of the present disclosure, the method described above can be implemented to be applied to data from a CT apparatus or scanner. Figure 7 shows an implementation of a radiation gantry included in a CT apparatus or CT scanner. As shown in Figure 7, the radiation gantry 750 is shown in a side view and further includes an X-ray tube 751, an annular frame 752, and a multi-row or two-dimensional array type X-ray detector 753. The X-ray tube 751 and the X-ray detector 753 are mounted on the annular frame 752 opposite each other with respect to the subject OBJ, and the annular frame 752 is supported so as to be rotatable around a rotation axis RA. A rotation unit 757 rotates the annular frame 752 at a high speed, such as 0.4 seconds / revolution, while the subject OBJ is moved along axis RA in the direction towards the back of the page or the front of the page as shown.

[0039] Embodiments of the X-ray CT apparatus of this disclosure are described below with reference to the accompanying drawings. It should be noted that X-ray CT apparatuses include various types of devices, such as rotary / rotating devices in which the X-ray tube and X-ray detector rotate together around the subject being examined, and fixed / rotating devices in which many detector elements are arranged in annular or planar manner, and only the X-ray tube rotates around the subject being examined. This disclosure is applicable to any of these types. Here, the currently dominant rotary / rotating type is given as an example.

[0040] The multislice X-ray CT apparatus further includes a high-voltage generator 759, which generates a tube voltage applied to the X-ray tube 751 through a slip ring 758 so that the X-ray tube 751 generates X-rays. The X-rays are irradiated toward the subject's obj, and the cross-sectional area of ​​the subject's obj is represented by a circle. For example, the X-ray tube 751 has an average X-ray energy during the first scan that is lower than the average X-ray energy during the second scan. In this way, two or more scans can be obtained corresponding to different X-ray energies. An X-ray detector 753 is located on the opposite side of the subject's obj from the X-ray tube 751 to detect the irradiated X-rays that have propagated through the subject's obj. The X-ray detector 753 further includes individual detector elements or units and may be a photon count detector. In the fourth-generation geometry system, the X-ray detector 753 may be one of several detectors arranged 360° around the subject's obj.

[0041] The CT scanner further includes other devices for processing the signals detected from the X-ray detector 753. The data acquisition circuit or data acquisition system (DAS) 754 converts the output signals from the X-ray detector 753 for each channel into voltage signals, amplifies those signals, and further converts those signals into digital signals. The X-ray detector 753 and DAS 754 are configured to process a predetermined total number of projections per rotation (TPPR).

[0042] The data described above is transmitted via a non-contact data transmitter 755 to a preprocessing device 756 housed in a console outside the radiation gantry 750. The preprocessing device 756 performs specific corrections, such as sensitivity corrections, on the raw data. A storage device 762 stores the resulting data, also called projection data, immediately before the reconstruction process. The storage device 762, along with the reconstruction device 764, input device 765, and display 766, is connected to the system controller 760 via a data / control bus 761. The system controller 760 controls a current regulator 763 that limits the current to a level sufficient to drive the CT system. In one embodiment, the system controller 760 implements optimized scan acquisition parameters.

[0043] In various generations of CT scanner systems, the detectors are rotated and / or fixed relative to the patient. In one embodiment, the CT system described above may be an example of a combination of a third-generation geometry system and a fourth-generation geometry system. In the third-generation system, the X-ray tube 751 and the X-ray detector 753 are mounted opposite each other on an annular frame 752, and rotate around the subject OBJ as the annular frame 752 rotates around the rotation axis RA. In the fourth-generation geometry system, the detectors are fixed around the patient, and the X-ray tubes rotate around the patient. In an alternative embodiment, the radiation gantry 750 has a number of detectors arranged on an annular frame 752 supported by a C-arm and a stand.

[0044] The storage device 762 can store measured values ​​indicating X-ray irradiance from the X-ray detector 753. Furthermore, the storage device 762 can store dedicated programs for performing CT image reconstruction, material decomposition, and PQR estimation methods, including the methods described herein.

[0045] The reconstruction device 764 can perform the methods described herein. The reconstruction device 764 may perform reconstruction according to one or more optimized image reconstruction parameters. Furthermore, the reconstruction device 764 may perform pre-reconstruction image processing, such as volume rendering and image subtraction, as needed.

[0046] The pre-reconstruction processing of projection data performed by the pre-processing device 756 may include, for example, detector calibration, correction for detector nonlinearity, and polarity effects.

[0047] The post-reconstruction process performed by the reconstruction device 764 may include, as necessary, filter generation and image smoothing, volume rendering, and image subtraction. The image reconstruction process may implement the optimal image reconstruction parameters derived above. The image reconstruction process can be performed using filtered back projection, iterative reconstruction, or probabilistic reconstruction.

[0048] The reconstruction device 764 can use memory to store, for example, projection data, forward projection training data, training images, unedited images, calibration data and parameters, as well as computer programs. The reconstruction device 764 may also include machine learning processing support, which includes calculating a reference dataset based on the obtained spatial distribution in the soft tissue region and generating filters by performing all or part of a machine learning process using the projection dataset as input data and the reference dataset as training data. Furthermore, it may generate one or more evaluation values ​​that represent image quality by applying machine learning, which may include the application of artificial neural networks.

[0049] The reconfiguration device 764 and denoising device described may be implemented individually in a single processor, or in a network or cloud of processors. The reconfiguration device 764 and denoising device may include a CPU (processing circuit) that can run as discrete logic gates, as an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other Complex Programmable Logic Device (CPLD). The FPGA or CPLD implementation may be coded in VDHL, Verilog, or other hardware description language, and the code may be stored directly in the electronic memory within the FPGA or CPLD, or in separate electronic memory. Furthermore, the storage device 62 may be non-volatile, such as ROM, EPROM, EEPROM, or flash memory. The storage device 762 may be volatile, such as static or dynamic RAM, and a processor, such as a microcontroller or microprocessor, may be provided to manage the electronic memory as well as the interaction between the FPGA or CPLD and the memory. In one embodiment, the reconstruction device 764 may include a CPU and a Graphics Processing Unit (GPU) for processing and generating the reconstructed image. The GPU may be a dedicated graphics card or an integrated graphics card that shares resources with the CPU, and may be one of various types of GPUs specialized for artificial intelligence, including NVIDIA Tesla and AMD FireStream.

[0050] For example, a trained DIP neural network is stored in memory device 762, and the processor retrieves the trained DIP neural network from memory device 762 and uses the retrieved trained DIP neural network to remove noise from the image. The processor is an example of a processing circuit. Memory device 762 is an example of a storage unit.

[0051] Alternatively, the CPU within the reconfigurable device 764 may execute a computer program comprising a set of computer-readable instructions that perform the functions described herein, the program being stored in any of the aforementioned non-temporary electronic memory and / or a hard disk drive, CD, DVD, flash drive, or any other known storage medium. Furthermore, the computer-readable instructions may be provided as a utility application, a background daemon, or an operating system component, or a combination thereof, and may be executed in cooperation with processors such as Intel® XEON® processors or AMD® OPTERON® processors, as well as Microsoft® 10, UNIX®, SOLARIS®, LINUX®, Apple MAC-OS®, and other operating systems known to those skilled in the art. Furthermore, the CPU within the reconfigurable device 764 may be implemented as multiple processors that operate concurrently and cooperatively to execute instructions.

[0052] In one embodiment, the reconstructed image can be displayed on a display 766. The display 766 may be an LCD display, a CRT display, a plasma display, an OLED, an LED, or any other display known in the art.

[0053] The storage device 762 may be a hard disk drive, a CD-ROM drive, a DVD drive, a flash drive, RAM, ROM, or other electronic storage device known in the art.

[0054] In addition to the other embodiments described above, additional embodiments are disclosed below in parentheses.

[0055] (1) A method for denoising an image, comprising: receiving a first medical image containing a first image of an anatomical structure; receiving a second medical image containing a second image of an anatomical structure; and training a Deep Image Prior (DIP) neural network to generate a denoised image by inputting the second medical image into the DIP neural network during training and combining convergence noise with the output of the DIP network, such that at the end of training, the convergence noise combined with the output of the DIP network approximates the first medical image, wherein the output of the DIP network represents the denoised image.

[0056] (2) The method of (1), wherein training the DIP neural network includes using a double-over parameterized training process for convergence noise.

[0057] (3) The method of either (1) or (2), wherein training a DIP neural network includes, but is not limited to, initializing first and second noise vectors, training the DIP neural network to generate a denoised image by training the first and second noise vectors such that a convolution-based function based on the first and second noise vectors is equal to a value that converges to the noise of the first medical image, and training the DIP neural network to approximate the first medical image with the convolution-based function subtracted.

[0058] (4) The method according to any one of (1) to (3), wherein the first medical image is a positron emission tomography (PET) image of the subject.

[0059] (5) The method according to (4), wherein the second medical image is a computed tomography (CT) image of the subject aligned with the PET image.

[0060] (6) The method according to (4), wherein the second medical image is a magnetic resonance imaging (MRI) image of the subject aligned with the PET image.

[0061] (7) The method according to any one of (1) to (3), wherein the first medical image is a single-photon emission computed tomography (SPECT) image of the subject.

[0062] (8) The method according to (7), wherein the second medical image is a computed tomography (CT) image of the subject aligned with the SPECT image.

[0063] (9) The method according to (7), wherein the second medical image is a magnetic resonance imaging (MRI) image of the subject aligned with the SPECT image.

[0064] (10) The method according to any one of (1) to (3), wherein the first medical image is an ungated cardiac computed tomography (CT) image of the subject.

[0065] (11) The method according to (10), wherein the second medical image is a gated cardiac CT image of the subject aligned to an ungated cardiac CT image.

[0066] (12) A medical image processing device comprising, but not limited to, a processing circuit configured to receive a first medical image including a first image of an anatomical structure, receive a second medical image including a second image of an anatomical structure, input the second medical image to a Deep Image Prior (DIP) neural network during training, and train the DIP neural network to generate a denoised image by combining convergence noise with the output of the DIP network so that at the end of training the convergence noise combined with the output of the DIP network approximates the first medical image, and the output of the DIP network represents the denoised image.

[0067] (13) The apparatus according to (12), wherein the processing circuit configured to train a DIP neural network is configured to use a double-over parameterized training process for convergence noise.

[0068] (14) The apparatus according to (12), comprising a processing circuit configured to train a DIP neural network, which initializes first and second noise vectors, trains the DIP neural network to generate a denoised image by training the first and second noise vectors such that a convolution-based function based on the first and second noise vectors is equal to a value that converges to the noise of a first medical image, and trains the DIP neural network to approximate the first medical image with the convolution-based function subtracted.

[0069] (15) The apparatus described in any one of (12) to (14), wherein the first medical image is a positron emission tomography (PET) image of the subject.

[0070] (16) The apparatus described in (15), wherein the second medical image is a computed tomography (CT) image of the subject aligned with the PET image.

[0071] (17) The apparatus described in (15), wherein the second medical image is a magnetic resonance imaging (MRI) image of the subject aligned with the PET image.

[0072] (18) The apparatus described in any one of (12) to (14), wherein the first medical image is a single-photon emission computed tomography (SPECT) image of the subject.

[0073] (19) The apparatus described in (18), wherein the second medical image is a computed tomography (CT) image of the subject aligned with the SPECT image.

[0074] (20) The apparatus described in (18), wherein the second medical image is a magnetic resonance imaging (MRI) image of the subject aligned with the SPECT image.

[0075] (21) The apparatus described in any one of (12) to (14), wherein the first medical image is an ungated cardiac computed tomography (CT) image of the subject.

[0076] (22) The apparatus according to (21), wherein the second medical image is a gated cardiac CT image of the subject aligned with an ungated cardiac CT image.

[0077] (23) A non-temporary computer-readable storage medium that stores computer-readable instructions that, when executed by a computer, cause the computer to perform one of the methods (1) to (11).

[0078] Clearly, numerous modifications and variations are possible in light of the above teachings. Therefore, it should be understood that the disclosures may be implemented in ways other than those specifically described herein, within the scope of the attached claims.

[0079] According to at least one embodiment described above, image noise can be removed with high accuracy.

[0080] While several embodiments of the present invention have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These embodiments can be carried out in a variety of other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims and their equivalents. [Explanation of symbols]

[0081] 762 Storage device 764 Reconfiguration Devices 870 processor 878 memory

Claims

1. 1. A method for training a Deep Image Prior (DIP) neural network for image denoising, comprising: receiving a first medical image including a first image of an anatomical structure; receiving a second medical image comprising a second image of the anatomical structure; training the at least one DIP neural network to generate a denoised image from the input image by inputting the second medical image into the at least one DIP neural network during training and combining converged noise with an output of the at least one DIP neural network such that at the end of the training, the converged noise combined with the output of the at least one DIP neural network approximates the first medical image; Including, the output of the at least one DIP neural network represents the denoised image; Training the at least one DIP neural network includes: initializing a first noise vector and a second noise vector; training the at least one DIP neural network to generate the denoised image by training the first noise vector and the second noise vector so that a convolution-based function based on the first noise vector and the second noise vector is equal to a value that converges to the noise of the first medical image; training the at least one DIP neural network to approximate the first medical image minus the convolution-based function; Including, Training the at least one DIP neural network further comprises: training a plurality of DIP neural networks as the at least one DIP neural network, each with different parameters; Including, method.

2. The method of claim 1 , wherein training the at least one DIP neural network comprises performing a double over-parameterized training process on the converged noise.

3. The method of claim 1 , wherein the first medical image is a Position Emission Tomography (PET) image of a subject.

4. The method of claim 3 , wherein the second medical image is a computed tomography (CT) image of the subject registered to the PET image.

5. The method of claim 3 , wherein the second medical image is a Magnetic Resonance Imaging (MRI) image of the subject registered to the PET image.

6. The method of claim 1 , wherein the first medical image is a Single-Photon Emission Computerized Tomography (SPECT) image of a subject.

7. The method of claim 6 , wherein the second medical image is a computed tomography (CT) image of the subject registered to the SPECT image.

8. The method of claim 6 , wherein the second medical image is a magnetic resonance imaging (MRI) image of the subject registered to the SPECT image.

9. The method of claim 1 , wherein the first medical image is a non-gated cardiac CT image of a subject.

10. 10. The method of claim 9, wherein the second medical image is a gated cardiac CT image of the subject registered to the non-gated cardiac CT image.

11. receiving a first medical image including a first image of an anatomical structure; receiving a second medical image comprising a second image of the anatomical structure; a processing circuit for training at least one Deep Image Prior (DIP) neural network to input the second medical image into the at least one DIP neural network during training and to generate a noise-removed image from the input image by combining converged noise with an output of the at least one DIP neural network such that at the end of the training, the converged noise combined with the output of the at least one DIP neural network approximates the first medical image; Equipped with the output of the at least one DIP neural network represents the denoised image; The processing circuitry, in training the at least one DIP neural network, initializing a first noise vector and a second noise vector; training the at least one DIP neural network to generate the denoised image by training the first noise vector and the second noise vector so that a convolution-based function based on the first noise vector and the second noise vector is equal to a value that converges to the noise of the first medical image; training the at least one DIP neural network to approximate the first medical image minus the convolution-based function; The processing circuitry further comprises: training a plurality of DIP neural networks as the at least one DIP neural network, each with different parameters; Medical imaging equipment.

12. The medical imaging apparatus of claim 11 , wherein the processing circuitry performs a double over-parameterized training process on the converged noise.

13. A program for causing a computer to execute a process of training a Deep Image Prior (DIP) neural network for removing noise from an image, comprising: The computer, receiving a first medical image comprising a first image of an anatomical structure; receiving a second medical image comprising a second image of the anatomical structure; training the at least one DIP neural network to generate a noise-removed image from the input image by inputting the second medical image into at least one DIP neural network during training and combining converged noise with an output of the at least one DIP neural network such that at the end of the training, the converged noise combined with the output of the at least one DIP neural network approximates the first medical image; Execute the output of the at least one DIP neural network represents the denoised image; The step of training the at least one DIP neural network comprises: initializing a first noise vector and a second noise vector; training the at least one DIP neural network to generate the denoised image by training the first noise vector and the second noise vector so that a convolution-based function based on the first noise vector and the second noise vector is equal to a value that converges to the noise of the first medical image; training the at least one DIP neural network to approximate the first medical image minus the convolution-based function; Including, The step of training the at least one DIP neural network further comprises: training a plurality of DIP neural networks as the at least one DIP neural network, each with different parameters; Including, program.

14. 14. The computer program product of claim 13, wherein training the at least one DIP neural network comprises performing a double-over-parameterization training process on the converged noise.

15. 1. A method for denoising an image using a Deep Image Prior (DIP) neural network for denoising an image, comprising: inputting the images into at least one DIP neural network that is trained to generate a denoised image from the input image, wherein during training, a second medical image including a second image of an anatomical structure is input, and by combining converged noise with an output, at the end of said training, said converged noise combined with said output approximates the first medical image including the first image of an anatomical structure; obtaining the image generated and output by the at least one DIP neural network, the image having noise removed from the input image; This includes: The at least one DIP neural network comprises: initializing a first noise vector and a second noise vector; training the at least one DIP neural network to generate the denoised image by training the first noise vector and the second noise vector so that a convolution-based function based on the first noise vector and the second noise vector is equal to a value that converges to the noise of the first medical image; training the at least one DIP neural network to approximate the first medical image minus the convolution-based function; and trained by a method comprising: The at least one DIP neural network further comprises: a plurality of DIP neural networks as the at least one DIP neural network, each trained using different parameters; method.

16. A medical image processing device that removes image noise using a Deep Image Prior (DIP) neural network, a storage unit that stores at least one DIP neural network that has been trained during training to receive a second medical image, the second medical image comprising a second image of an anatomical structure, and to generate a noise-removed image from the input image by combining converged noise with an output such that, at the end of the training, the converged noise combined with the output approximates the first medical image, the first image of an anatomical structure; a processing circuit for inputting an image to the at least one DIP neural network stored in the storage unit, and obtaining the image generated and output by the at least one DIP neural network, with noise removed from the input image; Equipped with The at least one DIP neural network comprises: initializing a first noise vector and a second noise vector; training the at least one DIP neural network to generate the denoised image by training the first noise vector and the second noise vector so that a convolution-based function based on the first noise vector and the second noise vector is equal to a value that converges to the noise of the first medical image; training the at least one DIP neural network to approximate the first medical image minus the convolution-based function; and trained by a method comprising: The at least one DIP neural network further comprises: a plurality of DIP neural networks as the at least one DIP neural network, each trained using different parameters; Medical imaging equipment.

17. A program for causing a computer to execute a process for removing image noise using a Deep Image Prior (DIP) neural network for removing image noise, comprising: The computer, obtaining at least one DIP neural network stored in a storage unit that stores at least one DIP neural network trained to generate a noise-removed image from the input image during training, the image comprising a second image of an anatomical structure, by combining converged noise with an output such that at the end of the training, the converged noise combined with the output approximates the first medical image comprising the first image of an anatomical structure; inputting an image into the at least one DIP neural network obtained, and obtaining an image generated and output by the DIP neural network in which noise has been removed from the input image; A program for executing The at least one DIP neural network comprises: initializing a first noise vector and a second noise vector; training the at least one DIP neural network to generate the denoised image by training the first noise vector and the second noise vector so that a convolution-based function based on the first noise vector and the second noise vector is equal to a value that converges to the noise of the first medical image; training the at least one DIP neural network to approximate the first medical image minus the convolution-based function; and trained by a method comprising: The at least one DIP neural network further comprises: a plurality of DIP neural networks as the at least one DIP neural network, each trained using different parameters; program.