Reducing artifacts in medical images using cascaded residual networks
By combining the residual network architecture and skip connection method, the problem that CNN has difficulty in removing artifacts of different scales in medical images is solved, achieving more efficient artifact reduction and improving image quality and diagnostic accuracy.
Patent Information
- Application Number
- CN202510251204.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-19
- Filing Date
- 2025-03-04
- Publication Date
- 2025-09-19
AI Technical Summary
Existing convolutional neural networks (CNNs) have difficulty in effectively reducing artifacts of different scales in medical images, especially local and global artifacts. Conventional residual networks suffer from performance degradation and training data dependence during training, resulting in insufficient artifact removal.
A series of residual network architectures are adopted to estimate and reduce artifacts of different scales separately through multiple levels of residual networks, skip connections are used to improve training efficiency, and different types of artifacts are gradually decoupled through joint optimization.
It significantly improves the diagnostic quality of medical images and the accuracy of image analysis, reduces the number and time of scans, and enhances the robustness and consistency of imaging.
Smart Images

Figure CN120672592A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the subject matter disclosed herein relate to medical imaging, and more particularly to systems and methods for removing artifacts from medical images. Background Art
[0002] Medical images, such as magnetic resonance (MR) images, may contain artifacts, which can degrade image quality and hinder diagnosis. Various approaches have been employed to reduce or remove artifacts. For example, convolutional neural networks (CNNs) can be trained to reduce artifacts in MR images. A CNN can be trained based on an image pair consisting of a first input image with artifacts and a second target (real) image without artifacts. The CNN learns to map an artifact-bearing image to an artifact-reduced image, and during training, the CNN outputs an artifact-reduced version of the MR image input to the CNN. However, a CNN trained in this manner may not be sufficient to reduce artifacts, with some artifacts remaining in the image after being processed by the CNN. In particular, a single CNN may struggle to reduce artifacts of varying scales, with a first CNN effectively reducing local artifacts in an image (such as noise) but not global artifacts (such as motion artifacts). A second CNN effectively reduces global artifacts but not local artifacts. Summary of the Invention
[0003] In one example, the above problem can be solved by an image processing system, which includes: a trained artifact estimation network, the trained artifact estimation network including multiple stages, the artifact estimation network being trained to estimate artifacts in a medical image; and a processor, which is capable of being communicatively coupled to a non-volatile memory storing the artifact estimation network, the memory including instructions that, when executed, cause the processor to: receive a medical image; generate an estimated artifact image from the medical image using the trained artifact estimation network; generate an artifact-reduced image by subtracting the estimated artifact image from the medical image, the artifact-reduced image being a version of the medical image that includes less artifacts than the medical image; and display the artifact-reduced image on a display device; wherein each stage of the trained artifact estimation network estimates artifacts of different scales in the medical image.
[0004] It should be understood that the above brief description is provided to introduce in a simplified form selected concepts that are further described in the detailed description. It is not meant to identify key features or essential features of the claimed subject matter, the scope of which is uniquely defined by the claims that follow the detailed description. Furthermore, the claimed subject matter is not limited to implementations that solve any disadvantages noted above or in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Various aspects of the present disclosure may be better understood by reading the following detailed description and referring to the accompanying drawings, in which:
[0006] Figure 1 shows a block diagram of an MRI system according to one or more embodiments of the present disclosure;
[0007] Figure 2 A block diagram illustrating an exemplary embodiment of an image processing system according to one or more embodiments of the present disclosure;
[0008] Figure 3 A block diagram illustrating an exemplary embodiment of an artifact estimation network training system for training an artifact estimation network according to one or more embodiments of the present disclosure is shown;
[0009] Figure 4 A flowchart illustrating an exemplary method for training an artifact estimation network according to one or more embodiments of the present disclosure is shown;
[0010] Figure 5 A flowchart illustrating an exemplary method for deploying a trained artifact estimation network to reduce artifacts in medical images according to one or more embodiments of the present disclosure is shown;
[0011] Figure 6 shows an exemplary architecture of an artifact estimation network according to one or more embodiments of the present disclosure;
[0012] Figure 7 shows an exemplary architecture of a first stage of an artifact estimation network according to one or more embodiments of the present disclosure;
[0013] Figure 8 shows an exemplary architecture of the second stage of an artifact estimation network according to one or more embodiments of the present disclosure;
[0014] Figure 9 shows a first exemplary output of a trained artifact estimation network according to one or more embodiments of the present disclosure;
[0015] Figure 10 shows a second exemplary output of a trained artifact estimation network according to one or more embodiments of the present disclosure; and
[0016] Figure 11 Exemplary outputs of different stages of a trained artifact estimation network are shown, according to one or more embodiments of the present disclosure.
[0017] The accompanying drawings illustrate specific aspects of the described systems and methods. Together with the following description, the drawings illustrate and explain the structures, methods, and principles described herein. In the drawings, the dimensions of components may be exaggerated or otherwise modified for clarity. Well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the described components, systems, and methods. DETAILED DESCRIPTION
[0018] Provided herein are methods and systems for reducing artifacts in medical image data, such as magnetic resonance (MR) images, computed tomography (CT) images, positron emission tomography (PET) images, or other types of medical images. Various methods have been developed to reduce or remove artifacts. In particular, deep learning (DL)-based methods have been developed for processing medical images to reduce artifacts in the images. For example, a neural network can be trained to detect and extract noise and artifacts from a medical image. A medical image including artifacts can be input into a first neural network, and the first neural network can output an artifact-reduced image, wherein the artifact-reduced image is a version of the medical image with reduced noise and artifacts. Alternatively, in some examples, the medical image including artifacts can be input into a second neural network, and the second neural network can output extracted noise and artifact image data. The extracted noise and artifact image data can be subtracted from the medical image to generate the artifact-reduced image.
[0019] Various types of neural networks can be used to remove artifacts from an image. The neural network can be a convolutional neural network (CNN) comprising a plurality of interconnected layers. At each subsequent layer of the CNN, more complex and abstract features can be extracted, which can improve the performance of the CNN. However, as the number of layers of the CNN increases above a threshold, the performance and accuracy of the CNN may decrease, and the CNN may stop learning. For example, the magnitude of the gradient of the loss function of the CNN may decrease as the number of layers of the CNN increases, wherein adjustments to the parameters (e.g., weights) of the CNN during backpropagation may become increasingly negligible. In other cases, the gradient may become larger, thereby causing instability during training.
[0020] In order to improve the performance and accuracy of CNNs containing multiple layers, these layers can be grouped into blocks (e.g., residual blocks), and skip connections can be included in the CNN, whereby the inputs to the block are added together as additional outputs of the block (e.g., identity mapping). The input added to the output bypasses the convolutional layers included in the block. When the number of layers of the CNN exceeds a threshold, the addition of skip connections can improve the efficiency of gradient descent during backpropagation. Therefore, the depth of the CNN can be increased without reducing performance, and the accuracy of the output can be improved. This architecture is generally referred to as a residual network. For example, a residual network can include 30 or more layers.
[0021] Conventional residual networks can be used to remove artifacts from medical images such as MR images. However, such medical images may include various types of artifacts. These artifacts may include both local features (e.g., noise, fine lines, ringing, etc.) and global features (motion, streaks, aliasing, etc.). Fine lines are typically caused by stimulated echoes, and ringing is often seen around edges due to truncation in the frequency domain. Streaks are typically observed in images acquired using a radial sampling pattern, while aliasing is typically caused by unsuppressed signal outside the specified field of view. The performance of a conventional residual network for the first type of artifact may be different from the performance of a conventional residual network for the second type of artifact. Training a conventional residual network to perform well on a variety of different artifact types can be difficult. Training a conventional residual network may rely on generating a large amount of training data that includes various types of artifacts in various combinations, which can be time-consuming and difficult to obtain. Furthermore, local and global artifacts may interact with each other, resulting in secondary features in the image domain. This additional complexity may make it more difficult for conventional residual networks to learn and separate different artifacts, resulting in inaccurate estimation and / or insufficient removal of these artifacts. As a result, conventional residual networks can remove one or more types of artifacts from medical images, but leave other types of artifacts behind.
[0022] Multiple conventional residual networks can be used in series to remove different types of artifacts. For example, a medical image can be input into a first conventional residual network, and the first conventional residual network can detect and reduce the first type of artifact. A second medical image with reduced artifacts of the first type, output by the first conventional residual network, can be input into a second conventional residual network. The second conventional residual network can detect and reduce the second type of artifact. A third medical image with reduced artifacts of the second type, output by the second conventional residual network, can be input into a third conventional residual network, and so on. By chaining conventional residual networks in this way, the generation of training datasets can be simplified, and the performance of each of these conventional residual networks can be improved individually.
[0023] However, the inventors herein have recognized a problem with chaining conventional residual networks in this manner, wherein the chained networks may not produce artifact-free images or images with a significantly reduced number or range of different types of artifacts. When a first conventional residual network reduces a first type of artifact, the first conventional residual network may also reduce or remove some of a second type of artifact, which may affect the ability of a second conventional residual network to learn to recognize the second type of artifact. Similarly, a second conventional residual network may remove image data of a third type of artifact, which may make it more difficult for a third conventional residual network to detect and remove the third type of artifact, and so on. As a result, remnants of the second artifact, the third artifact, and / or other artifacts may be present in the final artifact-reduced image generated by the final chained conventional residual network. Additionally, the order in which the conventional residual networks are chained is important, wherein different orders of the conventional residual networks may result in different artifact-reduced images of different quality.
[0024] Another problem with chaining residual networks is that each chained residual network depends on the output from the previous residual network. Therefore, if or when one chained model is adjusted, each of the other chained models must also be adjusted. This reduces the robustness and flexibility of the solution and may have regulatory implications.
[0025] Alternatively, a first conventional residual network, a second conventional residual network, and a third conventional residual network may be trained in parallel on the same training data set. The first conventional residual network, the second conventional residual network, and the third conventional residual network may each be trained to output noise and artifact data of a specific type of artifact of the medical image. The noise and artifact data output by the first conventional residual network, the second conventional residual network, and the third conventional residual network may then be added together to generate an artifact image (e.g., an artifact mask), and the artifact image may be subtracted from the medical image to remove different types of artifacts. However, due to the variety of artifacts included in the training data, it may be difficult for the first conventional residual network, the second conventional residual network, and the third conventional residual network to each achieve good performance on a single type of artifact.
[0026] To more significantly reduce various types of artifacts in medical images, this paper proposes a residual network architecture and training method that can more effectively remove different types of artifacts from medical images using a single trained residual network. The proposed residual network can be trained based on training data that includes various types of artifacts and noise distributions, and can use joint optimization to gradually decouple and extract different types of artifacts. According to the proposed method, different parts or levels of the proposed residual network architecture can estimate artifacts of different scales. That is, smaller-scale local artifacts (e.g., noise, ringing, etc.) can be first estimated via a first set of layers of the proposed residual network architecture and reduced or removed from the input image. Once the local noise and artifacts have been substantially reduced, larger-scale global artifacts (e.g., streaks, motion artifacts, etc.) can be estimated via a second set of layers of the proposed residual network architecture and removed from the input image. Additional levels can be used to further distinguish between scales. For example, the first stage may be used to reduce artifacts at a first scale (e.g., low-level noise); the second stage may be used to reduce artifacts at a second scale (e.g., ringing artifacts); the third stage may be used to reduce artifacts at a third scale (e.g., streak artifacts); the fourth stage may be used to reduce artifacts at a fourth scale (e.g., motion artifacts); and so on.
[0027] By training the proposed residual network as described herein, different types of artifacts can be removed from medical images more effectively than by using conventional residual networks (including when multiple conventional residual networks are chained or trained in parallel). An additional advantage of the proposed method and model is that multiple constraints and dependencies on training set data can be reduced, which can include a wider range of images with varying degrees of noise, a wider range of signal-to-noise ratios (SNRs), and more types of artifacts without increasing training time or degrading model performance due to over-representation of edge cases. In this way, the diagnostic image quality of medical images and the accuracy of image analysis and disease staging can be improved. In addition, the robustness and consistency of medical imaging can be improved, which can reduce the number and / or duration of scans performed on a subject.
[0028] First refer to the attached figure, Figure 1 An exemplary imaging system that can be used to acquire medical imaging data is illustrated. Figure 1 A magnetic resonance imaging (MRI) system is illustrated, but it will be appreciated that other medical imaging systems may be used without departing from the scope of the present disclosure. Figure 1A magnetic resonance imaging (MRI) apparatus 10 is illustrated, which includes a static magnetic field magnet unit 12, a gradient coil unit 13, an RF coil unit 14, an RF body or volume coil unit 15, a transmit / receive (T / R) switch 20, an RF driver unit 22, a gradient coil driver unit 23, a data acquisition unit 24, a controller unit 25, a patient couch or table 26, a data processing unit 31, an operation console unit 32, and a display unit 33. In some embodiments, the RF coil unit 14 is a surface coil, which is a local coil typically placed near an anatomical structure of interest on a subject 16. Here, the RF body coil unit 15 is a transmit coil that transmits RF signals, and the local surface RF coil unit 14 receives MR signals. Therefore, the transmit body coil (e.g., the RF body coil unit 15) and the surface receive coil (e.g., the RF coil unit 14) are independent but electromagnetically coupled components. The MRI apparatus 10 transmits electromagnetic pulse signals to a subject 16 placed in an imaging space 18 forming a static magnetic field to perform scanning to obtain magnetic resonance signals from the subject 16. One or more images of the subject 16 can be reconstructed based on the magnetic resonance signals thus obtained by the scanning.
[0029] The static magnetic field magnet unit 12 includes, for example, an annular superconducting magnet installed in an annular vacuum container. The magnet defines a cylindrical space surrounding the subject 16 and generates a constant main static magnetic field B0.
[0030] The MRI apparatus 10 also includes a gradient coil unit 13, which generates a gradient magnetic field in an imaging space 18 to provide three-dimensional position information for the magnetic resonance signals received by the RF coil array. The gradient coil unit 13 includes three gradient coil systems, each of which generates a gradient magnetic field along one of three spatial axes perpendicular to one another. The gradient coil unit 13 also generates gradient fields in each of the frequency encoding direction, the phase encoding direction, and the slice selection direction, depending on imaging conditions. More specifically, the gradient coil unit 13 applies a gradient field in the slice selection direction (or scanning direction) of the subject 16 to select a slice. The RF body coil unit 15 or the local RF coil array can then transmit RF pulses to the selected slice of the subject 16. The gradient coil unit 13 also applies a gradient field in the phase encoding direction of the subject 16 to phase-encode the magnetic resonance signals from the slice excited by the RF pulse. The gradient coil unit 13 then applies a gradient field in the frequency encoding direction of the subject 16 to frequency-encode the magnetic resonance signals from the slice excited by the RF pulse.
[0031] The RF coil unit 14 is positioned, for example, to surround a region of the subject 16 to be imaged. In some examples, the RF coil unit 14 may be referred to as a surface coil or a receiving coil. In the static magnetic field space or imaging space 18, where the static magnetic field B0 is formed by the static magnetic field magnet unit 12, the RF coil unit 15 transmits RF pulses, which are electromagnetic waves, to the subject 16 based on a control signal from the controller unit 25, thereby generating a high-frequency magnetic field B1. This excites the proton spins in the slice of the subject 16 to be imaged. The RF coil unit 14 receives the electromagnetic waves generated when the proton spins excited in the slice of the subject 16 to be imaged return to alignment with the initial magnetization vector as magnetic resonance signals. In some embodiments, the RF coil unit 14 can both transmit RF pulses and receive MR signals. In other embodiments, the RF coil unit 14 may be used only to receive MR signals, without transmitting RF pulses.
[0032] The RF body coil unit 15 is disposed, for example, to surround the imaging space 18 and generates RF magnetic field pulses orthogonal to the main magnetic field B0 generated by the static magnetic field magnet unit 12 within the imaging space 18 to excite nuclei. In contrast to the RF coil unit 14, which can be disconnected from the MRI apparatus 10 and replaced with another RF coil unit, the RF body coil unit 15 is fixedly attached and connected to the MRI apparatus 10. Furthermore, while local coils (such as the RF coil unit 14) can transmit or receive signals to or from a local region of the subject 16, the RF body coil unit 15 typically has a larger coverage area. For example, the RF body coil unit 15 can be used to transmit or receive signals to or from the entire body of the subject 16. Using a receive-only local coil and a transmit body coil provides uniform RF excitation and good image uniformity, at the expense of higher RF power deposited in the subject. With a transmit-receive local coil, the local coil provides RF excitation to the region of interest and receives MR signals, thereby reducing RF power deposited in the subject. It should be understood that the specific use of the RF coil unit 14 and / or the RF body coil unit 15 depends on the imaging application.
[0033] The T / R switch 20 can selectively electrically connect the RF body coil unit 15 to the data acquisition unit 24 when operating in the receive mode, and can selectively electrically connect the RF body coil unit 22 when operating in the transmit mode. Similarly, the T / R switch 20 can selectively electrically connect the RF coil unit 14 to the data acquisition unit 24 when the RF coil unit 14 operates in the receive mode, and can selectively electrically connect the RF coil unit 14 to the RF driver unit 22 when the RF coil unit operates in the transmit mode. When both the RF coil unit 14 and the RF body coil unit 15 are used for a single scan, for example, if the RF coil unit 14 is configured to receive MR signals and the RF body coil unit 15 is configured to transmit RF signals, the T / R switch 20 can direct control signals from the RF driver unit 22 to the RF body coil unit 15 while directing received MR signals from the RF coil unit 14 to the data acquisition unit 24. The coils of the RF body coil unit 15 may be configured to operate in a transmit-only mode or a transmit-receive mode. The coils of the local RF coil unit 14 may be configured to operate in a transmit-receive mode or a receive-only mode.
[0034] The RF driver unit 22 includes a gate modulator (not shown), an RF power amplifier (not shown), and an RF oscillator (not shown), which are used to drive an RF coil (e.g., the RF coil unit 15) and form a high-frequency magnetic field in the imaging space 18. The RF driver unit 22 modulates the RF signal received from the RF oscillator into a signal with a predetermined envelope and predetermined timing using the gate modulator based on a control signal from the controller unit 25. The RF signal modulated by the gate modulator is amplified by the RF power amplifier and then output to the RF coil unit 15.
[0035] The gradient coil driver unit 23 drives the gradient coil unit 13 based on a control signal from the controller unit 25 and thereby generates a gradient magnetic field in the imaging space 18. The gradient coil driver unit 23 includes three driver circuits (not shown) corresponding to the three gradient coil systems included in the gradient coil unit 13.
[0036] The data acquisition unit 24 includes a preamplifier (not shown), a phase detector (not shown), and an analog / digital converter (not shown) for acquiring magnetic resonance signals received by the RF coil unit 14. In the data acquisition unit 24, the phase detector uses the output of the RF oscillator from the RF driver unit 22 as a reference signal to perform phase detection on the magnetic resonance signals received from the RF coil unit 14 and amplified by the preamplifier. The phase-detected analog magnetic resonance signals are then output to the analog / digital converter for conversion into digital signals. The resulting digital signals are then output to the data processing unit 31.
[0037] The MRI apparatus 10 includes a table 26 for placing the subject 16 thereon. By moving the table 26 based on a control signal from the controller unit 25, the subject 16 can be moved inside and outside the imaging space 18.
[0038] The controller unit 25 includes a computer and a recording medium on which a program to be executed by the computer is recorded. When executed by the computer, the program causes the various components of the device to perform operations corresponding to predetermined scans. The recording medium may include, for example, a ROM, a floppy disk, a hard disk, an optical disk, a magneto-optical disk, a CD-ROM, or a non-volatile memory card. The controller unit 25 is connected to the operation console unit 32 and processes operation signals input to the operation console unit 32. It also controls the examination table 26, the RF driver unit 22, the gradient coil driver unit 23, and the data acquisition unit 24 by outputting control signals to these units. Based on the operation signals received from the operation console unit 32, the controller unit 25 controls the data processing unit 31 and the display unit 33 to obtain the desired image.
[0039] The operation console unit 32 includes user input devices such as a touch screen, a keyboard, and a mouse. The operation console unit 32 is used by the operator to input data such as imaging protocols and to set the area where the imaging sequence will be performed. Data on the imaging protocol and the imaging sequence execution area are output to the controller unit 25.
[0040] The data processing unit 31 includes a computer and a recording medium on which a program to be executed by the computer to perform predetermined data processing is recorded. The data processing unit 31 is connected to the controller unit 25 and performs data processing based on control signals received from the controller unit 25. The data processing unit 31 is also connected to the data acquisition unit 24 and generates spectrum data by applying various image processing operations to the magnetic resonance signals output from the data acquisition unit 24.
[0041] The display unit 33 includes a display device and displays an image on a display screen of the display device based on a control signal received from the controller unit 25. The display unit 33 displays, for example, images regarding input items for operation data input by the operator from the operation console unit 32. The display unit 33 also displays a two-dimensional (2D) slice image or a three-dimensional (3D) image of the subject 16 generated by the data processing unit 31.
[0042] Although an MRI system is described by way of example, it should be understood that the present technology may also be useful when applied to images acquired using other imaging modalities, such as CT, tomosynthesis, PET, C-arm angiography, etc. Discussion of the present invention with respect to the MRI imaging modality is provided merely as an example of one suitable imaging modality.
[0043] Now refer to Figure 2 , illustrates an image processing system 202 of a medical imaging system 200 according to an embodiment. In some embodiments, at least a portion of the image processing system 202 is disposed at a device (e.g., an edge device, a server, etc.) that is communicatively coupled to the medical imaging system 200 via a wired connection and / or a wireless connection. In some embodiments, at least a portion of the image processing system 202 is disposed at a separate device (e.g., a workstation) that can receive images from the medical imaging system 200 or from a storage device that stores images / data generated by the medical imaging system 200.
[0044] The image processing system 202 includes a processor 204 configured to execute machine-readable instructions stored in a non-transitory memory 206. The processor 204 can be a single-core or multi-core processor, and the programs executed thereon can be configured for parallel processing or distributed processing. In some embodiments, the processor 204 can optionally include separate components spread across two or more devices, which can be located at a distance and / or configured for collaborative processing. In some embodiments, one or more aspects of the processor 204 can be virtualized and performed by a remotely accessible networked computing device configured in a cloud computing configuration.
[0045] The non-transitory memory 206 may store a neural network module 208, a network training module 210, an inference module 212, and medical image data 214. The neural network module 208 may include one or more DL networks and instructions for implementing the DL network to reduce or optionally remove noise from the medical images of the medical image data 214, as described in more detail below. The neural network module 208 may include one or more trained and / or untrained neural networks and may also include various data or metadata related to the one or more neural networks stored therein. In particular, the neural network module may store an artifact estimation network 209, which will be referred to below. Figures 3 to 10 Describe in more detail.
[0046] The training module 210 may include instructions for training one or more neural networks that implement the artifact estimation network 209 and / or other DL models stored in the neural network module 208. In particular, the training module 210 may include instructions that, when executed by the processor 204, cause the image processing system 202 to perform one or more of the steps of the method 400 for training the artifact estimation network 209 at a training level, which will be described below with reference to Figure 3 and Figure 4 In some embodiments, the training module 210 includes instructions for implementing one or more gradient descent algorithms, applying one or more loss functions, and / or training routines for adjusting parameters of one or more neural networks of the neural network module 208. The inference module 212 may include instructions for using the trained DL model to reduce the amount of artifacts in new image data.
[0047] In some embodiments, non-transitory memory 206 may include components located at two or more devices, which may be remotely located and / or configured for collaborative processing. In some embodiments, one or more aspects of non-transitory memory 206 may include a remotely accessible networked storage device configured in a cloud computing configuration.
[0048] The image processing system 202 can be operably / communicatively coupled to a user input device 232 and a display device 234. The user input device 232 can include one or more of a touch screen, a keyboard, a mouse, a trackpad, or other devices configured to enable a user to interact with and manipulate data within the image processing system 202. The display device 234 can include one or more display devices utilizing virtually any type of technology. In some embodiments, the display device 234 can include a computer monitor and can display medical images. The display device 234 can be combined with the processor 204, the non-volatile memory 206, and / or the user input device 232 in a shared housing, or can be a peripheral display device and can include a monitor, a touch screen, a projector, or other display device known in the art that enables a user to view medical images generated by the medical imaging system and / or interact with various data stored in the non-volatile memory 206. In some examples, the display device 234 can be Figure 1 The display unit 33 and the user input device 232 may be Figure 1 at least a portion of the operating console unit 32 .
[0049] The non-transitory memory 206 also stores medical image data 214. The medical image data 214 may include, for example, medical images acquired via a scanner 236, which may be an MR scanner, a CT scanner, a scanner for spectral imaging, or a different imaging modality. The image processing system 202 may be operatively / communicatively coupled to the scanner 236. The scanner 236 may be, for example, Figure 1 The image processing system 202 may be any imaging device, such as the MRI apparatus 10, configured to image a subject, such as a patient, an inanimate object, one or more manufactured parts, and / or a foreign object, such as a dental implant, a stent, and / or a contrast agent, that is present in the body. The image processing system 202 may receive imaging data from the scanner 236, process the received imaging data via the processor 204 based on instructions stored in one or more modules of the non-transitory memory 206, and / or store the received imaging data in the medical image data 214.
[0050] It should be understood that Figure 1 The image processing system 202 shown is for illustration and not for limitation. Another suitable image processing system may include more, fewer, or different components.
[0051] refer to Figure 3 , shows an example of an artifact estimation network training system 300 that can be used to train a neural network, such as an artifact estimation network 302. Figure 4By performing one or more operations described in more detail in the method 400, the artifact estimation network 302 can be trained to estimate artifacts in a two-dimensional (2D) or three-dimensional (3D) medical image. The estimated artifacts can then be extracted from the 2D or 3D image to obtain an artifact-reduced image. The artifact estimation network training system 300 can be used by, for example, Figure 2 The image processing system 202 may be implemented to train the artifact estimation network 302 to estimate artifacts in MR images or different types of medical images, which artifacts may then be reduced or removed.
[0052] In some embodiments, the artifact estimation network 302 may be a deep neural network having multiple hidden layers. In one embodiment, the artifact estimation network 302 is a convolutional neural network (CNN) such as a residual neural network, as described in more detail below. The artifact estimation network 302 may be stored within a neural network module 301 of the image processing system, which may be a Figure 2 2 is a non-limiting example of a neural network module 208 of the image processing system 202.
[0053] The artifact estimation network training system 300 includes a training module 304, which may be Figure 2 2 is a non-limiting example of a training module 210 of the image processing system 202. The training module 304 includes a training dataset including a plurality of training data pairs (such as image pairs) divided into training image pairs 306 and test image pairs 308, which are used to train the artifact estimation network 302. The plurality of training image pairs 306 and test image pairs 308 can be selected to ensure that sufficient training data is available to prevent overfitting, whereby the artifact estimation network 302 learns to map specific features of training set samples that are not present in the test set.
[0054] Each image pair of training image pairs 306 and test image pairs 308 includes an input image and a target image. In various embodiments, the input image may be a noisy MR image 316 generated from a high-quality MR image 312 by combining one or more artifact images 314 with the high-quality MR image 312. Combining the one or more artifact images 314 with the high-quality MR image 312 may include, for each pixel of the high-quality MR image 312, adding the pixel intensity value of the corresponding pixel in each of the one or more artifact images 314 to the pixel intensity value of the pixel. For example, the first pixel intensity value of a pixel at a first location in the first high-quality MR image 312 may be added to the first pixel intensity value of a first pixel at the same location in the first artifact image 314 and the second pixel intensity value of a second pixel at the same location in the second artifact image 314, and so on for each pixel. The high-quality MR image may be a real MR image collected from a patient using a real scanner with few or no artifacts. In some embodiments, the high-quality MR image may be a synthetic image. The artifact image 314 may include a synthetic image with various types of artifacts. The various types of artifacts may include local noise, ringing artifacts, streaks, aliasing, etc. The artifact images 314 may be combined with the high-quality MR images 312 to create a corresponding set of noisy MR images 316 .
[0055] For example, a first set of one or more synthetic artifact images 314 can be added to a first high-quality MR image 312 to generate a first noisy MR image 316, which can be an input image for a first training image pair 306; and the first set of one or more synthetic artifact images 314 can be combined to generate a combined artifact target image for the first training image pair 306. A second set of one or more synthetic artifact images 314 can be added to a second high-quality MR image 312 to generate a second noisy MR image 316, which can be an input image for a second training image pair 306; and the second set of one or more synthetic artifact images 314 can be combined to generate a target image for the second training image pair 306. A third set of one or more synthetic artifact images 314 can be added to a third high-quality MR image 312 to generate a third noisy MR image 316, which can be an input image for a third training image pair 306; and the third set of one or more synthetic artifact images 314 can be combined to generate a target image for a third training image pair 306; and so on. The first, second, and third sets of one or more synthesized artifact images 314 may include the same, similar, or different artifacts, artifact types, and / or artifact quantities. In this way, a robust set of training data can be generated by exploiting the 1:1 correspondence between the input image and the target image. In some embodiments, each synthesized artifact image can be saved separately as a target image in a training image pair.
[0056] In other embodiments, the artifact estimation network 302 can be trained using different input and target image pairs. For example, in some embodiments, each training image pair 306 can include a noisy MR image 316 as an input image and a corresponding high-quality image 312 as a target image, where the noisy MR image 316 is generated by combining one or more artifact images 314 with the corresponding high-quality image 312. In such embodiments, the artifact estimation network 302 can be trained according to a similar training process as described herein to output an artifact-reduced image 338 instead of an artifact image 334.
[0057] The artifact estimation network training system 300 may include a training data generator 310, which may be used to generate image pairs. A noisy MR image 316 may be paired with an artifact image 314 via the training data generator 310, as described above. Once each image pair is generated, it may be assigned to either a training image pair 306 or a test image pair 308. In one embodiment, the image pairs may be randomly assigned to either a training image pair 306 or a test image pair 308 in a pre-established ratio. The artifact estimation network 302 may be trained based on the training image pairs to output one or more artifact images 314 associated with each noisy MR image 316. That is, the artifact estimation network 302 may be trained to extract artifacts from the noisy MR image 316 and output multiple artifact images or a combined artifact image including the extracted artifacts. The extracted artifact images may then be subtracted from the noisy MR image 316 at a subtraction module 336 to obtain an artifact-reduced image that is identical or similar to the high-quality MR image 312 used to generate the noisy MR image 316. Subtracting the artifact image from the noisy MR image 316 may include subtracting a first pixel intensity of each pixel of the artifact image from a corresponding second pixel intensity of the pixel at the same pixel location in the noisy MR image 316. Additionally, in various embodiments, an artifact-reduced image may be generated that extracts different artifacts at different stages of the plurality of stages of the artifact estimation network 302. As described below with reference to Figure 4 、 Figure 5 and Figure 6 Describing in more detail, at the end of each stage, an artifact image extracting a single artifact type can be output and subtracted from the noisy MR image 316 to produce a partially clean MR image in which the single artifact type has been reduced or removed. This partially clean MR image can then be input into a subsequent stage of the artifact estimation network 302. In this manner, artifacts can be sequentially output and removed until the last type of artifact is estimated or a final artifact-reduced image in which various types of artifacts have been removed or reduced is obtained.
[0058] The artifact estimation network training system 300 may include a validator 320 that validates the performance of the artifact estimation network 302 against the test image pairs 308. The validator 320 may take as input the partially trained artifact estimation network 302 and a dataset of test image pairs 308, and may output an assessment of the performance of the partially trained artifact estimation network 302 based on the dataset of test image pairs 308.
[0059] Once the artifact estimation network 302 has been validated, the trained artifact estimation network 322 (e.g., the validated artifact estimation network 302) may be used to generate a set of artifact images 334 from a set of acquired MR images 332. That is, for each MR image 332, the trained artifact estimation network 322 may output one or more corresponding artifact images at one or more stages of the trained artifact estimation network 322 that include estimated artifacts extracted from the MR image 332 (e.g., and without anatomical image data of the subject). The MR image 332 may include local and global artifacts. For example, the MR image 332 may be acquired by an MR imaging device 330, which may be a Figure 2 The trained artifact estimation network 322 may be stored in an inference module 321 of an image processing system (e.g., Figure 2 within the inference module 212).
[0060] For each MR image in the MR images 332 , one or more artifact images 334 output by the trained artifact estimation network 322 may then be subtracted from the corresponding input MR image 332 to generate an artifact-reduced image 338 , where the artifact-reduced image 338 is a version of the MR image 332 with artifacts reduced or removed.
[0061] The artifact estimation network 302 and the trained artifact estimation network 322 may be residual CNNs, wherein the layers of the artifact estimation network 302 may be grouped into residual blocks, and skip connections are used to propagate input data around one or more residual blocks. Compared to conventional residual networks, the artifact estimation network may have a cascaded multi-stage network architecture, as described below with reference to Figures 6 to 8 As stated.
[0062] Now refer to Figure 6 , shows a high-level architecture diagram of a residual neural network 600, which can be an artifact estimation network described herein. The residual neural network 600 can be used to extract the artifacts from the image data generated by, for example, Figure 1 The residual neural network 600 can be used in a MRI system such as an MRI apparatus 10 to estimate artifacts in an MR image acquired by an MR imaging system such as an MRI apparatus 10. Figure 3 The artifact estimation network training system 300 and other artifact estimation network training systems are trained.
[0063] The residual neural network 600 may be a cascaded residual neural network including a first stage 650 and a second stage 652, wherein each of the first stage 650 and the second stage 652 may have an input layer, a plurality of convolutional layers, and a level output layer. Figure 6 Two stages are shown in FIG6 , but in other embodiments, the residual neural network 600 may include one or more additional stages. Each of the first stage 650, the second stage 652, and any additional stages may focus on estimating artifacts of a particular scale from the input image. For example, the first stage 650 may detect and reduce local artifacts such as noise, ringing artifacts, etc. from the input image. The second stage 652 may detect and reduce global artifacts such as streak artifacts, motion artifacts, etc. from the input image. In other examples, the first stage may focus on estimating and extracting noise; the second stage may focus on estimating and extracting ringing artifacts; the third stage may focus on estimating and extracting streak artifacts; and the fourth stage may focus on estimating and extracting motion artifacts. In still other examples, different or additional stages may be included in the residual network 600.
[0064] If the embodiment of the residual neural network 600 includes additional stages, the artifact types can be classified based on scale, and each of the three stages can focus on detecting and reducing artifact data at a different scale from the input image. For example, the first stage can estimate and reduce the most localized type of artifact (e.g., noise); the second stage can estimate and reduce the next most localized type of artifact (e.g., ringing artifact); the third stage can estimate and reduce the next most localized type of artifact (e.g., streak artifact); and so on, until the final stage estimates and reduces the most global type of artifact. At each stage, the residual neural network 600 estimates and removes noise or artifacts at the corresponding scale while leaving as much image data as possible related to higher-level artifacts, so that subsequent stages can be more efficiently trained to estimate higher-level artifacts.
[0065] For example, in the depicted embodiment, the first stage 650 can reduce noise and ringing artifacts in medical images while removing little or no artifact data corresponding to streak artifacts and motion artifacts. Due to the removal of noise and ringing artifact data, the second stage 652 can achieve higher performance in detecting and reducing streak artifacts and motion artifacts. In contrast, due to the presence of noise and ringing artifacts in medical images, conventional residual neural networks may result in lower performance in detecting noisy and ringing artifact data.
[0066] During training of the residual neural network 600, an input image 602 of an image pair (e.g., the training image pair 306) may be input to a first set of convolutional layers 604 of the first stage 650. The input image 602 may be a medical image, such as an MR image, including anatomical features of the subject of the input image 602. The image data of the input image 602 may be propagated through the first set of convolutional layers 604 to one or more stage 1 output layers of the first set of convolutional layers 604. The stage 1 output layer may include a plurality of stage 1 output nodes, wherein the output of each stage 1 output node may represent a feature map of artifact data. Each artifact may be represented by a single stage 1 output layer. That is, each feature map may include artifact data of a particular type of artifact corresponding to the scale of the artifact data processed by the first stage 650. For example, the first level 1 output layer may output a first set of feature maps 606 comprising artifact data corresponding to noise (e.g., one feature map for each node of the first level 1 output layer); and the second level 1 output layer may output a second set of feature maps 608 comprising artifact data corresponding to ringing artifacts (e.g., one feature map for each node of the second level 1 output layer).
[0067] A first-level artifact-specific image can be generated from one or more feature maps of one or more level 1 output layers, each feature map representing a local artifact type. Thus, the first-level artifact-specific image can include artifact data corresponding to local artifacts (e.g., noise and ringing artifacts) of the input image 602, but not image data for anatomical features of the input image 602 or image data for other types of global artifacts. The one or more first-level artifact-specific images can then be subtracted from the input image 602 (e.g., the pixel intensity value of each pixel of the level-specific artifact image can be subtracted from the pixel intensity value of the corresponding pixel of the input image 602) to generate a partially cleaned image 610, where the partially cleaned image 610 can be a version of the input image 602 in which local artifacts have been reduced. However, other types of artifacts, such as global artifacts, may be present in the partially cleaned image 610.
[0068] The partially cleaned image 610 can then be input to a second set of convolutional layers 612 of the second stage 652. The second set of convolutional layers 612 can be configured differently than the first set of convolutional layers 604, and specifically configured to estimate global artifact features. For example, the nodes of the second set of convolutional layers 612 can be configured to have a larger receptive field than the nodes of the first set of convolutional layers 604. The image data of the partially cleaned image 610 can be propagated through the second set of convolutional layers 612 to one or more stage 2 output layers of the second set of convolutional layers 612. The stage 2 output layers can each include a plurality of stage 2 output nodes. The output of each stage 2 output node can represent a feature map of artifact data corresponding to a particular type of artifact at a scale of the artifact data processed by the second stage 652. Each artifact can be represented by a single stage 2 output layer. For example, the first level 2 output layer may output a first set of feature maps 614 including artifact data corresponding to streaks in the partially clean image 610; the second level 2 output layer may output a second set of feature maps 616 including artifact data corresponding to motion artifacts in the partially clean image 610; the third level 2 output layer may include artifact data corresponding to a third type of global artifact; and so on.
[0069] The feature maps of each stage 2 output node may then be added together to create a global artifact image, wherein the global artifact image includes global artifact data for the input image 602 (and the partially clean image 610) without including local artifact image data or image data of anatomical features of the input image 602. The global artifact data of the global artifact image may include artifact data for different types of global artifacts estimated at each stage 2 output node of the stage 2 output layer, and may not include artifact data for other types of artifacts. The global artifact image may be subtracted from the partially clean image 610 to generate an artifact-reduced image 620, wherein the artifact-reduced image 620 may be a version of the input image 602 in which both global and local artifacts have been reduced.
[0070] In various embodiments, the global artifact image can be an output of the residual neural network 600 comprising a first 2D matrix of values, wherein each value of the first 2D matrix of values corresponds to a pixel of the input image 602. The artifact-reduced image 620 is then generated by subtracting the first 2D matrix of values from a second 2D matrix of values corresponding to different intensities of pixels of the partially clean image 610. In other embodiments, the artifact-reduced image 620 can be an output of the residual neural network 600, wherein the artifact-reduced image 620 is a second 2D matrix of values, wherein each value corresponds to a pixel of the input image 602, wherein the different intensities of each pixel of the artifact-reduced image 620 generate a reconstruction of the input image 602, wherein the amount of artifacts in one or more regions of the artifact-reduced image 620 is lower than the amount of artifacts in one or more regions of the input image 602.
[0071] The global artifact image can then be compared to a target image of the image pair, which can be a true artifact image. A loss function can be used to calculate the difference, or loss, between the global artifact image and the true artifact image. This loss can then be backpropagated first through the second stage 652 and then through the first stage 650. As the loss is backpropagated, the parameters of the nodes of each convolutional layer of the first set of convolutional layers 604 and the second set of convolutional layers 612 can be adjusted using techniques known in the art.
[0072] Additionally, in some embodiments, different weights can be assigned to different types of artifacts to preferentially improve or reduce the relative performance of the residual network 600 for each different type of artifact. That is, before the feature maps of each output node of the first set of convolutional layers 604 are summed to create a combined local artifact image and subtracted from the input image 602 to generate the partially clean image 610, the image data included in each feature map can be multiplied by a weight value to increase or decrease the feature map's relative contribution to the local artifact image. Similarly, before the feature maps of each stage-2 output node of the second set of convolutional layers 612 are summed to create a combined global artifact image and subtracted from the partially clean image 610 to generate the artifact-reduced image 620, the image data included in each stage-2 feature map can be multiplied by a weight value to increase or decrease the stage-2 feature map's relative contribution to the global artifact image. For example, if a user prefers images with a medium signal-to-noise ratio (SNR), the first set of feature maps corresponding to noise in the image can be assigned a lower weight than the second set of feature maps corresponding to streak artifacts. In this way, the effectiveness of artifact removal can be tuned to meet the user's preferences.
[0073] After training, during a subsequent inference stage, the trained residual neural network 600 can be used to extract information from new medical images (such as those obtained by an imaging device (e.g., Figure 3 The MR imaging device 330) reduces or removes artifacts of various types and sizes in images generated during an examination of a patient. Figure 5 The method of FIG. 6 describes in more detail the application of the trained residual neural network 600 during the inference stage.
[0074] Figure 7A first architectural diagram of a stage 700 of a residual neural network 600 is shown. Stage 700 may correspond to either or both of the first stage 650 and / or the second stage 652 of the residual neural network 600. As described above, stage 700 may have an input layer, multiple convolutional layers, and an output layer. In the input layer, an input image 602 may be input to stage 700 and mapped to a set of features. Stage 700 may include a series of mappings from the input image 602 to multiple iteration images (e.g., trainable blocks). For example, the input image 602 may be mapped to a first set of iteration images 704. The first set of iteration images 704 may be mapped to a second set of iteration images 706 and a third set of iteration images 708. Each of the first set of iteration images 704, the second set of iteration images 706, and the third set of iteration images 708 may have a convolution kernel of a different size.
[0075] The second set of iterative images 706 may also be mapped to a third set of iterative images 708, as indicated by solid arrows 720. The third set of iterative images 708 may be mapped to an output image 710. In some examples, the images may also have residual connections, as indicated by dashed lines in the figure. For example, a first residual connection 712 may exist between the input image 602 and the output image 710, a second residual connection 714 may exist between each iterative image in the first set of iterative images 704, a third residual connection 716 may exist between each iterative image in the second set of iterative images 706, and a fourth residual connection 718 may exist between each iterative image in the third set of iterative images 708. Each iterative image may be connected to the previous image and the next image, where each iterative image receives input from the previous iterative image and transforms / maps the received input to an output to produce the next iterative image. In some examples, the convolution kernel size may increase or decrease between iterative images, while in other examples, the convolution kernel size may remain constant between iterative images.
[0076] Output image 710 may be an estimated image based on input image 602 and multiple iterated images based on the input and output layers of stage 700. Output image 710 may approximate a reference image based on the training of stage 700, where a reference image is a target image of a pair of training images. Thus, stage 700 illustrates a mapping of the transformations that occur as input image 602 propagates through the layers of the network.
[0077] Figure 8A second architectural diagram 800 of stage 700 of the residual neural network 600 is shown, wherein the input image 602 is shown mapped to a set of features, as described above. The second architectural diagram 800 illustrates the mapping of transformations that occur as the input image 602 propagates through the layers of stage 700 and the plurality of iteration images 706. Each iteration image in the iteration images 706 may be connected to a previous image and a next image, wherein each iteration image receives input from the previous iteration image and transforms / maps the received input to an output to produce the next iteration image. Furthermore, residual connections 812 may exist between non-adjacent iteration images. The nth iteration image may be mapped to an output image 808 (e.g., partially cleaned image 610 or artifact-reduced image 620). The output image 808 may be an estimated image based on the input image 602 and the plurality of iteration images 706 based on the input and output layers of the stage 700. The output image 808 may approximate a reference image 810 based on the training of the stage 700, where the reference image 810 is a target image of a pair of training images.
[0078] Now turn Figure 4 , shows a flow chart illustrating a method 400 for training an artifact estimation network. The artifact estimation network may be Figure 3 The artifact estimation network training system 300 includes the artifact estimation network 302 and / or Figures 6 to 8 Non-limiting example of a residual neural network 600. The method 400 may be performed by a method such as Figure 2 In one embodiment, some operations of method 400 may be stored in non-transitory memory of an image processing system (e.g., stored in a training module such as training module 210 of image processing system 202) and executed by a processor of the image processing system (e.g., processor 204 of image processing system 202).
[0079] Method 400 begins at 402, where method 400 includes receiving an MR image for training an artifact estimation network. In various embodiments, the MR image may be stored in a medical image dataset (such as a Figure 2 It should be understood that although method 400 and other methods included in the present disclosure are described herein with respect to MRI imaging and MR imaging, method 400 and other methods may be applied to other imaging modalities without departing from the scope of the present disclosure.
[0080] At 404, method 400 includes generating a dataset of training image pairs from the received images. An artifact estimation network may be trained based on the training data including the sets of image pairs. Each image pair in the sets of image pairs may include a set of target artifact images (e.g., Figure 3The present invention provides an artifact image 314) and a noisy MR image (e.g., noisy MR image 316) as an input image, wherein the noisy MR image is a combination of a high-quality MR image with little or no artifacts (e.g., high-quality MR image 312) and an artifact image. As described above, the artifact image may be a combination of various artifact images including different types of artifacts.
[0081] At 406, method 400 includes training an artifact estimation network based on the training pairs. More specifically, training the artifact estimation network based on the image pairs includes training the artifact estimation network to learn to extract artifacts from noisy MR images. The artifact estimation network may include one or more convolutional layers, which in turn include one or more convolutional filters (e.g., a convolutional neural network architecture). Additionally, the artifact estimation network may be a cascaded residual network having multiple stages, as described above with reference to Figure 6 In some embodiments, the artifact estimation network may include a residual neural network having a U-net architecture.
[0082] The convolutional filters of the artifact estimation network may include a plurality of weights, wherein the values of these weights are learned during a training procedure. The convolutional filters may correspond to one or more visual features / patterns, thereby enabling the artifact estimation network to recognize and extract features from medical images.
[0083] Training an artifact estimation network based on image pairs may include iteratively inputting the input image of each training image pair into an input layer of the artifact estimation network. In some embodiments, each pixel intensity value of the input image may be input into a different neuron of the input layer of the artifact estimation network. The artifact estimation network input image may be propagated from the input layer through multiple hidden layers to the output layer of the artifact estimation network. The multiple hidden layers may be organized into multiple residual blocks, wherein each residual block includes one or more convolutional layers, and residual (e.g., jump) connections are used to bypass one or more residual blocks during the forward propagation phase or the backward propagation phase of training.
[0084] In various embodiments, the output of the artifact estimation network may be a residual image that includes the extracted artifact data and does not include features of the input image. The image data of the residual image may be subtracted from the input image to generate an artifact-reduced image (as described above with reference to FIG. Figure 3
[0014] The artifact estimation network may further output an artifact image generated at each stage of the artifact estimation network, and each artifact image may be subtracted from the input image to generate the artifact-reduced image.
[0085] The artifact estimation network can be configured to iteratively adjust one or more of a plurality of weights of the artifact estimation network during backpropagation based on an evaluation of the difference between the input image and the target image included in each image pair of the training image pairs so as to minimize a loss function. In one embodiment, the loss function is a mean absolute error (MAE) loss function, in which the difference between the input image and the target image is compared and summed on a pixel-by-pixel basis. In another embodiment, the loss function can be a structural similarity index (SSIM) loss function. In other embodiments, the loss function can be a minimum loss function or a Wasserstein loss function. It should be understood that the examples provided herein are for illustrative purposes and that other types of loss functions may be used without departing from the scope of this disclosure.
[0086] The weights and biases of the artifact estimation network can be adjusted based on the difference between the output image of the relevant image pair and the target (e.g., true) image. The difference (or loss), as determined by the loss function, can be backpropagated through the neural learning network to update the weights (and biases) of the convolutional layers. The loss can be backpropagated through each stage of the artifact estimation network in the reverse order of the stages. The losses for each stage can also be combined into a single backpropagation to jointly optimize the network. Backpropagation can also be a combination of the two methods described above, for example, by first backpropagating the joint loss and then backpropagating each loss separately. For example, at the end of the first stage, a first loss can be calculated based on the first set of target artifact images for the training pair, and at the end of the second stage, a second loss can be calculated based on the second set of target artifact images for the training pair. In some embodiments, backpropagation of each loss or the joint loss can occur according to a gradient descent algorithm, where the gradient (first-order derivative or an approximation of the first-order derivative) of the loss function is determined for each weight and bias of the deep neural network. Each weight (and bias) of the artifact estimation network is then updated by adding the negative of the product of the gradient determined (or approximated) for the weight (or bias) with a predetermined step size. The weight and bias updates may be repeated until the weights and biases of the artifact estimation network converge, or the rate of change of the weights and / or biases of the deep neural network is below a threshold for each iteration of weight adjustment.
[0087] To avoid overfitting, training of the artifact estimation network can be periodically interrupted to verify the performance of the artifact estimation network relative to the test image pairs. In one embodiment, training of the artifact estimation network can be terminated when the performance of the artifact estimation network relative to the test image pairs converges (e.g., when the error rate on the test set converges to or within a threshold of a minimum value). In this way, the artifact estimation network can be trained to extract artifact data from the input image.
[0088] In some embodiments, the evaluation of the performance of the artifact estimation network can include a combination of a minimum error rate and a quality assessment, or a different function of the minimum error rate achieved on each of the test image pairs and / or one or more quality assessments, or another factor for evaluating the performance of the artifact estimation network. In other examples, other loss functions, error rates, quality assessments, and / or performance assessments can be used during training.
[0089] Now refer to Figure 5 , shows a method for deploying an artifact estimation network such as Figure 3 The artifact estimation network 302 and / or Figures 6 to 8 Flowchart of a method 500 for estimating and reducing the amount of different types of artifacts in medical images (such as MR images) using a residual neural network 600. The method 500 may be performed by, for example, Figure 2 Some operations of method 500 may be stored in a non-transitory memory of an image processing system (e.g., in inference module 212) and executed by a processor of the image processing system (e.g., processor 204). In various embodiments, the artifact estimation network may be as described above with reference to Figure 4 Training is performed as described in method 400.
[0090] Method 500 begins at 502, where method 500 includes receiving MR imaging data acquired from a scanned subject. As an example, the MR imaging data may include an MR image of the subject. For example, a MR image may be obtained using a MR image such as Figure 3 MR images are acquired using an MRI device, such as an MRI imaging device 330, or the scanner 236 of the image processing system 202. The acquired MR images may have the same region of interest and / or may include the same set of anatomical structures as the training image set on which the artifact estimation network is trained. In some embodiments, the subject of the acquired MR images may be similar to the subject of the training image set. In some examples, multiple artifact estimation networks may be trained based on different types of subjects or different anatomical structures of the subjects, and an artifact estimation network may be selected from the multiple artifact estimation networks based on the same set of anatomical structures and / or subject type. For example, the subject may be a child, in which case the acquired MR images may be input to a first artifact estimation network trained based on RGB images of the child; the subject may be a female, in which case the acquired MR images may be input to a second artifact estimation network trained based on reference images of the female subject; the subject may be a male, in which case the acquired MR images may be input to a third artifact estimation network trained based on reference images of the male subject; and so on.
[0091] At 504, the acquired MR imaging data is input into a trained artifact estimation network. In various embodiments, inputting the acquired MR imaging data into the trained artifact estimation network includes inputting image data for each pixel of the acquired MR image into a corresponding node of an input layer of the artifact estimation network. The value of the image data may be multiplied by a weight at the corresponding node and propagated through various hidden layers (e.g., convolutional layers) of each residual block of each stage of the artifact estimation network to a final output layer of the artifact estimation network. The output layer may include a node corresponding to each pixel of an output image, wherein the output image is based on the image data output by each node. At 506, method 500 includes receiving an output image from the artifact estimation network. In various embodiments, the output image may be an artifact image including artifacts extracted from the acquired MR (e.g., input) image. In other embodiments, the output image may be an artifact-reduced image having fewer artifacts than the input image, wherein the amount of artifacts in the input image is reduced or removed by the trained artifact estimation network.
[0092] At 507, method 500 optionally includes subtracting the output image from the input image to generate an artifact-reduced image. As described above, in some examples, the artifact estimation network can be trained to output an artifact image, wherein the artifact image includes image data of artifacts of the input image but does not include image data of anatomical structures and / or features of the input image. In this case, the image data of the artifacts can be subtracted from the image data of the input image, and the image data of the anatomical structures and / or features of the input image remains in the artifact-reduced image. In some examples, an artifact image can be output by each stage of the artifact estimation network, and the artifact image can be subtracted from the input image.
[0093] As mentioned above Figure 3 and Figure 4 As described above, a partially clean image may be generated at each stage of the trained artifact estimation network, wherein the partially clean image may be a version of the input image with artifacts of a given scale (and / or artifacts of a specific type) reduced or removed by subtracting one or more stage-specific artifact images output by the artifact estimation network. Figure 6As described above, in some scenarios, tuning parameters can be applied to the image data of the stage-specific artifact image generated at each stage to increase or decrease the contribution of that image data to subsequent stages of the trained artifact estimation network, thereby adjusting the amount of artifact type reduction associated with that stage-specific artifact image. For example, a user may wish to extract certain artifact types while leaving some noise in the resulting artifact-reduced image. To achieve this, the user can specify a tuning parameter between 0 and 1 to be multiplied by the intensity value of each pixel in the stage-specific artifact image before subtracting the stage-specific artifact image from the image input to the associated stage. For example, the tuning parameters can be set by the user in a preference file of the image processing system, or the image processing system can assign the tuning parameters based on a desired artifact profile specified by the user in a different manner. In this way, the tuning parameters can be used to weight different artifact types in the artifact image to adjust the overall level of artifact reduction / removal to suit different user preferences.
[0094] In some examples, the artifact weighting scheme may be submitted to the image processing system by a user via a selection from a menu displayed on a display screen, e.g., within a software application used to display the input image and / or the artifact-reduced image on the display screen. For example, a user may select a menu item for adjusting the weights of different types of artifacts desired to be removed. A lookup table may be used to determine one or more nodes of the artifact estimation network that generate a feature map associated with the selected artifact type, and the artifact weighting scheme may be applied to the node. In one example, the artifact weighting scheme may be applied when a stage-specific artifact image is subtracted from the input image at the relevant stage (e.g., stage 650, 652) of the artifact estimation network according to the following equation:
[0095]
[0096] Among them I final is an artifact-reduced image of the relevant level (eg, partially cleaned image 610), I input is the input image into the relevant stage, and w i is the i-th predicted artifact Res i The preferred weight of .
[0097] For example, a trained artifact estimation network can estimate two artifacts (noise and ringing) in the first stage and estimate streak artifacts in the second stage. That is, a first artifact image can be generated at the first output layer of the first stage, which can be subtracted from the input image to reduce noise artifacts, and a second artifact image can be generated at the second output layer of the first stage, which can be subtracted from the input image to reduce ringing artifacts. A user may prefer to remove all ringing and streak artifacts, while only removing 75% of the noise. To achieve this, a first tuning parameter of 0.75 can be multiplied by the intensity value of each pixel of the first artifact image (noise) generated at the first output layer of the first stage. A second tuning parameter of 1.0 can be multiplied by the intensity value of each pixel of the second artifact image (ringing) generated at the second output layer of the first stage. A third tuning parameter of 1.0 can be multiplied by the intensity value of each pixel of the third artifact image (streaks) generated at the output layer of the second stage. The first, second, and third artifact images may be output by the trained artifact estimation network and then subtracted from the input image (module 336) to generate an artifact-reduced image (e.g., artifact-reduced image 338), where 25% of the noise remains in the artifact-reduced image and ringing and streak artifacts are removed. The first and second artifact images may be fully subtracted from the input image to generate a partially clean image (e.g., partially clean image 610), where 100% of the noise and 100% of the ringing are removed. The partially clean image may then be propagated to the second stage.
[0098] It should be understood that the weighting of the artifact image and the creation of the partially clean image can be independent processes. That is, the artifact image can be subtracted from the input image to generate a partially clean image to be propagated to the next stage. The artifact image can be output by the network and separately weighted so as to be subtracted from the input image along with the artifact image generated at a subsequent stage of the network.
[0099] Brief Reference Figure 11, shows examples of partially cleaned images and artifact-reduced images generated as described above. An example input image 1100 can be input into a trained artifact estimation network. A partially cleaned image 1102 can be generated by removing artifacts predicted by the first stage of the trained artifact estimation network, and an artifact-reduced image 1104 can be generated by further removing artifacts estimated by the second stage of the trained artifact estimation network, as described above. Input image 1100 exhibits various artifacts, including noise, ringing, and streak artifacts. These artifacts are particularly visible in region 1101 of input MR image 1100. An expanded view of region 1101 is shown below input MR image 1110. As described above, the first stage estimates and reduces noise and ringing in input MR image 1110. Consequently, partially cleaned image 1102 may include streak artifacts, as seen in expanded view 1112 of corresponding region 1103 of partially cleaned image 1102. The second stage estimates and reduces streak artifacts in input MR image 1110. Thus, the artifact-reduced image 1104 may not include noise, ringing, and streak artifacts, as seen in the expanded view 1114 of the corresponding region 1105 of the artifact-reduced image 1104 .
[0100] At 508, method 500 includes displaying a display screen (e.g., Figure 2 The artifact-reduced image output by the trained artifact estimation network is displayed on a display device 234 of the image processing system. During an examination of the subject, the artifact-reduced image can be displayed in real time on the display screen so that an operator of the image processing system (e.g., a caregiver) can view the artifact-reduced image during the examination. The artifact-reduced image can also be output to a storage device or a picture archiving and communication system (PACS) for subsequent retrieval and / or remote viewing. For example, the artifact-reduced image can be used to diagnose a condition in the subject. By reducing the amount of artifacts in the acquired MR images, the caregiver can more clearly see the subject's anatomical features, thereby making it easier to diagnose the condition.
[0101] Figure 9 shows the use of a trained artifact estimation network such as Figure 3 The artifact estimation network 302 and / or Figures 6 to 8 The example artifact-reduced MR image 904 generated by the residual neural network 600 of FIG. 1 is compared with the input MR image 900 used to generate the artifact-reduced MR image 904 and the ... Figures 6 to 8904. The artifact-reduced MR image 904 is compared to a second exemplary artifact-reduced MR image 902 generated by a conventional (e.g., non-cascaded) residual neural network with the described architecture. The artifact-reduced MR image 904 shows fewer artifacts than the input MR image 900. For example, streak artifacts can be seen in portion 910 of the input MR image 900. In the artifact-reduced MR image 904, the streak artifacts have been largely eliminated in portion 910. Additionally, although the streak artifacts are reduced in the second exemplary artifact-reduced MR image 902, the streak artifacts are more noticeable in the second exemplary artifact-reduced MR image 902 than in the artifact-reduced MR image 904. Thus, in contrast to conventional residual neural network architectures, by using a CNN with Figures 6 to 8 The CNN architecture described in
[15] can remove a larger amount of artifacts.
[0102] Figure 10 A second example is shown in FIG. 1 , which shows an artifact-reduced MR image 1004 generated using a trained artifact estimation network, compared to an input MR image 1000 used to generate the artifact-reduced MR image 1004 and using a trained MR image without the above reference. Figures 6 to 8 A second exemplary artifact-reduced MR image 1002 generated by a conventional (e.g., non-cascaded) residual neural network using the described architecture is compared. Artifact-reduced MR image 1004 shows a reduced amount of both streak artifacts and noise compared to input MR image 1000. For example, streak artifacts can be seen in portion 1014 of input MR image 1000. In artifact-reduced MR image 1004, streak artifacts have been substantially eliminated in portion 1014. Although streak artifacts are reduced in second exemplary artifact-reduced MR image 1002, streak artifacts are more noticeable in second exemplary artifact-reduced MR image 1002 than in artifact-reduced MR image 1004. Additionally, a first amount of noise can be seen in portion 1012 of input MR image 1000. A second, lower amount of noise can be seen in portion 1012 of artifact-reduced MR image 1004, where the noise has been reduced by the trained artifact estimation network. The second exemplary artifact-reduced MR image 1002 shows a third amount of noise at portion 1012, wherein the third amount of noise is less than the first amount of noise of the input MR image 100, but greater than the second amount of noise of the artifact-reduced MR image 1004 generated by a conventional (e.g., non-cascaded) residual neural network.
[0103] Therefore, described herein are systems and methods for improving the performance of artifact reduction neural networks in reducing various types of artifacts in medical images. The disclosed artifact reduction residual network has a multi-stage architecture comprising a plurality of residual blocks organized into cascaded stages, each of which may include multiple convolutional layers. The residual blocks and convolutional layers at each stage are configured and trained to detect and reduce different types of artifacts occurring at different scales in the medical image. The initial stages of the artifact reduction neural network first remove local-scale artifacts. Subsequent stages of the artifact reduction neural network then remove global-scale artifacts. During training, the loss calculated between the artifact-reduced image output by the artifact reduction neural network and the target image is backpropagated through each stage of the multi-stage architecture. By training the artifact reduction neural network in this manner, artifacts occurring at lower scales are successively removed, leaving an image with higher-quality artifact data with artifacts at higher scales, enabling the artifact reduction neural network to more effectively learn to detect artifacts occurring at higher scales. Thus, different types of artifacts can be removed from medical images more effectively than by using conventional residual networks (including when multiple conventional residual networks are chained or trained in parallel). Additionally, the disclosed artifact reduction neural network can be trained on a single dataset that includes a wider range of images with varying degrees of noise, a wider range of signal-to-noise ratios (SNRs), and more types of artifacts than alternative artifact reduction neural networks without the disclosed architecture and training. A technical effect of training and using the disclosed artifact reduction neural network to reduce artifacts in medical images is that the quality of medical images can be improved, thereby enabling more accurate diagnoses and more effective patient treatments.
[0104] The present invention also provides support for an image processing system, the image processing system comprising: a trained artifact estimation network, the trained artifact estimation network comprising a plurality of stages, the artifact estimation network being trained to estimate artifacts in a medical image; and a processor communicatively coupled to a non-transitory memory storing the artifact estimation network, the memory comprising instructions that, when executed, cause the processor to: receive a medical image; generate an estimated artifact image from the medical image using the trained artifact estimation network; generate an artifact-reduced image by subtracting the estimated artifact image from the medical image, the artifact-reduced image being a version of the medical image that includes fewer artifacts than the medical image; and display the artifact-reduced image on a display device; wherein each stage of the trained artifact estimation network estimates artifacts of different scales in the medical image. In a first example of the system, each stage of the artifact estimation network comprises a first plurality of convolutional layers organized into a second plurality of residual blocks, and residual connections are used to bypass one or more convolutional layers within the stage. In a second example of the system, optionally including the first example, the artifact estimation network includes at least a first stage and a second stage, each of the first stage and the second stage including an input layer, a plurality of convolutional layers, and a stage output layer; the first stage estimates and reduces local artifact data from the medical image; and the second stage estimates and reduces global artifact data from the medical image. In a third example of the system, optionally including one or both of the first and second examples, the local artifact data includes noise and ringing artifacts, and the global artifact data includes streak artifacts and motion artifacts. In a fourth example of the system, optionally including one or more or each of the first to third examples, additional instructions are stored in the memory, which when executed cause the processor to: during generation of the artifact-reduced image from the medical image using the trained artifact estimation network: combine the feature maps of each output node of the first level to create a first-level artifact-specific image, the first-level artifact-specific image including local artifact data of the medical image and not including image data of anatomical features of the medical image; and subtract the first-level artifact-specific image from the medical image to generate a partially clean image, which is a version of the medical image in which local artifacts have been reduced. In a fifth example of the system, optionally including one or more or each of the first to fourth examples, additional instructions are stored in the memory, which, when executed, cause the processor to: input the partially clean image into the input layer of the second stage; combine the feature maps of each output node of the second stage to create a second-level specific artifact image, the second-level specific artifact image including global artifact data of the medical image and excluding image data of the anatomical feature of the medical image or the local artifact data; and subtract the second-level specific artifact image from the partially clean image to generate the artifact-reduced image.In a sixth example of the system, optionally including one or more or each of the first to fifth examples, the nodes of the plurality of convolutional layers of the second stage are configured to have a larger receptive field than the nodes of the plurality of convolutional layers of the first stage. In a seventh example of the system, optionally including one or more or each of the first to sixth examples, additional instructions are stored in the memory, which, when executed, cause the processor to: during training of the artifact estimation network: input a noisy medical image into the artifact estimation network, the noisy medical image being a combination of a high-quality medical image and one or more artifact images; backpropagate a loss between an artifact image output by the artifact estimation network and the one or more synthesized artifact images; and adjust parameters of both the first stage and the second stage of the artifact estimation network based on the backpropagated loss. In an eighth example of the system, optionally including one or more or each of the first to seventh examples, wherein the second-level artifact-specific image is an output of the artifact estimation network, the second-level artifact-specific image comprising a first set of 2D matrix of pixel intensity values, each value of the first set of 2D matrix corresponding to a pixel of the medical image. In a ninth example of the system, optionally including one or more or each of the first to eighth examples, the first-level artifact-specific image is an additional output of the artifact estimation network, the first-level artifact-specific image comprising a second set of 2D matrix of values, each value of the second 2D matrix corresponding to a pixel of the medical image, and wherein the artifact-reduced image is generated by subtracting the first set of 2D matrix of values and the second set of 2D matrix of values from the medical image. In a tenth example of the system, optionally including one or more or each of the first to ninth examples, further instructions are stored in the memory that, when executed, cause the processor to: receive relative weight values for different types of artifacts of the medical image from a user of the image processing system; and apply the relative weight values to at least one of the first-level specific artifact image and the second-level specific artifact image to preferentially adjust the amount of the different types of artifacts to be reduced from the medical image. In an eleventh example of the system, optionally including one or more or each of the first to tenth examples, the relative weight values are received via an artifact weighting scheme that is submitted to the image processing system by the user via selection from a menu, the menu being displayed on the display device. In a twelfth example of the system, optionally including one or more or each of the first to eleventh examples, the one or more artifact images are synthesized artifact images. In a thirteenth example of the system, optionally including one or more or each of the first to twelfth examples, wherein the artifact-reduced image is displayed in real time on the display device during examination of a subject of the medical image.
[0105] The present disclosure also provides support for a method for training a residual neural network to reduce the amount of artifacts in a medical image, the method comprising: receiving a set of training image pairs, each training image pair comprising a plurality of real target artifact images, and a noisy medical image comprising a high-quality medical image combined with the plurality of real target artifact images as an input image; inputting the input images of the training image pairs of the set of training image pairs into a first stage of the residual neural network; estimating a first set of local artifact images of the plurality of real target artifact images at the first stage of the residual neural network; subtracting the first set of estimated local artifact images from the noisy medical image to generate a partially clean image, the partially clean image comprising artifacts of a reduced amount at a local scale; inputting the partially clean image into a second stage of the residual neural network; and The invention relates to a method for obtaining a plurality of real target artifact images of an image processing unit (IMU) by using a convolutional neural network (CNN) to obtain a plurality of real target artifact images of the image processing unit (IMU). The method comprises the following steps: estimating a second set of global artifact images of the plurality of real target artifact images at the second stage of the residual neural network; combining the first set of estimated local artifact images with the second set of estimated global artifact images to create a combined artifact image; subtracting the second set of estimated global artifact images from the partially clean image to generate an artifact-reduced image, the artifact-reduced image including artifacts of reduced amounts at both local and global scales; backpropagating a loss between the combined artifact image and the plurality of real target artifact images of the training image pair through the second plurality of convolutional layers of the second stage and the first plurality of convolutional layers of the first stage; and adjusting a first set of parameters at the first plurality of nodes of the first plurality of convolutional layers and a second set of parameters at the second plurality of nodes of the second plurality of convolutional layers based on the backpropagated loss. In a first example of the method, inputting the input image into the first stage of the residual neural network to estimate the first set of local artifact images and generate the partially clean image further includes: propagating the image data of the input image through the first plurality of convolutional layers of the first stage to generate a first plurality of feature maps at corresponding first plurality of output nodes of the first stage, each feature map of the first plurality of feature maps including artifact data of different types of local artifacts, and generating the first set of artifact images based on the first plurality of feature maps. In a second example of the method, optionally including the first example, inputting the partially clean image into the second stage of the residual neural network to estimate the second set of global artifact images further includes: propagating the image data of the partially clean image through the second plurality of convolutional layers of the second stage to generate a second plurality of feature maps at corresponding second plurality of output nodes of the second stage, each feature map of the second plurality of feature maps including artifact data of different types of global artifacts, and generating the second set of global artifact images based on the second plurality of feature maps. In a third example of the method, optionally including one or both of the first and second examples, the nodes of the convolutional layer of the second stage are configured to have a larger receptive field than the nodes of the convolutional layer of the first stage.In a fourth example of the method, optionally including one or more or each of the first to third examples, the first set of local artifact images includes at least one of noise and ringing artifacts, and the second set of global artifact images includes at least one of streak artifacts and motion artifacts.
[0106] The present disclosure also provides support for a residual neural network trained to reduce the amount of artifacts in a medical image, the residual neural network comprising multiple stages, the multiple stages including at least: a first stage that takes the medical image as input and generates a first version of the medical image with a reduced number of artifacts at a first scale; and a second stage that takes the first version of the medical image as input and generates a second version of the medical image with a reduced number of artifacts at both the first scale and the second scale.
[0107] When introducing the elements of the various embodiments of the present disclosure, the articles "one", "a kind of" and "the" are intended to mean that there are one or more such elements. The terms "first", "second" etc. do not represent any order, amount or importance, but are used to distinguish one element from another. The terms "comprise", "comprising" and "having" are intended to be inclusive and mean that in addition to the listed elements, additional elements may also be present. As used herein, the terms "connected to", "coupled to" etc., an object (e.g., a material, element, structure, member, etc.) may be connected to or coupled to another object, regardless of whether the object is directly connected or coupled to another object, or whether there are one or more intervening objects between the object and another object. In addition, it should be understood that reference to "one embodiment" or "embodiment" of the present disclosure is not intended to be interpreted as excluding the existence of additional embodiments that also combine the cited features.
[0108] In addition to any modifications previously indicated, those skilled in the art may devise numerous other variations and alternative arrangements without departing from the spirit and scope of this specification, and the appended claims are intended to cover such modifications and arrangements. Thus, although the information has been described above with particularity and detail in connection with what are presently considered to be the most practical and preferred aspects, it will be apparent to those skilled in the art that many modifications, including but not limited to form, function, mode of operation, and use, may be made without departing from the principles and concepts set forth herein. Likewise, as used herein, the examples and embodiments are intended to be illustrative in all respects only and should not be construed as limiting in any way.
Claims
1. An image processing system (202), comprising: a trained artifact estimation network (302, 209, 322), the trained artifact estimation network comprising a plurality of stages, the artifact estimation network (302, 209) being trained to estimate artifacts in a medical image; and a processor (204) communicatively coupled to a non-transitory memory (206) storing the artifact estimation network (302, 209), the memory comprising instructions that, when executed, cause the processor (204): receiving medical images; generating an estimated artifact image (314) from the medical image using the trained artifact estimation network (302, 209, 322); generating an artifact-reduced image (1104, 338, 620) by subtracting the estimated artifact image (314) from the medical image, the artifact-reduced image (1104, 338, 620) being a version of the medical image that includes a smaller amount of artifacts than the medical image; and displaying the artifact-reduced image (1104, 338, 620) on a display device (234); Each stage (650, 652) of the trained artifact estimation network (302, 209, 322) estimates artifacts of different scales in the medical image.
2. The image processing system (202) of claim 1, wherein each stage (650, 652) of the artifact estimation network (302, 209) comprises a first plurality of convolutional layers (612, 604) organized into a second plurality of residual blocks, and residual connections (812) are used to bypass one or more convolutional layers (612, 604) within the stage (650, 652).
3. The image processing system (202) according to claim 1, wherein: The artifact estimation network (302, 209) includes at least a first stage (650) and a second stage (652), each of the first stage (650) and the second stage (652) including an input layer, a plurality of convolutional layers (612, 604) and a stage output layer; The first stage (650) estimates and reduces local artifact data from the medical image; and The second stage (652) estimates and reduces global artifact data from the medical image.
4. The image processing system (202) of claim 3, wherein the local artifact data comprises noise and ringing artifacts, and the global artifact data comprises streak artifacts and motion artifacts.
5. The image processing system (202) of claim 3, wherein further instructions are stored in the memory, which when executed cause the processor (204) to: During generation of the artifact-reduced image (1104, 338, 620) from the medical image using the trained artifact estimation network (302, 209, 322): combining the feature maps (614, 606, 608, 616) of each output node of the first stage (650) to create a first stage specific artifact image (314), the first stage specific artifact image (314) including local artifact data of the medical image and excluding image data of anatomical features of the medical image; and The first level specific artifact image (314) is subtracted from the medical image to generate a partially clean image (610, 1102), the partially clean image (610, 1102) being a version of the medical image in which local artifacts have been reduced.
6. The image processing system (202) of claim 5, wherein further instructions are stored in the memory, which when executed cause the processor (204) to: inputting the partially cleaned image (610, 1102) into the input layer of the second stage (652); combining the feature maps (614, 606, 608, 616) of each output node of the second stage (652) to create a second stage specific artifact image (314), the second stage specific artifact image (314) including global artifact data of the medical image and not including image data of the anatomical feature of the medical image or the local artifact data; and The second level specific artifact image (314) is subtracted from the partially clean image (610, 1102) to generate the artifact-reduced image (1104, 338, 620).
7. The image processing system (202) of claim 3, wherein the nodes of the plurality of convolutional layers (612, 604) of the second stage (652) are configured to have a larger receptive field than the nodes of the plurality of convolutional layers (612, 604) of the first stage (650).
8. The image processing system (202) of claim 3, wherein further instructions are stored in the memory, the further instructions, when executed, causing the processor (204): During training of the artifact estimation network (302, 209): Inputting a noisy medical image into the artifact estimation network (302, 209), the noisy medical image being a combination of a high-quality medical image and one or more artifact images (314, 334); back-propagating a loss between an artifact image (314) output by the artifact estimation network (302, 209) and the one or more synthesized artifact images (314, 334, 314); and Parameters of both the first stage (650) and the second stage (652) of the artifact estimation network (302, 209) are adjusted based on the back-propagated loss.
9. The image processing system (202) according to claim 6, wherein the second-level specific artifact image (314) is the output of the artifact estimation network (302, 209), and the second-level specific artifact image (314) includes a first set of 2D pixel intensity value matrices, and each value of the first set of 2D matrices corresponds to a pixel of the medical image.
10. The image processing system (202) of claim 9, wherein a first level specific artifact image (314) is an additional output of the artifact estimation network (302, 209), the first level specific artifact image (314) comprising a second set of 2D value matrices, each value of the second 2D matrix corresponding to a pixel of the medical image, and wherein the artifact-reduced image (1104, 338, 620) is generated by subtracting the first set of 2D value matrices and the second set of 2D value matrices from the medical image.
11. The image processing system (202) of claim 10, wherein further instructions are stored in the memory, which when executed cause the processor (204) to: receiving relative weight values of different types of artifacts of the medical image from a user of the image processing system (202); and The relative weight values are applied to at least one of the first level specific artifact image (314) and the second level specific artifact image (314) to preferentially adjust the amounts of the different types of artifacts to be reduced from the medical image.
12. The image processing system (202) of claim 11, wherein the relative weight values are received via an artifact weighting scheme, the artifact weighting scheme being submitted to the image processing system (202) by the user via selection from a menu, the menu being displayed on the display device (234).
13. The image processing system (202) of claim 8, wherein the one or more artifact images (314, 334) are synthesized artifact images.
14. The image processing system (202) of claim 1, wherein the artifact-reduced image (1104, 338, 620) is displayed in real time on the display device (234) during examination of a subject (16) of the medical image.
15. A method for training a residual neural network to reduce the amount of artifacts in a medical image, the method comprising: receiving a set of training image pairs (402), each training image pair comprising a plurality of real target artifact images, and a noisy medical image comprising a high-quality medical image combined with the plurality of real target artifact images as an input image; Inputting an input image of a training image pair of the set of training image pairs into a first stage of the residual neural network (406); estimating a first set of local artifact images of the plurality of real object artifact images at the first stage of the residual neural network; subtracting the first set of estimated local artifact images from the noisy medical image to generate a partially clean image, the partially clean image comprising artifacts of reduced local scale; Inputting the partially cleaned image into the second stage of the residual neural network; estimating a second set of global artifact images of the plurality of real object artifact images at the second stage of the residual neural network; combining the first set of estimated local artifact images with the second set of estimated global artifact images to create a combined artifact image; subtracting the second set of estimated global artifact images from the partially clean image to generate an artifact-reduced image, the artifact-reduced image including reduced amounts of artifacts at both local and global scales; back-propagating a loss between the combined artifact image and the plurality of true target artifact images of the training image pairs through a second plurality of convolutional layers of the second stage and a first plurality of convolutional layers of the first stage; as well as A first set of parameters at a first plurality of nodes of the first plurality of convolutional layers and a second set of parameters at a second plurality of nodes of the second plurality of convolutional layers are both adjusted based on the backpropagated loss.