Method and system for converting a magnetic resonance image into a pseudo-computed tomography image
Converting MRI images into pseudo-CT images through deep multitasking neural networks solves the problem that it is difficult to obtain accurate electron density information in the MRI workflow using only, especially in the high dynamic range areas of the bones, achieving efficient and accurate acquisition of in vivo electron density information.
Patent Information
- Application Number
- CN202111016987.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-25
- Filing Date
- 2021-08-31
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-08-31
AI Technical Summary
The prior art has difficulty obtaining accurate in vivo electron density information in vivo intra-body using magnetic resonance imaging (MRI) workflows, especially in high dynamic range areas of bones, resulting in errors in radiation therapy planning and positron emission tomography (PET) imaging.
Electronic density information similar to CT imaging is provided by converting MRI images into pseudo-CT images using deep multitasking neural networks. The neural network is configured with multiple tasks, including the entire image transformation, accurate segmentation of the region of interest, and image value estimation to improve the accuracy of the bone region.
Accurate electron density information of high dynamic range bone areas in MRI-only use-MRI workflows is achieved, reducing errors in radiation therapy planning and PET imaging, and avoiding CT dose exposure in patients.
Smart Images

Figure CN114255289B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the subject matter disclosed herein relate to magnetic resonance imaging and, more particularly, to converting magnetic resonance images into computed tomography-like images. Background Art
[0002] Electron density information within the body is essential for accurate dose calculations in radiotherapy planning and for calculating attenuation correction maps in positron emission tomography (PET) imaging. In conventional radiotherapy planning and PET imaging, computed tomography (CT) images provide the necessary information on the electron density and attenuation characteristics of tissues. Specifically, CT imaging enables the simultaneous and accurate depiction of internal anatomical structures such as bone, soft tissue, and blood vessels. Summary of the Invention
[0003] In one embodiment, a method includes: acquiring a magnetic resonance (MR) image; generating a pseudo-CT image corresponding to the MR image using a multi-task neural network; and outputting the MR image and the pseudo-CT image. In this way, the benefits of CT imaging with respect to accurate density information (especially in sparse regions of bone that exhibit a high dynamic range) can be obtained in an MR-only workflow, thus realizing the benefits of enhanced soft tissue contrast in the MR image while eliminating the patient's CT dose exposure.
[0004] It should be understood that the above summary is provided to introduce in a simplified form selected concepts that are further described in the detailed description. This is not meant to identify key or essential features of the claimed subject matter, the scope of which is uniquely defined by the claims that follow the detailed description. Moreover, the claimed subject matter is not limited to embodiments that solve any disadvantages noted above or in any part of this disclosure. Brief Description of the Drawings
[0005] The present disclosure will be better understood by reading the following description of non-limiting embodiments with reference to the accompanying drawings, in which:
[0006] Figure 1 is a block diagram of an MRI system according to an embodiment of the present disclosure;
[0007] Figure 2 is a schematic diagram of an image processing system for converting an MR image into a pseudo-CT image according to an embodiment of the present disclosure;
[0008] Figure 3 is a schematic diagram showing the layout of an embodiment of a multi-task neural network for converting an MR image into a pseudo-CT image according to an embodiment of the present disclosure;
[0009] Figure 4is a schematic diagram showing the layout of a deep multi-task neural network that can be used in an Figure 2 image processing system;
[0010] Figure 5 is a high-level flowchart showing an exemplary method for training a deep multi-task neural network to generate a pseudo CT image from an MR image with accuracy in a focused region of interest;
[0011] Figure 6 is a high-level flowchart showing an exemplary method for using a deep multi-task neural network to generate a pseudo CT image from an MR image;
[0012] Figure 7 shows a set of images showing exemplary pseudo CT images generated according to different techniques compared to an input MR image and a ground truth CT image;
[0013] Figure 8 shows a set of graphs showing normalized histograms of soft tissue and bone regions of pseudo CT and CT images for multiple cases; and
[0014] Figure 9 shows a set of graphs showing graphs of Dice coefficients of pseudo CT bone regions at different bone density thresholds for multiple cases. DETAILED DESCRIPTION
[0015] The following description relates to various embodiments for converting MR images into pseudo CT or CT-like images. Electron density information in the body is essential for accurate dose calculation in radiotherapy (RT) planning and for calculating attenuation correction maps in positron emission tomography (PET) imaging. In traditional RT treatment planning and PET / CT imaging, CT images provide the necessary information on the electron density and attenuation characteristics of tissues. However, there is an increasing trend to use only MR clinical workflows to take advantage of the enhanced soft tissue contrast in MR images. To replace CT images, it is necessary to infer density maps for RT dose calculation and PET / MR attenuation correction from MRI. One way to replace CT using MRI can include mapping the MR image to a corresponding CT image to provide CT Hounsfield unit (HU) values as a pseudo CT (pCT) image. In this way, it is possible to use only an MRI system (such as Figure 1The depicted MR device) to obtain certain benefits of CT imaging. However, in CT images, bone values can range from 250 Hu to over 2000 HU, while only occupying a portion of the body area. Therefore, previous methods for generating pCT images based on machine learning models tend to be biased towards training the spatial principal values from soft tissue and background regions, for example, due to the large dynamic range and spatial sparsity of the bone region, resulting in reduced accuracy within the bone region. Additionally, the bone region becomes sparser at higher densities and contributes even less to network optimization. Inaccurate bone value assignment in the pseudo-CT image can lead to a series of errors, for example, in dose calculation for RT treatment planning. To better utilize MRI for RT treatment planning, an image processing system (such as Figure 2 The depicted image processing system) may include a deep multi-task neural network module configured to generate a pCT image with accurate bone value assignment. Specifically, the deep multi-task neural network module takes an MR image as input and outputs a pCT image with accurate HU value assignment across different densities and tissue classes. In one example, to increase the accuracy of bone estimation in synthetic CT, the tasks of whole image transformation, accurate segmentation of the region of interest, and image value estimation within the region of interest are assigned to the multi-task neural network. For example, as Figure 3 Depicted, the multi-task neural network thus outputs a pseudo-CT image, a bone mask, and a bone HU image or bone density map for three corresponding tasks. The multi-task neural network can be implemented as a two-dimensional U-Net convolutional neural network with multiple output layers, for example, as Figure 4 Depicted. The multi-task neural network can be trained, for example, according to a training method (such as Figure 5 The depicted method) to perform multiple tasks simultaneously, such that the related tasks improve the normalization of the network. A method for implementing such a multi-task neural network after training (such as Figure 6 The depicted method) may include generating a pseudo-CT image from the MR image and updating the pseudo-CT image with the bone HU image also generated by the multi-task neural network. By constructing a neural network as described herein and training the neural network to perform multiple related tasks, the accuracy of pCT generation is improved relative to other methods and even other deep neural network-based methods, as Figures 7 to 9 Depicted in the qualitative and quantitative comparisons in
[0016] Now turning to the pictures, Figure 1An MRI (Magnetic Resonance Imaging) apparatus 10 is shown, which includes a static magnetic field magnet unit 12, a gradient coil unit 13, an RF coil unit 14, an RF body or volume coil unit 15, a transmit / receive (T / R) switch 20, an RF driver unit 22, a gradient coil driver unit 23, a data acquisition unit 24, a controller unit 25, a patient examination table or bed 26, a data processing unit 31, an operation console unit 32, and a display unit 33. In some embodiments, the RF coil unit 14 is a surface coil, which is a local coil typically placed near the anatomical structure of interest of the subject 16. Here, the RF body coil unit 15 is a transmit coil for transmitting RF signals, and the local surface RF coil unit 14 receives MR signals. Thus, the transmit body coil (e.g., the RF body coil unit 15) and the surface receive coil (e.g., the RF coil unit 14) are independent but electromagnetically coupled components. The MRI apparatus 10 transmits electromagnetic pulse signals to the subject 16 placed in the imaging space 18, where a static magnetic field is formed to perform a scan to obtain magnetic resonance signals from the subject 16. One or more images of the subject 16 can be reconstructed based on the magnetic resonance signals obtained by the scan.
[0017] The static magnetic field magnet unit 12 includes, for example, a toroidal superconducting magnet installed in a toroidal vacuum vessel. The magnet defines a cylindrical space around the subject 16 and generates a constant main static magnetic field B0.
[0018] The MRI apparatus 10 further includes a gradient coil unit 13 that forms a gradient magnetic field in the imaging space 18 to provide three-dimensional position information for the magnetic resonance signals received by an RF coil array (e.g., the RF coil unit 14 and / or the RF body coil unit 15). The gradient coil unit 13 includes three gradient coil systems, each of which generates a gradient magnetic field along one of three mutually perpendicular spatial axes and generates a gradient field in each of the frequency encoding direction, the phase encoding direction, and the slice selection direction according to the imaging conditions. More specifically, the gradient coil unit 13 applies a gradient field in the slice selection direction (or the scan direction) of the subject 16 to select a slice; and the RF body coil unit 15 or the local RF coil array can transmit an RF pulse to the selected slice of the subject 16. The gradient coil unit 13 also applies a gradient field in the phase encoding direction of the subject 16 to phase-encode the magnetic resonance signals from the slice excited by the RF pulse. Then the gradient coil unit 13 applies a gradient field in the frequency encoding direction of the subject 16 to frequency-encode the magnetic resonance signals from the slice excited by the RF pulse.
[0019] The RF coil unit 14 is arranged to surround, for example, the region of the subject 16 to be imaged. In some examples, the RF coil unit 14 may be referred to as a surface coil or a receive coil. In the static magnetic field space or imaging space 18 where the static magnetic field B0 is formed by the static magnetic field magnet unit 12, the RF coil unit 15 transmits an RF pulse, which is an electromagnetic wave, to the subject 16 based on a control signal from the controller unit 25, and thereby generates a high-frequency magnetic field B1. This excites the proton spins in the slice of the subject 16 to be imaged. The RF coil unit 14 receives the electromagnetic wave generated when the proton spins excited thereby in the slice of the subject 16 to be imaged return to alignment with the initial magnetization vector as a magnetic resonance signal. In some embodiments, the RF coil unit 14 may transmit the RF pulse and receive the MR signal. In other embodiments, the RF coil unit 14 may be used only for receiving the MR signal and not for transmitting the RF pulse.
[0020] The RF body coil unit 15 is arranged to surround, for example, the imaging space 18 and generates an RF magnetic field pulse orthogonal to the main magnetic field B0 generated by the static magnetic field magnet unit 12 within the imaging space 18 to excite nuclei. In contrast to the RF coil unit 14, which can be disconnected from the MRI apparatus 10 and replaced with another RF coil unit, the RF body coil unit 15 is fixedly attached and connected to the MRI apparatus 10. Further, although local coils such as the RF coil unit 14 may transmit or receive signals only from a local region of the subject 16, the RF body coil unit 15 generally has a larger coverage area. For example, the RF body coil unit 15 can be used to transmit or receive signals to / from the whole body of the subject 16. Using only receive local coils and a transmit body coil provides uniform RF excitation and good image uniformity at the cost of higher RF power deposited in the subject. For transmit-receive local coils, the local coil provides RF excitation to the region of interest and receives the MR signal, thereby reducing the RF power deposited in the subject. It should be understood that the specific use of the RF coil unit 14 and / or the RF body coil unit 15 depends on the imaging application.
[0021] When operating in the receive mode, the T / R switch 20 can selectively electrically connect the RF body coil unit 15 to the data acquisition unit 24, and when operating in the transmit mode, the T / R switch 20 can selectively electrically connect the RF body coil unit 15 to the RF driver unit 22. Similarly, when the RF coil unit 14 operates in the receive mode, the T / R switch 20 can selectively electrically connect the RF coil unit 14 to the data acquisition unit 24, and when operating in the transmit mode, the T / R switch 20 can selectively electrically connect the RF coil unit 14 to the RF driver unit 22. When both the RF coil unit 14 and the RF body coil unit 15 are used for a single scan, for example, if the RF coil unit 14 is configured to receive MR signals and the RF body coil unit 15 is configured to transmit RF signals, the T / R switch 20 can direct the control signal from the RF driver unit 22 to the RF body coil unit 15 while directing the received MR signals from the RF coil unit 14 to the data acquisition unit 24. The coils of the RF body coil unit 15 can be configured to operate in a transmit-only mode or a transmit-receive mode. The coils of the local RF coil unit 14 can be configured to operate in a transmit-receive mode or a receive-only mode.
[0022] The RF driver unit 22 includes a gate modulator (not shown), an RF power amplifier (not shown), and an RF oscillator (not shown), which are used to drive an RF coil (e.g., the RF coil unit 15) and form a high-frequency magnetic field in the imaging space 18. The RF driver unit 22 modulates the RF signal received from the RF oscillator into a signal with a predetermined envelope and a predetermined timing based on the control signal from the controller unit 25 and using the gate modulator. The RF signal modulated by the gate modulator is amplified by the RF power amplifier and then output to the RF coil unit 15.
[0023] The gradient coil driver unit 23 drives the gradient coil unit 13 based on the control signal from the controller unit 25, thereby generating a gradient magnetic field in the imaging space 18. The gradient coil driver unit 23 includes three driver circuit systems (not shown) corresponding to the three gradient coil systems included in the gradient coil unit 13.
[0024] The data acquisition unit 24 includes a preamplifier (not shown), a phase detector (not shown), and an analog / digital converter (not shown) for acquiring the magnetic resonance signals received by the RF coil unit 14. In the data acquisition unit 24, the phase detector uses the output of the RF oscillator from the RF driver unit 22 as a reference signal to detect the magnetic resonance signals received from the RF coil unit 14 and amplified by the preamplifier, and outputs the phase-detected analog magnetic resonance signals to the analog / digital converter to be converted into digital signals. The digital signals thus obtained are output to the data processing unit 31.
[0025] The MRI apparatus 10 includes an examination table 26 on which a subject 16 is placed. By moving the examination table 26 based on a control signal from the controller unit 25, the subject 16 can be moved inside and outside the imaging space 18.
[0026] The controller unit 25 includes a computer and a recording medium on which a program to be executed by the computer is recorded. When the program is executed by the computer, it causes the respective parts of the apparatus to perform operations corresponding to a predetermined scan. The recording medium may include, for example, a ROM, a floppy disk, a hard disk, an optical disk, a magneto-optical disk, a CD-ROM, or a non-volatile memory card. The controller unit 25 is connected to the operation console unit 32 and processes operation signals input to the operation console unit 32, and also controls the examination table 26, the RF driver unit 22, the gradient coil driver unit 23, and the data acquisition unit 24 by outputting control signals to them. The controller unit 25 also controls the data processing unit 31 and the display unit 33 based on the operation signals received from the operation console unit 32 to obtain a desired image.
[0027] The operation console unit 32 includes user input devices such as a touch screen, a keyboard, and a mouse. The operator uses the operation console unit 32, for example, to input data such as an imaging protocol and to set the region where an imaging sequence is to be executed. Data regarding the imaging protocol and the imaging sequence execution region is output to the controller unit 25.
[0028] The data processing unit 31 includes a computer and a recording medium on which a program to be executed by the computer to perform predetermined data processing is recorded. The data processing unit 31 is connected to the controller unit 25 and performs data processing based on a control signal received from the controller unit 25. The data processing unit 31 is also connected to the data acquisition unit 24 and generates spectral data by applying various image processing operations to the magnetic resonance signals output from the data acquisition unit 24.
[0029] The display unit 33 includes a display device and displays an image on the display screen of the display device based on a control signal received from the controller unit 25. The display unit 33 displays, for example, an image of an input item for which the operator inputs operation data from the operation console unit 32. The display unit 33 also displays a two-dimensional (2D) slice image or a three-dimensional (3D) image of the subject 16 generated by the data processing unit 31.
[0030] Reference Figure 2, shows a medical image processing system 200 according to an exemplary embodiment. In some embodiments, the medical image processing system 200 is incorporated into a medical imaging system, such as an MRI system, a CT system, an X-ray system, a PET system, an ultrasound system, etc. In some embodiments, the medical image processing system 200 is disposed at a device (e.g., an edge device, a server, etc.) communicatively coupled to the medical imaging system via a wired connection and / or a wireless connection. In some embodiments, the medical image processing system 200 is disposed at a separate device (e.g., a workstation) that can receive images from the medical imaging system or from a storage device storing images generated by the medical imaging system. The medical image processing system 200 may include an image processing system 202, a user input device 216, and a display device 214.
[0031] The image processing system 202 includes a processor 204 configured to execute machine-readable instructions stored in a non-transitory memory 206. The processor 204 can be single-core or multi-core, and the program executed thereon can be configured for parallel or distributed processing. In some embodiments, the processor 204 may optionally include separate components distributed in two or more devices, which may be remotely located and / or configured for coordinated processing. In some embodiments, one or more aspects of the processor 204 may be virtualized and executed by a remotely accessible networked computing device configured in a cloud computing configuration.
[0032] The non-transitory memory 206 may store a deep multi-task neural network module 208, a training module 210, and medical image data 212, such as magnetic resonance image data. The deep multi-task neural network module 208 may include one or more deep multi-task neural networks, the one or more deep multi-task neural networks including a plurality of parameters (including weights, biases, activation functions), and instructions for implementing the one or more deep multi-task neural networks to receive MR images and map the MR images to an output, where a pseudo CT image corresponding to the MR image can be generated from the output. For example, the deep multi-task neural network module 208 may store instructions for implementing a multi-task neural network (such as Figure 4 the multi-task convolutional neural network (CNN) of the CNN architecture 400 shown). The deep neural network module 208 may include trained and / or untrained multi-task neural networks, and may also include various data or metadata related to the one or more multi-task neural networks stored therein.
[0033] The non-transitory memory 206 may also store a training module 210 that includes instructions for training one or more of the deep neural networks stored in the deep multi-task neural network module 208. The training module 210 may include instructions that, when executed by the processor 204, cause the image processing system 202 to perform one or more of the steps of method 500 discussed in detail below. In some embodiments, the training module 210 includes instructions for: implementing one or more gradient descent algorithms; applying one or more loss functions for each task and a composite loss function based on the one or more loss functions for each task; and / or training routines for adjusting the parameters of one or more of the deep multi-task neural networks of the deep multi-task neural network module 208. In some embodiments, the training module 210 includes instructions for intelligently selecting a training data set from the medical image data 212. In some embodiments, the training data set includes corresponding MR and CT medical image pairs of the same anatomical region of the same patient. Additionally, in some embodiments, the training module 210 includes instructions for generating a training data set by generating a bone mask and a bone HU image based on the CT images in the medical image data 212. In some embodiments, the training module 210 is not provided at the image processing system 202. The deep multi-task neural network module 208 includes networks for training and validation.
[0034] The non-transitory memory 206 also stores medical image data 212. The medical image data 212 includes, for example, MR images captured from an MRI system, CT images acquired by a CT imaging system, etc. For example, the medical image data 212 may store corresponding MR images and CT images of a patient. In some embodiments, the medical image data 212 may include a plurality of training data pairs that include MR image and CT image pairs.
[0035] In some embodiments, the non-transitory memory 206 may include components disposed on two or more devices that may be remotely located and / or configured for coordinated processing. In some embodiments, one or more aspects of the non-transitory memory 206 may include a networked storage device configured in a cloud computing configuration that is remotely accessible.
[0036] The image processing system 200 may also include a user input device 216. The user input device 216 may include a touch screen, a keyboard, a mouse, a touchpad, a motion sensing camera, or one or more of other devices configured to enable a user to interact with and manipulate the data within the image processing system 202. For example, the user input device 216 may enable a user to select a medical image (such as an MR image) to be converted into a pseudo CT image.
[0037] The display device 214 may include one or more display devices utilizing almost any type of technology. In some embodiments, the display device 214 may include a computer monitor and may display unprocessed and processed MR images and / or pseudo-CT images. The display device 214 may be combined with the processor 204, the non-transitory memory 206, and / or the user input device 216 in a shared housing or may be a peripheral display device and may include a monitor, a touch screen, a projector, or other display devices known in the art, which may enable a user to view medical images and / or interact with various data stored in the non-transitory memory 206.
[0038] It should be understood that Figure 2 the illustrated image processing system 200 is for illustration and not limitation. Another suitable image processing system may include more, fewer, or different components.
[0039] Turning to Figure 3 , a schematic diagram of a first embodiment of a conversion process 300 for converting an MR image 310 into a pseudo-CT or pCT image 330 is shown. Specifically, the conversion process 300 includes inputting the MR image 310 into a multi-task neural network 320, which in turn outputs a pseudo-CT image 330, a bone mask 340, and a bone HU image 350 corresponding to the input MR image 310.
[0040] Thus, the multi-task neural network 320 maps the MR image 310 to its corresponding pseudo-CT image 330 that matches the ground truth CT image (not shown). The CT image (I CT ) can be considered as a spatially non-overlapping set having three different density classes:
[0041] I CT =(I 空气 ∪I 组织 ∪I 骨 ),
[0042] where I 空气 corresponds to air, I 组织 corresponds to tissue, and I 骨 corresponds to bone. Assuming that the MR image I MR (e.g., the MR image 310) and the CT image I CT are spatially aligned, which means the spatial alignment of the pseudo-CT image I pCT (e.g., the pseudo-CT image 330), the error between the CT image and the pCT image 330 can be defined as
[0043] e = I CT - I pCT .
[0044] A smaller value of e results in a density deviation within a certain category. However, a larger value of e results in pixels being classified differently; such errors are more likely to occur at the boundary positions between two categories and can lead to cumulative classification errors. Therefore, the error e can be regarded as including both the classification error between different categories and the image value estimation error within each category. The overall objective of the network is to map the MR image 310 to the pCT image 330 by minimizing the error e between the ground truth CT image and the pCT image 330.
[0045] The multi-task neural network 320 is configured with multiple tasks such that the tasks of classification and regression are separated, rather than configuring a neural network with a single task of mapping the MR image 310 to the pCT image 330. By separating the tasks of classification and regression and by optimizing the multi-task neural network 320 to simultaneously reduce the two errors, implicit enhancement can be achieved for each relevant task in the related tasks. Although the tasks are related, it is expected that the multi-task neural network 320 learns them in different ways from each other, and in order to optimize the tasks individually, each task is driven by a dedicated loss function. As further described herein, the multi-task neural network 320 is configured with three tasks: whole image transformation, accurate segmentation of the region of interest, and image value estimation within the region of interest. Each task is driven by a loss function that is modulated to minimize a specific error, thereby contributing to the overall optimal state of the multi-task neural network 320.
[0046] The mean absolute error (MAE) is a suitable loss function for image regression. However, the MAE is a global measure that neither takes into account the imbalance between the regional volumes of each category in the image nor can it focus on regions of the image as needed. The MAE can be adapted to include the ability of spatial focusing by positively weighting the regional loss compared to the rest of the image, where the relative volume of the region can be used as an implicit weight factor. For example, for a given region k with N k samples, the mean absolute error (MAE) within region k is calculated as:
[0047]
[0048] where y i is the true value, and is the estimated value. Then the weighted MAE for an image with two complementary spatial categories {k, k'} including the first category k and the second category k' can be defined as:
[0049]
[0050] where
[0051] N k +N k ′ = N
[0052] is the volume of the entire image. In the volume N of the first category - k is much smaller than the volume N of the second category k' in the scenario of class imbalance, the mean absolute error MAE of the first category - k is emphasized by the volume N of the second category k' such that the mean absolute error MAE of the first category - k is equivalent to the mean absolute error MAE of the second category k' . This result can be regarded as spatial focusing on the region within the image represented by the first category being k. When the volume N of the first category k is equal to the volume N of the second category k' , for example, when each volume is equal to half of the total volume N of the image, then the above weighted mean absolute error (e.g., wMAE k ) becomes the global mean absolute error MAE.
[0053] For the segmentation task, the smooth Dice coefficient loss is usually the preferred loss function. Between a given pair of segmentation probability maps, the Dice loss is defined as:
[0054]
[0055] where x i and are the true bone probability value and the predicted bone probability value in the image respectively.
[0056] As mentioned above, the multi - task neural network 320 is configured to learn multiple tasks with the main purpose of generating the pseudo - CT image 330. The tasks include the first task of generating the pseudo - CT image 330, the second task of generating the bone mask 340, and the third task of generating the bone HU image 350 (e.g., the image values within the bone region of interest in terms of HU). To generate the pseudo - CT image I pCT , the main task of the multi - task neural network 320 is the regression of the entire image corresponding to the entire range of CT values (HU) of different classes. Therefore, the first task or the pCT image task is driven by the regression loss of the body region:
[0057]
[0058] To generate the bone mask X 骨 , the auxiliary task of the multi - task neural network is to segment the bone region from the rest of the image. Specifically, the loss of the second task regularizes the shape of the bone region by penalizing the misclassification of other regions as bone. For this purpose, the second task or the bone mask task is thus driven by the segmentation loss, which may include the Dice loss L discussed above D:
[0059]
[0060] To generate a bone HU value map or a bone HU image I 骨 , the auxiliary task of the multi-task neural network is to generate a continuous density value map within the bone region. Although the third task is a subset of the first task, considering the large dynamic range of the target, the loss of the third task regularizes the regression in the region of interest (e.g., the bone region) explicitly. To focus on the bone region, the rest of the body region together with the background is regarded as a complementary class. Thus, the third task or the bone HU image task is driven by a regression loss focusing on a sub-range of values, defined by the following formula:
[0061]
[0062] The overall objective of the multi-task neural network is defined by the composite task of generating a pseudo-CT image I MR from the input MR image I pCT , a bone map X 骨 , and a bone HU image I 骨 :
[0063] I MR → {I pCT ; X 骨 ; I 骨}.
[0064] To this end, the multi-task neural network is optimized by minimizing the composite loss function L of the multi-task neural network 320:
[0065]
[0066] where the loss coefficient weights w1, w2, and w3 can be selected empirically according to the importance of the corresponding tasks or by modeling the uncertainty of each task. As an illustrative example, the loss coefficient weights can be selected empirically by setting the weight w1 of the main task to one and up-weighting the bone segmentation and regression losses. For example, w1 can be set to 1.0, w2 can be set to 1.5, and w3 can be set to 1.3.
[0067] Although in Figure 3A single input MR image 310 is depicted and described above, but it should be understood that in some examples, more than one MR image 310 may be input into the multi-task neural network 320. The input MR image 310 may thus include a plurality of MR images 310. For example, the input MR image 310 may include a single two-dimensional MR image slice and / or a three-dimensional MR image volume (e.g., including a plurality of two-dimensional MR image slices). Thus, it should be understood that the term "image" as used herein with respect to MR images may thus refer to a two-dimensional slice or a three-dimensional volume.
[0068] As referred to herein Figure 4 Further, as an illustrative and non-limiting example, the multi-task neural network 320 may include a two-dimensional U-Net convolutional neural network having a plurality of output layers. In some illustrative and non-limiting examples, the encoder network may include four levels having two convolutional-batch normalization-exponential linear unit (ELU)-max pooling operation blocks. These layers may be followed by two convolutional-batch normalization-ELU operation blocks in a bottleneck layer. The decoder network may include four levels having two upsampling-convolution-batch normalization-dropout operation blocks. The input MR image 310 is encoded by a convolutional layer block that operates at two different scales (e.g., filter sizes of 7×7 and 5×5) at each resolution. This scale space feature pyramid at each resolution encodes features better than a single feature at a single scale. The decoder path is designed to have shared layers until the final layer. At the final layer, the pCT image 330 is obtained via a convolutional-linear operation, the bone mask 340 is obtained via a convolutional-sigmoid operation, and the bone HU value map or bone image 350 is obtained via a convolutional-rectified linear unit (ReLU) operation, where each output layer operates with a 1×1 filter size.
[0069] As further described herein, the specific implementation-specific parameters described herein (such as the number of filters, U-Net layers, filter size, max pooling size, and learning rate) are illustrative and non-limiting. In fact, any suitable neural network configured for multi-task learning may be implemented. One or more specific embodiments of the present disclosure are described herein in order to provide a thorough understanding. Those skilled in the art will understand that the specific details described in the embodiments may be modified during implementation without departing from the essence of the present disclosure.
[0070] Turning to Figure 4, showing the architecture of an exemplary multi-task convolutional neural network (CNN) 400. The multi-task CNN 400, simply and interchangeably referred to herein as CNN 400, represents an example of a machine learning model according to the present disclosure, where the parameters of the CNN 400 can be learned using training data generated according to one or more methods disclosed herein. The CNN 400 includes a U-net architecture, which can be divided into an autoencoder part (down part, elements 402-430) and an auto-decoder part (up part, elements 432-458). The CNN 400 is configured to receive an MR image including a plurality of pixels / voxels and map the input MR image to a plurality of predetermined types of outputs. The CNN 400 includes a series of mappings that start from the input image patches 402 that can be received by the input layer, pass through a plurality of feature maps, and finally map to the output layers 458a to 458c. Although two-dimensional inputs are described herein, it should be understood that the multi-task neural networks described herein (including the multi-task CNN 400) can be configured to additionally or alternatively accept three-dimensional images as inputs. In other words, the multi-task CNN 400 can accept two-dimensional MR image slices and / or three-dimensional MR image volumes as inputs.
[0071] Various elements including the CNN 400 are labeled in the legend 460. As indicated in the legend 460, the CNN 400 includes a plurality of feature maps (and / or replicated feature maps), where each feature map can receive input from an external file or a previous feature map and can transform / map the received input into an output to generate the next feature map. Each feature map can include a plurality of neurons, where in some embodiments, each neuron can receive input from a subset of neurons in the previous layer / feature map and can calculate a single output based on the received input, where the output can be propagated to a subset of neurons in the next layer / feature map. The feature maps can be described using spatial dimensions such as length, width, and depth, where the dimensions refer to the number of neurons included in the feature map (e.g., how many neurons long, how many neurons wide, and how many neurons deep a specified feature map is).
[0072] In some embodiments, the neurons of the feature map can calculate the output by performing a dot product of the received input using a set of learned weights (each set of learned weights can be referred to herein as a filter), where each received input has a unique corresponding learned weight, and the learned weight is learned during the training of the CNN.
[0073] The transformations / mappings performed by each feature map are indicated by arrows, where each type of arrow corresponds to a different transformation, as indicated in legend 460. The solid black arrow pointing to the right indicates a 3×3 convolution with a stride of 1, where the output of a 3×3 grid of feature channels from the immediately previous feature map is mapped to a single feature channel of the current feature map. Each 3×3 convolution can be followed by an activation function, where in one embodiment, the activation function includes a rectified linear unit (ReLU).
[0074] The hollow arrow pointing down indicates 2×2 max pooling, where the maximum value of a 2×2 grid from a feature channel propagates from the immediately previous feature map to a single feature channel of the current feature map, resulting in a 4-fold reduction in the spatial resolution of the immediately previous feature map.
[0075] The hollow arrow pointing up indicates 2×2 transposed convolution, which includes mapping the output of a single feature channel from the immediately previous feature map to a 2×2 grid of feature channels in the current feature map, thereby increasing the spatial resolution of the immediately previous feature map by 4 times.
[0076] The dashed arrow pointing to the right indicates copying and cropping a feature map to concatenate with another later-occurring feature map. Cropping enables the dimensions of the copied feature map to match the dimensions of the feature channels to which the copied feature map is to be concatenated. It should be understood that when the size of the first feature map being copied is equal to the size of the second feature map to which the first feature map is to be concatenated, cropping may not be performed.
[0077] The arrow pointing to the right with a hollow elongated triangular head indicates a 1×1 convolution, where each feature channel in the immediately previous feature map is mapped to a single feature channel of the current feature map, or in other words, where a 1-to-1 mapping of feature channels occurs between the immediately previous feature map and the current feature map. As depicted, the other arrows pointing to the right with hollow triangular heads indicate convolutions with different activation functions, including a linear activation function, a rectified linear unit (ReLU) activation function, and a sigmoid activation function.
[0078] In addition to the operations indicated by the arrows within legend 460, CNN 400 also includes feature maps represented by solid-filled rectangles in Figure 4 where the feature maps include height (the length from top to bottom as shown, which corresponds to the y spatial dimension in the x-y plane), width( Figure 4 not shown in, assumed to have a magnitude equal to the height, and corresponding to the x spatial dimension in the x-y plane), and depth (as Figure 4 not shown, assumed to have a magnitude equal to the height, and corresponding to the x spatial dimension in the x-y plane), and depth (as Figure 4The left - right length shown, which corresponds to the number of features within each feature channel). Similarly, CNN 400 includes replicated and cropped feature maps represented by hollow (unfilled) rectangles in Figure 4 where the replicated feature map includes a height (the top - to - bottom length as Figure 4 shown, which corresponds to the y - spatial dimension in the x - y plane), a width ( Figure 4 not shown in, assumed to be equal in magnitude to the height and corresponding to the x - spatial dimension in the x - y plane), and a depth (the left - to - right length as Figure 4 shown, which corresponds to the number of features within each feature channel).
[0079] Starting at the input image tile 402 (also referred to herein as the input layer), data corresponding to the MR image can be input and mapped to a first set of features. In some embodiments, the input data is pre - processed (e.g., normalized) before being processed by the neural network. The weights / parameters of each layer of CNN 400 can be learned during the training process, where a matching pair of input and expected output (ground truth output) is fed into CNN 400. The parameters can be adjusted based on a gradient descent algorithm or other algorithms until the output of CNN 400 matches the expected output (ground truth output) within a threshold accuracy.
[0080] As indicated by the solid black right - pointing arrow immediately to the right of the input image tile 402, a 3×3 convolution of the feature channels of the input image tile 402 is performed to produce the feature map 404. As discussed above, the 3×3 convolution involves mapping the input from a 3×3 grid of the feature channel to a single feature channel of the current feature map using learned weights, where the learned weights are referred to as convolution filters. Each 3×3 convolution in the CNN architecture 400 can include a subsequent activation function, which in one embodiment includes passing the output of each 3×3 convolution through ReLU. In some embodiments, activation functions other than ReLU can be employed such as Softplus (also known as SmoothReLU), leaky ReLU, noisy ReLU, exponential linear unit (ELU), Tanh, Gaussian, Sinc, Bent identity, logistic function, and other activation functions known in the machine learning field.
[0081] As indicated by the solid black right - pointing arrow immediately to the right of the feature map 404, a 3×3 convolution is performed on the feature map 404 to produce the feature map 406.
[0082] As indicated by the downward-pointing arrow below the signature map 406, a 2×2 max pooling operation is performed on the signature map 406 to produce the signature map 408. Briefly, the 2×2 max pooling operation involves determining the maximum eigenvalue from a 2×2 grid of feature channels of the immediately preceding signature map, and setting a single feature in a single feature channel of the current signature map to the maximum value so determined. Additionally, the signature map 406 is copied and concatenated with the output from the signature map 448 to produce the signature map 450, as indicated by the dashed right-pointing arrow immediately to the right of the signature map 406.
[0083] As indicated by the solid black right-pointing arrow immediately to the right of the signature map 408, a 3×3 convolution with a stride of 1 is performed on the signature map 408 to produce the signature map 410. As indicated by the solid black right-pointing arrow immediately to the right of the signature map 410, a 3×3 convolution with a stride of 1 is performed on the signature map 410 to produce the signature map 412.
[0084] As indicated by the downward-pointing hollow-headed arrow below the signature map 412, a 2×2 max pooling operation is performed on the signature map 412 to produce the signature map 414, where the signature map 414 is one-quarter of the spatial resolution of the signature map 412. Additionally, the signature map 412 is copied and concatenated with the output from the signature map 442 to produce the signature map 444, as indicated by the dashed right-pointing arrow immediately to the right of the signature map 412.
[0085] As indicated by the solid black right-pointing arrow immediately to the right of the signature map 414, a 3×3 convolution with a stride of 1 is performed on the signature map 414 to produce the signature map 416. As indicated by the solid black right-pointing arrow immediately to the right of the signature map 316, a 3×3 convolution with a stride of 1 is performed on the signature map 416 to produce the signature map 418.
[0086] As indicated by the downward-pointing arrow below the signature map 418, a 2×2 max pooling operation is performed on the signature map 418 to produce the signature map 420, where the signature map 420 is half of the spatial resolution of the signature map 419. Additionally, the signature map 418 is copied and concatenated with the output from the signature map 436 to produce the signature map 438, as indicated by the dashed right-pointing arrow immediately to the right of the signature map 418.
[0087] As indicated by the solid black right-pointing arrow immediately to the right of the signature map 420, a 3×3 convolution with a stride of 1 is performed on the signature map 420 to produce the signature map 422. As indicated by the solid black right-pointing arrow immediately to the right of the signature map 422, a 3×3 convolution with a stride of 1 is performed on the signature map 422 to produce the signature map 424.
[0088] As indicated by the downward-pointing arrow below the signature map 424, a 2×2 max pooling operation is performed on the signature map 424 to produce the signature map 426, where the signature map 426 is one-fourth of the spatial resolution of the signature map 424. Additionally, the signature map 424 is copied and concatenated with the output from the signature map 430 to produce the signature map 432, as indicated by the dashed right-pointing arrow immediately to the right of the signature map 424.
[0089] As indicated by the solid black right-pointing arrow immediately to the right of the signature map 426, a 3×3 convolution is performed on the signature map 426 to produce the signature map 428. As indicated by the solid black right-pointing arrow immediately to the right of the signature map 428, a 3×3 convolution with a stride of 1 is performed on the signature map 428 to produce the signature map 430.
[0090] As indicated by the upward-pointing arrow immediately above the signature map 430, a 2×2 transposed convolution is performed on the signature map 430 to produce the first half of the signature map 432, while using the copied features from the signature map 424 to produce the second half of the signature map 432. In short, a 2×2 transposed convolution with a stride of 2 (also referred to as deconvolution or upsampling in this document) involves distributing the features from a single feature channel in the previous signature map to four features in four feature channels in the current signature map (i.e., the output from a single feature channel is treated as the input to four feature channels). Transposed convolution / deconvolution / upsampling involves projecting the feature values from a single feature channel through a deconvolution filter (also referred to as a deconvolution kernel in this document) to produce multiple outputs.
[0091] As indicated by the solid black right-pointing arrow immediately to the right of the signature map 432, a 3×3 convolution is performed on the signature map 432 to produce the signature map 434.
[0092] As Figure 4As indicated, a 3×3 convolution is performed on signature map 434 to produce signature map 436 and a 2×2 up-convolution is performed on signature map 436 to produce half of signature map 438, while replicated features from signature map 418 produce the second half of signature map 438. Additionally, a 3×3 convolution is performed on signature map 438 to produce signature map 440, a 3×3 convolution is performed on signature map 440 to produce signature map 442, and a 2×2 up-convolution is performed on signature map 442 to produce the first half of signature map 444, while using replicated and cropped features from signature map 412 to produce the second half of signature map 444. A 3×3 convolution is performed on signature map 444 to produce signature map 446, a 3×3 convolution is performed on signature map 446 to produce signature map 348, and a 2×2 up-convolution is performed on signature map 448 to produce the first half of signature map 450, while using replicated features from signature map 406 to produce the second half of signature map 450. A 3×3 convolution is performed on signature map 450 to produce signature map 452, a 3×3 convolution is performed on signature map 452 to produce signature map 454, and a 1×1 convolution is performed on signature map 454 to produce output layer 456. In short, a 1×1 convolution involves a one-to-one mapping of feature channels in a first feature space to feature channels in a second feature space, where no reduction in spatial resolution occurs.
[0093] As depicted by the multiple arrows pointing away from output layer 456, multiple outputs can be obtained from output layer 456 by performing two-dimensional convolutions with different activation functions. First, a two-dimensional convolution with a linear activation function is performed on output layer 456 to produce a pseudo-CT image output 458a. Second, a two-dimensional convolution with a sigmoid activation function is performed on output layer 456 to produce a bone mask output 458b. Third, a two-dimensional convolution with a ReLU activation function is performed on output layer 456 to produce a bone image output 458c.
[0094] The pseudo-CT image output layer 458a includes an output layer of neurons, where the output of each neuron corresponds to a pixel of the pseudo-CT image. The bone mask output layer 458b includes an output layer of neurons, where the output of each neuron corresponds to a pixel of the bone mask or bone mask image. The bone HU image output layer 458c includes an output layer of neurons, where the output of each neuron corresponds to a pixel that includes the HU value within the bone region and is empty outside the bone region.
[0095] In this way, the multi-task CNN 400 enables the MR image to be mapped to multiple outputs. Figure 4The architecture of the CNN 400 shown in includes feature map transformation, which occurs as the input image patches propagate through the neuron layers of the convolutional neural network to produce multiple prediction outputs.
[0096] The weights (and biases) of the convolutional layers in the CNN 400 are learned during training, which will be discussed in detail below with reference to Figure 5 This will be discussed in detail below. Briefly, a loss function is defined to reflect the difference between each prediction output and each ground truth output. The composite difference / loss of the multi-task CNN 400 based on the loss functions for each task can be back-projected 8 into the CNN 400 to update the weights (and biases) of the convolutional layers. Multiple training datasets including MR images and corresponding ground truth outputs can be used to train the CNN 400.
[0097] It should be understood that the present disclosure includes neural network architectures that include one or more regularization layers, which include batch normalization layers, dropout layers, Gaussian noise layers, and other regularization layers known in the machine learning field. These can be used during training to mitigate overfitting and improve training efficiency while reducing training time. The regularization layers are used during the training of the CNN and are deactivated or removed during the post-training implementation of the CNN. These layers can be scattered Figure 4 between the layers / feature maps shown, or can replace one or more of the layers / feature maps shown.
[0098] It should be understood that Figure 4 the architecture and configuration of the CNN 400 shown in are for illustration and not limitation. Any suitable multi-task neural network can be used. The above describes one or more specific embodiments of the present disclosure to provide a thorough understanding. Those skilled in the art will understand that the specific details described in the embodiments can be modified during implementation without departing from the essence of the present disclosure.
[0099] Figure 5 is a high-level flowchart showing an exemplary method 500 for training a deep multi-task neural network (such as Figure 4 the CNN 400 shown in ) to generate pseudo-CT images from MR images with focused region of interest accuracy. The method 500 can be implemented by a training module 210.
[0100] Method 500 begins at 505. At 505, method 500 feeds a training data set including an MR image, a ground truth CT image, a ground truth bone mask, and a ground truth bone HU image into a multi-task neural network. The MR image and the ground truth CT image include medical images of the same region of interest of the same patient acquired via MR and CT imaging modalities, respectively, such that the MR image and the ground truth CT image correspond to each other. The ground truth bone mask and the ground truth bone HU image are generated from the ground truth CT image. For example, the ground truth CT image can be segmented to obtain segments of the ground truth CT image that contain bone, and the ground truth bone mask can include the segments of the ground truth CT image that contain bone. The ground truth bone mask thus includes an image mask that indicates the locations of the ground truth CT image corresponding to bone and further indicates the locations of the ground truth CT image that do not correspond to bone, for example, by representing bone segments as black pixels and non-bone segments as white pixels, or vice versa. Similarly, although the ground truth bone mask includes an image mask indicating bone segments, the ground truth bone HU value map or the ground truth bone HU image includes HU values within the bone segments.
[0101] The ground truth can include an expected, ideal, or "correct" result obtained from the multi-task neural network based on the input of the MR image. The ground truth output including the ground truth CT image, the ground truth bone mask, and the ground truth bone HU image corresponds to the MR image, such that the multi-task neural network described herein can be trained on multiple tasks, including generating a pseudo CT image corresponding to the MR image, generating a bone mask indicating the location of bone within the pseudo CT image, and generating a bone HU image indicating the bone HU values within the pseudo CT image. The training data set and multiple training data sets including the training data set can be stored in an image processing system, such as in the medical image data 212 of the image processing system 202.
[0102] At 510, method 500 inputs the MR image into the input layer of the multi-task neural network. For example, the MR image is input into the input layer 402 of the multi-task CNN 400. In some examples, each voxel or pixel value of the MR image is input into a different node / neuron of the input layer of the multi-task neural network.
[0103] At 515, method 500 determines the current output of the multi-task neural network including a pCT image, a bone mask, and a bone HU image. For example, the multi-task neural network maps the input MR image to the pCT image, the bone mask, and the bone HU image by propagating the input MR image through one or more hidden layers from the input layer until reaching the output layer of the multi-task neural network. The pCT image, the bone mask, and the bone HU image include the output of the multi-task neural network.
[0104] At 520, method 500 calculates a first loss between the pCT image and the ground truth CT image. Method 500 can calculate the first loss by computing the difference between the pCT image output by the multi-task neural network and the ground truth CT image. For example, since the first task of the multi-task neural network is whole-image regression over the entire CT value (HU) range corresponding to different classes, the first loss can be calculated according to the following formula
[0105]
[0106] where MAE 身体 includes the mean absolute error over the entire body region, which includes the bone region, tissue region, etc., as referred to above Figure 3 as described
[0107] At 525, method 500 calculates a second loss between the bone mask and the ground truth bone mask. Method 500 can calculate the second loss by computing the difference between the bone mask output by the multi-task neural network and the ground truth bone mask. For example, since the second task of the multi-task neural network is to segment the bone region of the MR image, the second loss regularizes the shape of the bone region by penalizing misclassifications of other regions as bone. To this end, the second loss can be calculated as:
[0108]
[0109] where the Dice loss L D can include the smoothed Dice coefficient loss, as referred to above Figure 3 as described
[0110] At 530, method 500 calculates a third loss between the bone HU image and the ground truth bone HU image. Method 500 can calculate the third loss by computing the difference between the bone HU image output by the multi-task neural network and the ground truth bone HU image. For example, the third loss explicitly regularizes the regression in the region of interest (e.g., the bone region), and to focus on the bone region, the rest of the body region along with the background is considered a complementary class. Thus, method 500 can calculate the third loss by computing a regression loss focused on a sub-range of values, defined by the following formula:
[0111]
[0112] where wMAE 骨 includes the weighted mean absolute error of the bone region as referred to above Figure 3 as described
[0113] At 535, method 500 calculates a composite loss based on a first loss, a second loss, and a third loss. For example, method 500 calculates a composite loss function L for a multi-task neural network:
[0114]
[0115] where the loss coefficient weights w1, w2, and w3 can be determined based on the importance of the corresponding tasks or by modeling the uncertainty of each task.
[0116] At 540, method 500 adjusts the weights and biases of the multi-task neural network based on the composite loss calculated at 535. The composite loss can be backpropagated through the multi-task neural network to update the weights and biases of the convolutional layers. In some examples, the backpropagation of the composite loss can occur according to a gradient descent algorithm, where the gradient (first derivative or approximation of the first derivative) of the composite loss function is determined for each weight and bias of the multi-task neural network. Then, each weight and bias is updated by adding the negative of the product of the determined (or approximated) gradient and a predetermined step size. Then, method 500 returns. It should be understood that method 500 can be repeated until the weights and biases of the multi-task neural network converge, or until the rate of change of the weights and / or biases of the multi-task neural network is below a threshold for each iteration of method 500.
[0117] In this way, method 500 enables the multi-task neural network to be trained to generate pseudo-CT images with increased structural and quantitative accuracy in regions with different electron densities, with a particular focus on accurate bone value prediction.
[0118] Once the multi-task neural network is trained as described above, the multi-task neural network can be deployed to generate pseudo-CT images, which can in turn be used to improve the clinical workflow with a single imaging modality. As an illustrative example, Figure 6 is a high-level flowchart showing an exemplary method 600 for generating pseudo-CT images from MR images using a deep multi-task neural network according to an embodiment of the present disclosure. Referring to Figures 1 to 4 the system and components of
[0119] Method 600 begins at 605. At 605, method 600 obtains an MR image. In an example where the medical image processing system 200 is integrated into an imaging system (such as the MRI device 10), for example, method 600 can control the MRI device 10 to perform a scan of a subject (such as a patient) by generating an RF signal and measuring an MR signal. In such examples, method 600 can also construct an MR image of the subject based on the measured MR signal as described above with reference to Figure 1 In other examples, where the image processing system 200 is provided at a separate device (e.g., a workstation) that is communicatively coupled to the imaging system (such as the MRI device 10) and is configured to receive the MR image from the imaging system, method 600 can obtain the MR image by retrieving or receiving the MR image from the imaging system. In other examples, method 600 can obtain the MR image by retrieving the MR image from a storage device (e.g., via a Picture Archiving and Communication System (PACS)).
[0120] At 610, method 600 inputs the MR image into a trained multi-task neural network. In some examples, the trained multi-task neural network includes a U-Net two-dimensional convolutional neural network architecture configured with multiple output layers, such as the CNN 400 described above with reference to Figure 4 having an autoencoder-auto decoder type architecture trained according to the training method 500 described above with reference to Figure 5 The trained multi-task neural network generates at least a pseudo CT (pCT) image corresponding to the input MR image, as well as additional outputs related to other tasks, such as a bone mask and a bone HU image as described above.
[0121] Thus, at 615, method 600 receives the pCT image, the bone HU image, and the bone mask corresponding to the MR image from the trained multi-task neural network. Since the bone HU image is generated using a specific objective of quantitative accuracy, the HU value of the bone region indicated by the bone mask may be more accurate than the HU value of the same region in the pCT image. Thus, at 620, method 600 updates the pCT image using the bone HU image. For example, method 600 can paste the bone HU image onto the pCT image guided by the bone mask in some examples, such that the bone HU values depicted in the bone HU image replace the corresponding pixels in the pCT image. Alternatively, the bone HU image can be blended with the pCT image to improve the quantitative accuracy of the pCT image without replacing the pixels of the pCT image.
[0122] At 620, method 600 outputs the MR image and the updated pCT image. For example, method 600 can display the MR image and the updated pCT image via a display device (such as the display device 214 or the display unit 33). Then, method 600 returns.
[0123] Figure 7 A set of images 700 showing exemplary pseudo CT images generated according to different techniques are shown compared to an input MR image 705 and a ground truth CT image 710. The input MR image 705 comprises a zero echo time (ZTE) MR image of a patient acquired by controlling an MR device during an MR scan using a ZTE protocol suitable for capturing bone information in a single contrast MRI. The ground truth CT image 710 comprises a CT image of a patient acquired by controlling a CT imaging system to perform a CT scan of the patient.
[0124] Image registration between the MR image 705 and the ground truth CT image 710 is also performed. For example, the CT image 710 is aligned to match the MR image space of the MR image 705 by applying an affine transformation to the CT image. As an illustrative example, the registration can be performed by minimizing a combination of mutual information and cross-correlation measures. Such registration can be performed specifically for MR-CT image training pairs to further improve the accuracy of pCT image regression from MR images, bone segmentation, and bone image regression performed by the multi-task neural network described herein.
[0125] As mentioned above, the set of images 700 also includes exemplary pseudo CT images generated according to different techniques. For example, the first pseudo CT image 720 includes a multi-task pseudo CT image 720 generated from the input MR image 705 using a multi-task neural network, as described above. The difference map 722 depicts the pixel-level difference or residual error (e.g., I CT –I pCT ).
[0126] In addition, the second pseudo CT image 730 includes a single-task pseudo CT image 730 generated from the input MR image 705 using a single-task neural network adapted to a similar architecture to the multi-task neural network described herein, but trained only for a single task of pseudo CT image regression. The difference graph 732 depicts the difference between the ground truth CT image 710 and the single-task pseudo CT image 730.
[0127] As another example, the third pseudo CT image 740 includes a standard pseudo CT image 740 generated from the input MR image 705 using a standard regression network, specifically a fully connected DenseNet56 neural network trained to perform pseudo CT image regression. The difference map 742 depicts the difference between the ground truth CT image 710 and the standard pseudo CT image 740.
[0128] As depicted, the residual error depicted by the difference map 722 for the multi-task pseudo-CT image 720 is lower than the residual errors depicted by the difference maps 732 and 742. The multi-task pseudo-CT image 720 and the single-task pseudo-CT image 730 look similar, but a comparison of the difference map 722 and the difference map 732 indicates lower errors for the multi-task pseudo-CT image 720 in the entire bone region of the image, particularly in the frontal and nasal bone regions of the skull, as depicted by the darker regions in the difference map 732. The difference map 742 indicates a more extensive residual error in the entire bone region (such as in the occipital bone region of the skull).
[0129] As another illustrative example of the qualitative differences between the use of the multi-task neural network provided herein and the use of standard single-task neural networks, Figure 8 a set of graphs 800 showing normalized histograms of soft tissue and bone regions of pseudo-CT and CT images for multiple cases is shown. Specifically, the set of graphs 800 includes: a first set of graphs 810 for a first case, including a graph 812 of the soft tissue region and a graph 814 of the bone region; a set of graphs 820 for a second case, including a graph 822 of the soft tissue region and a graph 824 of the bone region; a set of graphs 830 for a third case, including a graph 832 of the soft tissue region and a graph 834 of the bone region; and a fourth set of graphs 840 for a fourth case, including a graph 842 of the soft tissue region and a graph 844 of the bone region.
[0130] Each graph shows a graph of the normalized histogram of the ground truth CT image and the pseudo-CT image obtained via the various techniques described above (including the multi-task neural network, the single-task neural network, and the standard DenseNet neural network described herein). Specifically, as depicted in the legend 880, the graph with the solid line corresponds to the measurement of the multi-task neural network, the graph with the longer dashed line corresponds to the measurement of the ground truth CT image, the graph with the short dashed line corresponds to the measurement of the single-task neural network, and the graph with the shortest dashed line corresponds to the measurement of the DenseNet56 neural network.
[0131] The proximity of the predicted image histogram to the CT histogram in each region is an indicator of the image similarity at different values within that range. As depicted in each graph of the set of graphs 800, for all HU values, the pseudo-CT histogram of the multi-task neural network (depicted by the solid line graph) more closely matches the ground truth CT histogram (depicted by the longer dashed line graph) compared to the other pseudo-CT histograms (depicted by the shorter dashed line graphs) for both the soft tissue region and the bone region.
[0132] Figure 7 and Figure 8The qualitative analysis depicted shows that the multi-task neural network described herein provides a substantial qualitative improvement in regression to pseudo-CT images over other neural network-based methods. Additionally, the quantitative analysis of different techniques also confirms that the multi-task neural network configured and trained for regression to pseudo-CT images as described herein achieves better performance relative to other methods. For example, for the proposed multi-task neural network evaluated using five cases, the mean absolute error (MAE) in the body region is 90.21 ± 8.06, the MAE in the soft tissue region is 61.60 ± 7.20, and the MAE in the bone region is 159.38 ± 12.48. In contrast, for the single-task neural network evaluated using the same five cases, the MAE in the body region is 103.05 ± 10.55, the MAE in the soft tissue region is 66.90 ± 8.24, and the MAE in the bone region is 214.50 ± 16.05. Similarly, for the DenseNet56 neural network evaluated using the same five cases, the MAE in the body region is 109.92 ± 12.56, the MAE in the soft tissue region is 64.88 ± 8.36, and the MAE in the bone region is 273.66 ± 24.88. The MAE in the body is an indication of the overall accuracy of the prediction. The MAE in the soft tissue region (in the range of -200 Hu to 250 HU) indicates the accuracy of the prediction in low density values. The MAE in the bone region (in the range of 250 Hu to 3000 HU) indicates the accuracy of the prediction in the bone value range. Thus, the specific advantage of the multi-task neural network over other networks is evident in the bone region error. The reduced error of the multi-task neural network can be attributed to the advantage of ROI focus loss via the additional tasks of bone segmentation and bone image regression, thus driving image regression in the bone region. 身体 was 90.21 ± 8.06, the MAE in the soft tissue region 软组织 was 61.60 ± 7.20, and the MAE in the bone region 骨 was 159.38 ± 12.48. In contrast, for the single-task neural network evaluated using the same five cases, the MAE in the body region 身体 was 103.05 ± 10.55, the MAE in the soft tissue region 软组织 was 66.90 ± 8.24, and the MAE in the bone region 骨 was 214.50 ± 16.05. Similarly, for the DenseNet56 neural network evaluated using the same five cases, the MAE in the body region 身体 was 109.92 ± 12.56, the MAE in the soft tissue region 软组织 was 64.88 ± 8.36, and the MAE in the bone region 骨 was 273.66 ± 24.88. The MAE in the body 身体 is an indication of the overall accuracy of the prediction. The MAE in the soft tissue region - 软组织 (in the range of -200 Hu to 250 HU) indicates the accuracy of the prediction in low density values. The MAE in the bone region 骨 (in the range of 250 Hu to 3000 HU) indicates the accuracy of the prediction in the bone value range. Thus, the specific advantage of the multi-task neural network over other networks is evident in the bone region error. The reduced error of the multi-task neural network can be attributed to the advantage of ROI focus loss via the additional tasks of bone segmentation and bone image regression, thus driving image regression in the bone region.
[0133] Additionally, as an additional illustrative example, Figure 9A set of graphs 900 showing graphs of Dice coefficients of pseudo-CT bone regions at different bone density thresholds in multiple cases is shown. The set of graphs 900 includes a graph 910 for a first case, a graph 920 for a second case, a graph 930 for a third case, and a graph 940 for a fourth case, where each graph depicts a graph (solid line) for a multi-task neural network, a graph (longer dashed line) for a single-task neural network, and a graph (shorter dashed line) for a DenseNet56 neural network, as depicted in legend 980. The deterioration of the curve at higher CT numbers or HU values indicates an underestimation of high-density bone values. Although the multi-task neural network shows a slight deterioration for high-density bone values, the single-task neural network and DenseNet56 show significantly worse estimations for the same bone values. Therefore, compared with the standard neural network method for pCT image regression, the multi-task neural network with enhanced focus on bone proposed and described herein ensures accurate prediction across all HU values.
[0134] To evaluate the utility of the multi-task neural network in radiotherapy planning, a comparative analysis of pCT dosimetry performance in radiotherapy planning was performed. After collecting MR and CT data of two patients with brain tumors, treatment plans were developed based on CT images using standard clinical guidelines and ROIs drawn by physicians. Then, the treatment plans were evaluated using a treatment planning system based on both CT and pCT data (generated by the multi-task neural network described herein), and the results were compared. The differences in the average dose of the planned target volume (PTV) relative to the prescribed dose were found to be 0.18% and -0.13% respectively. Therefore, generating accurate pseudo-CT images from MR images using the multi-task neural network enables replacing CT imaging to obtain density maps for radiotherapy dose calculation, and thus enables an MR-only clinical workflow for radiotherapy.
[0135] The technical effects of the present disclosure include generating pseudo-CT images from MR images. Another technical effect of the present disclosure includes generating CT-like images from MR images with enhanced accuracy in regions containing bone. Yet another technical effect of the present disclosure includes generating pseudo-CT images, bone masks, and bone images using a multi-task neural network based on input MR images.
[0136] In one embodiment, a method includes: acquiring a magnetic resonance (MR) image; generating a pseudo-CT image corresponding to the MR image using a multi-task neural network; and outputting the MR image and the pseudo-CT image.
[0137] In a first example of the method, a focal loss of an interested region including bone in an MR image is used to train a multi-task neural network. In a second example of the method optionally including the first example, the method further includes generating a bone mask and a bone image corresponding to the MR image using the multi-task neural network. In a third example of the method optionally including one or more of the first example and the second example, the multi-task neural network is trained using an entire image regression loss of a pseudo-CT image, a segmentation loss of the bone mask, and a regression loss focused on bone fragments of the bone image. In a fourth example of the method optionally including one or more of the first example to the third example, the method further includes training the multi-task neural network using a composite loss including an entire image regression loss, a segmentation loss, and a regression loss focused on bone fragments, wherein each loss is weighted in the composite loss. In a fifth example of the method optionally including one or more of the first example to the fourth example, the method further includes updating the pseudo-CT image using the bone image and outputting the updated pseudo-CT image using the MR image. In a sixth example of the method optionally including one or more of the first example to the fifth example, the multi-task neural network includes a U-Net convolutional neural network configured with multiple output layers, wherein one of the multiple output layers outputs the pseudo-CT image.
[0138] In another embodiment, a magnetic resonance imaging (MRI) system includes an MRI scanner; a display device; a controller unit communicatively coupled to the MRI scanner and the display device; and a memory storing executable instructions that, when executed, cause the controller unit to: acquire a magnetic resonance (MR) image via the MRI scanner; generate a pseudo-CT image corresponding to the MR image using a multi-task neural network; and output the MR image and the pseudo-CT image to the display device.
[0139] In a first example of an MRI system, a multi-task neural network is trained using a focal loss of a region of interest including bone in an MR image. In a second example of an MRI system optionally including the first example, the memory further stores executable instructions that, when executed, cause the controller unit to generate a bone mask and a bone image corresponding to the MR image using the multi-task neural network. In a third example of an MRI system optionally including one or more of the first example and the second example, the multi-task neural network is trained using an overall image regression loss of a pseudo-CT image, a segmentation loss of the bone mask, and a regression loss focused on bone segments of the bone image. In a fourth example of an MRI system optionally including one or more of the first example to the third example, the memory further stores executable instructions that, when executed, cause the controller unit to train the multi-task neural network using a composite loss including an overall image regression loss, a segmentation loss, and a regression loss focused on bone segments, wherein each loss is weighted in the composite loss. In a fifth example of an MRI system optionally including one or more of the first example to the fourth example, the memory further stores executable instructions that, when executed, cause the controller unit to update the pseudo-CT image using the bone image and output the updated pseudo-CT image to a display device using the MR image.
[0140] In yet another implementation, a non-transitory computer-readable medium includes instructions that, when executed, cause a processor to: obtain a magnetic resonance (MR) image; generate a pseudo-CT image corresponding to the MR image using a multi-task neural network; and output the MR image and the pseudo-CT image to a display device.
[0141] In a first example of a non-transitory computer-readable medium, a multi-task neural network is trained using a focal loss of a region of interest including bone in an MR image. In a second example of the non-transitory computer-readable medium optionally including the first example, when the instructions are executed, the processor is further caused to generate a bone mask and a bone image corresponding to the MR image using the multi-task neural network. In a third example of the non-transitory computer-readable medium optionally including one or more of the first example and the second example, the multi-task neural network is trained using an entire image regression loss of a pseudo CT image, a segmentation loss of the bone mask, and a regression loss focused on bone fragments of the bone image. In a fourth example of the non-transitory computer-readable medium optionally including one or more of the first example to the third example, when the instructions are executed, the processor is further caused to train the multi-task neural network using a composite loss including an entire image regression loss, a segmentation loss, and a regression loss focused on bone fragments, where each loss is weighted in the composite loss. In a fifth example of the non-transitory computer-readable medium optionally including one or more of the first example to the fourth example, when the instructions are executed, the processor is further caused to update the pseudo CT image using the bone image and output the updated pseudo CT image to a display device using the MR image. In a sixth example of the non-transitory computer-readable medium optionally including one or more of the first example to the fifth example, the multi-task neural network includes a U-Net convolutional neural network configured with multiple output layers, where one of the multiple output layers outputs the pseudo CT image.
[0142] As used herein, an element or step recited in the singular and preceded by the word "a" or "an" should be understood as not excluding a plurality of the elements or steps, unless expressly stated to the contrary. Further, a reference to "one embodiment" of the present invention is not intended to be construed as excluding the existence of additional embodiments that also incorporate the recited features. Additionally, unless expressly stated to the contrary, an embodiment that "comprises," "includes," or "has" an element or elements with a particular property may include additional such elements that do not have that property. The terms "comprises" and "comprising" are used as a concise verbal equivalent of the corresponding terms "includes" and "including." Further, the terms "first," "second," and "third," etc. are used merely as labels and are not intended to impose numerical requirements or a particular positional order on their objects.
[0143] The written description uses examples to disclose the invention, including the best mode, and also enables one of ordinary skill in the relevant art to practice the invention, including making and using any device or system and performing any included method. The scope of the invention for which protection is sought is defined by the claims and may include other examples that occur to one of ordinary skill in the art. Such other examples are intended to fall within the scope of the claims if they have structural elements that do not differ from the literal language of the claims or if they include equivalent structural elements with insubstantial differences from the literal language of the claims.
Claims
1. A method for converting a magnetic resonance image into a pseudo-computed tomography image, comprising: Obtaining a magnetic resonance (MR) image; Using a multi-task neural network to generate a pseudo-CT image, a bone mask, and a bone HU image corresponding to the MR image, wherein the pseudo-CT image includes a set of three density classes, the density classes including air, tissue, and bone, and wherein the bone HU image includes image values within a bone region of interest in terms of HU; and Outputting the MR image and the pseudo-CT image, wherein the multi-task neural network is trained using an overall image regression loss of the pseudo-CT image, a segmentation loss of the bone mask, and a regression loss focused on bone segments of the bone HU image.
2. The method according to claim 1, wherein the multi-task neural network is trained using a focus loss of a region of interest including bone in the MR image.
3. The method according to claim 1, further comprising training the multi-task neural network using a composite loss including the overall image regression loss, the segmentation loss, and the regression loss focused on the bone segments, wherein each loss is weighted in the composite loss.
4. The method according to claim 1, further comprising updating the pseudo-CT image using the bone HU image and outputting the updated pseudo-CT image using the MR image.
5. The method according to claim 1, wherein the multi-task neural network comprises a U-Net convolutional neural network configured with multiple output layers, and one of the multiple output layers outputs the pseudo-CT image.
6. A magnetic resonance imaging (MRI) system, comprising: An MRI scanner; A display device; A controller unit communicatively coupled to the MRI scanner and the display device; and A memory storing executable instructions that, when executed, cause the controller unit to: Obtain a magnetic resonance (MR) image via the MRI scanner; Use a multi-task neural network to generate a pseudo-CT image, a bone mask, and a bone HU image corresponding to the MR image, wherein the pseudo-CT image includes a set of three density classes, the density classes including air, tissue, and bone, and wherein the bone HU image includes image values within a bone region of interest in terms of HU; and Output the MR image and the pseudo-CT image to the display device, wherein the multi-task neural network is trained using an overall image regression loss of the pseudo-CT image, a segmentation loss of the bone mask, and a regression loss focused on bone segments of the bone HU image.
7. The magnetic resonance imaging (MRI) system according to claim 6, wherein the multi-task neural network is trained using a focus loss of a region of interest including bone in the MR image.
8. The magnetic resonance imaging (MRI) system according to claim 6, wherein the memory further stores executable instructions that, when executed, cause the controller unit to train the multi-task neural network using a composite loss that includes the whole-image regression loss, the segmentation loss, and the regression loss focused on the bone fragments, wherein each loss is weighted in the composite loss.
9. The magnetic resonance imaging (MRI) system according to claim 6, wherein the memory further stores executable instructions that, when executed, cause the controller unit to update the pseudo-CT image using the bone HU image and output the updated pseudo-CT image to the display device using the MR image.
10. A non-transitory computer-readable medium comprising instructions that, when executed, cause a processor to: Obtain a magnetic resonance (MR) image; Generate a pseudo-CT image, a bone mask, and a bone HU image corresponding to the MR image using a multi-task neural network, where, The pseudo-CT image includes a set of three density classes including air, tissue, and bone, and wherein the bone HU image includes image values within the bone region of interest in terms of HU; and Output the MR image and the pseudo-CT image to a display device, wherein the multi-task neural network is trained using the whole-image regression loss of the pseudo-CT image, the segmentation loss of the bone mask, and the regression loss focused on the bone fragments of the bone HU image.
11. The non-transitory computer-readable medium according to claim 10, wherein the multi-task neural network is trained using a focus loss of the region of interest including bone in the MR image.
12. The non-transitory computer-readable medium according to claim 10, wherein the instructions, when executed, further cause the processor to train the multi-task neural network using a composite loss that includes the whole-image regression loss, the segmentation loss, and the regression loss focused on the bone fragments, wherein each loss is weighted in the composite loss.
13. The non-transitory computer-readable medium according to claim 10, wherein the instructions, when executed, further cause the processor to update the pseudo-CT image using the bone HU image and output the updated pseudo-CT image to the display device using the MR image.
14. The non-transitory computer-readable medium according to claim 10, wherein the multi-task neural network includes a U-Net convolutional neural network configured with multiple output layers, and wherein one of the multiple output layers outputs the pseudo-CT image.
15. A method for converting a magnetic resonance image into a pseudo-computed tomography image, comprising: Obtain a magnetic resonance (MR) image; Generate a pseudo-CT image, a bone mask, and a bone HU image corresponding to the MR image using a multi-task neural network, wherein the pseudo-CT image includes a set of three density classes including air, tissue, and bone, and wherein the bone mask is obtained by a convolutional sigmoid operation and the bone HU image is obtained by a ReLU operation; and Output the MR image and the pseudo-CT image, wherein, the multi-task neural network is trained using the overall image regression loss of the pseudo-CT image, the segmentation loss of the bone mask, and the regression loss focusing on bone fragments of the bone HU image.