Method and system for characterizing anatomical features in medical images
Automatically characterize HCC lesions through deep learning and image processing technology, solving the complexity and subjectivity of liver imaging in multiphase CT examinations, and achieving efficient and accurate detection and evaluation of HCC tumors.
Patent Information
- Application Number
- CN202011644062.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-23
- Filing Date
- 2020-12-31
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-08-01
AI Technical Summary
Prior art In the early detection and evaluation of hepatocellular carcinoma (HCC), especially during multiphase CT examination, there is complexity and subjectivity of liver imaging processing, resulting in misdiagnosis and variability, making it difficult to accurately track changes in lesions and evaluate treatment responses.
Deep learning methods and image processing technology are used to automatically characterize HCC lesions by segmenting, registering and resegrating medical images, reducing subjectivity and improving the accuracy of detection and evaluation.
It improves the efficiency and accuracy of HCC tumor detection and characterization, reduces misdiagnosis, and enhances the reliability of physician diagnosis and patient prognosis.
Smart Images

Figure CN113240719B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the subject matter disclosed herein relate to medical imaging systems and, more particularly, to image feature analysis using deep neural networks. Background Art
[0002] Liver cancer is a leading cause of cancer-related death worldwide, and hepatocellular carcinoma (HCC) is the most common type of liver cancer. Early tumor detection and evaluation can favorably impact patient prognosis. As an example, contrast-enhanced computed tomography (CT) scans can be used as a non-invasive method for assessing liver tissue in patients at risk for HCC. CT scans use differential transmission of x-rays through a target volume (which includes the liver when a patient is being evaluated for HCC) to acquire image data and construct tomographic images (e.g., three-dimensional representations of the interior of the human body). The Liver Imaging Reporting and Data System or LI-RADS provides standards for interpreting and reporting the resulting images from CT scans.
[0003] In some examples, follow-up scans can be performed after treatment to evaluate the effectiveness of the treatment. For example, during the analysis of the follow-up scans, a user / physician can compare the gray-scale values of relevant tissue regions to assess the differences before and after treatment. For example, if the identified lesion shows a significant reduction in viable tissue in the follow-up scan after the course of therapy, this can indicate a positive response to the therapy. Otherwise, no response or progressive disease can be indicated. Summary of the Invention
[0004] In one embodiment, a method includes: acquiring a plurality of medical images over time during an examination; registering the plurality of medical images; after registering the plurality of medical images, segmenting an anatomical structure into one of the plurality of medical images; creating and characterizing a reference region of interest (ROI) in each of the plurality of medical images; determining a characteristic of the anatomical structure by tracking pixel values of the segmented anatomical structure over time; and outputting the determined characteristic on a display device. In this way, the anatomical structure can be characterized quickly and accurately.
[0005] The above advantages and other advantages and features of the present specification will be apparent from the following detailed description when taken alone or in connection with the accompanying drawings. It should be understood that the above summary of the invention is provided to introduce in a simplified form a selection of concepts that are further described in the detailed description. This does not mean identifying key or essential features of the claimed subject matter, the scope of which is uniquely defined by the claims that follow the detailed description. Furthermore, the claimed subject matter is not limited to embodiments that solve any disadvantages mentioned above or in any part of this disclosure. Brief Description of the Drawings
[0006] Aspects of the present disclosure can be better understood by reading the following detailed description and referring to the accompanying drawings, in which:
[0007] Figure 1 is a schematic diagram showing an image processing system for segmenting and tracking anatomical features across multiple images according to an embodiment;
[0008] Figure 2 shows a pictorial view of an imaging system that can utilize an image processing system (such as Figure 1 the image processing system) according to one embodiment;
[0009] Figure 3 shows a block schematic diagram of an exemplary imaging system according to one embodiment;
[0010] Figure 4 is a schematic diagram showing the architecture of a deep learning network that can be used in the Figure 1 system according to an embodiment;
[0011] Figure 5 is a flowchart showing a method for training a segmentation model according to an embodiment;
[0012] Figure 6 is a high-level flowchart showing a method for analyzing and characterizing hepatocellular carcinoma (HCC) lesions in a multi-phase examination according to an embodiment;
[0013] Figure 7 shows a schematic diagram showing an exemplary deep learning method for aligning images from a multi-phase examination according to an embodiment;
[0014] Figure 8 shows a block schematic diagram of a workflow for creating and analyzing a reference region of interest (ROI);
[0015] Figure 9 shows a block schematic diagram of a workflow for performing HCC lesion re-segmentation and characterization according to an embodiment;
[0016] Figure 10 shows an example of lesion re-segmentation and characterization according to an embodiment; and
[0017] Figure 11 shows an example of feature extraction from liver images obtained during a multi-phase examination according to an embodiment. DETAILED DESCRIPTION
[0018] The following description relates to systems and methods for characterizing and tracking changes in anatomical structures, such as cancerous lesions. Specifically, systems and methods are provided for determining the characteristics of hepatocellular carcinoma (HCC) lesions based on images acquired during a multiphase computed tomography (CT) examination. One such multiphase CT examination includes four sequential patient CT scans: a non-contrast (or unenhanced) image acquired before injection of a contrast agent, and additional images acquired after injection of the contrast agent corresponding to the arterial, portal venous, and delayed phases of contrast agent uptake and clearance. In some examples, a physician may rely on the images obtained during a multiphase CT examination to identify potential cancerous tissue and classify the lesion according to a standardized scoring process (e.g., the Liver Imaging Reporting and Data System or LI-RADS), which gives an indication of the malignancy of the lesion. In other examples, a physician may use the images obtained during a multiphase CT examination to track changes in a known lesion over time in order to evaluate response to treatment and further develop treatment strategies. However, accurate processing of liver images obtained during multiphase CT remains challenging due to the complex liver background, respiratory motion, blurred tumor boundaries, heterogeneous appearance, and highly variable shapes of the tumor and the liver. In addition, LI-RADS scoring and post-treatment assessment are complex and subjective for the physician performing the analysis, which can lead to variability and / or misdiagnosis.
[0019] Accordingly, in accordance with the embodiments disclosed herein, a workflow including image processing and deep learning methods is provided to improve the efficiency and accuracy of HCC tumor detection and characterization. The workflow can be used to identify the characteristics of an original HCC lesion (e.g., prior to treatment) in order to classify the lesion according to LI-RADS characteristics and / or evaluate changes in the HCC lesion over time (such as after treatment). Such follow-up studies can indicate, for example, the physiological response to a particular medical treatment and can assist in determining a treatment plan. For example, the workflow can segment the liver and the lesions within the liver in each image acquired during a multiphase CT examination; register the segmented liver and lesions between the images; and determine pre-treatment or post-treatment characteristics of the lesions by re-segmenting the tissue within the lesions, which can be output to the physician for evaluation. By using an automated workflow that combines image processing and deep learning methods to characterize cancerous lesions, variability between and within physicians can be reduced. In addition, misdiagnosis can be reduced by reducing the amount of subjectivity in tumor assessment. In this way, the accuracy of cancer diagnosis and prognosis can be improved while reducing the mental burden on the physician.
[0020] Figure 1 An exemplary image processing system for implementing a deep learning model is shown, which can be Figure 4The deep learning model shown. The deep learning model may include multiple neural network models, each neural network model being trained with a training set according to, for example, Figure 5 An imaging system, such as Figure 2 and Figure 3 The CT imaging system shown, can be used to generate images analyzed by a Figure 1 image processing system. Figure 6 An exemplary method of a workflow for using a Figure 4 exemplary deep learning model to determine and / or track lesion characteristics is shown. Specifically, Figure 6 The method of Figure 7 can analyze the contrast agent kinetics in multiple examination phases to segment the liver from other anatomical features, segment the lesion from other liver tissues, and re-segment the lesion based on various tissues found in the lesion. To use images from multiple examination phases, Figure 6 An exemplary deep learning method for image alignment that can be used in the Figure 8 method is shown. Figure 9 Schematically shows a workflow that can be used to determine and analyze a reference region of interest (ROI) within healthy liver tissue around a lesion. Figure 6 Schematically shows a workflow that can be used to classify re-segmented lesion tissue in the Figure 10 method of Figure 11 An exemplary frame of a re-segmented lesion is shown, and
[0021] See Figure 1 , an exemplary medical image processing system 100 is shown. In some embodiments, the medical image processing system 100 is incorporated into a medical imaging system (such as, a magnetic resonance imaging (MRI) system, a CT system, a single photon emission computed tomography (SPECT) system, etc.). In some embodiments, at least a portion of the medical image processing system 100 is disposed at a device (e.g., an edge device or a server) that is communicatively coupled to the medical imaging system via a wired and / or wireless connection. In some embodiments, the medical image processing system 100 is disposed at a separate device (e.g., a workstation) that can receive images from the medical imaging system or from a storage device that stores images generated by the medical imaging system. The medical image processing system 100 may include an image processing system 31, a user input device 32, and a display device 33. For example, the image processing system 31 may be operably / communicatively coupled to the user input device 32 and the display device 33.
[0022] The image processing system 31 includes a processor 104 configured to execute machine-readable instructions stored in a non-transitory memory 106. The processor 104 can be single-core or multi-core, and the program executed by the processor 104 can be configured for parallel or distributed processing. In some embodiments, the processor 104 can optionally include separate components distributed across two or more devices, which can be remotely located and / or configured to coordinate processing. In some embodiments, one or more aspects of the processor 104 can be virtualized and executed by a remotely accessible networked computing device configured in a cloud computing configuration. In some embodiments, the processor 104 can include other electronic components capable of performing processing functions, such as a digital signal processor, a field programmable gate array (FPGA), or a graphics board. In some embodiments, the processor 104 can include multiple electronic components capable of performing processing functions. For example, the processor 104 can include two or more electronic components selected from a plurality of possible electronic components, which include: a central processing unit, a digital signal processor, a field programmable gate array, and a graphics board. In additional embodiments, the processor 104 can be configured as a graphics processing unit (GPU), including a parallel computing architecture and parallel processing capabilities.
[0023] The non-transitory memory 106 can store a segmentation module 112 and medical image data 114. The segmentation module 112 can include one or more machine learning models (such as deep learning networks), the one or more machine learning models including a plurality of weights and biases, activation functions, loss functions, gradient descent algorithms, and instructions for implementing the one or more deep neural networks to process input medical images. For example, the segmentation module 112 can store instructions for implementing a neural network (such as Figure 4 the convolutional neural network (CNN) 400 shown and described below). The segmentation module 112 can include trained and / or untrained neural networks and can also include training routines or parameters (such as weights and biases) associated with one or more neural network models stored therein.
[0024] The image processing system 31 can be communicatively coupled to a training module 110, which includes instructions for training one or more of the machine learning models stored in the segmentation module 112. The training module 110 can include instructions that, when executed by the processor, cause the processor to perform the Figure 5Steps of method 500. In one example, training module 110 includes instructions for receiving a training data set from medical image data 114, the training data set including a collection of medical images, associated ground truth labels / images, and associated model outputs for training one or more machine learning models stored in segmentation module 112. Training module 110 may receive medical images, associated ground truth labels / images, and associated model outputs for training one or more machine learning models from sources other than medical image data 114 (such as other image processing systems, cloud, etc.). In some embodiments, one or more aspects of training module 110 may include a remotely accessible networked storage device configured in a cloud computing configuration. Additionally, in some embodiments, training model 110 is included in non-transitory memory 106. Additionally or alternatively, in some embodiments, training model 110 may be used to generate segmentation module 112 offline and away from image processing system 100. In such embodiments, training module 110 may not be included in image processing system 100, but may generate data stored in image processing system 100.
[0025] Non-transitory memory 106 also stores medical image data 114. Medical image data 114 includes, for example, functional imaging images captured by functional imaging modalities (such as SPECT and PET systems), anatomical images captured by MRI systems or CT systems, etc. For example, medical image data 114 may include initial and subsequent medical scan images stored in non-transitory memory 106. As an example, medical image data 114 may include a series of CT images acquired during a multi-phase contrast-enhanced examination of a patient's liver.
[0026] In some embodiments, non-transitory memory 106 may include components located at two or more devices, which may be remotely located and / or configured for coordinated processing. In some embodiments, one or more aspects of non-transitory memory 106 may include a remotely accessible networked storage device configured in a cloud computing configuration.
[0027] Image processing system 100 may also include a user input device 32. User input device 32 may include one or more of a touch screen, keyboard, mouse, touchpad, motion-sensing camera, or other devices configured to enable a user to interact with and manipulate data within image processing system 31. For example, user input device 32 may enable a user to analyze and sort imaging structures.
[0028] The display device 33 may include one or more display devices utilizing any type of display technology. In some embodiments, the display device 33 may include a computer monitor and may display unprocessed images, processed images, parametric maps, and / or examination reports. The display device 33 may be combined with the processor 104, the non-transitory memory 106, and / or the user input device 32 in a shared housing or may be a peripheral display device. The display device 33 may include a monitor, a touch screen, a projector, or another type of display device that enables a user to view medical images and / or interact with various data stored in the non-transitory memory 106.
[0029] It should be understood that Figure 1 the illustrated image processing system 100 is a non-limiting embodiment of an image processing system, and other imaging processing systems may include more, fewer, or different components without departing from the scope of the present disclosure.
[0030] Figure 2 An exemplary CT system 200 configured for CT imaging is shown. Specifically, the CT system 200 is configured to image a subject 212 (such as a patient, an inanimate object, one or more manufactured parts) and / or a foreign object (such as a dental implant, a stent, and / or a contrast agent present in the body). The CT system 200 can be used to generate, for example, medical images processed by Figure 1 the image processing system 100. In one embodiment, the CT system includes a gantry 202, which in turn may also include at least one x-ray source 204 configured to project an x-ray radiation beam 206 (see Figure 3 ) for imaging a subject 212 lying on an examination table 214. Specifically, the x-ray source 204 is configured to project the x-ray radiation 206 towards a detector array 208 positioned on the opposite side of the gantry 202. Although Figure 2 only a single x-ray source 204 is depicted, in certain embodiments, multiple x-ray sources and detectors may be employed to project multiple x-ray radiation beams 206 for acquiring projection data at different energy levels corresponding to the patient. In some embodiments, the x-ray source 204 may implement dual-energy Gemstone Spectral Imaging (GSI) through rapid peak kilovoltage (kVp) switching. In some embodiments, the x-ray detector employed is a photon-counting detector capable of distinguishing x-ray photons of different energies. In other embodiments, two sets of x-ray sources and detectors are used to generate dual-energy projections, with one set of x-ray sources and detectors set to a low kVp and the other set set to a high kVp. It should thus be understood that the methods described herein can be implemented using single-energy acquisition techniques as well as dual-energy acquisition techniques.
[0031] In some embodiments, the CT system 200 further includes an image processor unit 210 that is configured to reconstruct an image of a target volume of a subject 212 using iterative or analytical image reconstruction methods. For example, the image processor unit 210 may use an analytical image reconstruction method such as filtered back projection (FBP) to reconstruct an image of the target volume of the patient. As another example, the image processor unit 210 may use an iterative image reconstruction method (such as advanced statistical iterative reconstruction (ASIR), conjugate gradient (CG), maximum likelihood expectation maximization (MLEM), model-based iterative reconstruction (MBIR), etc.) to reconstruct an image of the target volume of the subject 212. As further described herein, in some examples, in addition to the iterative image reconstruction method, the image processor unit 210 may also use an analytical image reconstruction method (such as FBP). In some examples, the image processor unit 210 may be included as Figure 1 part of or communicatively coupled to an image processing system 31.
[0032] In some configurations of the CT system 200, the x-ray source 204 projects a cone-shaped x-ray radiation beam that is collimated to lie within the X-Y-Z plane of a Cartesian coordinate system and is commonly referred to as the "imaging plane". The x-ray radiation beam 206 passes through an object being imaged, such as a patient 212. The x-ray radiation beam impinges on an array of detector elements at the detector array 208 after being attenuated by the object. The intensity of the attenuated x-ray radiation beam received at the detector array 208 depends on the attenuation of the x-ray radiation beam by the object. Each detector element of the array generates a separate electrical signal that is a measurement of the x-ray beam attenuation at the detector location. The attenuation measurements from all detector elements are acquired separately to produce a transmission profile.
[0033] In some configurations of the CT system 200, the gantry 202 is used to rotate the x-ray source 204 and the detector array 208 around the object to be imaged within the imaging plane such that the angle at which the x-ray radiation beam 206 intersects the object is constantly changing. A set of x-ray radiation attenuation measurements (e.g., projection data) from the detector array 208 at one gantry angle is referred to as a “view”. A “scan” of an object includes a set of views obtained at different gantry angles or perspectives during one rotation of the x-ray source and detector. It is contemplated that the benefits of the methods described herein are derived from medical imaging modalities other than CT, and thus as used herein, the term “view” is not limited to the use described above with respect to projection data from one gantry angle. The term “view” is used to mean a single data acquisition whenever there are multiple data acquisitions from different angles (whether from CT, positron emission tomography (PET), or single photon emission CT (SPECT) acquisitions), and / or any other modality (including modalities yet to be developed) and their combinations in fused (e.g., hybrid) embodiments.
[0034] The projection data is processed to reconstruct an image corresponding to a two-dimensional slice acquired through the object, or in some examples where the projection data includes multiple views or scans, an image corresponding to a three-dimensional rendering of the object. A method for reconstructing an image from a set of projection data is called the filtered backprojection technique. Transmission and emission tomography reconstruction techniques also include statistical iterative methods such as maximum likelihood expectation maximization (MLEM) and ordered subset expectation maximization techniques and iterative reconstruction techniques. This method converts the attenuation measurements from the scan into integers called “CT numbers” or “Hounsfield units (HU)”, which are used to control the brightness (or intensity) of the corresponding pixels on a display device.
[0035] To reduce the total scan time, “helical” scans can be performed. To perform a helical scan, the patient is moved while data for a specified number of slices is acquired. Such systems produce a single helix from a cone beam helical scan. The helix mapped out by the cone beam produces projection data from which an image in each of the specified slices can be reconstructed.
[0036] As used herein, the phrase “reconstructed image” is not intended to exclude embodiments of the present invention in which data representing an image is generated rather than a visual image. Thus, as used herein, the term “image” broadly refers to both a visual image and data representing a visual image. However, many embodiments generate (or are configured to generate) at least one visual image.
[0037] Figure 3 is shown similar to Figure 2Exemplary imaging system 300 of CT system 200. In accordance with aspects of the present disclosure, imaging system 300 is configured to image a subject 304 (e.g., Figure 2 subject 212). In one embodiment, imaging system 300 includes detector array 208 (see Figure 2 ). Detector array 208 also includes a plurality of detector elements 302 that together sense an x-ray radiation beam 206 passing through subject 304 (such as a patient) (see Figure 3 ) to acquire corresponding projection data. Thus, in one embodiment, detector array 208 is fabricated in a multi-slice configuration including multiple rows of cells or detector elements 302. In such a configuration, one or more additional rows of detector elements 302 are arranged in a parallel configuration for acquiring projection data.
[0038] In certain embodiments, imaging system 300 is configured to traverse different angular positions around subject 304 to acquire the desired projection data. Thus, gantry 202 and the components mounted thereon can be configured to rotate about a center of rotation 306 to acquire projection data, for example, at different energy levels. Alternatively, in embodiments where the projection angle relative to subject 304 varies over time, the mounted components can be configured to move along a generally curved path rather than along an arc of a circle.
[0039] Thus, as x-ray source 204 and detector array 208 rotate, detector array 208 collects data of the attenuated x-ray beam. Then, the data collected by detector array 208 undergoes preprocessing and calibration to condition the data to represent a line integral of the attenuation coefficient of the scanned subject 304. The processed data is generally referred to as a projection. In some examples, individual detectors or detector elements 302 in detector array 208 may include photon counting detectors that bin interactions of individual photons into one or more energy bins. It should be understood that the methods described herein may also be implemented using energy integrating detectors.
[0040] The acquired projection data set can be used for basis material decomposition (BMD). During BMD, the measured projections are converted into a set of material density projections. The material density projections can be reconstructed to form a pair or a set of material density maps or images (such as bone, soft tissue, and / or contrast agent maps) of each corresponding basis material. The density maps or images can then be correlated to form a volume rendering of the basis materials (e.g., bone, soft tissue, and / or contrast agent) in the imaged volume.
[0041] Once reconstructed, the underlying material image generated by the imaging system 300 reveals the internal characteristics of the subject 304 represented by the densities of the two underlying materials. Density images can be displayed to showcase these characteristics. In traditional methods of diagnosing medical conditions such as disease states, and more generally medical events, a radiologist or physician would consider a hard copy or display of the density image to discern characteristic features of interest. Such features can include lesions, size, and shape of specific anatomical structures or organs, as well as other features that should be distinguishable in the image based on the skills and knowledge of the individual practitioner.
[0042] In one embodiment, the imaging system 300 includes a control mechanism 308 to control the movement of components, such as the rotation of the gantry 202 and the operation of the x-ray source 204. In certain embodiments, the control mechanism 308 further includes an x-ray controller 310 that is configured to provide power and timing signals to the x-ray source 204. Additionally, the control mechanism 308 includes a gantry motor controller 312 that is configured to control the rotational speed and / or position of the gantry 202 based on imaging requirements.
[0043] In certain embodiments, the control mechanism 308 further includes a data acquisition system (DAS) 314 that is configured to sample the analog data received from the detector elements 302 and convert the analog data into digital signals for subsequent processing. The DAS 314 can also be configured to selectively aggregate analog data from a subset of the detector elements 302 into so-called macro detectors, as further described herein. The data sampled and digitized by the DAS 314 is transmitted to a computer or computing device 316. In one example, the computing device 316 stores the data in a storage device or mass storage device 318. For example, the storage device 318 can include a hard disk drive, a floppy disk drive, a compact disc-read / write (CD-R / W) drive, a digital versatile disc (DVD) drive, a flash drive, and / or a solid-state storage drive.
[0044] Additionally, the computing device 316 provides commands and parameters to one or more of the DAS 314, the x-ray controller 310, and the gantry motor controller 312 to control system operations, such as data acquisition and / or processing. In certain embodiments, the computing device 316 controls system operations based on operator input. The computing device 316 receives operator input via an operator console 320 that is operably coupled to the computing device 316, and the operator input includes, for example, commands and / or scan parameters. The operator console 320 can include, for example, a keyboard (not shown) or a touch screen to allow the operator to specify commands and / or scan parameters.
[0045] Although Figure 3Only one operator console 320 is shown, but more than one operator console may be coupled to the imaging system 300, for example, for inputting or outputting system parameters, requesting examinations, plotting data, and / or viewing images. Additionally, in some embodiments, the imaging system 300 may be coupled to a plurality of locally or remotely located displays, printers, workstations, and / or similar devices, such as within an institution or hospital or at a completely different location, via one or more configurable wired and / or wireless networks (such as the Internet and / or virtual private networks, wireless telephone networks, wireless local area networks, wired local area networks, wireless wide area networks, wired wide area networks, etc.). For example, the imaging system 300 may be coupled to Figure 1 the image processing system 100.
[0046] In one embodiment, for example, the imaging system 300 includes a picture archiving and communication system (PACS) 324 or is coupled to the PACS. In one exemplary implementation, the PACS 324 is further coupled to remote systems (such as a radiology information system, a hospital information system) and / or is coupled to an internal or external network (not shown) to allow operators at different locations to supply commands and parameters and / or obtain access to image data.
[0047] The computing device 316 uses operator-supplied and / or system-defined commands and parameters to operate the examination table motor controller 326, which in turn may control the examination table 214, which may be an electric examination table. Specifically, the examination table motor controller 326 may move the examination table 214 to properly position the subject 304 within the gantry 202 to acquire projection data corresponding to the target volume of the subject 304.
[0048] As previously described, the DAS 314 samples and digitizes the projection data acquired by the detector elements 302. Subsequently, the image reconstructor 330 uses the sampled and digitized x-ray data to perform high-speed reconstruction. Although Figure 3 the image reconstructor 330 is shown as a separate entity, in some embodiments, the image reconstructor 330 may form part of the computing device 316. Alternatively, the image reconstructor 330 may not be present in the imaging system 300, and instead the computing device 316 may perform one or more functions of the image reconstructor 330. Additionally, the image reconstructor 330 may be locally or remotely located and may be operably connected to the imaging system 300 using a wired or wireless network. Specifically, one exemplary embodiment may use computing resources in a "cloud" network cluster for the image reconstructor 330. Further, in some examples, the image reconstructor 330 is included as Figure 2 part of the image processor unit 210.
[0049] In one embodiment, the image reconstructor 330 stores the reconstructed image in the storage device 318. Alternatively, the image reconstructor 330 may transmit the reconstructed image to the computing device 316 to generate available patient information for diagnosis and evaluation. In some embodiments, the computing device 316 may transmit the reconstructed image and / or patient information to a display or display device 332 that is communicatively coupled to the computing device 316 and / or the image reconstructor 330. In one embodiment, the display 332 allows an operator to evaluate the imaged anatomical structure. The display 332 may also allow the operator to select a volume of interest (VOI) and / or request patient information, for example via a graphical user interface (GUI), for subsequent scans or processing.
[0050] In some embodiments, the reconstructed image may be transmitted from the computing device 316 or the image reconstructor 330 to the storage device 318 for short-term or long-term storage. Additionally, in some embodiments, the computing device 316 may be coupled to or may be operatively coupled to Figure 1 processor 104. Thus, the raw data and / or the image reconstructed from the data acquired by the imaging system 300 may be transmitted to the image processing system 100 (see Figure 1 ) for further processing and analysis. Additionally, the various methods and processes further described herein (such as the methods described hereinafter with reference to Figure 6 ) may be stored as executable instructions in a non-transitory memory on a computing device (or controller). At least some of the instructions may be stored in the non-transitory memory in the imaging system 300. In one embodiment, the image reconstructor 330 may include such executable instructions in the non-transitory memory to reconstruct an image from the scan data. In another embodiment, the computing device 316 may include instructions in the non-transitory memory and may at least partially apply the methods described herein to the reconstructed image after receiving the reconstructed image from the image reconstructor 330. In yet another embodiment, the methods and processes described herein may be distributed between the image reconstructor 330 and the computing device 316. Additionally or alternatively, the methods and processes described herein may be distributed between the imaging system 300 (e.g., in the image reconstructor 330 and / or the computing device 316) and Figure 1 the medical image processing system 100 (e.g., in the processor 104 and / or the non-transitory memory 106).
[0051] Go to Figure 4, an exemplary convolutional neural network (CNN) architecture 400 for segmenting anatomical features in medical images is shown. For example, CNN 400 can segment an anatomical feature of interest from other anatomical features in the image. For example, segmentation can define the boundary between the anatomical feature of interest and other tissues, organs, and structures. The liver will be used as the anatomical feature of interest for description. Figure 4 , where the CNN architecture 400 is used to segment the liver from other anatomical features. However, it should be understood that in other examples, the CNN architecture 400 can be applied to segment other anatomical features. CNN 400 represents an example of a machine learning model according to the present disclosure, where the parameters of CNN 400 can be learned using training data generated according to one or more methods disclosed herein (such as the method described below). Figure 5 ).
[0052] The CNN architecture 400 represents a U-net architecture, which can be divided into an encoder part (down part, elements 402 - 430) and a decoder part (up part, elements 432 - 456). The CNN architecture 400 is configured to receive medical images, which can be, for example, magnetic resonance (MR) images, computed tomography (CT) images, SPECT images, etc. In one embodiment, the CNN architecture 400 is configured to receive data from a CT image (such as the input medical image 401) including a plurality of pixels / voxels, and map the input image data to a segmented image of the liver (such as the output segmented medical image 460) based on the output of the acquisition parameter transformation. The CNN architecture 400 will be described as a 3D network, but a 2D network can include a similar architecture. Thus, the input medical image 401 is an input 3D volume including a plurality of voxels.
[0053] The CNN architecture 400 includes a series of mappings that extend from the input image tile 402 (which can be received by the input layer from the input medical image 401) through a plurality of feature maps and ultimately reach the output segmented medical image 460. The output segmented medical image 460, which is an output 3D volume, can be generated based on the output from the output layer 456.
[0054] Various elements including the CNN architecture 400 are marked in the legend 458. As indicated in the legend 458, the CNN architecture 400 includes a plurality of feature maps (and / or replicated feature maps) connected by one or more operations indicated by arrows. The arrows / operations receive input from an external file or a previous feature map, and transform / map the received input to produce the next feature map. Each feature map may include a plurality of neurons. In some embodiments, each neuron may receive input from a subset of neurons in the previous layer / feature map and may compute a single output based on the received input. The output may be propagated / mapped to a subset or all of the neurons in the next layer / feature map. The feature maps may be described using spatial dimensions such as length, width, and depth, where the dimensions refer to the number of neurons included in the feature map (e.g., how many neurons long, how many neurons wide, and how many neurons deep a specified feature map is).
[0055] In some embodiments, the neurons of the feature map may compute the output by performing a convolution of the received input using a set of learned weights (each set of learned weights may be referred to herein as a filter), where each received input has a unique corresponding learned weight, and where the learned weight is learned during the training of the CNN.
[0056] The transformations / mappings performed between the respective feature maps are indicated by the arrows. Each different type of arrow corresponds to a different type of transformation, as shown in the legend 458. The solid black arrow pointing to the right indicates a 3×3×3 convolution with a stride of 1, where the output of a 3×3×3 grid from the feature channels of the immediately previous feature map is mapped to a single feature channel of the current feature map. Each 3×3×3 convolution may be followed by an activation function, where in one embodiment, the activation function includes a rectified linear unit (ReLU).
[0057] The arrow pointing down indicates a 2×2×2 max pooling operation, where the maximum value of a 2×2×2 grid from the feature channels at a single depth is propagated from the immediately previous feature map to a single feature channel of the current feature map, resulting in an output feature map with a spatial resolution reduced by a factor of 8 compared to the immediately previous feature map.
[0058] The hollow arrow pointing up indicates a 2×2×2 transposed convolution, which includes mapping the output of a single feature channel from the immediately previous feature map to a 2×2×2 grid of feature channels in the current feature map, thereby increasing the spatial resolution of the immediately previous feature map by a factor of 8.
[0059] The right-pointing dashed arrow indicates the copying and cropping of a feature map for concatenation with another, later-appearing feature map. Cropping enables the dimensions of the copied feature map to match the dimensions of the feature channels to be concatenated with the copied feature map. It should be understood that no cropping may be performed when the dimensions of the first feature map being copied and the dimensions of the second feature map to be concatenated with the first feature map are equal.
[0060] A right-pointing arrow with a hollow head indicates a 1×1×1 convolution, in which each feature channel in the immediately previous feature map is mapped to a single feature channel in the current feature map, or in other words, a 1-to-1 mapping of feature channels occurs between the immediately previous feature map and the current feature map. The processing at each feature map may include the convolution and deconvolution described above, as well as activation, where the activation function is a nonlinear function that limits the output value of the processing to within a bounded range.
[0061] In addition to the operations indicated by the arrows in the legend 458, the CNN architecture 400 also includes a solid filled rectangle representing a feature map, where the feature map includes height (e.g., Figure 4 The length from top to bottom shown corresponds to the y spatial dimension in the xy plane), the width ( Figure 4 , assuming that the magnitude is equal to the height and corresponds to the x spatial dimension in the xy plane) and the depth (as Figure 4 , which corresponds to the number of features in each feature channel). Similarly, CNN architecture 400 includes a hollow (unfilled) rectangle representing a copied and cropped feature map, where the copied feature map includes a height (e.g., Figure 4 The length from top to bottom shown corresponds to the y spatial dimension in the xy plane), the width ( Figure 4 , assuming that the magnitude is equal to the height and corresponds to the x spatial dimension in the xy plane) and the depth (as Figure 4 The length from left to right is shown, which corresponds to the number of features within each feature channel).
[0062] Starting from the input image tile 402 (also referred to herein as the input layer), data corresponding to the medical image 401 is input and mapped to a first set of features. In some embodiments, the input medical image 401 is pre-processed (e.g., normalized) before being processed by the neural network. The weights / parameters of each layer of the CNN 400 can be learned during the training process, where a matching pair of input and expected output (ground truth output) is fed into the CNN 400. The parameters can be adjusted based on a gradient descent algorithm or other algorithms until the output of the CNN 400 matches the expected output (ground truth output) within a threshold accuracy. The input medical image 401 can include a two-dimensional (2D) or three-dimensional (3D) image / map of the liver (or in other examples, another patient anatomical region).
[0063] As indicated by the solid black right-pointing arrow immediately to the right of the input image tile 402, a 3×3×3 convolution of the feature channels of the input image tile 402 is performed to produce the feature map 404. As discussed above, the 3×3×3 convolution includes mapping the input from a 3×3×3 grid of feature channels to a single feature channel of the current feature map using learned weights, where the learned weights are referred to as convolution filters. Each 3×3×3 convolution in the CNN architecture 400 can include a subsequent activation function, which in one embodiment includes passing the output of each 3×3×3 convolution through ReLU. In some embodiments, activation functions other than ReLU can be employed such as Softplus (also known as SmoothReLU), leaky ReLU, noisy ReLU, exponential linear unit (ELU), Tanh, Gaussian, Sinc, Bent identity, logistic function, and other activation functions known in the field of machine learning.
[0064] As indicated by the solid black right-pointing arrow immediately to the right of the feature map 404, a 3×3×3 convolution is performed on the feature map 404 to produce the feature map 406.
[0065] As indicated by the downward-pointing arrow below the feature map 406, a 2×2×2 max pooling operation is performed on the feature map 406 to produce the feature map 408. Briefly, the 2×2×2 max pooling operation includes determining the maximum feature value from a 2×2×2 grid of feature channels of the immediately previous feature map, and setting a single feature in a single feature channel of the current feature map to the determined maximum feature value. Additionally, the feature map 406 is copied and concatenated with the output from the feature map 448 to produce the feature map 450, as indicated by the dashed right-pointing arrow immediately to the right of the feature map 406.
[0066] As indicated by the solid black right - pointing arrow immediately to the right of signature map 408, a 3×3×3 convolution with a stride of 1 is performed on signature map 408 to produce signature map 410. As indicated by the solid black right - pointing arrow immediately to the right of signature map 410, a 3×3×3 convolution with a stride of 1 is performed on signature map 410 to produce signature map 412.
[0067] As indicated by the downward - pointing hollow - headed arrow below signature map 412, a 2×2×2 max - pooling operation is performed on signature map 412 to produce signature map 414, where signature map 414 is one - quarter of the spatial resolution of signature map 412. Additionally, signature map 412 is copied and concatenated with the output from signature map 442 to produce signature map 444, as indicated by the dashed right - pointing arrow immediately to the right of signature map 412.
[0068] As indicated by the solid black right - pointing arrow immediately to the right of signature map 414, a 3×3×3 convolution with a stride of 1 is performed on signature map 414 to produce signature map 416. As indicated by the solid black right - pointing arrow immediately to the right of signature map 416, a 3×3×3 convolution with a stride of 1 is performed on signature map 416 to produce signature map 418.
[0069] As indicated by the downward - pointing arrow below signature map 418, a 2×2×2 max - pooling operation is performed on signature map 418 to produce signature map 420, where signature map 420 is half of the spatial resolution of signature map 418. Additionally, signature map 418 is copied and concatenated with the output from signature map 436 to produce signature map 438, as indicated by the dashed right - pointing arrow immediately to the right of signature map 418.
[0070] As indicated by the solid black right - pointing arrow immediately to the right of signature map 420, a 3×3×3 convolution with a stride of 1 is performed on signature map 420 to produce signature map 422. As indicated by the solid black right - pointing arrow immediately to the right of signature map 422, a 3×3×3 convolution with a stride of 1 is performed on signature map 422 to produce signature map 424.
[0071] As indicated by the downward-pointing arrow below the signature map 424, a 2×2×2 max pooling operation is performed on the signature map 424 to produce a signature map 426, where the signature map 426 is one-fourth of the spatial resolution of the signature map 424. Additionally, the signature map 424 is copied and concatenated with the output from the signature map 430 to produce a signature map 432, as indicated by the right-pointing arrow with a dashed line immediately to the right of the signature map 424.
[0072] As indicated by the solid black right-pointing arrow immediately to the right of the signature map 426, a 3×3×3 convolution is performed on the signature map 426 to produce a signature map 428. As indicated by the solid black right-pointing arrow immediately to the right of the signature map 428, a 3×3×3 convolution with a stride of 1 is performed on the signature map 428 to produce a signature map 430.
[0073] As indicated by the upward-pointing arrow directly above the signature map 430, a 2×2×2 transposed convolution is performed on the signature map 430 to produce the first half of the signature map 432, while using the copied features from the signature map 424 to produce the second half of the signature map 432. In short, a 2×2×2 transposed convolution with a stride of 2 (also referred to as deconvolution or upsampling in this document) involves distributing the features in a single feature channel of the immediately preceding signature map to four features that are distributed among four feature channels in the current signature map (i.e., the output from a single feature channel is treated as the input for four feature channels). Transposed convolution / deconvolution / upsampling involves projecting the feature values from a single feature channel through a deconvolution filter (also referred to as a deconvolution kernel in this document) to produce multiple outputs.
[0074] As indicated by the solid black right-pointing arrow immediately to the right of the signature map 432, a 3×3×3 convolution is performed on the signature map 432 to produce a signature map 434.
[0075] As Figure 4As indicated, a 3×3×3 convolution is performed on feature map 434 to produce feature map 436, and a 2×2×2 up-convolution is performed on feature map 436 to produce the first half of feature map 438, while the replicated and cropped features from feature map 418 produce the second half of feature map 438. Additionally, a 3×3×3 convolution is performed on feature map 438 to produce feature map 440, a 3×3×3 convolution is performed on feature map 440 to produce feature map 442, and a 2×2×2 up-convolution is performed on feature map 442 to produce the first half of feature map 444, while using the replicated and cropped features from feature map 412 to produce the second half of feature map 444. A 3×3×3 convolution is performed on feature map 444 to produce feature map 446, a 3×3×3 convolution is performed on feature map 446 to produce feature map 448, and a 2×2×2 up-convolution is performed on feature map 448 to produce the first half of feature map 450, while using the replicated features from feature map 406 to produce the second half of feature map 450. A 3×3×3 convolution is performed on feature map 450 to produce feature map 452, a 3×3×3 convolution is performed on feature map 452 to produce feature map 454, and a 1×1×1 convolution is performed on feature map 454 to produce output layer 456. In short, the 1×1×1 convolution includes a one-to-one mapping of the feature channels in the first feature space to the feature channels in the second feature space, where no reduction in spatial resolution occurs.
[0076] Output layer 456 may include an output layer of neurons, where each neuron may correspond to a voxel of the segmented medical image, and where the output of each neuron may correspond to a predicted anatomical feature or characteristic (or lack thereof) at a given location in the input medical image. For example, the output of a neuron may indicate whether the corresponding voxel of the segmented medical image is part of the liver or another anatomical feature.
[0077] In some embodiments, the output layer 456 can be fed back to the input layer of the CNN 400. For example, the output layer from a previous iteration of the CNN 400 can serve as a feedback layer and be applied as an input to the current iteration of the CNN 400. The feedback layer can be included as another layer of the input image (at the same resolution) and thus can be included as part of the input image tile 402. For example, the input medical image 401 and the output layer from a previous iteration of the CNN 400 (e.g., a buffered output, where buffering means the output is stored in a buffer until used as an input to the CNN 400) can be formed into a vector that is fed in as an input to the CNN 400. In some examples, the input medical image that was used as an input in a previous iteration of the CNN 400 can also be included in the input layer.
[0078] In this way, the CNN 400 can implement a mapping from a medical image to an output for segmenting an anatomical feature of interest (e.g., the liver). Figure 4 The architecture of the CNN 400 shown includes feature map transformations that occur as the input image tile propagates through the neuron layers of the convolutional neural network to produce a predicted output. The weights (and biases) of the convolutional layers in the CNN 400 are learned during training, which will be discussed in detail below with reference to Figure 5 Briefly, a loss function is defined to reflect the difference between the predicted output and the ground truth output. The difference / loss can be backprojected into the CNN to update the weights (and biases) of the convolutional layers. The CNN 400 can be trained using multiple training data sets that include medical images and their corresponding ground truth outputs.
[0079] It should be understood that the present disclosure includes neural network architectures that include one or more regularization layers, which include batch normalization layers, dropout layers, Gaussian noise layers, and other regularization layers known in the machine learning field, and they can be used during training to mitigate overfitting and improve training efficiency while reducing training time. The regularization layers are used during the training of the CNN and are deactivated or removed during the post-training implementation of the CNN 400. These layers can be interspersed Figure 4 between the layers / feature maps shown, or can replace one or more of the layers / feature maps shown.
[0080] It should be understood that Figure 4The architecture and configuration of the CNN 400 shown are illustrative and not restrictive. Any suitable neural network such as ResNet, recurrent neural network, generalized regression neural network (GRNN), etc. may be used. One or more specific embodiments of the present disclosure have been described above in order to provide a thorough understanding. Those skilled in the art will understand that specific details described in the embodiments may be modified during implementation without departing from the essence and scope of the present disclosure.
[0081] Refer to Figure 5 , a flowchart of a method 500 for training a segmentation model (such as Figure 4 the CNN 400 shown) according to an exemplary embodiment is shown. For example, the method 500 may be implemented by Figure 1 the training module 110.
[0082] At 502, a training data set is fed into the segmentation model. The training data set may be selected from a plurality of training data sets and may include current medical images, previous model outputs, and corresponding ground truth labels. The previous model output may be determined based on a previous medical image acquired immediately before the current medical image or according to a previous image acquired before the current medical image but having one or more intermediate medical images acquired between the current medical image and the previous medical image. For example, the previous medical image may be the first frame of medical data collected by an imaging system, and the current medical image may be the fifth frame of medical data collected by the imaging system, where, for the purpose of the training data set, the second, third, and fourth frames of medical data collected by the imaging system are discarded. As an example, the training data may include at least several hundred data sets in order to cover high variability in image acquisition parameters (e.g., injection period variability, artifacts, peak x-ray source kilovoltage), anatomical structures, and pathologies.
[0083] The ground truth may include the expected, ideal, or "correct" result obtained from a machine learning model based on the input of the current medical image. In one example, in a machine learning model trained to identify anatomical structures (e.g., the liver) and / or features of the anatomical structure (e.g., cancerous lesions within the liver) in a medical image, the ground truth output corresponding to a specific medical image may include an expert-curated segmentation map of the medical image, which may include the anatomical structure segmented from the background and labels identifying each different tissue type of the anatomical structure. In another example, the ground truth output may be generated by an analysis method / algorithm. In this way, the ground truth labels can identify the identity and location of each anatomical feature in each image for each image. In some embodiments, the training data set (and the plurality of training data sets) may be stored in an image processing system, such as stored in Figure 1In the medical image data 114 of the illustrated image processing system 31. In other embodiments, the training data set may be obtained via a communication coupling between the image processing system and an external storage device (such as via an Internet connection to a remote server). As an example, the training data may include at least 600 data sets in order to cover a high variability in image acquisition parameters (e.g., injection period variability, artifacts, peak x-ray source kilovoltage), anatomical structures, and pathologies.
[0084] At 504, the current image of the training data set is input into the input layer of the model. In some embodiments, the current image is input into the input layer of a CNN having an encoder-decoder type architecture (such as Figure 4 the illustrated CNN 400). In some embodiments, each voxel or pixel value of the current image is input into a different node / neuron of the input layer of the model.
[0085] At 506, the current model output is determined using the current image and the model, the current model output indicating the identity and location of one or more anatomical structures and / or features in the current image. For example, the model may map the input current image to the identity and location of anatomical features by propagating the input current image from the input layer, through one or more hidden layers, and to the output layer of the model. In some embodiments, the output of the model includes a matrix of values, where each value corresponds to the identified anatomical feature (or lack of an identified feature) at the corresponding pixel or voxel of the input current image.
[0086] At 508, the image processing system calculates the difference between the current output of the model and the ground truth label corresponding to the current image. In some embodiments, the difference between each output value of the predicted anatomical feature corresponding to the input current image and the anatomical feature indicated by the ground truth label is determined. The difference may be calculated according to a loss function, such as:
[0087]
[0088] where S is the ground truth label and T is the predicted anatomical feature. That is, for each pixel or voxel of the input current image, the output of the model may include an indication of which anatomical feature the pixel is part of (or lacks an anatomical feature). The ground truth label may similarly include an indication of which identified anatomical feature each pixel of the current image is part of. Then, the difference between each output value and the ground truth label may be determined.
[0089] At 510, the weights and biases of the model are adjusted based on the difference calculated at 508. The difference (or loss) as determined by the loss function can be backpropagated through the model (e.g., a neural learning network) to update the weights (and biases) of the convolutional layers. In some embodiments, the backpropagation of the loss can occur according to a gradient descent algorithm, where the gradient (first derivative or approximation of the first derivative) of the loss function is determined for each weight and bias of the model. Then, each weight (and bias) of the model is updated by adding the negative of the product of the gradient determined (or approximated) for the weight (or bias) and a predetermined step size. Then method 500 can return. For example, method 500 can be repeated until the weights and biases of the model converge, or until the rate of change of the weights and / or biases of the model is below a threshold for each iteration of method 500.
[0090] In this way, method 500 enables the model to be trained to predict the location and / or other attributes (e.g., tissue characteristics) of one or more anatomical features from a current medical image, thereby facilitating the automatic determination of the identified anatomical feature characteristics in subsequent medical scans. Specifically, method 500 can be used to train Figure 4 the CNN 400 to segment the liver within the input medical image.
[0091] Now referring to Figure 6 , a flowchart of an exemplary method 600 for analyzing and characterizing HCC lesions in a multiphase examination using image processing and a deep learning model is shown. Method 600 can be implemented by one or more of the systems disclosed above, such as Figure 1 the medical image processing system 100, Figure 2 the CT system 200, and / or Figure 3 the imaging system 300. Specifically, method 600 provides a workflow for accelerating and improving the accuracy of HCC tumor detection and characterization in both pre-treatment and post-treatment patients.
[0092] At 602, method 600 includes obtaining a liver image. As an example, the liver image may be obtained during a contrast-enhanced computed tomography (CT) scan (or examination). A contrast-enhanced CT scan can be used as a non-invasive method to evaluate liver tissue in patients, for example, at risk of HCC. Thus, obtaining a liver image may include acquiring a series of CT images of the liver during different phases of the CT scan. For example, CT images may be acquired during various phases before and after injecting a contrast agent into the patient. Specifically, a first image may be acquired before injecting the contrast agent (e.g., at 0 seconds after injection) and may be referred to as an unenhanced (or non-contrast) image. A second image may be acquired within a short duration after injecting the contrast agent (e.g., at 20 seconds after injection) and may be referred to as an arterial phase image. A third image may be acquired within another short duration after the second image (e.g., at 50 seconds after injection) and may be referred to as a portal venous phase image. A fourth image may be acquired within a duration after the third image (e.g., at 150 seconds after injection) and may be referred to as a delayed phase image. However, in other examples, a delayed phase image may not be obtainable.
[0093] As detailed above with respect to Figure 2 and Figure 3 each image may be acquired by activating an x-ray source (e.g., Figure 2 x-ray source 204 of Figure 2 configured to project an x-ray radiation beam through the patient, specifically through the patient's liver onto a detector array (e.g., Figure 2 detector array 208 of
[0094] The detector array measures the x-ray attenuation caused by the patient, which can be used (e.g., by an image processor, such as Figure 2 image processor unit 210 of
[0094] to reconstruct the image. In some examples, additional processing, such as various corrections and normalizations, may be performed on the liver image before the following processing.
[0094] At 604, method 600 includes performing liver and lesion registration. Because a contrast-enhanced CT scan is multi-phase, the images acquired from different phases may be aligned at 604 to account for patient movement during the scan. For example, even small movements, such as those caused by patient breathing, may result in displacement of the liver and lesions within the scan view. As will be further described below with respect to Figure 7 registration may include first performing a rigid affine transformation (three rotations and three translations), and then a non-rigid transformation (dense displacement field). Thus, rigid and non-rigid registration may be performed to more directly compare the lesions on the respective images in order to more accurately track the contrast agent kinetics within the lesions, as detailed below.
[0095] At 606, method 600 includes performing liver segmentation on one of the liver images in the liver image set. For example, a fully convolutional neural network based on a 3D Unet architecture with ResNet connections and depth monitoring can be used to separate the liver from other anatomical features in one of the acquired liver image sets. Additionally, stochastic gradient descent with an adaptive learning rate method can be used as an optimizer along with dice loss. Further, affine transformation, elastic deformation, noise addition, and gray level deformation can be performed on the acquired liver images. As an example, the liver segmentation can be performed using the CNN400 shown in Figure 4 . In some examples, a liver mask can be created to separate the segmented liver from the remaining pixels / voxels of the image. Since all liver images are aligned (e.g., at 604), the created liver mask can be applied to all other liver images in the set. Thus, the liver segmentation performed on one of these liver images is valid for all liver images in the set. By performing liver segmentation on one of these liver images after registering the liver images, processing time and resources can be reduced. However, in an alternative embodiment, liver segmentation can be performed on each of multiple images.
[0096] At 608, method 600 includes performing lesion segmentation on one of the liver images. As an example, the lesion segmentation can be a semi-automatic process where the user / physician provides the maximum axial diameter of a tumor within the segmented liver, such as via a user input device (e.g., Figure 1The user input device 32). Then, texture features within a given maximum axial diameter are extracted (e.g., using mean, median, standard deviation, and edge detection filters), and clustering is performed using the k-means method, which can output a texture map. This texture map can be used to define "object" and "background" labels (or seeds), which are further used as inputs to a random walk algorithm. The random walk algorithm can define the edge between a lesion and healthy liver tissue (e.g., parenchyma) based on the similarity or difference between adjacent pixels. Thus, the lesion can be segmented from the surrounding parenchyma in each of a series of acquired liver images. In some examples, a lesion mask can be created to separate the segmented lesion from the remaining pixels / voxels of the image. Since all liver images are aligned (e.g., at 604), the created lesion mask can be applied to all other liver images in the series. Thus, the lesion segmentation performed on one of these liver images is valid for all liver images in the series. By performing lesion segmentation on one of these liver images after registering the liver images, processing time and resources can be reduced. However, in an alternative embodiment, lesion segmentation can be performed on each of multiple images.
[0097] At 610, method 600 includes creating and characterizing a reference region of interest (ROI) on each liver image. Once the different phases of the contrast-enhanced CT examination are registered and resampled to the same resolution, the kinetics of the contrast agent within the tumor can be tracked. To characterize the kinetics of the contrast agent within the tumor, reference measurements extracted from the liver parenchyma surrounding the tumor are used. Thus, a reference ROI is created and characterized for each liver image in the series. Creating the reference ROI includes applying a distance transform from the lesion mask to create an ROI at a distance between 10 mm and 30 mm from the tumor. The ROI is constrained to be within the liver mask. By creating an ROI at a distance away from the tumor, the likelihood that parts of the lesion are inadvertently included in the reference ROI is reduced. Thus, it is expected that the reference ROI reflects the contrast agent uptake kinetics of the liver parenchyma.
[0098] Characterizing the reference ROI includes extracting statistical values from the parenchyma, which include the mean, standard deviation, noise autocorrelation, and count of pixels (or voxels). However, some structures can introduce some bias (other tumors, blood vessels, etc.), so robust measurements are used. The mean can be estimated from the mode of the ROI histogram regularized by a Gaussian filter. For the standard deviation, the half full width half maximum can be used as a robust estimator. Also, the noise autocorrelation can be calculated and used for lesion analysis, as will be detailed below.
[0099] By automatically placing the reference ROI, without relying on the user for correct placement, the position of the reference ROI around the tumor is precisely controlled, and the size of the reference ROI is much larger than the spherically shaped ROI that can be manually defined. Accordingly, the accuracy and repeatability of reference ROI measurements are increased, which also increases the accuracy and repeatability of tumor measurements made therefrom. The reference Figure 8 Describes illustrative examples of creating and characterizing a reference ROI.
[0100] At 612, method 600 includes performing lesion re-segmentation. HCC lesions can have heterogeneous tissue, which includes viable tissue (both arterial and portal veins), necrotic tissue, chemoembolized tissue (if treated with chemoembolization products), or tissue of undetermined nature. The kinetics of the contrast agent in each tissue are different over time. For example, necrotic tissue has an average gray value that remains below that of the reference ROI throughout the multi-phase examination, while typical viable tissue is similar to, higher than, and lower than the reference ROI during the unenhanced, arterial, and portal venous phases, respectively. Accordingly, pixels / voxels will be assigned to different tissue classes according to ad hoc rules defined by a physician (e.g., a hepatologist) regarding the expected contrast agent kinetics in each tissue class (or type). As will be referenced Figure 9 below, the image can undergo noise reduction and filtering in order to reduce noise while retaining contrast information. The contrast information can enable generation of a temporal profile for each pixel (or voxel) to classify the tissue contained in that pixel (or voxel). For example, the average gray value of each pixel / voxel can be tracked across the various examination phases to determine how the gray value (e.g., pixel / voxel intensity) changes relative to the reference ROI during each phase. Then, the temporal profile of a given pixel / voxel can be compared to criteria defined for each tissue class in order to determine the type of tissue in that pixel / voxel. Additionally, in some examples, performing the re-segmentation can include calculating the proportion of each tissue type within the lesion in order to give an overall indication of the tissue composition. Exemplary HCC lesions that have undergone segmentation and re-segmentation are shown in Figure 10 and will be described below.
[0101] At 614, it is determined whether the lesion is in a pre-treatment state. The pre-treatment state corresponds to the original lesion that has not undergone chemotherapy (such as transarterial chemoembolization (TACE) or drug-eluting bead transarterial chemoembolization (DEB-TACE)) or any other interventional procedure (such as radioembolization or radiofrequency ablation). For example, TACE involves injecting an embolic agent such as lipiodol (e.g., ethylated oil) into the artery (e.g., the hepatic artery) that directly supplies blood to the tumor to block the blood supply to the tumor, thereby inducing cell death (e.g., necrosis). Thus, a tumor that has undergone TACE, DEB-TACE, radioembolization, and / or radiofrequency ablation is in a post-treatment state, while a tumor that has not undergone TACE, DEB-TACE, radioembolization, or radiofrequency ablation may be in a pre-treatment state. In some examples, the user may manually indicate whether the lesion is in a pre-treatment state. For example, when entering patient information for a CT scan, the user may be prompted to select one of "pre-treatment lesion" and "post-treatment lesion". As another example, the system may automatically determine whether the lesion is in a pre-treatment state based on the type of ordered assessment (e.g., LI-RADS feature extraction for pre-treatment tumors or LI-RADS treatment response feature extraction for post-treatment tumors). As yet another example, the system may infer whether the lesion is in a pre-treatment state or a post-treatment state based on data from an electronic health record (EHR) associated with the patient. For example, the EHR may include previously performed examinations, diagnoses, and current treatments, which can be used to determine whether the lesion is in a pre-treatment state.
[0102] If the lesion is in a pre-treatment state, method 600 proceeds to 616 and includes extracting LI-RADS features. LI-RADS features include: a 2D lesion diameter, which can be measured from the lesion segmentation and represents the size of the lesion; arterial phase hyperenhancement (APHE), which occurs during the arterial phase; washout (WO), which occurs during the portal venous phase and / or the delayed phase; and the presence of a capsule enhancement (CAP), which occurs during the portal venous phase and / or the delayed phase. APHE, WO, and CAP are detected based on the contrast between the lesion and its surrounding environment (e.g., the reference ROI created and characterized at 610), which is used to assess whether the two distributions are similar or different. The probability that the two distributions are different takes into account changes in Hounsfield units, noise standard deviation and autocorrelation, and lesion / reference ROI size. These parameters are estimated for both the reference ROI and the tumor biopsy as determined by the resegmentation at 612.
[0103] The probabilities of APHE, WO, and CAP are calculated by the confidence interval of the mean. Under the normal distribution assumption, the upper limit (e.g., the boundary) (μ, σ) of the distribution N is defined by Equation 1:
[0104]
[0105] where μ is the mean of the given distribution, σ is the standard deviation of the given distribution, α is the confidence level, t is the Student's one-tailed t-distribution, and N eq is the number of independent voxels in the data (N eq taking into account ROI size, noise autocorrelation, and mode statistical efficiency, which is approximately ⅕). Equation 1 represents the probability α% that the true distribution mean is included in the interval ]-∞, μ sup .
[0106] Similarly, the lower limit of the distribution is defined by Equation 2:
[0107]
[0108] where μ is the mean of the given distribution, σ is the standard deviation of the given distribution, α is the confidence level, t is the Student's one-tailed t-distribution, and N eq is the number of independent voxels in the data.
[0109] Next, the confidence level α% is calculated such that μ 1sup (α) = μ 2inf (α). In this example, μ 1sup is the upper limit of the first distribution (e.g., reference ROI) and μ 2inf is the lower limit of the second distribution (e.g., the viable part of the lesion). Thus, α 2 corresponds to the probability of having APHE washout or enhancing capsule. When the probability is at least the threshold percentage, the corresponding feature (e.g., APHE, WO, or CAP) is determined to be present, while when the probability is less than the threshold percentage, the corresponding feature is determined not to be present. The threshold percentage is the probability above which the corresponding feature can be assumed to be present with high confidence. As a non-limiting example, the threshold percentage is 90%. An example of LI-RADS feature extraction is shown in Figure 11 and will be described below.
[0110] At 618, method 600 includes determining a LI-RADS score based on the LI-RADS features extracted at 616. In at least some examples, the determined LI-RADS score can be a preliminary score or a score range that helps guide the physician in determining the final LI-RADS score. Thus, if needed, the physician can change the score. The LI-RADS score can be determined using known criteria that correlate LI-RADS features (including lesion size and the presence or absence of each of APHE, WO, and CAP) with the relative risk of HCC. The scores range from LR-1 (benign, non-HCC) to LR-5 (definitely HCC). As an illustrative example, a lesion with APHE and WO that is 15 mm in size and has no enhancing capsule corresponds to a LI-RADS score of LR-4 (possibly HCC). By determining the LI-RADS score based on automatically extracted LI-RADS features, the mental burden on the physician is reduced, and the repeatability and accuracy of the score are improved.
[0111] At 620, method 600 includes outputting the image of the re-segmented HCC lesion and the LI-RADS report. For example, the re-segmented HCC lesion including tissue type annotations can be output to a display device (e.g., Figure 1 display device 33) to be shown to the user / physician. The LI-RADS report can include one or more of the determined APHE probability, the determined washout probability, the size and tissue composition of the lesion, and the LI-RADS score, and the LI-RADS report can also be output to the display device. As another example, outputting the image of the re-segmented HCC lesion and the LI-RADS report can include saving the image and the LI-RADS report to a specified storage location. As described above, the physician can use the LI-RADS report to diagnose the patient, including the malignancy of the lesion. Then method 600 can end.
[0112] Return to 614. If the lesion is not in a pre-treatment state, then the lesion is in a post-treatment state (e.g., after TACE or another treatment), and method 600 proceeds to 622 and includes comparing the current HCC lesion re-segmentation with the HCC lesion re-segmentation from a previously acquired one. For example, the proportion of remaining viable tissue, the proportion of necrotic tissue, and the proportion of chemoembolized tissue can be compared with the values determined for the same HCC lesion before TACE (or the current treatment round). However, in other examples, the current HCC lesion re-segmentation may not be compared with the previous HCC lesion re-segmentation, such as when no previous data is available. Additionally, in some examples, LI-RADS features can also be extracted post-treatment to compare the pre-treatment and post-treatment lesion sizes and APHE, WO, and CAP probabilities.
[0113] At 624, method 600 includes determining a treatment response based on the comparison performed at 622. For example, the treatment response can include a response score determined based on predefined criteria, such as the percentage change in the proportion of viable tissue in the tumor and / or the percentage change in the proportion of necrotic tissue in the tumor after a treatment (e.g., TACE). When prior data is not available for a given lesion, the treatment response can be inferred based on data from other HCC lesions with similar post-treatment tissue composition. The response score can give an indication of the patient's prognosis. For example, the response score can be used to determine whether additional treatment with the same chemoembolization agent is expected to further reduce the proportion of viable tissue in the tumor, reduce the tumor size, and / or clear the tumor. For example, a higher response score can indicate a high response of the tumor to the current treatment course. As another example, the response score can suggest considering other treatment options, such as when the response score is low (e.g., the tumor is relatively non-responsive to the current treatment course). Thus, the determined treatment response can help guide the physician in determining the treatment course for the patient.
[0114] At 626, method 600 includes outputting the image of the re-segmented HCC lesion and a treatment report. For example, the re-segmented HCC lesion and the treatment report can be output to a display device and / or a storage location, as described above at 620. The treatment report can include one or more of the determined treatment response, the response score, and treatment recommendations. Then method 600 can end.
[0115] In this way, the method uses image processing and deep learning methods to accelerate and improve the efficiency of tumor detection and characterization. Overall, HCC tumors can be characterized more accurately, and variability can be reduced. Thus, positive patient prognosis can be increased.
[0116] Next, Figure 7 An exemplary workflow is shown for a deep learning network 700 to align images obtained during a multi-phase imaging examination using rigid registration and non-rigid registration. Deep learning network 700 includes a first network portion 702 and a second network portion 704. In the example shown, the first network portion 702 performs rigid registration through an affine transformation, and the second network portion 704 performs non-rigid registration through a voxel deformation diffeomorphic architecture.
[0117] Legend 799 illustrates various elements included in deep learning network 700. As shown in Legend 799, deep learning network 700 includes a plurality of convolutional layers / feature maps connected by one or more operations. The operations receive an input from an external file (e.g., an input image) or a previous layer / feature map, and transform / map the received input to produce the next layer / feature map. Each layer / feature map may include a plurality of neurons. In some embodiments, each neuron may receive an input from a subset of neurons in the previous layer / feature map and may compute a single output based on the received input. The output may be propagated / mapped to a subset or all of the neurons in the next layer / feature map. The layers / feature maps may be described using spatial dimensions such as length, width, and depth, where the dimensions refer to the number of neurons included in the feature map (e.g., how many neurons long, wide, and deep a specified layer / feature map is).
[0118] As shown in Legend 799, the diagonally shaded feature maps (e.g., feature maps 710, 712, 714, 730, 732, 734, and 736) include 3D convolution (CONV 3D) with batch normalization (BN) and leaky ReLu activation (LEAKY RELU). The lighter dot-shaded feature maps (e.g., feature maps 736, 738, 740, and 742) include 3D transposed convolution (CONV 3D transpose) with concatenation (CONCAT) and 3D convolution (CONV 3D + leaky RELU) with leaky ReLu activation. The diamond-shaded feature maps (e.g., feature maps 716a, 716b, 718a, and 718b) include 3D convolution. The darker dot-shaded feature maps (e.g., feature maps 720a and 720b) include global average pooling. The vertically shaded feature map (e.g., feature map 744) includes a velocity field. The non-shaded feature map (e.g., feature map 746) includes an integration layer.
[0119] The source image (e.g., mobile imaging) 706a and the target image (e.g., fixed image) 708a are inputs for the first network portion 702. The source cropped image 706b and the target cropped image 708b are outputs of the first network portion 702 and inputs for the second network portion 704. The source image 706a and the target image 708a are optionally selected from a series of liver images obtained during a multiphase CT scan, as described above with respect to Figure 6As described above. For example, the target image 708a can be used as a template (e.g., a reference configuration) to align with the source image 706a. Thus, the first network part 702 finds the best rigid transformation to be applied to the active image (source image) to match the fixed image (target image). The same image as the target image 708a can be used to align all the remaining images in the series of liver images. One of the remaining images (e.g., not the target image 708a) can be selected as the source image 706a until all the remaining images have undergone registration.
[0120] The source image 706a and the target image 708a are input into the feature map 710. The resulting global average pooling feature map 720a undergoes six parametric transformations 722, including three rotation transformations and three translations (e.g., in x, y, and z). The resulting global average pooling feature map 720b undergoes six parametric evaluations 724, which produce values for the bounding box center ("BB_Center") and produce values for the bounding box size ("BB_size"). The resulting rotated and translated layer is input together with the source image 706a into the spatial transformation 726. The outputs of the spatial transformation 726, the bounding box center, the bounding box size, and the target image 708a are input into the cropping function 728 to crop the source image and the target image with the same bounding box. The cropping function 728 outputs the source cropped image 706b and the target cropped image 708b.
[0121] The source cropped image 706b and the target cropped image 708b are input into the feature map 730 of the second network part 704. The second network part 704 finds the best non-rigid deformation field to be applied to the active image to match the fixed image. The velocity field layer feature map 744 and the integration layer feature map 746 are used to generate the deformation field 748, which provides a matrix of displacement vectors for each voxel (or pixel for 2D images) in the source cropped image 706b relative to each similar voxel in the target cropped image 708b. The deformation field 748 and the source cropped image 706b are input into the spatial transformation 750, which outputs the deformed image 752. The deformed image 752 includes the source cropped image 706b aligned with the target cropped image 708b. Thus, the deformed image 752 is the output of the second network part 704 and the overall output of the deep learning network 700, and has undergone both rigid registration (e.g., via the first network part 702) and non-rigid registration (e.g., via the second network part 704).
[0122] Figure 8 An exemplary workflow 800 for creating a reference ROI characterizing a medical image is schematically shown. Specifically, a reference ROI is created within the medical image 802 obtained during a CT scan of a patient's liver. As described above with respect toFigure 6 As described, the medical image 802 undergoes processing and analysis to segment the liver and lesions within the liver, thereby generating a liver mask 804 and a lesion mask 806. The liver mask 804 shows the liver portion of the medical image 802 in white and the non-liver portion of the medical image in black. The liver mask 804 separates the pixels / voxels corresponding to the non-liver portion of the medical image 802 from the pixels / voxels within the liver. In this way, the liver mask 804 hides the non-liver portion of the medical image 802 from the rest of the workflow 800, such that these pixels / voxels are not considered for the creation and characterization of the reference ROI. Similarly, the lesion mask 806 separates the pixels / voxels corresponding to the lesions from other liver tissues, and shows the lesions in white and the rest of the medical image (including other portions of the liver) in black.
[0123] A distance transform 808 is applied to the lesion mask 806 to place a distance between the edges of the lesions and the reference ROI. This distance can be in a range, for example, between 10 mm and 30 mm. This creates a preliminary reference ROI 810, which is shown as a 3D volume. The preliminary reference ROI 810 does not include the voxels within the distance transform 808, which includes all the voxels within the lesion mask 806. The liver mask 804 and the preliminary reference ROI 810 are input into an intersection function 812, which refines the preliminary reference ROI by constraining the preliminary reference ROI 810 to be located within the liver mask 804. The intersection function 812 outputs a reference ROI 814, which is shown as a 3D volume. Since the reference ROI 814 excludes the voxels within the lesion mask 806 (and the distance from the lesion mask 806) and is constrained to the liver mask 804, the reference ROI 814 includes voxels corresponding to healthy liver tissue.
[0124] Statistics 816 are performed on the created reference ROI 814. The statistics can include (but are not limited to) estimating the average intensity value of the pixels / voxels in the ROI, the standard deviation of the intensity values, the noise autocorrelation, and the quantity. The average value can be estimated by the mode of the ROI histogram, which can be normalized by a Gaussian filter. The relatively large size of the reference ROI (e.g., compared to a volume that can be manually selected by a user / physician) provides a robust estimate of the average value and the standard deviation, because the size of the reference ROI makes the reference ROI less susceptible to structures other than the liver parenchyma (e.g., other tumors, blood vessels) and registration errors. The workflow 800 can output a reference ROI analysis 818, which can include the statistics 816. The reference ROI analysis 818 can be used for subsequent HCC lesion re-segmentation and characterization, as described above with respect to Figure 6 described and as detailed below.
[0125] Figure 9 An exemplary workflow 900 for performing HCC lesion re-segmentation and characterization in images obtained during contrast-enhanced multi-phase CT scans of the liver is schematically illustrated. For example, workflow 900 may be executed by an image processing system (e.g., Figure 1 image processing system 100) as part of Figure 6 method 600.
[0126] At 902, a plurality of images are input into workflow 900. Each image includes one image acquired during one phase of a multi-phase examination, as detailed above with respect to Figure 6 . In the Figure 9 example, three images are shown: an unenhanced (UE) image 904a, an arterial phase image (Art) 906a, and a portal venous phase image (Port) 908a. At 910, each image is averaged, and at 912, denoising and filtering are performed on each averaged image using a guided filter 911. The guided filter 911 may include a joint bilateral filter that has the averaged image as a guidance image, enabling noise reduction while maintaining contrast. The resulting filtered images are shown at 914. Specifically, 914 shows the filtered unenhanced image 904b, the filtered arterial phase image 906b, and the filtered portal venous phase image 908b.
[0127] At 916, a per-pixel (or per-voxel) temporal analysis is performed. That is, the intensity of each pixel (or voxel) may be tracked through a series of filtered images to generate a temporal profile, which is shown as a graph 918. Graph 918 includes time as the horizontal axis, where the time points at which different images are acquired are marked on this horizontal axis. For example, the vertical axis represents the change (e.g., Δ Hounsfield units or HU) in the intensity of a given pixel relative to a reference ROI created and analyzed using, for example, Figure 8 the workflow. Each line on the graph shows a different pixel.
[0128] At 920, each pixel (or voxel) is classified as a specific type of tissue based on its temporal profile. Possible tissue types include necrotic, active arterial, active portal venous, chemoembolized (e.g., after transarterial chemoembolization with a chemotherapeutic agent such as lipiodol), peripherally enhanced, indeterminate, and parenchyma. The kinetics of the contrast agent over time are different in each tissue type. For example, the average gray value of necrotic tissue remains less than the average gray value of the reference ROI over time, while the average gray value of typical living tissue is similar to, higher than, and lower than the average gray value of the reference ROI during the unenhanced, arterial, and portal venous phases, respectively. Thus, at 920, pixels (or voxels) are assigned to different classes based on specific classification rules created by a qualified physician (e.g., a hepatologist or oncologist). In this way, the lesion can be re-segmented based on the tissue type within the lesion. The resulting re-segmented lesion can be output as an image and / or report containing quantitative and / or qualitative information about the tissue composition of the lesion.
[0129] Figure 10 An exemplary image 1000 of an HCC lesion that has been segmented and re-segmented, such as according to Figure 6 the method and using the workflow outlined in Figure 9 is shown. Specifically, image 1000 shows the post-treatment lesion. The first image 1002 shows the segmented lesion, including the lesion boundary 1006. The lesion boundary 1006 separates the lesion from the liver tissue surrounding the lesion. The second image 1004 shows the re-segmented lesion. The re-segmented lesion includes different tissue classifications for different parts of the segmented lesion. The viable region 1008 is bounded by a dashed boundary, the chemoembolized region 1010 is bounded by a longer dashed boundary, and the necrotic region 1012 is bounded by a shorter dashed boundary. As described above with respect to Figure 6 , in some examples, an image processing system (e.g., Figure 1 image processing system 100) can use the re-segmented lesion to determine the percentage of the tumor that remains viable and / or estimate the effectiveness of the treatment.
[0130] Figure 11 An exemplary set of multiple liver images 1100 obtained during a multi-phase CT scan is shown, along with the use of image processing and deep learning methods by an image processing system (such as Figure 1 image processing system 100) according to Figure 6Features extracted by analyzing these images using, for example, the method of [reference]. Specifically, a plurality of exemplary liver images 1100 are arranged in a grid, where each row corresponds to a series of four images obtained during a multiphase CT scan and the resulting feature analysis. Each column represents a liver image from a single phase of the CT scan or a given feature analysis. For example, the first row 1102 shows the first series of liver images and the resulting feature analysis, the second row 1104 shows the second series of liver images and the resulting feature analysis, the third row 1106 shows the third series of liver images and the resulting feature analysis, and the fourth row 1108 shows the fourth series of liver images and the resulting feature analysis. For each liver image in the series of liver images, the first column 1110 shows the non-contrast (NC) image, the second column 1112 shows the arterial phase image, the third column 1114 shows the portal venous phase image, the fourth column 1116 shows the delayed phase image, the fifth column 1118 shows lesion re-segmentation, the sixth column 1120 shows the probability of viable area and arterial phase hyper-enhancement (APHE), the seventh column 1122 shows the probability of viable area and washout (WO), and the eighth column 1124 shows the probability of the peritumoral enhancement area and capsule (CAP). In some examples, the image processing system may display all or some of the exemplary liver images 1100 to the user, such as via a display device (e.g., Figure 1 display device 33).
[0131] The APHE probability in column 1120 is determined from the arterial phase image (column 1112), the washout probability in column 1122 is determined from the portal venous phase image (column 1114), and the capsule probability in column 1124 is determined from the delayed phase image (column 1116), such as Figure 6 described at 616 of [reference]. These probabilities are affected by the difference in HU (e.g., pixel intensity) between the lesion and the surrounding tissue, the noise standard deviation, the noise autocorrelation, and the region size. For example, when the difference in HU is higher, the probability is higher. Also, as the noise increases, the probability decreases. Also, a higher noise autocorrelation results in a smaller probability. Also, the probability decreases as the region size decreases. In Figure 11 an example, a probability of at least 90% triggers a positive identification of the corresponding feature, while a probability less than 90% indicates the absence of the corresponding feature.
[0132] First, looking at the first series of liver images in the first row 1102, the non-contrast image (column 1110), the arterial phase image (column 1112), the portal venous phase image (column 1114), and the delayed phase image (column 1116) are registered with each other, such as with respect to Figure 7 described. The boundary 1126 of the lesion is shown in the arterial phase image (column 1112), which can be determined by extracting texture features and using a random walk segmentation algorithm, as described above with respect to Figure 6As described. As shown, the pixel intensity of the segmented lesion is more similar to the healthy tissue surrounding the lesion in the arterial phase image (column 1112) than in the non-contrast image (column 1110). Additionally, the portal venous phase image (column 1114) shows that the lesion darkens (e.g., pixel intensity decreases) relative to the healthy tissue surrounding the image, thereby creating a strong contrast between the lesion and the surrounding tissue. The pixel intensity in the lesion remains relatively low in the delayed phase image (column 1116). In the portal venous phase image and the delayed phase image, the periphery of the lesion is brighter than the surrounding tissue.
[0133] The image processing system generates a time profile for each pixel in the segmented lesion, such as described above with respect to Figure 9 and uses this time profile to re-segment the lesion. The resulting re-segmented lesion shown in column 1118 includes a lesion boundary 1126, an arterial viable region within the dashed boundary 1128, a portal venous viable region within the dashed boundary 1130, a peripheral enhancement region within the long dashed boundary 1132, and a necrotic region within the short dashed boundary 1134. The arterial viable region is also shown in column 1120, which shows the probability of 75% APHE. Therefore, APHE is determined to be absent because the probability is less than 90%. The portal venous viable region is also shown in column 1122, which shows a clearance probability of 100% (e.g., there is WO). The peripheral enhancement region is also shown in column 1124, which shows a capsule probability of 100% (e.g., there is a capsule). The image processing system can determine the APHE, clearance, and capsule probabilities, as described above with respect to Figure 6 Because the diameter of the lesion shown in the first row 1102 < 20 mm, there is no APHE, and it has WO and an enhancing capsule, the image processing system can score the lesion as LR-4.
[0134] The second series of liver images in the second row 1104 are arranged as described above. The boundary 1134 of the lesion is shown in the arterial phase image (column 1112). As shown, the pixel intensity of the segmented lesion is more similar to the healthy tissue surrounding the lesion in the arterial phase image (column 1112) than in the non-contrast image (column 1110). Additionally, in the arterial phase image (column 1112), the pixel intensity of the segmented lesion in the second series of liver images (second row 1104) is more similar to the healthy tissue surrounding the lesion than the segmented lesion in the first series of liver images (first row 1102). Similar to the first series of liver images (first row 1102), the portal venous phase image (column 1114) shows that the lesion darkens (e.g., pixel intensity decreases) relative to the healthy tissue surrounding the image, thereby creating a contrast between the lesion and the surrounding tissue and a brightening of the lesion periphery. The pixel intensity in the lesion remains relatively low in the delayed phase image (column 1116).
[0135] The resulting re-segmented lesion shown in column 1118 includes a lesion boundary 1134, an arterial viable region within the dashed boundary 1136, a portal venous viable region within the dashed boundary 1138, a peripheral enhancement region within the long dashed boundary 1140, and a necrotic region within the short dashed boundary 1142. The arterial viable region is also shown in column 1120, which shows a 100% probability of APHE. The portal venous viable region is also shown in column 1122, which shows a 100% clearance probability. The peripheral enhancement region is also shown in column 1124, which shows a 100% capsule probability. Because the lesion shown in the second row 1104 has a diameter between 10 mm and 20 mm and has APHE, WO, and an enhancing capsule, the image processing system can score the lesion as LR-5.
[0136] Now looking at the third series of liver images in the third row 1106, the boundary 1144 of the lesion is shown in the arterial phase image (column 1112). As shown, the segmented lesion has a higher (e.g., brighter) pixel intensity in the arterial phase image (column 1112) compared to the non-contrast image (column 1110). Additionally, the pixel intensity of the segmented lesion in the arterial phase image (column 1112) of the third series of liver images (third row 1106) is higher than the pixel intensity of the segmented lesion in the arterial phase image (column 1112) of the first series of liver images (first row 1102) and the segmented lesion in the second series of liver images (second row 1104). In the third series of liver images, in the portal venous phase image (column 1114) and the delayed phase image (column 1116), there is no significant contrast between the lesion and the parenchyma. Additionally, in the portal venous phase image (column 1114) and the delayed phase image (column 1116), there is no significant bright contrast between the periphery of the lesion and the parenchyma.
[0137] The resulting re-segmented lesion shown in column 1118 includes a lesion boundary 1144; an arterial viable region within the dashed boundary 1146, which substantially overlaps with the lesion boundary 1144; and a portal venous viable region within the dashed boundary 1148. The arterial viable region is also shown in column 1120, which shows a 100% probability of APHE. The portal venous viable region is also shown in column 1122, which shows a 60% clearance probability. There is no visible peripheral enhancement region in column 1124, providing a 0% capsule probability. Because the lesion shown in the third row 1106 has a diameter less than 10 mm and has APHE but no WO and no enhancing capsule, the image processing system can score the lesion as LR-4.
[0138] The fourth series of liver images in the fourth row 1108 shows the lesion boundary 1150 in the arterial phase image (column 1112). As shown, the lesion is substantially indistinguishable from the parenchyma in the non-contrast image (column 1110), and the pixel intensity in the arterial phase image (column 1112) is higher than that in the non-contrast image (column 1110). In addition, the pixel intensity of the segmented lesion in the arterial phase image (column 1112) of the fourth series of liver images (fourth row 1108) is higher than that of the segmented lesions in the arterial phase images (column 1112) of the first series of liver images (first row 1102), the second series of liver images (second row 1104), and the third series of liver images (third row 1106). In the fourth series of liver images, the lesion has a periphery with higher intensity pixels and a central region with lower intensity pixels (e.g., relative to the periphery) in the portal venous phase image (column 1114) and the delayed phase image (column 1116).
[0139] The resulting re-segmented lesion shown in column 1118 includes the lesion boundary 1150; an arterial viable region within the dashed boundary 1152 that substantially overlaps with the lesion boundary 1150; a portal venous viable region within the dashed boundary 1154; and a peripheral enhancement region within the long dashed boundary 1156. The arterial viable region is also shown in column 1120, which shows the probability of 100% APHE. The portal venous viable region is also shown in column 1122, which shows a clearance probability of 45%. The peripheral enhancement region is also shown in column 1122, which shows a capsule probability of 100%. Since the lesion shown in the fourth row 1108 has a diameter between 10 mm and 20 mm and has APHE but no WO and has an enhancing capsule, the image processing system can score the lesion as LR-4 or LR-5.
[0140] The technical effect of automatically characterizing liver cancer lesions using deep learning methods is to improve the accuracy of characterization, thereby improving the accuracy and timeliness of patient treatment decisions.
[0141] One example provides a method that includes acquiring a plurality of medical images over time during an examination; registering the plurality of medical images; after registering the plurality of medical images, segmenting an anatomical structure in one of the plurality of medical images; creating and characterizing a reference region of interest (ROI) in each of the plurality of medical images; determining characteristics of the anatomical structure by tracking pixel values of the segmented anatomical structure over time; and outputting the determined characteristics on a display device.
[0142] In one example, registering the plurality of medical images includes applying rigid and non-rigid registration to a source image and a target image selected from the plurality of medical images.
[0143] In one example, the method further includes, after registering the plurality of medical images and segmenting the anatomical structure, segmenting a lesion within the anatomical structure. In one example, segmenting the anatomical structure includes using a convolutional neural network, and segmenting the lesion includes extracting and clustering texture features based on a maximum axial diameter of the received lesion, the texture features including at least one of an average value, a median value, and a standard deviation of pixel values within the segmented anatomical structure in each of the plurality of medical images. In some examples, the examination is a multi-phase contrast-enhanced computed tomography (CT) examination, and segmenting the anatomical structure includes segmenting the liver. For example, the multi-phase contrast-enhanced CT examination includes a non-enhanced phase, an arterial phase after the non-enhanced phase, a portal venous phase after the arterial phase, and a delayed phase after the portal venous phase, and the plurality of medical images includes a first CT image of the liver obtained during the non-enhanced phase, a second CT image of the liver obtained during the arterial phase, a third CT image of the liver obtained during the portal venous phase, and a fourth CT image of the liver obtained during the delayed phase.
[0144] In an example, the reference ROI includes tissue outside the segmented lesion and within the segmented liver, wherein the characteristic of the anatomical structure includes a tissue classification of each part of the segmented lesion, and wherein determining the characteristic of the anatomical structure by tracking the pixel values of the segmented anatomical structure over time includes determining the tissue classification of each part based on the value of each pixel within each part of the segmented lesion in each of the plurality of medical images relative to the reference ROI. For example, the tissue classification is one of necrosis, viable, chemoembolization, indeterminate, peripheral enhancement, and parenchyma.
[0145] In one example, the characteristic of the anatomical structure further includes a treatment effect, and determining the characteristic of the anatomical structure by tracking the pixel values of the segmented anatomical structure over time further includes determining the treatment effect by comparing the tissue classification of each part of the segmented lesion from a first examination, performed before the treatment, with the tissue classification of each part of the segmented lesion from a second examination, performed after the treatment.
[0146] In one example, the characteristic of the anatomical structure further includes a Liver Imaging Reporting and Data System (LI-RADS) score, and determining the characteristic of the anatomical structure by tracking the pixel values of the segmented anatomical structure over time further includes: determining an arterial phase high enhancement (APHE), washout, and capsule probability of a contrast agent used in the multi-phase contrast-enhanced CT examination based on the pixel values of the segmented lesion in each of the plurality of medical images relative to the reference ROI.
[0147] Exemplary methods include receiving a series of liver images from a multiphase examination; aligning the series of liver images; segmenting the liver and lesions within the series of aligned liver images; re-segmenting the lesions based on changes in pixel values of the segmented lesions on the series of aligned liver images; and outputting an analysis of the re-segmented lesions. In an example, the series of liver images includes one liver image from each phase of the multiphase examination, and segmenting the liver includes: using a convolutional neural network to define a liver boundary in one image of the series of aligned liver images; and applying the liver boundary to each additional image of the series of aligned liver images. In one example, segmenting the lesions includes: based on texture features of one image of the series of aligned liver images, defining a boundary of the lesion within the liver boundary in the one image, the texture features including at least one of an average value, a median value, and a standard deviation of pixel values within the liver boundary in the one image; and applying the boundary of the lesion to each additional image of the series of aligned liver images.
[0148] In an example, re-segmenting the lesions based on changes in pixel values of the segmented lesions on the series of aligned liver images includes: in each image of the series of aligned liver images, comparing the pixel values of the segmented lesions with the pixel values in a reference region that is within the boundary of the liver and outside the boundary of the lesion; and based on the change in pixel values of each sub-segment of the lesion relative to the pixel values in the reference region in the series of aligned liver images, determining the classification of the sub-segment as one of viable tissue, necrotic tissue, indeterminate tissue, and chemoembolized tissue. In some examples, outputting an analysis of the re-segmented lesions includes outputting the classification of each sub-segment of the lesion as an annotated image of the re-segmented lesion. In one example, the method further includes determining a diagnostic score of the lesion based on the intensity distribution of each pixel within the viable tissue sub-segment on the series of aligned liver images, and wherein outputting an analysis of the re-segmented lesion includes outputting the diagnostic score.
[0149] An exemplary system includes: a computed tomography (CT) system; a memory that stores instructions; and a processor communicatively coupled to the memory and configured to, when executing the instructions: receive a plurality of images acquired by the CT system during a multiphase examination enhanced with a contrast agent; identify boundaries of anatomical features in at least one of the plurality of images via a deep learning model; determine a diagnostic score for the anatomical feature based on low enhancement and high enhancement of the contrast agent within regions of the anatomical feature throughout the multiphase examination; and output the diagnostic score. In an example, the anatomical feature is the liver, and the processor is further configured to, when executing the instructions: identify boundaries of cancerous lesions within the liver via a random walk algorithm. In an example, the region includes a region of viable tissue within the lesion, and wherein the processor is further configured to, when executing the instructions: create a reference region of interest (ROI) outside the lesion and within the liver; statistically analyze the reference ROI to determine an average value in each of the plurality of images; and re-segment the region of viable tissue within the lesion by generating a temporal profile of the change in brightness in each pixel of the lesion compared to the reference ROI over the plurality of images, the re-segmented region including a peripheral enhancement zone. In one example, the multiphase examination includes a non-enhanced phase, an arterial phase, a portal venous phase, and a delayed phase, and the processor is further configured to, when executing the instructions: calculate a first probability of high enhancement of the contrast agent in the arterial phase within the re-segmented region of viable tissue within the lesion; calculate a second probability of low enhancement of the contrast agent in the portal venous phase within the re-segmented region of viable tissue within the lesion; calculate a third probability of high enhancement of the contrast agent in the portal venous phase within the peripheral enhancement zone; and determine the diagnostic score based on the first probability, the second probability, and the third probability.
[0150] When introducing elements of various embodiments of the present disclosure, the words "a", "an", and "the" are intended to mean that there is one or more of these elements. The terms "first", "second", etc. do not denote any order, quantity, or importance, but are used to distinguish one element from another. The terms "comprising", "including", and "having" are intended to be inclusive and mean that additional elements may exist in addition to the listed elements. As used herein, terms such as "connected to", "coupled to", etc., an object (e.g., a material, element, structure, component, etc.) may be connected to or coupled to another object, regardless of whether the one object is directly connected or coupled to the other object, or whether there is one or more intervening objects between the one object and the other object. Further, it should be understood that references to "one embodiment" or "an embodiment" of the present disclosure are not to be construed as excluding the existence of additional embodiments that also incorporate the recited features.
[0151] Except for any previously indicated modifications, many other variations and alternative arrangements can be devised by those skilled in the art without departing from the substance and scope of this description, and the appended claims are intended to cover such modifications and arrangements. Thus, although the information has been specifically and detailedly described above in connection with the presently considered most practical and preferred aspects, it will be apparent to those of ordinary skill in the art that many modifications can be made without departing from the principles and concepts set forth herein, including but not limited to form, function, manner of operation, and use. Similarly, as used herein in all respects, the examples and embodiments are only intended to be illustrative and should not be construed in any way as restrictive.
Claims
1. A method for characterizing anatomical features in medical images, the method comprising: Acquiring a plurality of medical images over time during an examination; Registering the plurality of medical images; After registering the plurality of medical images, segmenting anatomical structures in one of the plurality of medical images; Creating and characterizing a reference region of interest (ROI) in each of the plurality of medical images; Determining characteristics of the anatomical structure by tracking pixel values of the segmented anatomical structure over time; Performing class-based re-segmentation of the viable tissue region within the lesion by generating a temporal profile of the brightness change in each pixel of the lesion compared to the reference ROI on the plurality of medical images, the re-segmented region including a peripheral enhancement zone; And Outputting the determined characteristics on a display device.
2. The method according to claim 1, wherein registering the plurality of medical images includes applying rigid and non-rigid registration to a source image and a target image selected from the plurality of medical images.
3. The method according to claim 1, the method further comprising: After registering the plurality of medical images and segmenting the anatomical structure, segmenting a lesion within the anatomical structure.
4. The method according to claim 3, wherein segmenting the anatomical structure includes using a convolutional neural network, and segmenting the lesion includes extracting and clustering texture features based on the maximum axial diameter of the received lesion, the texture features including at least one of the mean, median, and standard deviation of the pixel values within the segmented anatomical structure in each of the plurality of medical images.
5. The method according to claim 3, wherein the examination is a multi-phase contrast-enhanced computed tomography (CT) examination, and segmenting the anatomical structure includes segmenting the liver.
6. The method according to claim 5, wherein the multi-phase contrast-enhanced CT examination includes a non-enhanced phase, an arterial phase after the non-enhanced phase, a portal venous phase after the arterial phase, and a delayed phase after the portal venous phase, and the plurality of medical images include a first CT image of the liver obtained during the non-enhanced phase, a second CT image of the liver obtained during the arterial phase, a third CT image of the liver obtained during the portal venous phase, and a fourth CT image of the liver obtained during the delayed phase.
7. The method according to claim 5, wherein the reference ROI includes tissue outside the segmented lesion and within the segmented liver, wherein the characteristics of the anatomical structure include the tissue classification of each part of the segmented lesion, and wherein determining the characteristics of the anatomical structure by tracking the pixel values of the segmented anatomical structure over time includes determining the tissue classification of the part based on the value of each pixel within each part of the segmented lesion in each of the plurality of medical images relative to the reference ROI.
8. The method according to claim 7, wherein the tissue is classified as one of necrosis, viable, chemoembolization, indeterminate, peripheral enhancement, and parenchyma.
9. The method according to claim 8, wherein the characteristic of the anatomical structure further includes a treatment effect, and determining the characteristic of the anatomical structure by tracking the pixel values of the segmented anatomical structure over time further includes determining the treatment effect by comparing the tissue classification of each part of the segmented lesion from a first examination, which is performed before the treatment, with the tissue classification of each part of the segmented lesion from a second examination, which is performed after the treatment.
10. The method according to claim 7, wherein the characteristic of the anatomical structure further includes a Liver Imaging Reporting and Data System (LI-RADS) score, and determining the characteristic of the anatomical structure by tracking the pixel values of the segmented anatomical structure over time further includes: Based on the pixel values of the segmented lesion in each of the plurality of medical images relative to the reference ROI, determining the arterial phase high enhancement (APHE), washout, and capsule probability of the contrast agent used in the multi-phase contrast-enhanced CT examination.
11. A method for characterizing anatomical features in medical images, the method comprising: Receiving a series of liver images from a multi-phase examination; Aligning the series of liver images; Segmenting the liver and lesions within the series of aligned liver images; Re-segmenting the lesions into categories based on changes in the pixel values of the segmented lesions on the series of aligned liver images; and Outputting an analysis of the re-segmented lesions.
12. A system for characterizing anatomical features in medical images, the system comprising: A computed tomography (CT) system; A memory that stores instructions; And A processor communicatively coupled to the memory and configured, when executing the instructions, to: Receive a plurality of images acquired by the CT system during a multi-phase examination enhanced with a contrast agent; Identify the boundaries of anatomical features in at least one of the plurality of images via a deep learning model; Determine a diagnostic score of the anatomical features based on low and high enhancement of the contrast agent within regions of the anatomical features throughout the multi-phase examination; And Output the diagnostic score, wherein the processor is further configured, when executing the instructions, to re-segment the viable tissue regions within the lesions into categories by generating a temporal profile of the brightness change in each pixel of the lesions compared to a reference region of interest (ROI) on the plurality of images, and the re-segmented regions include a peripherally enhanced region.
13. The system according to claim 12, wherein the anatomical feature is the liver, and the processor is further configured, when executing the instructions, to: Identify the boundaries of cancerous lesions within the liver via a random walk algorithm.
14. The system according to claim 13, wherein the region includes a viable tissue region within the lesion, and wherein the processor, when executing the instructions, is further configured to: create the reference ROI outside the lesion and within the liver; and statistically analyze the reference ROI to determine an average value in each of the plurality of images.
15. The system according to claim 14, wherein the multi-phase examination includes an unenhanced phase, an arterial phase, a portal venous phase, and a delayed phase, and wherein the processor, when executing the instructions, is further configured to: calculate a first probability of arterial phase hyper-enhancement of the contrast agent within the re-segmented region of the viable tissue within the lesion; calculate a second probability of portal venous phase hypo-enhancement of the contrast agent within the re-segmented region of the viable tissue within the lesion; calculate a third probability of portal venous phase hyper-enhancement of the contrast agent within the peripheral enhancement region; and determine the diagnostic score based on the first probability, the second probability, and the third probability.
Citation Information
Patent Citations
Method and system for lesion segmentation
US20110158491A1
System for computing quantitative biomarkers of texture features in tomographic images
US9092691B1