Lesion Detection Artificial Intelligence Pipeline Computing System
Through a multi-level machine learning computer model pipeline, liver lesions are automatically identified and classified, solving the problems of low efficiency and large errors in liver lesion detection in existing technologies, and achieving efficient and accurate lesion detection and classification.
Patent Information
- Application Number
- CN202111267667.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-30
- Filing Date
- 2021-10-29
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2041-10-29
AI Technical Summary
In existing technologies, the detection and classification of liver lesions mainly rely on manual evaluation by human experts, which is inefficient and prone to errors. Especially when processing a large number of medical images, it is difficult to achieve accurate and efficient lesion detection and classification.
A multi-level machine learning computer model pipeline is used, including phase classification, anatomical structure detection, lesion detection, segmentation and classification. The AI pipeline automatically identifies and classifies liver lesions, generates a lesion list and provides accurate lesion contours, reducing false positive detections.
It achieves automated and accurate detection and classification of liver lesions, reduces errors, improves detection efficiency, and provides reliable lesion information for subsequent treatment decisions.
Smart Images

Figure CN114529501B_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to an improved data processing apparatus and method, and more particularly to mechanisms for providing an artificial intelligence pipeline computing system capable of identifying anatomical features of interest (such as lesions, etc.) in electronic medical image data. Background Art
[0002] Liver lesions (liver lesion) are the groups of abnormal cells in the liver of a biological entity, and may also be referred to as masses or tumors. Noncancerous or benign liver lesions are common and do not spread to other areas of the body. Such benign liver lesions do not cause any health problems usually. However, some liver lesions are formed due to cancer. Patients with specific medical conditions (medical condition) may be more likely to have cancerous liver lesions than other patients. These medical conditions include, for example, hepatitis B or hepatitis C, cirrhosis, iron storage disease (hemochromatosis), obesity, or exposure to toxic chemicals (such as arsenic or aflatoxin).
[0003] Liver lesions can typically only be identified by having a medical imaging test, such as, for example, ultrasound, magnetic resonance imaging (MRI), computerized tomography (CT), or positron emission tomography (PET) scan. Such medical imaging tests must be reviewed by a human medical imaging subject matter expert (SME), who must use their own knowledge and expertise, as well as human ability, to review patterns in the images to determine whether the medical imaging test shows any lesions. If a human SME identifies a potential cancerous lesion, the patient's physician may perform a biopsy to determine whether the lesion is cancerous.
[0004] Abdominal contrast-enhanced (CE) CT is the current standard for evaluating various abnormalities (e.g., lesions) in the liver. These lesions can be assessed by human SMEs as malignant (hepatocellular carcinoma, cholangiocarcinoma, angiosarcoma, metastasis and other malignant lesions) or benign (hemangioma, focal nodular hyperplasia, adenoma, cyst or lipoma, granuloma, etc.). Manual evaluation of such images by human SMEs is important for guiding subsequent interventions. Many times, in order to appropriately evaluate lesions in CE CT, a multi-stage study is performed, in which the multi-stage study provides medical imaging of different levels of enhancement of healthy liver parenchyma and comparison with lesion enhancement to determine differential detection. The human SME can then determine the diagnosis of the lesion based on these differences. Summary of the Invention
[0005] This Summary is provided to introduce a selection of concepts in a simplified form that are further described herein in the Detailed Description. This Summary is not intended to identify key factors or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0006] In one illustrative embodiment, a method is provided in a data processing system comprising at least one processor and at least one memory, the at least one memory comprising instructions executed by the at least one processor to implement a lesion detection and classification artificial intelligence (AI) pipeline comprising a plurality of trained machine learning computer models, wherein the AI pipeline execution comprises the following method. The AI pipeline execution comprises the following method: processing an input volume of a medical image by one or more first machine learning computer models of the AI pipeline to determine whether the input volume depicts a predetermined amount of anatomical structure of interest. The method further comprises determining, by logic of the AI pipeline, whether the output of the one or more first machine learning computer models satisfies one or more predetermined criteria. The one or more predetermined criteria include a predetermined amount of anatomical structure of interest depicted in the input volume.
[0007] In response to determining that the output of the one or more first machine learning computer models meets the one or more predetermined criteria, the method performs a lesion processing operation. The lesion processing operation includes: processing the input volume by one or more second machine learning computer models of the AI pipeline to detect lesions corresponding to the anatomical structure of interest present in the medical image of the input volume. The lesion processing operation also includes: processing the detected lesions by one or more third machine learning computer models of the AI pipeline to perform lesion segmentation, and combining lesion contours of different medical images associated with the same lesion in the input medical image to generate a lesion list and lesion contours. The lesion processing operation also includes: processing the lesion list and contours associated with the lesions in the lesion list by one or more fourth machine learning computer models of the AI pipeline to classify the lesions in the lesion list into one or more predetermined lesion categories corresponding to the anatomical structure of interest.
[0008] The method also includes outputting, by the AI pipeline, the list of lesions and the classifications associated with the lesions for processing by a downstream computing system. Using this approach, improved lesion detection and classification are enabled in images of anatomical structures of interest. The method provides an automated computing tool that automatically identifies lesions in anatomical structures and generates outlines of these lesions while minimizing false positives and simultaneously associating lesion outlines associated with the same lesion. This enables more accurate detection and classification of lesions compared to prior art mechanisms.
[0009] In some demonstrative embodiments, the one or more first machine learning computer models of the AI pipeline include a body part detection machine learning model executed on the input volume, detecting a portion of a biological entity's body represented in the input volume, and determining whether the portion of the biological entity's body corresponds to a body part in which the anatomical structure of interest is present. In some demonstrative embodiments, the body part is a human abdomen and the anatomical structure of interest is a human liver. This operation allows the AI pipeline mechanism to first ensure that the image volume being processed provides an image of a portion of the biological entity's body (e.g., the abdomen) in which a particular anatomical structure of interest (e.g., the liver) is present, so that processing resources are not applied to an image volume in which the anatomical structure is not present.
[0010] In some demonstrative embodiments, the one or more first machine learning computer models of the AI pipeline include a phase classification machine learning model and an anatomical structure detection machine learning model that operate on the input volume, wherein the phase classification machine learning model determines the phase of medical imaging represented in each of the medical images of the input volume, and wherein the anatomical structure detection machine learning model determines the amount of the anatomical structure of interest represented in the input volume. These operations ensure that the images of the volume all represent the same phase, and thus, differences between images due to different phases will not cause errors in lesion detection. In addition, determining whether a specific amount of the anatomical structure of interest is present in the input volume provides another mechanism for ensuring that the images will provide useful results by representing at least a minimum amount of the anatomical structure of interest, so that a reliable result of whether a patient has a lesion can be made by the AI pipeline.
[0011] In some demonstrative embodiments, the one or more criteria further include an input volume having medical imaging of a single phase present in the input volume, and wherein a minimum threshold amount of anatomical structures of interest present in the input volume is determined based on the results of the operation of the phase classification machine learning model on the input volume and based on the results of the operation of the anatomical structure detection machine learning model on the input volume. In some demonstrative embodiments, in response to the results of the operation of the phase classification machine learning model indicating the presence of medical imaging of more than one phase in the input volume, a subset of medical images in the input volume corresponding to the phase of interest is selected, and the lesion processing operation is performed only on the selected subset of medical images in the input volume. Similarly, these operations ensure the reliability of the results of the AI pipeline by eliminating potential sources of error due to the presence of multiple phases in the input volume. By providing the ability to select a subset of volumes associated with a single phase, this allows for adaptive processing of the input volume without necessarily rejecting the entire volume. In some demonstrative embodiments, the phase of interest is one of a pre-contrast phase, an arterial phase, a portal phase, or a delayed phase.
[0012] In some demonstrative embodiments, determining whether the output of one or more first machine learning computer models satisfies one or more predetermined criteria further comprises: performing, by the axial scoring logic of the AI pipeline, an axial scoring operation on the medical images in the input volume, wherein the axial scoring operation comprises scoring the images in the input volume according to a predefined scoring algorithm to infer an axial score of a most inferior slice in the volume (MISV) and a most superior slice in the volume (MSSV) based on scores associated with the images, and identifying a score of an anatomical structure of interest present in the input volume based on the axial scores of the MISV and the MSSV. These operations provide a mechanism for determining the amount of anatomical structure present in the input volume. The scoring provides a measurement mechanism based on the range of slices present in the input volume.
[0013] In some demonstrative embodiments, the one or more second machine learning computer models of the AI pipeline comprise an ensemble of machine learning computer models, wherein each machine learning computer model in the ensemble is trained through a machine learning process to process an input volume differently than other machine learning computer models in the ensemble to generate corresponding lesion detection predictions. By integrating differently trained machine learning computer models, the machine learning computer models can compensate for loss / error function focus and / or potential weaknesses in training of other computer models in the ensemble, thereby improving overall lesion detection.
[0014] In some demonstrative embodiments, the integration of the machine learning computer models includes a mask generation machine learning model, an input volume processing machine learning model, and a masked input volume processing machine learning model, wherein the mask generation machine learning model is trained by a machine learning process to generate a mask corresponding to the anatomical structure of interest, and the mask is applied to the input volume to generate a masked input volume, wherein the input volume processing machine learning model is trained by a machine learning process to generate a first lesion prediction output based on the input volume, and wherein the masked input volume processing machine learning model is trained by a machine learning process to generate a second lesion prediction output based on the masked input volume. By generating the mask and having the machine learning model operate on the masked input volume, the processing performed by the machine learning model can be focused on portions of the image corresponding to the anatomical structure of interest. Combining the operations of this machine learning model with other machine learning models that operate on non-masked inputs allows for better lesion detection.
[0015] In some demonstrative embodiments, a first machine learning model in the ensemble is trained using a first loss function that penalizes errors in false-negative classification of lesions, and a second machine learning model in the ensemble, different from the first machine learning model, is trained using a second loss function that penalizes errors in false-positive classification of lesions. By having multiple machine learning models that implement different loss functions, each machine learning model can compensate for weaknesses in sensitivity and specificity of the other machine learning models.
[0016] In some demonstrative embodiments, the ensemble logic applies a third loss function that compares a first lesion detection output of the first machine learning model with a second lesion detection output of the second machine learning model and operates to reconcile the first lesion detection output with the second lesion detection output. The third loss function ensures that the outputs of the ensemble's differently trained machine learning models are consistent with one another during training, thereby improving the training of the different machine learning models, even if they were trained using different loss functions.
[0017] In some demonstrative embodiments, the one or more third machine learning computer models of the AI pipeline include: a lesion segmentation machine learning model that segments the image of the input volume into contours corresponding to lesions to generate lesion segmentations; a z-connectivity machine learning model that combines a subset of lesion segmentations based on the segmentations generated by the lesion segmentation machine learning model that are determined to be associated with the same lesion in three-dimensional space; and a contour refinement machine learning model that separates lesions in the input volume and refines the contours of lesions. These models allow for identification of lesions in three-dimensional space by connecting two-dimensional lesion contours associated with the same lesion and then refining the contours of these lesions. Thus, a more accurate representation of detected lesions in three dimensions is enabled.
[0018] In some demonstrative embodiments, one or more false positive removal machine learning computer models are provided that operate on the lesion list to remove false positive detections of lesions from the lesion list. False positive detection helps eliminate those contours and areas of the image in the input volume that are detected as lesions but are not actually lesions.
[0019] In some demonstrative embodiments, the one or more false positive removal machine learning computer models include one or more false positive removal machine learning computer models trained by a machine learning process based on two different operating points, wherein a first operating point of the two different operating points corresponds to a patient-level operating point and a second operating point of the two different operating points corresponds to a lesion-level operating point. The dual operating point implementation of false positive removal allows for more accurate false positive removal by first determining at the patient level whether the patient has a lesion and, if so, performing a more sensitive evaluation of the lesion list to remove lesions from the lesion list that are not actual lesions.
[0020] In some demonstrative embodiments, the one or more fourth machine learning computer models classify each lesion in the list of lesions into one of a plurality of predetermined categories, wherein the one or more predetermined categories include at least one of a benign category, a malignant category, or an indeterminate category or lesion type. By classifying the lesions, medical personnel are informed of the lesions in the patient requiring treatment. Furthermore, in some demonstrative embodiments, this lesion classification can be used to process the lesion list and the output of the classification associated with the lesions by a downstream computing system to generate a visual representation of the lesions in the lesion list in a graphical user interface. In this way, medical personnel are provided with a visual representation of the lesions that require their attention in order to provide treatment to the patient.
[0021] In some demonstrative embodiments, the input volume includes an image corresponding to at least one of a magnetic resonance imaging modality or a computed tomography modality.Thus, the illustrative embodiments may be applied to known medical imaging techniques as well as later developed medical imaging techniques in which lesion detection and classification are to be performed.
[0022] In other illustrative embodiments, a computer program product is provided that includes a computer-usable or readable medium having a computer-readable program. When executed on a computing device, the computer-readable program causes the computing device to perform various or combinations of the operations outlined above with respect to the method illustrative embodiment.
[0023] In yet another illustrative embodiment, a system / apparatus is provided. The system / apparatus may include one or more processors and a memory coupled to the one or more processors. The memory may include instructions that, when executed by the one or more processors, cause the one or more processors to perform various operations and combinations of operations outlined above with respect to the method illustrative embodiment.
[0024] These and other features and advantages of the present invention will be described in or will become apparent to those skilled in the art in view of the following detailed description of exemplary embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The invention, its preferred mode of use, and further objects and advantages will be best understood by reference to the following detailed description of illustrative embodiments when read in conjunction with the accompanying drawings in which:
[0026] Figure 1 is an example block diagram of an AI pipeline implementing multiple specially configured and trained ML / DL computer models to perform anatomical structure recognition and lesion detection in input medical image data, according to one illustrative embodiment;
[0027] Figure 2 is an example flowchart overview of example operations of an AI pipeline according to one illustrative embodiment;
[0028] Figure 3A is an example diagram illustrating an example input volume of a slice (medical image) of the abdomen of a human patient according to one illustrative embodiment;
[0029] Figure 3B Shown Figure 3A Another depiction of the input volume, where a slice is presented along with its corresponding axial scores s'inf and s' sup Together we express;
[0030] Figure 3C yes Figure 3A Example diagram of an input volume where the volume is divided axially into n completely overlapping segments;
[0031] Figures 4A-4C is an example diagram of an illustrative embodiment of an ML / DL computer model configured and trained to estimate s′ for a segment of an input volume of a medical image, according to an illustrative embodiment. sup and s'inf value;
[0032] Figure 5 is a flow chart outlining example operation of liver detection and predetermined amount of anatomical structure determination logic of an AI pipeline according to one illustrative embodiment;
[0033] Figure 6 is an example diagram of an ensemble of ML / DL computer models for performing lesion detection in an anatomical structure of interest (e.g., liver) according to an illustrative embodiment;
[0034] Figure 7 is a flow chart outlining example operation of liver / lesion detection logic in an AI pipeline according to one illustrative embodiment;
[0035] Figure 8 depicts a block diagram illustrating aspects of lesion segmentation according to an illustrative embodiment;
[0036] Figure 9 depicts results of lesion detection and slice-wise partitioning according to an illustrative embodiment;
[0037] Figures 10A-10D Seed positioning according to one illustrative embodiment is described;
[0038] Figure 11A is a block diagram illustrating a mechanism for lesion splitting according to one illustrative embodiment;
[0039] Figure 11B is a block diagram illustrating a mechanism for seed relabeling according to one illustrative embodiment;
[0040] Figure 12 is a flow chart outlining example operations for lesion splitting according to one illustrative embodiment;
[0041] Figures 13A-13C illustrates z-connectivity of a lesion according to an illustrative embodiment;
[0042] Figure 14A and Figure 14B illustrates results of a training model for z-direction lesion connectivity according to an illustrative embodiment;
[0043] Figure 15 is a flow chart outlining example operations of a mechanism for connecting two-dimensional lesions along the z-axis according to one illustrative embodiment;
[0044] Figure 16 illustrates an example of the outlines of two lesions in the same image according to one illustrative embodiment;
[0045] Figure 17 is a flow chart outlining an example operation of a mechanism for slice-wise contour refinement according to one illustrative embodiment;
[0046] Figure 18A is an example of an ROC curve determined for patient-level and lesion-level operating points according to one illustrative embodiment;
[0047] Figure 18B is an example flow chart of operations for performing false positive removal based on patient-level and lesion-level operating points in accordance with one illustrative embodiment;
[0048] Figure 18C is an example flow chart of operations for performing voxel-wise false positive removal based on input volume-level and voxel-level operating points, according to one illustrative embodiment;
[0049] Figure 19 is a flow chart outlining an example operation of false positive removal logic of an AI pipeline according to one illustrative embodiment;
[0050] Figure 20 is an exemplary diagram of a distributed data processing system in which aspects of the illustrative embodiments may be implemented; and
[0051] Figure 21 is an example block diagram of a computing device in which aspects of the illustrative embodiments may be implemented. DETAILED DESCRIPTION
[0052] The detection of lesions or abnormal cell groups is a primarily manual process in modern medicine. Because this manual process is fraught with sources of error, it is subject to artificial limitations on an individual's ability to detect portions of a digital medical image that display such lesions, particularly given the increasing demand for such individuals to evaluate an increasing number of images in a shorter amount of time. While some automated image analysis mechanisms have been developed, there remains a need to improve such automated image analysis mechanisms to provide more efficient and accurate analysis of medical image data for detecting lesions in imaged anatomical structures (e.g., the liver or other organs).
[0053] The illustrative embodiments are specifically directed to providing improved computational tools for providing specifically trained automated computer-driven artificial intelligence medical image analysis that, through machine learning / deep learning computer processes, detect anatomical structures, detect lesions or other biological structures of interest within or associated with such anatomical structures, perform specialized segmentation of the detected lesions or other biological structures, perform false positive removal based on the specialized segmentation, and perform classification of the detected lesions or other biological structures, and provide the results of the lesion / biological structure detection to a downstream computing system for performing additional computer operations. The following description of the illustrative embodiments will assume specific reference to embodiments of the illustrative embodiments' mechanisms specifically trained with respect to liver lesions as biological structures of interest, however, the illustrative embodiments are not limited thereto. Rather, one of ordinary skill in the art will recognize that the illustrative embodiments' machine learning / deep learning-based artificial intelligence mechanisms can be implemented with respect to a wide variety of other types of biological structures / lesions within or associated with other anatomical structures represented in medical imaging data without departing from the spirit and scope of the present invention. Furthermore, the illustrative embodiments may be described with respect to the medical imaging data being computed tomography (CT) medical imaging data, however, the illustrative embodiments may be implemented with any digital medical imaging data from various types of medical imaging techniques including, but not limited to, positron emission tomography (PET) and other nuclear medicine imaging, ultrasound, magnetic resonance imaging (MRI), elastography, photoacoustic imaging, echocardiography, magnetic particle imaging, functional near-infrared spectroscopy, elastography, various radiographic imaging including fluoroscopy, and the like.
[0054] In general, the illustrative embodiments provide an improved artificial intelligence (AI) computer pipeline that includes multiple specifically configured and trained AI computer tools (e.g., neural networks, cognitive computing systems, or other AI mechanisms trained to perform specified tasks based on limited data sets). Each of the configured and trained AI computer tools is specifically configured / trained to perform a specific type of artificial intelligence processing on an input medical image volume, represented as one or more sets of data and / or metadata defining a medical image captured by a medical imaging technique. Typically, these AI tools employ machine learning (ML) / deep learning (DL) computer models (or simply ML models) to perform tasks that mimic human thought processes regarding generated outcomes while using distinct computer processes specific to the computer tool, and in particular, the ML / DL computer models. These ML / DL computer models learn patterns and relationships between data that represent specific outcomes (e.g., image classifications or labels, data values, medical treatment recommendations, etc.). ML / DL computer models are essentially functions of elements including a machine learning algorithm, configuration settings for the machine learning algorithm, features of the input data identified by the ML / DL computer model, and labels (or outputs) generated by the ML / DL computer model. The functions of these elements are specifically adjusted through the machine learning process to generate a specific ML / DL computer model instance. Different ML models can be specifically configured and trained to perform different AI functions on the same or different input data.
[0055] Since artificial intelligence (AI) pipelines implement multiple ML / DL computer models, it should be understood that these ML / DL computer models are trained through ML / DL processes for specific purposes. Thus, as an overview of the ML / DL computer model training process, it should be understood that machine learning involves the design and development of techniques that take empirical data (such as medical image data) as input and identify complex patterns in the input data. A common pattern in machine learning techniques is to use an underlying computer model M, the parameters of which are optimized to minimize a cost function associated with M given the input data. For example, in the context of classification, model M can be a straight line that separates the data into two classes (e.g., labels), such that M = a*x+b*y+c, and the cost function is the number of misclassified points. The learning process then operates by adjusting the parameters a, b, and c to minimize the number of misclassified points. After this optimization phase (or learning phase), model M can be used to classify new data points. Typically, given the input data, M is a statistical model, and the cost function is inversely proportional to the likelihood of M. This is merely a simple example to provide a general explanation of machine learning training and other types of machine learning using different models, cost (or loss) functions, and optimizations may be used with the mechanisms of the illustrative embodiments without departing from the spirit and scope of the invention.
[0056] For the purpose of anatomical structure detection and / or lesion detection (where lesions are "abnormal" in medical imaging data), a learning machine can construct an ML / DL computer model of a normal structure representation to detect data points in medical images that deviate from this ML / DL computer model of normal structure representation. For example, a given ML / DL computer model (e.g., a supervised, unsupervised, or semi-supervised model) can be used to generate an anomaly score and report it to another device, generating a classification output indicating one or more categories into which the input is classified, the probabilities or scores associated with different categories, etc. Example machine learning techniques that can be used to construct and analyze such ML / DL computer models can include, but are not limited to, nearest neighbor (NN) techniques (e.g., k-NN models, replicator NN models, etc.), statistical techniques (e.g., Bayesian networks, etc.), clustering techniques (e.g., k-means, etc.), neural networks (e.g., reservoir networks, artificial neural networks, etc.), support vector machines (SVMs), etc.
[0057] The artificial intelligence (AI) pipeline implemented by the processor of the illustrative embodiments typically includes one or both of machine learning (ML) and deep learning (DL) computer models. In some cases, one or the other of ML and DL can be used or implemented to achieve a specific result. Traditional machine learning may include or use algorithms such as Bayesian decision making, regression, decision trees / forests, support vector machines, or neural networks. Deep learning can be based on deep neural networks and can use multiple layers, such as convolutional layers. Such DL (such as using a layered network) can be effective in their implementation and can provide enhanced accuracy relative to traditional ML techniques. Traditional ML can generally be distinguished from DL because DL models can outperform classic ML models, however, DL models can consume relatively large amounts of processing and / or power resources. In the context of the illustrative embodiments, references to one or the other of ML and DL herein can be understood to cover one or both forms of AI processing.
[0058] With respect to the illustrative embodiment, after being configured and trained by the ML / DL training process, the ML / DL computer model of the AI pipeline is executed and performs complex computer medical imaging analysis to detect anatomical structures in the input medical image and generate outputs that specifically identify target biological structures of interest (hepatic lesions, for the purpose of describing the example embodiment, hereinafter assumed to be CT medical image data), classify the target biological structures of interest, delineate where these target biological structures of interest (e.g., liver lesions) are present in the input medical image (hereafter assumed to be CT medical image data), and other information that helps human subject matter experts (SMEs) (such as radiologists, doctors, etc.) understand the patient's medical condition from the perspective of the captured input medical image. In addition, the outputs can be provided to other downstream computer systems to perform additional artificial intelligence operations (such as treatment recommendations and other decision support operations based on classification, delineation, etc.).
[0059] Initially, an artificial intelligence (AI) pipeline of the illustrative embodiments receives an input volume of computed tomography (CT) medical imaging data and detects which portion of a biological entity's body is depicted in the CT medical imaging data. A "volume" of a medical image is a three-dimensional representation of the internal anatomical structure of a biological entity composed of a stack of two-dimensional slices, where the slices may be individual medical images captured by medical imaging technology. A stack of slices may also be referred to as a "slab," and differs from a slice itself in that the stack represents a portion of an anatomical structure having thickness, wherein the stack of slices or slabs generates a three-dimensional representation of the anatomical structure.
[0060] For the purposes of this specification, it will be assumed that the biological entity is a human, however, the present invention can operate on medical images for various types of biological entities. For example, in veterinary medicine, the biological entity can be different types of small animals (e.g., pets such as dogs, cats, etc.) or large animals (e.g., horses, cows, or other farm animals). For embodiments in which the AI pipeline is specifically trained to detect liver lesions, the AI pipeline determines whether the input CT medical imaging data represents an abdominal scan present in the CT medical imaging data, and if not, the operation of the AI pipeline terminates with respect to the input CT medical imaging data because it is not directed to the correct part or portion of the human body. It should be understood that according to the illustrative embodiments, there can be different AI pipelines trained to process input medical images for different parts of the body and different target biological structures, and the input CT medical image can be input to each of the AI pipelines, or routed to the AI pipeline based on the body part or body part classification depicted in the input CT medical image. For example, the input CT medical image can first be classified with respect to the body part or body part represented in the input CT medical image, and then a corresponding trained AI pipeline can be selected from multiple trained AI pipelines of the type described herein to process the input CT medical image. For the purposes of the following description, a single AI pipeline trained to detect liver lesions will be described, but given this description, its extension to a suite or collection of AI pipelines will be apparent to those of ordinary skill in the art.
[0061] Assuming that the input CT medical image volume comprises a medical image of the abdomen of a human body (for the purpose of liver lesion detection), further processing of the input CT medical image is performed in two primary stages, which may be performed substantially in parallel with each other and / or sequentially, depending on the desired implementation. The two primary stages include a phase classification stage and an anatomical structure detection stage (e.g., a liver detection stage in the case where the AI pipeline is configured to perform liver lesion detection).
[0062] The phase classification stage determines whether the volume of the input CT medical image includes a single imaging phase or multiple imaging phases. "Phase" in medical imaging is an indication of contrast agent uptake. For example, in some medical imaging techniques, the phase can be defined based on when the contrast agent is introduced into a biological entity, which allows the capture of medical images, including capturing the path of the contrast agent. For example, the phase can include a pre-contrast phase, an arterial phase, a portal phase, and a delayed phase, where the medical image is captured in any or all of these phases. The phase is generally related to the timing after injection and the characteristics of the structural enhancement within the image. The timing information can be taken into account to "classify" the potential phase (for example, the delayed phase will always be acquired after the portal phase) and estimate the potential phase of a given image. Regarding the use of the enhancement characteristics of structures within the image, an example of using this type of information to determine the phase is described in commonly assigned and co-pending U.S. patent application serial number 16 / 926,880, filed on July 13, 2020, and entitled "Method of Determining Contrast Phase of a Computerized Tomography Image." Additionally, the timing information can be used in combination with other information (sampling, reconstruction kernel, etc.) to pick the best representation for each phase (a given acquisition can be reconstructed in several ways).
[0063] Once the images in the input volume are assigned or classified to their corresponding phases based on enhanced timing and / or features, it can be determined based on the phase classification whether the volume includes images of a single phase (e.g., the presence of a portal vein phase but no arterial phase) or a multi-phase examination (e.g., portal vein and artery). If the phase classification indicates that a single phase is present in the volume of the input CT medical image, further processing of the AI pipeline is performed as described below. If multiple phases are detected, the volume is not further processed by the AI pipeline. However, in some illustrative embodiments, while this single / multiple phase-based volume filtering only accepts volumes with images from a single phase and rejects multiple phase volumes, in other illustrative embodiments, the AI pipeline processing described herein can filter out images of volumes that are not classified into the target phase of interest, for example, the portal vein phase images in the volume can be retained while filtering out images of the volume that are not classified as part of the portal vein phase, thereby modifying the input volume into a modified volume having only a subset of images classified as the target phase. Furthermore, as previously discussed, different AI pipelines can be trained for different types of volumes. In some illustrative embodiments, phase classification of images within an input volume can be used to route or distribute images of the input volume to corresponding AI pipelines that are trained and configured to process images of different phases, such that the input volume can be subdivided into constituent subvolumes and routed to their corresponding AI pipelines for processing, e.g., a first subvolume corresponding to a portal vein phase image is sent to a first AI pipeline, while a second subvolume corresponding to an arterial phase image is sent to a second AI pipeline for processing. If the volume of the input CT medical image comprises a single phase, or after filtering the subvolumes and optionally routing them to corresponding AI pipelines so that the AI pipelines process images of the input volume or subvolumes of a single phase, the volume (or subvolume) is then passed to the next stage of the AI pipeline for further processing.
[0064] The second main stage is the anatomical structure of interest (in the exemplary embodiment, the liver) detection stage, in which volumetric portions depicting the anatomical structure of interest are identified and passed to the next downstream stage of the AI pipeline. The anatomical structure of interest detection stage (hereinafter referred to as the liver detection stage, according to the exemplary embodiment) includes a machine learning (ML) / deep learning (DL) computer model that is specifically trained and configured to perform computerized medical image analysis to identify portions of an input medical image corresponding to the anatomical structure of interest (e.g., the liver). Such medical image analysis can include training the ML / DL model on labeled training medical image data as input to determine whether the input medical image (the training image during training) includes the anatomical structure of interest (e.g., the liver). Based on the ground truth of the image labels, the operating parameters of the ML / DL model can be adjusted to reduce the loss or error in the results generated by the ML / DL model until convergence is achieved (i.e., the loss is minimized). Through this process, the ML / DL model is trained to recognize patterns in the medical image data that indicate the presence of the anatomical structure of interest (in this example, the liver). Thereafter, once trained, the ML / DL model can be executed on new input data to determine whether the new input medical image data has a pattern that indicates the presence of an anatomical structure, and if the probability is greater than a predetermined threshold, it can be determined that the medical image data includes the anatomical structure of interest.
[0065] Therefore, at the liver detection stage, the AI pipeline uses a trained ML / DL computer model to determine whether the volume of the input CT medical image includes an image depicting the liver. The portion of the volume depicting the liver, along with the results of the phase classification stage, is passed to the determination stage of the AI pipeline. This determination stage determines whether medical imaging of a single phase exists and whether at least a predetermined amount of the anatomical structure of interest (e.g., the liver) is present in the portion of the volume depicting the anatomical structure of interest. The determination of the presence of the predetermined amount of the anatomical structure of interest can be based on known measurement mechanisms for determining measurements of structures from medical images (e.g., calculating the size of the structure from differences in pixel positions within the image). The measurements can be compared to predetermined sizes (e.g., average sizes) of anatomical structures of similar patients with similar demographics, so that if the measurements represent at least a predetermined amount or portion of the anatomical structure, further processing by the AI pipeline is performed. In one illustrative embodiment, for example, the determination determines whether at least one-third of the liver is present in the portion of the volume of the input CT medical image determined to depict the liver. While one-third is used in the exemplary embodiment, any predetermined amount of structure determined to be appropriate for a particular implementation may be used without departing from the spirit and scope of the present invention.
[0066] In one illustrative embodiment, to determine whether a predetermined amount of an anatomical structure of interest is present in a volume of an input CT medical image, an axial score (axial score) is defined such that a slice of the medical image corresponding to the first representation of the anatomical structure of interest (e.g., liver) in the volume, i.e., the first slice (FSL) containing the liver is given a slice score of 0, and the last slice (LSL) containing the liver has a score of 1. Assuming a human biological entity, the first and last slices are defined from the most inferior slice (MISV) in the volume (closest to the lower limb, e.g., the foot) to the most superior slice (MSSV) in the volume (closest to the head). The liver axial score estimate (LAE) is calculated by a pair of slice scores s sup and s inf To define, the slice score s sup and s inf Slice scores corresponding to MSSV and MISV slices, respectively. As will be described in more detail below, the ML / DL computer model is specifically configured and trained to determine the slice score s of the volume of the input CT medical image. sup and s inf Knowing these slice scores and knowing from the above definition that the liver extends from 0 to 1, the mechanisms of the illustrative embodiments are able to determine the fraction of the liver that is in the field of view of the volume of the input CT medical image.
[0067] In some demonstrative embodiments, the slice score s may be found indirectly by first dividing the volume of the input CT medical image into a plurality of segments, and then, for each segment, executing the configured and trained ML / DL computer model on the slices of the segment to estimate the height of each slice. sup and s inf , in order to determine the sup and the uppermost (closest to the head) and lowermost (closest to the feet) liver slices in s'inf. Given s' sup and s'inf, find s by extrapolation sup and s inf Since it is known how the segments are positioned relative to the entire volume of the input CT medical image, the method is based on a robust estimator of the height of any slice from the input volume (or a subvolume associated with the target phase). Such an estimator can be obtained by learning a regression model, for example by using a deep learning model that performs the estimation of the height from a chunk (a set of consecutive slices). For example, long short-term memory (LSTM) type artificial neural network models are suitable for these tasks because they have the ability to encode the ordering of slices containing the liver and abdominal anatomical structures. It should be noted that for each volume, there will be n s sup and s infWhere n is the number of segments per volume. In one illustrative embodiment, the final estimate is obtained by taking an unweighted average of these n estimates, however, in other illustrative embodiments, other functions of the n estimates may be used to generate the final estimate.
[0068] The volume s of the input CT medical image has been determined sup and s inf The final estimate of the value of , based on which the fraction of the anatomical structure of interest (e.g., the liver) is calculated. This task is made possible by an estimate of the height of each slice. Based on the estimate of the height of the first slice (h1) and the height of the last slice (h2) of the liver in the input volume, assuming that the heights of the actual first and last slices of the liver (whether they are contained in the input volume or not) are H1 and H2, the portion of the liver visible in the input volume can be expressed as (min(h1, H1)-max(h2, H2)) / (H1-H2). This calculated fraction can then be compared with a predetermined threshold to determine whether a predetermined minimum amount of the anatomical structure of interest is present in the volume of the input CT medical image, for example, whether at least 1 / 3 of the liver is present in the volume of the input CT medical image.
[0069] If the determination results in a determination that multiple phases are present and / or a predetermined amount of anatomical structure of interest depicting the anatomical structure is not present in the portion of the volume of the input CT medical image, further processing of the volume may be discontinued. If the determination results in a determination that the volume of the input CT medical image includes a single phase and at least a predetermined amount of anatomical structure of interest (e.g., 1 / 3 of the liver is shown in the image), the portion of the volume of the input CT medical image depicting the anatomical structure is forwarded to the next stage of the AI pipeline for processing.
[0070] At the next level of the AI pipeline, the AI pipeline performs lesion detection on a portion of the volume of the input CT medical image that represents an anatomical structure of interest (e.g., the liver). This liver and lesion detection stage of the AI pipeline uses an integration of ML / DL computer models to detect the liver and lesions in the liver as represented in the volume of the input CT medical image. The integration of the ML / DL computer models performs liver and lesion detection using differently trained ML / DL computer models, where the ML / DL computer models are trained and use a loss function to balance false positives and false negatives in lesion detection. In addition, the integrated ML / DL computer models are configured such that a third loss function causes the outputs of the ML / DL computer models to be consistent with each other.
[0071] Assuming that liver detection and lesion detection are performed at this stage of the AI pipeline, a first ML / DL computer model is executed on a volume of an input CT medical image to detect the presence of the liver. This ML / DL computer model can be the same ML / DL computer model employed in a previous AI pipeline stage for anatomical structure detection of interest, and thus can leverage previously obtained results. Multiple (two or more) other ML / DL computer models are configured and trained to perform lesion detection in portions of the medical image depicting the liver. The first ML / DL computer model is configured with two loss functions. The first loss function penalizes errors in false negatives, i.e., classifications that incorrectly indicate the absence of a lesion (normal anatomical structure). The second loss function penalizes errors in false positives, i.e., classifications that incorrectly indicate the presence of a lesion (abnormal anatomical structure). The second ML / DL is trained to detect lesions using an adaptive loss function that penalizes false positives in slices of the liver containing normal tissue and penalizes false negatives in slices of the liver containing lesions. The detection outputs from the two ML / DL models are averaged to produce a final lesion detection.
[0072] The results of the liver / lesion detection stage of the AI pipeline include one or more contours (outlines) of the liver and a detection map that identifies portions of medical imaging data elements corresponding to the detected lesions, e.g., a voxel-wise map of the liver lesions detected in the volume of the input CT medical image. The image map is then input to the lesion segmentation stage of the AI pipeline. As described in more detail below, the lesion segmentation stage uses a watershed technique to partition the detection map to generate partitions of image elements (e.g., voxels) of the input CT medical image. Based on the partitions, the liver lesion segmentation stage identifies all contours corresponding to lesions present in slices of the volume of the input CT medical image and performs an operation to identify which contours correspond to the same lesion in three dimensions. Lesion segmentation aggregates related lesion contours to generate a three-dimensional partition of the lesion. Lesion segmentation uses lesion image elements (e.g., voxels) represented in the medical image and patches of non-liver tissue to focus on each lesion individually and perform active contour analysis. In this way, individual lesions can be identified and processed without biasing the analysis due to other lesions in the medical image or due to portions of the image outside the liver.
[0073] The result of lesion segmentation is a list of lesions with their corresponding outlines or contours in the volume of the input CT medical image. These outputs may include findings that are not actual lesions. In order to minimize the impact of those false positives, the output is provided to the next stage of the AI pipeline involving false positive removal using a trained false positive removal model. This false positive removal model of the AI pipeline acts as a classifier to identify what outputs are actual lesions and what are false positives from the detected findings. The input consists of the image volume (VOI) around the detected findings associated with the mask generated from the lesion segmentation refinement. The false positive removal model is trained using data that is the result of the detection / segmentation stage: objects detected by the detection algorithm as lesions from the ground truth are used to represent the lesion class during training, while detections that do not match any lesions from the ground truth are used to represent the non-lesion (false positive) class.
[0074] To further improve the overall performance, a dual operating point strategy is adopted on lesion detection and false positive models. The idea is to note that the output of the AI pipeline can be interpreted at different levels. First, the output of the AI pipeline can be used to discriminate between the examination volume, i.e., the input volume or image volume (VOI) with or without lesions. Second, the output of the AI pipeline aims to maximize the detection of lesions, regardless of whether they are contained in the same patient / examination / volume. For clarity, the measurements made for the examination will be referred to as "patient level" here, and the measurements made for the lesions will be referred to as "lesion level" here. Maximizing the sensitivity at the "lesion level" will reduce the specificity at the "patient level" (for a patient, one detection is sufficient to say that the lesion is contained). This may ultimately be suboptimal for clinical use, as one must choose between having poor specificity at the patient level or having low sensitivity at the lesion level.
[0075] Given this, the illustrative embodiment uses a dual operating point approach for both lesion detection and false positive removal. The principle is to first run the process using a first operating point that gives reasonable performance at the patient level. Then, for patients with at least one detected lesion from the first run, the detected lesions are reinterpreted / processed using a second operating point. This second operating point is chosen to be more sensitive. While the specificity of this second operating point is lower than the first, this loss of specificity is accounted for at the patient level because all patients without lesions detected using the first operating point remain unchanged, regardless of whether the second operating point would have detected additional lesions. Therefore, patient-level specificity is determined solely by the first operating point. Patient-level sensitivity is between that of either the first or second operating point taken alone (a false-negative case from the first operating point can be turned into a true positive by the second operating point). On the lesion side, the actual lesion-level sensitivity is improved compared to the first operating point alone. Lesion specificity is better than that obtained from the less specific second operating point alone because there are no false positives from cases treated using only the first operating point.
[0076] While the illustrative embodiments will assume a specific configuration and use of a dual operating point approach, it should be understood that the dual operating point approach can be used with other configurations and for other purposes, where one is interested in measuring performance at both the group level (in the illustrative embodiment, this group level is the "patient level") and the element level (in the illustrative embodiment, this element level is the "lesion level"). While in the illustrative embodiments, the dual operating point approach is applied to both lesion detection and false positive removal, it should be understood that the dual operating point approach can be extended beyond these stages of the AI pipeline. For example, lesion detection can be performed at the voxel level (element) versus the volume level (group), rather than at both the patient level and the lesion level. As another example, the voxel or lesion level can be used for the element level, and the slab (a collection of slices) can be used as the group level. In yet another example, the entire examined volume can be used as the group level, rather than a single volume. It should be understood that the approach can also be applied to two-dimensional images (e.g., 2D X-rays of the chest, mammography, etc.) rather than three-dimensional volumes of the image to be analyzed. Specificity (such as the average number of false positives per patient / group) can be used to select the operating point. Additionally, while the illustrative embodiments are described as applied to lesion detection and classification, the dual operating point based approach may be applied to other structures (clips, stents, implants, etc.) and beyond medical imaging.
[0077] The results of dual-operating point detection and false positive removal lead to the identification of a final filtered list of lesions to be further processed by the lesion classification stage of the AI pipeline. At the lesion classification stage of the AI pipeline, a configured and trained ML / DL computer model is executed on the lesion list and its corresponding contour data to classify the lesions into one of a plurality of predetermined lesion classifications. For example, each lesion in the final filtered lesion list and its attributes (e.g., contour data) can be input into a trained ML / DL computer model, which then operates on this data to classify the lesion as a specific type of lesion. Classification can be performed using a classifier previously trained on ground truth data (e.g., a trained neural network computer model) in combination with the results of previous processing steps of the AI pipeline. The classification task can be more or less complex, for example, it can provide a label between benign, malignant, or uncertain, or in another example, provide the actual lesion type (e.g., cyst, metastasis, hemangioma, etc.). The classifier can be, for example, a neural network-based computer model classifier (e.g., SVM, decision tree, etc.) or a deep learning computer model. The actual input to the classifier is a patch around the lesion, which in some embodiments may be augmented with a lesion mask or outline (contour).
[0078] After the lesions are classified by the lesion classification stage of the AI pipeline, the AI pipeline outputs a list of lesions and their classification, along with any contour attributes of the lesions. In addition, the AI pipeline may also output liver contour information for the liver. The information generated by the AI pipeline may be provided to further downstream computing systems for further processing and generating representations of the anatomical structures of interest and any detected lesions present in the anatomical structures. For example, a graphical representation of the volume of the input CT medical image may be generated in a medical image viewer or other computer application, wherein the contour information generated by the AI pipeline is used to overlay or otherwise highlight the anatomical structures and detected lesions in the graphical representation. In other illustrative embodiments, downstream processing of the information generated by the AI pipeline may include diagnostic decision support operations, automated medical imaging report generation based on the list of detected lesions, classifications, and contours. In other illustrative embodiments, based on the classification of the lesions, different treatment recommendations may be generated for review and consideration by practicing physicians.
[0079] In some demonstrative embodiments, a list of lesions, their classifications, and outlines may be stored in a historical data structure associated with a patient corresponding to a volume of an input CT medical image, such that multiple executions of the AI pipeline on different volumes of the input CT medical image associated with the patient may be stored and evaluated over time. For example, differences between the list of lesions and / or their associated classifications and outlines may be determined to assess the progression of a disease or medical condition in the patient, and such information may be presented to a medical professional for use in assisting with the patient's treatment.
[0080] Other downstream computing systems and processing of the specific anatomical structure and lesion detection information generated by the AI mechanism of the illustrative embodiments may be implemented without departing from the spirit and scope of the present invention. For example, the output of the AI pipeline may be used by another downstream computing system to process the anatomical structure and lesion information in the output of the AI pipeline to identify discrepancies with other information sources (e.g., radiology reports) in order to make clinical staff aware of potentially overlooked findings.
[0081] Thus, the illustrative embodiments provide mechanisms for providing an automated AI pipeline that includes multiple configured and trained ML / DL computer models that implement various artificial intelligence operations for each stage of the AI pipeline to identify anatomical structures and lesions associated with these anatomical structures in an input medical image volume, determine contours associated with such anatomical structures and lesions, determine classifications for such lesions, and generate a list of such lesions and the contours of the lesions and the anatomical structures for further downstream computer processing of the AI-generated information from the AI pipeline. The operations of the AI pipeline are automated such that no human intervention is present at any stage of the AI pipeline. Instead, specifically configured and trained ML / DL computer models trained through machine learning / deep learning computer processes are employed to perform the designated AI analysis at each stage. The only points at which human intervention may occur are before the input of the input medical image volume (e.g., during medical imaging of a patient) and after the output of the AI pipeline (e.g., when viewing an enhanced medical image presented via a computer image viewing application based on the output of the lesion list and contours generated by the AI pipeline). Thus, AI pipelines perform operations that humans cannot perform as mental processes and do not organize any human activities because AI pipelines specifically involve automated computer tools that are implemented as improvements to artificial intelligence using specified machine learning / deep learning processes that exist only within a computer environment.
[0082] Before proceeding to discuss various aspects of the illustrative embodiments and the improved computer operations performed by the illustrative embodiments, it should first be recognized that throughout this specification, the term "mechanism" will be used to refer to elements of the present invention that perform various operations, functions, etc. As used herein, the term "mechanism" can be an implementation of a function or aspect of the illustrative embodiments in the form of a device, program, or computer program product. In the case of a process, the process is implemented by one or more devices, apparatuses, computers, data processing systems, etc. In the case of a computer program product, the logic represented by the computer code or instructions implemented in or on the computer program product is executed by one or more hardware devices to implement the function or perform the operation associated with the specified "mechanism." Thus, the mechanisms described herein can be implemented as dedicated hardware, software executed on hardware to configure the hardware to perform the specialized functions of the present invention that the hardware is not otherwise capable of performing, software instructions stored on a medium so that the instructions can be readily executed by hardware to specifically configure the hardware to perform the functionality and specific computer operations described herein, a process or method for performing the functions, or any combination thereof.
[0083] This specification and claims may utilize the terms "a," "at least one," and "one or more" with respect to particular features and elements of the illustrative embodiments. It should be understood that these terms and phrases are intended to state that there is at least one of the particular features or elements present in a particular illustrative embodiment, but more than one may be present. That is, these terms / phrases are not intended to limit the specification or claims to the presence of a single feature / element or to require the presence of a plurality of such features / elements. Rather, these terms / phrases require only at least one of a single feature / element, while multiple such features / elements are possible within the scope of the specification and claims.
[0084] Furthermore, it should be understood that if the term "engine" is used herein with respect to describing embodiments and features of the present invention, it is not intended to limit any particular implementation for implementing and / or performing actions, steps, processes, etc. that are attributable to and / or performed by an engine. An engine may be, but is not limited to, software, hardware, and / or firmware, or any combination thereof, that performs specified functions, including but not limited to any use of a general-purpose and / or special-purpose processor in combination with appropriate software loaded or stored in a machine-readable memory and executed by the processor. Further, unless otherwise specified, any name associated with a particular engine is for ease of reference and is not intended to be limiting of a particular implementation. Additionally, any functionality attributed to an engine may be performed equally by multiple engines, incorporated into and / or combined with another engine of the same or different type, or distributed across one or more engines in various configurations.
[0085] In addition, it should be understood that the following description uses a number of different examples of different elements of the illustrative embodiments to further illustrate exemplary implementations of the illustrative embodiments and to aid in understanding the mechanisms of the illustrative embodiments. These examples are intended to be non-limiting and not an exhaustive list of every possibility for implementing the mechanisms of the illustrative embodiments. In view of this description, it will be clear to one of ordinary skill in the art that there are many other alternative implementations for these different elements that can be utilized in addition to or in place of the examples provided herein without departing from the spirit and scope of the present invention.
[0086] The present invention may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon, the computer-readable program instructions being used to cause a processor to perform various aspects of the present invention.
[0087] Computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. Computer readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as a punch card or a raised structure in a groove with instructions recorded thereon), and any suitable combination of the foregoing. As used herein, computer readable storage medium should not be interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (for example, a light pulse by a fiber optic cable), or an electrical signal transmitted by a wire.
[0088] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0089] The computer-readable program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, and traditional procedural programming languages, such as "C" programming language or similar programming languages. The computer-readable program instructions can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (for example, by using the Internet of an Internet service provider). In some embodiments, an electronic circuit (including, for example, a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA)) can execute the computer-readable program instructions to personalize the electronic circuit by utilizing the state information of the computer-readable program instructions, so as to perform various aspects of the present invention.
[0090] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer-readable program instructions.
[0091] These computer-readable program instructions can be provided to a processor of a computer or other programmable data processing device to produce a machine, such that the instructions, executed by the processor of the computer or other programmable data processing device, create a device for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that can direct a computer, programmable data processing device, and / or other device to function in a specific manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0092] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0093] The flow charts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the systems, methods and computer program products according to different embodiments of the present invention. To this end, each box in the flow chart or block diagram can represent a part of a module, segment or instruction, and a part of a module, segment or instruction includes one or more executable instructions for realizing a specified logical function. In some alternative embodiments, the functions marked in the box may not occur in the order marked in the figure. For example, depending on the functions involved, the two boxes shown in succession can actually be executed substantially simultaneously, or these boxes can sometimes be executed in the opposite order. It will also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a system based on special-purpose hardware, which performs a specified function or action or performs a combination of special-purpose hardware and computer instructions.
[0094] Overview of the Lesion Detection and Classification AI Pipeline
[0095] Figure 1 is an example block diagram of a lesion detection and classification artificial intelligence (AI) pipeline (referred to herein as the "AI pipeline") that implements multiple specifically configured and trained ML / DL computer models to perform anatomical structure recognition and lesion detection in input medical image data in accordance with an illustrative embodiment. For illustrative purposes only, the depicted AI pipeline is specifically described as being directed to liver detection and liver lesion detection in medical image data. As described above, the illustrative embodiments are not limited thereto and may be applied to any anatomical structure of interest and lesions associated with such anatomical structure of interest that may be represented in image elements of medical image data captured by medical imaging techniques and corresponding computing systems. For example, the mechanisms of the illustrative embodiments may be applied to other anatomical structures (such as the lungs, heart, etc.), as well as detection, contour recognition, classification, etc. of lesions associated with the lungs, heart, or other anatomical structures of interest.
[0096] Furthermore, it should be understood that the following description provides Figure 1, and subsequent sections of this description will go into additional detail regarding the various stages of the AI pipeline. In some illustrative embodiments, each stage of the AI pipeline is implemented as a configured and trained ML / DL computer model, such as a neural network of a deep learning neural network, as represented by the symbol 103 in each stage of the AI pipeline 100. These different ML / DL computer models are specifically configured and trained to perform the specific AI operations described herein (e.g., body part recognition, liver detection, phase classification, liver minimum amount detection, liver / lesion detection, lesion segmentation, false positive removal, lesion classification, etc.). Although these additional sections described below will explain specific embodiments for implementing each stage of the AI pipeline that provides novel techniques, mechanisms, and methods for performing the AI operations at different stages, it should be understood that in the context of the AI pipeline as a whole, other equivalent techniques, mechanisms, or methods may be used without departing from the spirit and scope of the illustrative embodiments. In view of this description, these other equivalent techniques, mechanisms, or methods will be apparent to those of ordinary skill in the art and are intended to be within the spirit and scope of the present invention.
[0097] like Figure 1 As shown, according to one illustrative embodiment, an artificial intelligence (AI) pipeline 100 receives an input medical image volume 105, which in the depicted example is a volume of an input computed tomography (CT) medical image represented as one or more data structures, as input that is then automatically processed by various stages of the AI pipeline 100 to ultimately generate an output 170 including a list of lesions and their classification and contour information, as well as contour information about an anatomical structure of interest (e.g., the liver in the depicted example). The input medical image volume 105 can be captured by a medical imaging technology 102 using any of a number of commonly known or later developed medical imaging techniques and devices that render images of the internal anatomical structure of a biological entity (i.e., a patient) as one or more medical image data structures. In some demonstrative embodiments, the volume of input medical images 105 includes two-dimensional slices of a portion of the anatomical structure of a portion of a patient's body (individual medical images), which are then combined to generate slabs (a combination of slices along an axis, thereby providing a collection of medical images having thickness along the axis), and the two-dimensional slices are combined to generate a three-dimensional representation (i.e., a volume of the anatomical structure of the body portion).
[0098] In the first stage logic 110 of the AI pipeline 100, the AI pipeline 100 determines 112 a portion of the patient's body corresponding to an input volume 105 of CT medical imaging data and determines, via body part of interest determination logic 114, whether the portion of the patient's body represents a portion of the patient's body corresponding to an anatomical structure of interest (e.g., an abdominal scan rather than a cranial scan, a lower body scan, etc.). This evaluation is performed using initial filters on the AI pipeline 100 with respect to only the volume 105 of input CT medical imaging data (hereinafter referred to as the "input volume" 105), for which the AI pipeline 100 is specifically configured and trained to perform anatomical structure recognition, as well as contour and lesion recognition, contouring, and classification. This detection of a body part represented in the input volume 105 may include looking at metadata associated with the input volume 105, which may include a field specifying the region of the patient's body being scanned, as may be specified by the source medical imaging technology computing system 102 when performing the medical imaging scan. Alternatively, the first-level logic 110 of the AI pipeline 100 may implement a specially configured and trained ML / DL computer model for body part detection 112 that performs medical image classification with respect to a specific portion of the patient's body, wherein the medical image classification performs computerized pattern analysis on the medical image data of the input volume 105 and predicts a classification of the medical imaging data with respect to one or more predetermined portions of the patient's body. In some illustrative embodiments, this assessment may be binary (e.g., either an abdominal medical imaging volume or not), or may be a more complex multi-class assessment (e.g., specifically identifying probabilities or scores with respect to multiple different body part classifications (e.g., abdominal, cranial, lower extremity, etc.)).
[0099] If the body part of interest determination logic 114 of the first stage logic 110 of the AI pipeline 100 determines that the input volume 105 does not represent a portion of the patient's body in which the anatomical structure of interest can be found (e.g., the abdominal portion of the body where the liver can be found), then processing of the AI pipeline 100 can be interrupted (a reject condition). If the body part of interest determination logic 114 of the first stage logic 110 of the AI pipeline 100 determines that the input volume 105 does represent a portion of the patient's body in which the anatomical structure of interest can be found, then further processing of the input volume 105 by the AI pipeline 100 is performed as described below. It should be appreciated that in some demonstrative embodiments, a plurality of different instances of the AI pipeline 100 may be provided, each instance being configured and trained to process input volumes 105 corresponding to different anatomical structures that may be present in different portions of the patient's body. Thus, the first-level logic 110 may be provided external to the AI pipeline 100 and may operate as routing logic to route the input volume 105 to a corresponding AI pipeline 100 that is specifically configured and trained to process a particular classification of the input volume 105, e.g., one AI pipeline instance for liver and liver lesion detection / classification, another AI pipeline instance for lung and lung lesion detection / classification, a third AI pipeline instance for heart and cardiac lesion detection / classification, etc. Thus, the first-level logic 110 may include routing logic that stores a mapping of which AI pipeline instance 100 corresponds to different body parts / anatomies of interest, and based on the detection of a body part represented in the input volume 105, the first-level logic 110 may automatically route the input volume 105 to a corresponding AI pipeline instance 100 that is specifically configured and trained to process the input volume 105 corresponding to the detected body part.
[0100] Assuming that the input volume 105 is detected as a portion of the patient's body representing the presence of an anatomical structure of interest (e.g., an abdominal scan is present in the input volume 105 for the purpose of liver lesion detection), further processing of the input volume 105 is performed by the AI pipeline 100 in the second stage logic 120. This second stage logic 120 includes two main sub-stages 122 and 124, which may be executed substantially in parallel with each other and / or sequentially, depending on the desired implementation (in the example of FIG. Figure 1 The two main sub-stages 122, 124 include a phase classification sub-stage 122 and an anatomical structure detection sub-stage 124 (eg, a liver detection sub-stage 124 if the AI pipeline 100 is configured to perform liver lesion detection).
[0101] The phase classification sub-stage 122 determines whether the input volume 105 includes a single imaging phase (e.g., a pre-contrast phase, an arterial phase, a portal phase, a delayed phase, etc.). Again, the phase classification sub-stage 122 can be implemented as logic that evaluates metadata associated with the input volume 105, which can include a field specifying the phase of the medical imaging study to which the medical image corresponds, as may be generated by the medical imaging technology computing system 102 when performing medical imaging. Alternatively, the illustrative embodiments can implement a configured and trained ML / DL computer model that is specifically trained to detect patterns in medical images that indicate different phases of a medical imaging study, and thereby can classify the medical images of the input volume 105 as to which phases they correspond. The output of the phase classification sub-stage 122 can be binary, indicating whether the input volume 105 includes one phase or multiple phases, or can be a classification for each phase represented in the input volume 105, which can then be used to determine whether a single phase or multiple phases are represented.
[0102] If the phase classification indicates the presence of a single phase in the input volume 105, the AI pipeline 100 performs further processing through downstream stages 130-170 as described below. If multiple phases are detected, the input volume 105 is not further processed by the AI pipeline 100, or, as previously described, can be filtered and / or divided into sub-volumes, each having an image corresponding to a single phase, such that the AI pipeline 100 processes only the sub-volumes corresponding to the target phase and / or routes the sub-volumes to corresponding AI pipelines configured and trained to process input volumes corresponding to their particular phase classification. It should be understood that an input volume can be rejected for a number of reasons (e.g., liver is not present in the image, a single-phase input volume is not present in the image, not enough liver is present in the image, etc.). Depending on the actual root cause of the rejection, the reason for the rejection can be communicated to the user via a user interface, etc. For example, in response to the rejection, an output of the AI pipeline 100 can indicate the reason for the rejection and can be utilized by a downstream computing system (e.g., a viewer or additional automated processing system) to communicate the reason for the rejection via the output. For example, in the event that no liver is detected in the input volume, the input volume may be silently ignored, e.g., without communicating this rejection to the user, while for an input volume that contains a liver, but including a multi-phase input volume, the rejection may be communicated to the user (e.g., a radiologist) by clearly stating in a user interface generated by a viewer's downstream computing system that the input volume was not processed by the AI pipeline 100 because it had images of more than one phase, e.g., so as not to be misinterpreted by an input volume that does not contain any findings.
[0103] The second main sub-stage 124 is a detection sub-stage for detecting the anatomical structure of interest (in the exemplary embodiment, the liver) in portions of the input volume 105. That is, slices, slabs, etc. in the input volume 105 that specifically depict the anatomical structure of interest (the liver) are identified and evaluated to determine whether a predetermined minimum amount of the anatomical structure of interest (the liver) is present in these slices, slabs, or in the input volume as a whole. As previously described, the detection sub-stage 124 includes an ML / DL computer model 125 that is specifically trained and configured to perform computerized medical image analysis to identify portions of the input medical image that correspond to the anatomical structure of interest (e.g., the human liver).
[0104] Thus, in the liver detection sub-stage 124, the AI pipeline 100 uses the trained ML / DL computer model 125 to determine whether the volume of the input CT medical image includes an image depicting the liver. The portion of the volume depicting the liver is passed along with the results of the phase classification sub-stage 122 to a determination sub-stage 126 of the AI pipeline 100. The determination sub-stage 126 includes single phase determination logic 127 and minimum structure amount determination logic 128. The determination sub-stage 126 determines whether medical imaging of a single phase 127 is present and whether at least a predetermined amount of anatomical structure of interest 128 is present in the portion of the volume depicting the anatomical structure of interest (e.g., the liver). As previously described, the presence of the predetermined amount of anatomical structure of interest can be determined based on known measurement mechanisms for determining structural measurements from medical images, for example, calculating the size of the structure from differences in pixel locations within the image, and comparing these measurements to one or more predetermined thresholds to determine whether a minimum amount of anatomical structure of interest (e.g., the liver) is present in the input volume 105, for example, 1 / 3 of the liver is present in the portion of the input volume 105 determined to depict the liver.
[0105] In one illustrative embodiment, to determine whether a predetermined amount of the anatomical structure of interest (liver) is present in the input volume 105, the previously described axial scoring mechanism may be used to evaluate the portion of the anatomical structure present in the input volume 105. As previously described, for the input volume 105, the ML / DL computer model may be configured and trained to estimate slice scores s corresponding to slice scores for MSSV and MISV slices, respectively. sup and s inf In some demonstrative embodiments, the s′ of the segment may be estimated by first dividing the input volume 105 into a plurality of segments and then, for each segment, executing a configured and trained ML / DL computer model on a slice of the segment. sup and the slice scores of the first and last slices in s'inf to indirectly find the slice score s sup and s infGiven s' sup and s' inf Estimates of s are found by extrapolation. sup and s inf Since it is known how the segments are positioned relative to the entire volume of the input CT medical image, it should be noted that for each input volume 105, there will be n s sup and s inf Where n is the number of segments per volume. In one illustrative embodiment, the final estimate is obtained by taking an unweighted average of these n estimates, however, in other illustrative embodiments, other functions of the n estimates may be used to generate the final estimate.
[0106] The volume s of the input CT medical image has been determined sup and S inf A final estimate of the anatomical structure of interest (e.g., liver) is calculated based on these values. This calculated score can then be compared with a predetermined threshold to determine whether a predetermined minimum amount of the anatomical structure of interest is present in the volume of the input CT medical image, for example, at least 1 / 3 of the liver is present in the volume of the input CT medical image.
[0107] If the determination logic 127 and 128 indicates that multiple phases are present and / or a predetermined amount of anatomical structure of interest is absent in the portion of the input volume 105 depicting the liver, further processing of the input volume 105 by the AI pipeline 100 for stages 130-170 may be discontinued (i.e., the input volume 105 is rejected). If the determination logic 127 and 128 results in a determination that the input volume 105 has an image of a single phase and depicts at least a predetermined amount of the liver, the portion of the input volume 105 depicting the anatomical structure is forwarded to the next stage 130 of the AI pipeline 100 for processing. While the exemplary illustrative embodiment forwards a sub-portion of the input volume containing the liver for further processing, in other illustrative embodiments, context surrounding the liver may also be provided. This may be accomplished by adding a predetermined amount of margin above and below the selected liver region. Depending on how much context is required for subsequent processing operations, this margin may be increased to fully cover the original input volume.
[0108] In the next stage 130 of the AI pipeline 100, the AI pipeline 100 performs lesion detection on a portion of the input volume 105 representing an anatomical structure of interest (e.g., the liver). This liver / lesion detection stage 130 of the AI pipeline 100 uses an ensemble of ML / DL computer models 132-136 to detect the liver and lesions in the liver as represented in the input volume 105. The ensemble of ML / DL computer models 132-136 performs liver and lesion detection using differently trained ML / DL computer models 132-136, wherein the ML / DL computer models 132-136 are trained and use a loss function to balance false positives and false negatives in lesion detection. In addition, the integrated ML / DL computer models 132-136 are configured such that a third loss function causes the outputs of the ML / DL computer models 132-136 to be consistent with each other.
[0109] In one illustrative embodiment, a configured and trained ML / DL computer model 132 is executed on an input volume 105 to detect the presence of a liver. This ML / DL computer model 132 can be identical to the ML / DL computer model 125 employed in the previous AI pipeline stage 120, thereby leveraging previously obtained results. Multiple (two or more) other ML / DL computer models 134-136 are configured and trained to perform lesion detection in portions of a medical image of the input volume 105 depicting a liver. A first ML / DL computer model 134 is configured and trained to operate directly on the input volume 105 and generate lesion predictions. A second ML / DL computer model 136 is configured with two different decoders implementing two different loss functions: one that penalizes false negatives (i.e., classifications that incorrectly indicate the absence of a lesion (normal anatomy)) and a second that penalizes false positives (i.e., classifications that incorrectly indicate the presence of a lesion (abnormal anatomy)). The first decoder of the ML / DL computer model 136 is trained to recognize patterns representing a relatively large number of different lesions at the expense of a high number of false positives. The second decoder of the ML / DL computer model 136 is trained to be less sensitive to lesion detection, but lesions that are detected are more likely to be accurately detected. A third loss function of the ensemble of ML / DL computer models as a whole compares the results of the decoders of the ML / DL computer models 136 with each other and reconciles them with each other. The lesion prediction results of the first and second ML / DL computer models 134, 136 are combined to generate a final lesion prediction for the ensemble, while the other ML / DL computer model 132 that generates a prediction of the liver mask provides an output representing the liver and its contours. Figure 6 Example architectures for these ML / DL computer models 132 - 136 are described in more detail.
[0110] The results of the liver / lesion detection stage 130 of the AI pipeline 100 include one or more contours (outlines) of the liver and a detection map (e.g., a voxel-wise map of the liver lesions detected in the input volume 105) that identifies the portions of the medical imaging data elements corresponding to the detected lesions 135. The detection map is then input to the lesion segmentation stage 140 of the AI pipeline 100. As described in more detail below, the lesion segmentation stage 140 partitions the detection map using a watershed technique and a corresponding ML / DL computer model 142 to generate image element (e.g., voxel) partitions of the medical images (slices) of the input volume 105. The liver lesion segmentation stage 140 provides other mechanisms (such as an ML / DL computer model 144) that, based on the partitions, identify all contours corresponding to lesions present in the slices of the input volume 105 and perform operations to identify which contours correspond to the same lesion in three dimensions. The lesion segmentation stage 140 further provides mechanisms (such as an ML / DL computer model 146) that aggregate related lesion contours to generate three-dimensional partitions of the lesions. Lesion segmentation uses image elements (e.g., voxels) representing lesions in medical images and in-painting of non-liver tissue to focus on each lesion individually and perform active contour analysis. In this way, individual lesions can be identified and processed without biasing the analysis due to other lesions in the medical image or due to portions of the image outside the liver.
[0111] The result of the lesion segmentation 140 is a list of lesions 148 having their corresponding outlines or contours in the input volume 105. These outputs 148 are provided to the false positive removal stage 150 of the AI pipeline 100. The false positive removal stage 150 uses a configured and trained ML / DL computer model that uses a dual operating point strategy to reduce false positive lesion detections in the list of lesions produced by the lesion segmentation stage 140 of the AI pipeline 100. By configuring the ML / DL computer model of the false positive removal stage 150 to remove as many lesions as possible, a first operating point is selected that is sensitive to false positives. After the sensitive false positive removal, it is determined whether a predetermined number or fewer lesions remain in the list. If so, a second operating point that is relatively less sensitive to false positives is used to reconsider the other lesions removed from the list. The results of these two methods identify the final filtered list of lesions to be further processed by the lesion classification stage of the AI pipeline.
[0112] After false positives have been removed from the list of lesions and their contours generated by the lesion segmentation stage 140, the resulting filtered lesion list 155 is provided as input to the lesion classification stage 160 of the AI pipeline 100, where a configured and trained ML / DL computer model is executed on the lesion list and its corresponding contour data to thereby classify the lesion into one of a plurality of predetermined lesion classifications. For example, each lesion in the final filtered lesion list and its attributes (e.g., contour data) can be input into the trained ML / DL computer model of the lesion classification stage 160, which then operates on this data to classify the lesion into a particular predetermined type or category of lesion.
[0113] After the lesions are classified by the lesion classification stage 160 of the AI pipeline 100, the AI pipeline 100 generates an output 170, which includes a finalized list of lesions and their classifications, as well as any contour attributes of the lesions. Furthermore, the AI pipeline 100 output 170 may also include liver contour information of the liver obtained from the liver / lesion detection stage 130. The output generated by the AI pipeline 100 may be provided to a further downstream computing system 180 for further processing and generation of representations of the anatomical structures of interest and any detected lesions present therein. For example, a graphical representation of the input volume may be generated in a medical image viewer or other computer application of the downstream computing system 180, where the contour information generated by the AI pipeline is superimposed or otherwise highlighted in the graphical representation. In other illustrative embodiments, downstream processing by the downstream computing system 180 may include diagnostic decision support operations and automated medical imaging report generation based on the list of detected lesions, their classifications, and contours. In other illustrative embodiments, based on the lesion classification, different treatment recommendations may be generated for review and consideration by a medical practitioner. In some illustrative embodiments, the lesion lists, their classifications, and profiles can be stored in a history data structure of a downstream computing system 180 in association with a patient identifier, so that multiple executions of the AI pipeline 100 on different input volumes 105 associated with the same patient can be stored and evaluated over time. For example, differences between the lesion lists and / or their associated classifications and profiles can be determined to assess the progression of a patient's disease or medical condition, and such information can be presented to a medical professional for use in assisting with the patient's treatment. Other downstream computing systems 180 of the illustrative embodiments and the processing of specified anatomical structures and lesion detection information generated by the AI pipeline 100 can be implemented without departing from the spirit and scope of the present invention.
[0114] Figure 2 is an example flow chart outlining example operations of an AI pipeline according to one illustrative embodiment. Figure 2The operations outlined in can be implemented through various logic levels, including configuration and training of ML / DL computer models, as in Figure 1 , and described above with reference to specific example embodiments described in the following separate sections of this specification. It should be understood that this operation is specifically directed to automated artificial intelligence pipelines implemented in one or more data processing systems having one or more computing devices specifically configured to implement these automated computer tool mechanisms. In addition to medical image volume creation time and when using output from downstream computing systems, Figure 1 and 2 There is no human intervention in the operations outlined. The present invention, inter alia, provides improved automated artificial intelligence computing mechanisms to perform the described operations that avoid human interaction and reduce potential errors due to previous manual processes by providing new and improved procedures that are specifically different from any previous manual processes and are specifically directed to providing logic and data structures that allow the improved artificial intelligence computing mechanisms of the present invention to be implemented in automated computing tools.
[0115] like Figure 2 As shown, the operation begins by receiving an input volume of a medical image from a medical imaging technology computing system (e.g., a computing system that provides computed tomography (CT) medical images) (step 210). The AI pipeline operates on the received input volume to perform body part detection (step 212), so that a determination can be made as to whether a body part of interest is present in the received input volume (step 214). If the body part of interest (e.g., the abdomen in the case of liver lesion detection and classification) is not present in the input volume, the operation terminates. If the body part of interest is present in the input volume, phase classification and minimal anatomical structure assessment are performed either sequentially or in parallel.
[0116] That is, Figure 2 As shown, phase classification is performed on the input volume (step 216) to determine whether the input volume includes a single phase for medical imaging (e.g., pre-contrast imaging, partial contrast imaging, delayed phase, etc.) or a medical image (slice) of multiple phases. A determination is then made as to whether the phase classification indicates a single phase or multiple phases (step 218). If the input volume includes a medical image pointing to multiple phases, the operation terminates; otherwise, if the input volume includes a medical image pointing to a single phase, the operation continues to step 220.
[0117] In step 220, a detection of the anatomical structure of interest (e.g., the liver in the depicted example) is performed to determine whether a minimum amount of the anatomical structure exists in the input volume to enable accurate execution of subsequent stages of the AI pipeline operation. A determination is made as to whether a minimum amount of the anatomical structure exists (e.g., representing at least 1 / 3 of the liver in the input volume) (step 222). If a minimum amount does not exist, the operation terminates; otherwise, the operation continues to step 224.
[0118] In step 224, liver / lesion detection is performed to generate outlines and detection maps of the lesions. These outlines and detection maps are provided to lesion segmentation logic, which performs lesion segmentation based on these outlines and detection maps (e.g., liver lesion segmentation in the depicted example) (step 226). Lesion segmentation results in the generation of a list of lesions and their outlines, as well as detection and contour information of anatomical structures (e.g., liver) (step 228). Based on this list of lesions and their outlines, a false positive removal operation is performed on the lesions in the list to remove false positives and generate a filtered list of lesions and their outlines (step 230).
[0119] The filtered list of lesions and their contours is provided to lesion classification logic, which performs lesion classification to generate a final list of lesions, their contours, and lesion classifications (step 232). This final list, along with the liver contour information, is provided to a downstream computing system (step 234), which can operate on this information to generate medical imaging views in a medical imaging viewer application, generate treatment recommendations based on the classification of the detected lesions, evaluate the historical progression of lesions over time in the same patient based on comparisons of the final lesion lists generated by the AI pipeline at different time points, and so on.
[0120] Thus, the illustrative embodiments, as outlined above, provide automated artificial intelligence mechanisms and ML / DL computer models that operate on an input volume of a medical image and generate a list of lesions, their contours, and classifications while minimizing false positives. The illustrative embodiments provide automated artificial intelligence computer tools that specifically identify, within a given set of image voxels of an input volume, which of the voxels correspond to a portion of an anatomical structure of interest (e.g., the liver), and which of these voxels correspond to lesions in the anatomical structure of interest (e.g., liver lesions). The illustrative embodiments provide a significant improvement over previous methods, both manual and automated, because the illustrative embodiments can be integrated into a fully automated computer tool within a clinician's workflow. In fact, based on the early stages of the AI pipeline design of the illustrative embodiments, which accept input volumes of only a single phase (e.g., abdominal scans) and reject input volumes that do not depict the anatomy of interest (e.g., liver) or do not depict a predetermined amount of the anatomy of interest (e.g., too small and too much liver), only meaningful input volumes are processed by the automated AI pipeline, thereby preventing radiologists from expending valuable manual resources on useless or flawed results when reviewing input volumes of non-anatomical structures of interest (e.g., non-liver cases). In addition to preventing radiologists from being overwhelmed with useless information, the automated AI pipeline of the illustrative embodiments also ensures smooth information technology integration by avoiding congestion of the AI pipeline and downstream computing systems (such as networks) with data associated with cases that do not correspond to the anatomy of interest or do not provide a sufficient amount of the anatomy of interest to archive and review the computing systems. Furthermore, as described above, the automated AI pipeline of the illustrative embodiments allows for the accurate detection, measurement, and characterization of lesions in a fully automated manner, which is technically possible through one or more of the automated AI pipeline structures of the illustrative embodiments and their corresponding automated ML / DL computer model-based components.
[0121] ML / DL computer models for detecting the presence of minimal anatomical structures in an input volume
[0122] As previously described, as part of the processing of the input volume 105, it is important to ensure that the input volume 105 represents a single phase of medical imaging and that at least a minimum amount of anatomical structure of interest is represented in the input volume 105. To determine that a minimum amount of anatomical structure of interest is present in the input volume 105, in one illustrative embodiment, the determination logic 128 implements a specially configured and trained ML / DL computer model that estimates a slice score for determining a portion of an anatomical structure (e.g., the presence of a liver in the input volume 105). The following description provides an example embodiment of such a configured and trained ML / DL computer model based on a defined axial scoring technique.
[0123] Figure 3A is an example diagram illustrating an example input volume (medical image) of the abdomen of a human patient according to one illustrative embodiment. Figure 3A In the description of , a two-dimensional representation of a three-dimensional volume is shown. The slice is Figure 3A Horizontal lines within the two-dimensional representation shown in , but would be represented as planes extending into and / or out of the page to represent flat two-dimensional slices of the human body, where stacking of these planes results in a three-dimensional image.
[0124] like Figure 3A As shown, the illustrative embodiment defines axial scores for slices ranging from 0 to 1. The axial scores are defined such that a slice corresponding to the first slice containing the liver (FSL) has a slice score of 0, while the last slice containing the liver (LSL) has a score of 1. In the depicted example, the first and last slices are defined as being associated with the most inferior slice in the volume (MISV) and the most superior slice in the volume (MSSV), where the axial scores along a given axis of the volume (e.g., Figure 3A The lowermost and uppermost slices are determined by the y-axis in the depicted example. Thus, in this depicted example, the MSSV is located at the slice with the highest y-axis value and the MISV is located at the slice with the lowest y-axis value. For example, the MISV may be closest to the lower limb of the biological entity (e.g., the foot of a human subject) and the MSSV may be closest to the upper part of the biological entity (e.g., the head of a human subject). The FLS is the slice that depicts the anatomical structure of interest (e.g., the liver) that is relatively closest to the MISV. The LSL is the slice that depicts the anatomical structure of interest that is relatively closest to the MSSV. In one illustrative embodiment, a trained ML / DL computer model (e.g., a neural network) can assign axial scores by taking a chunk of slices as input and outputting the height (axial score) of the center slice in the chunk. The trained ML / DL computer model is trained with a cost function that minimizes the error of the actual height (e.g., minimum squared error). The trained ML / DL computer model is then applied to all chunks covering the input volume (there may be some overlap between chunks).
[0125] Liver axial score estimation (LAE) is calculated from a pair of slice scores s sup and s inf To define, the slice score s sup and s inf Slice scores corresponding to MSSV and MISV slices, respectively. Figure 1 The ML / DL computer model of the determination logic 128 in the input volume 105 is specifically configured and trained to determine the slice score s sup and s inf, and knowing these slice scores, the mechanisms of the illustrative embodiments are able to determine the fraction of the liver that is in the field of view of the input volume 105 .
[0126] In some demonstrative embodiments, the slice score s may be found indirectly by first dividing input volume 105 into a plurality of segments, eg, segments comprising X slices (eg, 20 slices). sup and s inf , then for each segment, the configured and trained ML / DL computer model is executed on the slices of that segment to estimate the slice scores s' for the first and last slices in the segment sup and s'inf, where "first" and "last" can be determined based on the direction of progression along the axis of the three-dimensional volume 105 (e.g., from the first slice to the last slice along the y-axis, from the smallest y-axis value slice to the highest y-axis value slice). Given s' sup and s'inf, find s by extrapolation sup and s inf Since it is known how the segments are positioned relative to the entire volume 105, it should be noted that for each volume there will be n s sup and s inf Where n is the number of segments per volume. In one illustrative embodiment, the final estimate is obtained by taking an unweighted average of these n estimates, however, in other illustrative embodiments, other functions of the n estimates may be used to produce the final estimate.
[0127] For example, Figure 3B Shown Figure 3A Another depiction of the input volume, where the slice segments are sliced together with their corresponding axial scores s' inf and s' sup Together. Figure 3B As shown, in this example, the segments are defined as 20 slices separated by 5 mm. For each 20-slice segment of the volume, the slice score s' is estimated by the ML / DL computer model. sup and s' inf , and from these s' along a given range (eg, a range from 0 to 1, a range from -0.5 to 1.2, or any other desired predetermined range suitable for a particular embodiment) sup and extrapolation of s'inf values to obtain s sup and s inf In this example, assuming a predetermined range of -0.5 to 1.2, if s is estimated by applying an ML / DL computer model sup and extrapolated to about 1.2 and s inf is estimated to be -0.5, indicating that the entire liver is contained within the volume. Similarly, if s is estimated to besup is 1.2 and s inf = 0.5, then these values indicate that approximately 50% of the upper axial liver extension is contained in the volume (e.g., coverage is (1.2-0.5) / (1.2-(-0.5))=0.7 / 1.7≈0.41). As another example, in another illustrative embodiment, where the liver starts at -2.0 and ends at 0.8 (i.e., s sup is estimated to be 0.8 and s inf The upper limit of the liver is below 1.2, so the liver is cut in the upper part, and the lower limit is below -0.5, so that the base of the liver is completely covered. This indicates that about 80% of the lower axial liver extension is contained in the volume (i.e., the coverage is (0.8-max(-2, -0.5)) / (1.2-(-0.5))=(0.8-(-0.5) / 1.7)≈0.76).
[0128] Figure 3C yes Figure 3A Example diagram of an input volume where the volume is divided axially into n completely overlapping segments. In the depicted example, there are 7 segments indicated by arrows. It should be noted that in this example, the last two segments (arrows at the top of the figure) are almost identical. As previously mentioned, s' sup and s' inf The values were estimated by the ML / DL computer model for each of these segments and used to extrapolate the s values for MSSV and MISV slices. sup and s inf value, then, s sup and s inf The value may be used to determine the amount of anatomical structure of interest present in the input volume 105 .
[0129] Thus, by first dividing the input volume 105 into segments and then estimating for each segment the slice score s′ of the first and last slice in the segment sup and s'inf to indirectly find the s of MSSV and MISV sup and s inf Given these estimates, estimate s by extrapolation sup and s inf Since it is known how the segments are positioned relative to the entire input volume 105, there are n s extrapolated from each segment. sup and s inf where n is the number of segments per volume. The final estimate may be obtained, for example, by estimating any suitable combining function of the n estimates, such as an unweighted average of the n estimates or any other suitable combining function.
[0130] Figures 4A-4CAn example diagram showing an illustrative embodiment of an ML / DL computer model configured and trained to estimate s′ for a segment of an input volume of a medical image, according to an illustrative embodiment. sup and s'inf values. Figures 4A-4C The ML / DL computer model of the embodiment is only one example of an architecture of an ML / DL computer model, and many modifications can be made to the architecture without departing from the spirit and scope of the present invention (such as changing the tensor size of the input slices of the input volume, changing the number of nodes in the layers of the ML / DL computer model, changing the number of layers, etc.). Given this description, one of ordinary skill in the art will recognize how to modify the ML / DL computer model of the illustrative embodiment to a desired implementation.
[0131] like Figure 4A As shown in FIG, a sequence of 20 slices representing a segment 410 or "thick slice" of the input volume 105 is provided as input to processing blocks (PBs) 420-430. In the illustrative embodiment depicted, the PBs 420-430 are logical blocks that mix convolutional layers and LSTM layers (e.g., Figure 4B and 4C ). Features are extracted from the convolutional layers of PBs 420, 430 and then fed as input to the LSTM layers of PBs 420, 430. This is a smart / light modeling of the fact that the slices have a specific order in the anatomical region of interest or anatomical structure (e.g., abdomen / liver) that is driven by the anatomy (e.g., the relative positions of the liver, kidneys, heart, etc. in addition to the liver anatomy itself). In the depicted example, the tensor size of the initial 20 slices 410 in this example is 128x128. In this example embodiment, the first processing block 420 reduces the size of the tensor by 8 to generate 20 slice segments with slices of size 16x16x32 (it should be understood that the number of slices in a segment is implementation specific and can be modified without departing from the spirit and scope of the invention), where 32 is the number of filters. The second processing block 430 converts the input segment slices into 20 slice segments with slices of dimensions 2x2x64, where 64 is the number of filters. A subsequent neural network 440 configured with flat layers, dense layers, and linear layers is configured and trained to generate s′ for the input segment 410 of the input volume 105. sup and s'inf estimates. Figure 4B shows the composition of a processing block (PB) with respect to a convolutional layer and an LSTM layer according to an illustrative embodiment, and Figure 4C An example configuration of each of these convolutional and LSTM layers for each PB is shown, according to one illustrative embodiment.
[0132] by Figures 4A-4C Taking the ML / DL computer model architecture of FIG1 as an example, during the training of the ML / DL computer model, in one illustrative embodiment, medical imaging data (e.g., Digital Imaging and Communications in Medicine (DICOM) data) is assembled into, for example, a i The input volume is a 3D array of size 512x512, with floating point numbers 32, having Hounsfield Unit (HU) values, which is a normalized physical value that describes the attenuation of X-rays by the material present at a given location (e.g., voxel). i is the number of slices in the i-th volume, where i is in the range of 0 to N-1, where N is the total number of volumes. Each input volume is processed by a body part detector and an approximate region corresponding to the abdomen is extracted as described above (in the case of liver detection). The abdomen is defined as a continuous region between, for example, the axial scores from the body part detector -30 and 23. Slices outside this continuous region are rejected, and the ground truth can be defined as the position of the appropriately adjusted FSL and LSL. For example, assuming the input volume ranges from a to b, if there is no overlap between [a:b] and [-30, 23], the input volume is rejected. In other words, if b>23 or if a<-30.
[0133] The input segments 410 or "slabs" are overlapped to a predetermined slice interval (eg, 5 mm). The input segments 410 are reshaped to 128x128 in the x, y dimensions, which produces N shapes of M i x128x128 segments 410. This is called downsampling of the data in the input volume. Since the ordering of slices within the input volume depends on coarse information (e.g., the size of the organ), the AI pipeline still operates well on the downsampled data, and both the processing and training time of the AI pipeline are improved due to the reduction in the size of the downsampled data.
[0134] Input segments 410 having less than a predetermined number of slices (eg, 20) or a pixel size less than a predetermined pixel size (eg, 55 mm) are rejected, resulting in N′ M i x128x128 segment. Use a linear transformation from its acquisition range (e.g., -1024, 2048) to the range (0, 1) to clip and normalize the values in the segment. At this point, as described above, the N' M i The x128x128 segments constitute the training set on which the neural network 440 is trained to generate s' for the input segments. sup and estimates of s'inf.
[0135] With respect to performing inference with the trained neural network 440, the above operations are used to process the input volume 105 through body part detection, slice selection corresponding to the body part of interest, re-slicing, reshaping, rejection of certain slices that do not meet predetermined requirements, and again performing the generation of cropped and normalized segments for the new segments of the input volume 105. After generating the cropped and normalized segments, the input volume 105 is divided into R-ceil(M-10) / 10 sub-volumes or segments containing 20 slices to thereby generate partitions of slices with overlapping chunks. For example, if there are N'=31 slice volumes (slice numbers 0-30), then there are three segments or sub-volumes containing the following overlapping slices: 0-19, 10-29, 11-30. These segments or sub-volumes will typically have an overlap of approximately at least 50%.
[0136] Thus, an ML / DL computer model is provided, configured and trained to calculate the s′ of a segment of a volume corresponding to a predetermined number of slices (medical images) given a defined axial score range from 0 to 1. sup and s'inf value estimation, estimate the input volume s sup and s inf Based on these estimates, a determination can be made as to whether the input volume includes medical slices that together constitute at least a predetermined amount of the anatomical structure of interest (e.g., the liver). As previously discussed, this determination can be part of the determination logic 128 of the AI pipeline 100 for determining whether there is sufficient representation of the anatomical structure in the input volume 105 to allow for accurate liver / lesion detection, lesion segmentation, etc. in further downstream stages of the AI pipeline 100.
[0137] Figure 5 is a flow chart outlining an example operation of the liver detection and predetermined amount of anatomical structure determination logic of the AI pipeline according to one illustrative embodiment. Figure 5As shown, the liver detection operation of the AI pipeline begins by receiving an input volume (step 510) and dividing the input volume into multiple overlapping segments with a predetermined number of slices for each segment (step 520). The slices of each segment are input into a trained ML / DL computer model, which estimates the axial scores of the first and last slices in each segment (step 530). The axial scores of the first and last slices are used to extrapolate the scores of the lowest slice in the volume (MISV) and the highest slice in the volume (MSSV) of the input volume (step 540). This results in multiple estimates of the axial scores of the MISV and MSSV, which are then combined by functions of the respective estimates to generate an estimate of the axial scores of the MISV and MSSV of the input volume (e.g., a weighted average, etc.) (step 550). Based on the estimates of the axial scores of the MISV and MSSV, the axial scores are compared with criteria for determining whether a predetermined amount of anatomical structure of interest (e.g., liver) is present in the input volume (step 560). Thereafter, the operation terminates.
[0138] Liver / lesion detection
[0139] As previously described, assuming that the input volume 105 is determined to have a single phase represented, and that the input volume 105 has a predetermined amount of anatomical structures of interest represented in slices of the input volume 105, liver / lesion detection is performed on the portion of the input volume 105 that includes the anatomical structures of interest. In one illustrative embodiment, the liver / lesion detection logic stage 130 of the AI pipeline 100 employs a configured and trained ML / DL computer model that operates to detect the anatomical structure of interest (e.g., liver) in slices of the input volume 105 (again, in some illustrative embodiments, this can be the same ML / DL computer model 125 used in stage 120 for liver detection). The liver / lesion detection logic stage 130 of the AI pipeline 100 also includes an integration of multiple other configured and trained ML / DL computer models to detect lesions in images of the anatomical structure of interest (liver).
[0140] Figure 6is an example diagram of an integration of ML / DL computer models for performing lesion detection in an anatomical structure of interest (e.g., liver) according to an illustrative embodiment. The integration of ML / DL computer models 600 includes a first ML / DL computer model 610 for detecting the anatomical structure of interest (e.g., liver) and generating a corresponding mask. The integration of ML / DL computer models 600 also includes a second ML / DL computer model 620 that is configured and trained to process a liver mask input implemented in two decoders of the second ML / DL computer model 620 and generate lesion predictions using two competing loss functions. One loss function is configured to penalize false positive errors (yielding low sensitivity, but high accuracy), and the other loss function is configured to penalize false negative errors (yielding high sensitivity, but lower accuracy). An additional loss function is employed (in Figure 6 The ML / DL computer model 620 is used to make the outputs generated by the two competing decoders similar (consistent) to each other. The ensemble of ML / DL computer models further includes a third ML / DL computer model 630 that is configured and trained to directly process the input volume 105 and generate lesion predictions.
[0141] like Figure 6 As shown and described above, the integration 600 includes a first configured and trained ML / DL computer model 610 that is specifically configured and trained to identify anatomical structures of interest in an input medical image. In some illustrative embodiments, the first ML / DL computer model 610 includes a U-Net neural network model 612 that is configured and trained to perform image analysis to detect the liver within the medical image. However, it should be understood that the illustrative embodiments are not limited to this particular neural network model and that any ML / DL computer model that can perform segmentation may be utilized without departing from the spirit and scope of the present invention. U-Net is a convolutional neural network developed at the Computer Science Department of the University of Freiburg, Germany for biomedical image segmentation. The U-Net neural network is based on a fully convolutional network with an architecture that has been modified and extended to work with fewer training images and produce more accurate segmentations. U-Net is generally known in the art and thus a more detailed explanation will not be provided herein.
[0142] like Figure 6As shown, in one illustrative embodiment, a first ML / DL computer model 610 can be trained to process a predetermined number of slices at a time, where the number is determined to be appropriate for the desired implementation (e.g., 3 slices are determined through an empirical process to produce good results). In one illustrative embodiment, the slices of the input volume are, for example, 512x512 pixel medical images, but other implementations may use different slice sizes without departing from the spirit and scope of the illustrative embodiments. The U-Net generates a segmentation of the anatomical structure in the input slice, thereby producing one or more segments corresponding to the anatomical structure of interest (e.g., the liver). As part of this segmentation, the first ML / DL computer model 610 generates a segmentation representing a liver mask 614. The liver mask 614 is provided as input to at least one of the other ML / DL computer models 620 of the ensemble 600 so that the processing performed by the ML / DL computer model 620 is focused only on the portion of the input slice of the input volume 105 that corresponds to the liver. By pre-processing the input of the ML / DL computer model with the liver mask 614, the processing performed by the ML / DL computer model can be focused on the portion of the input slice corresponding to the anatomical structure of interest, rather than on "noise" in the input image. Other ML / DL computer models (e.g., ML / DL computer model 630) receive the input volume 105 directly without using the liver mask 614 generated by the first ML / DL computer model 610 for liver masking.
[0143] In the depicted illustrative embodiment of the integrated 600, the third ML / DL computer model 630 is composed of an encoder portion 634-636 and a decoder portion 638. The ML / DL computer model 630 is configured to receive 9 thick slices of the input volume 105, which are then separated into groups 631-633 of 3 slices each, wherein each group 631-633 is input into a corresponding encoder network 634-636. Each encoder 634-636 is a convolutional neural network (CNN) (such as DenseNet-121 (D121)) that has been pre-trained to recognize different types of objects (e.g., lesions) present in the input slice without a fully connected head and output a classification output (e.g., as an output classification vector, etc.) indicating the type of detected object present in the input slice. CNNs 634-636 may, for example, operate on three channels of the input slice, and the resulting output features of CNNs 634-636 are provided to the cascaded NHWC logic 637, where NHWC refers to the number of images in the batch (N), the image's height (H), the image's width (W), and the image's number of channels (C). The architecture of the original DenseNet network includes many convolutional layers and skip-true connections that downsample the 3-slice full-resolution input to many feature channels of smaller resolution. From there, a fully connected head aggregates all features and maps them to multiple categories in the final output of the DenseNet. Because the DenseNet network is used as the encoder in the depicted architecture, the head is removed and only the downsampled features are retained. Then, in the cascaded NHWC logic 637, all feature channels are concatenated to pass them to the decoder stage 638, which has the function of upsampling the image until the desired output probability map resolution (e.g., 512x512) is reached.
[0144] The encoders 634-636 share the same parameters that are optimized through the training process (e.g., weights, sampling of lesion types during training, weights of losses, types of augmentation, etc.). The training of the ML / DL computer model 630 uses two different loss functions. The primary loss function is an adaptive loss that is specifically configured to penalize false positive errors in slices that do not have lesions in the ground truth, and also to penalize false negative errors in slices that have lesions in the ground truth. The loss function is a modified version of the Tversky loss as follows:
[0145] For each output slice:
[0146] TP = sum(prediction * target)
[0147] FP = sum((1-target)*prediction)
[0148] FN = sum((1-prediction)*target)
[0149] LOSS=1-((TP+1) / (TP+1+α*FN+β*FP))
[0150] Where "prediction" is the output probability of the ML / DL computer model 630 and "target" is the ground truth lesion mask. The output probability value ranges between 0 and 1. For each pixel in the slice, the target has 0 or 1. For slices without lesions, the "α" term is small (e.g., zero) and "β" is large (e.g., 10). For slices with lesions, "α" is large (e.g., 10) and "β" is small (e.g., 1).
[0151] The second loss function 639 is a function connected to the output of the encoders 634-636. Because the input to this loss comes from the middle of the ML / DL computer model 630, it is called "deep supervision" 639. Deep supervision has been shown to enable the encoder neural network 634-636 to learn better representations of the input data during training. In an illustrative embodiment, this second loss is a simple mean squared error of predicting whether there is a lesion in the slice. Therefore, a mapping network is used to map the output features of the encoders 634-636 to 9 values between 0 and 1, which represent the probability of having a lesion in each of the 9 slice inputs. The decoder 638 generates an output of a probability map of the lesions detected in the specified input image.
[0152] The second ML / DL computer model 620 receives a pre-processed input of 3 slices from the input volume, which has been pre-processed with the liver mask 614 generated by the first ML / DL computer model 610 to identify the parts of the 3 slices corresponding to the liver mask 614. The resulting pre-processed input slices (having a size of 192x192x3 in the depicted example illustrative embodiment) are provided to the second ML / DL computer model 620, which includes a DenseNet-169 (D169) encoder 621 connected to two decoders (2D DEC - meaning the decoder consists of 2-dimensional neural network layers). The D169 encoder 621 is a neural network feature extractor that is widely used in computer vision applications. It consists of a series of convolutional layers, where the features extracted from each layer are connected in a feed-forward manner to any other layer. The features extracted in the encoder 621 are passed to two independent decoders 622, 623, where each decoder 622, 623 consists of a 2D convolutional layer and an upsampling layer (in Figure 6The decoder 622 is composed of a plurality of decoders 622 and 623 (referred to as 2D DEC in the text). Each decoder 622, 623 is trained to detect lesions (e.g., liver lesions) in the input slice. As discussed above and below, although the two decoders 622, 623 are trained to perform the same task (i.e., lesion detection), the key difference in their training is that the two decoders 622, 623 each utilize a different loss function in order to drive the detection training into two competing directions. The final detection map of the second ML / DL model 620 is combined with the final detection map of the third ML / DL model 630 by an averaging operation 640. This process is applied to all input thick slices of the input volume 105 to generate a final detection map (e.g., liver lesions).
[0153] As described above, the second ML / DL computer model 620 is trained using two different loss functions that attempt to achieve opposite detection operating point performance. That is, while one of the decoders 622 uses a loss function for training that penalizes errors in false-negative lesion detection, resulting in high-sensitivity detection with relatively low accuracy, the other encoder in the decoder 623 uses a loss function for training that penalizes errors in false-positive lesion detection, resulting in low-sensitivity detection but high accuracy. An example of such a loss function can be a focal Tversky loss (see Abraham et al., "A Novel Focal Tversky Lossfunction with Improved Attention U-Net for Lesion Segmentation," arXiv:1810.07842[cs], October 2018), where parameters are adjusted for either high or low penalties for false positives and false negatives according to an illustrative embodiment. A third loss function (consistency loss 627) is used to enforce consistency between the predicted detections of each decoder 622, 623. The consistency loss logic 627 compares the outputs 624, 625 of the two decoders 622, 623 to each other and makes these outputs similar to each other. This loss can be, for example, a mean square error loss between two predicted detections, a structural similarity loss, or any other loss that enforces consistency / similarity between the compared predicted detections.
[0154] In operation, using these relative operating point decoders 622, 623, the second ML / DL computer model 620 generates two lesion outputs 624, 625, which are input to slice averaging (SLC AVG) logic 626 that generates an average of the lesion outputs. This average of the lesion outputs is then resampled to generate an output that is dimensionally commensurate with the output of the third ML / DL computer model 630 for comparison (note that this process includes restoring the liver masking operation, and therefore the lesion outputs are calculated at the original 512x512x3 resolution).
[0155] At runtime, slice average (SLC AVG) logic 626 operates on the lesion prediction outputs 624 and 625 of the decoders 622 and 623 to generate the final detection map for the ML / DL model 620. It should be appreciated that while a consistency loss 627 was applied during training to drive each decoder 622 and 623 to learn consistent detections, this consistency loss is no longer utilized at runtime. Instead, the ML / DL model 620 outputs two detection maps that need to be aggregated by the SLC AVG module 626. The results of the SLC AVG logic 626 are resampled to generate an output with dimensions commensurate with the input slab (512x512x3). All detections generated by the ML / DL model 620 for each slab of the input volume 105 are combined with the detections generated by the ML / DL model 630 via volume average (VOL AVG) logic 640. This logic calculates the average of the two detection masks at the voxel level. The result is a final lesion mask 650 corresponding to the lesions detected in the input volume 105.
[0156] Thus, after training the ML / DL computer models 620, 630, when presented with a new slice of a new input volume 105, the first ML / DL computer model 610 generates a liver mask 614 that is used to pre-process the input of the second ML / DL computer model 620, and both ML / DL computer models 620, 630 process the input slice to generate a lesion prediction that is averaged over the volume by volume averaging logic 640. The result is a final lesion output 650 and a liver mask output 660 based on the operation of the first ML / DL computer model 610. These outputs can be provided as outputs of the liver / lesion detection logic stage 130 of the AI pipeline 100, which are provided to the lesion segmentation logic stage 140 of the AI pipeline 100, as previously discussed above and described in more detail below. Thus, the mechanisms of the illustrative embodiments provide an integrated 600 approach for anatomical structure recognition and lesion detection in an input volume 105 of a medical image (slice).
[0157] Through Figure 6The illustrated integrated architecture achieves improved performance over that achieved using a single ML / DL computer model. Specifically, it has been observed that by using the integrated architecture, by combining the detection outputs of multiple ML / DL computer models in the integration, improved detection specificity is achieved at the same level of sensitivity as a single ML / DL computer model. That is, while errors (false positives) are generated at different locations using ML / DL models 620 and 630, when the detection outputs of the different locations are averaged, the signal from the false positives is reduced, while the signal from the true positive lesions prevails, resulting in improved performance.
[0158] Figure 7 is a flow chart outlining an example operation of the liver / lesion detection logic in the AI pipeline according to one illustrative embodiment. Figure 7 As shown, the operation begins by receiving an input volume (step 710) and performing anatomical structure detection (e.g., liver detection) using a first trained ML / DL computer model (such as a U-Net computer model configured and trained to recognize anatomical structures (e.g., liver)) (step 720). The result of the anatomical structure detection is a segmentation of the input volume to identify a mask of the anatomical structure (e.g., liver mask) (step 730). The input volume is also processed by the integrated first trained ML / DL computer model, which is specifically configured and trained to perform lesion detection (step 740). The first trained ML / DL computer model generates a first set of lesion detection prediction outputs based on its processing of the input volume (step 750).
[0159] The integrated second trained ML / DL computer model receives a masked input generated by applying the generated anatomical structure mask to the input volume, and thereby identifies the portion of the medical image in the input volume that corresponds to the anatomical structure of interest (step 760). The second trained ML / DL computer model processes the masked input via two different decoders with two different and competing loss functions (e.g., one loss function penalizes errors in false positive lesion detections, while the other loss function penalizes errors in false negative lesion detections) (step 770). The result is two sets of lesion prediction outputs, which are then combined by combinational logic to generate a lesion prediction output for the second ML / DL computer model (step 780). If necessary, the second lesion prediction output is resampled and combined with the first lesion prediction output generated by the integrated first ML / DL computer model to generate a final lesion prediction output (step 790). The final lesion prediction output is then output together with the anatomical structure mask (step 795), and the operation terminates.
[0160] Lesion segmentation
[0161] As previously described, lesion prediction outputs are generated through the operation of various ML / DL computer models and stages of the AI pipeline logic including body part detection, determination of body parts of interest, phase classification, identification of anatomical structures of interest, and anatomical structure / lesion detection. Figure 1 In the illustrated AI pipeline 100, the results of the liver / lesion detection stage 130 of the AI pipeline 100 include one or more contours (outlines) of the liver, and a detection map (e.g., a voxel-wise map of liver lesions detected in the input volume 105) that identifies portions of the medical imaging data elements corresponding to the detected lesions 135. The detection map is then input to the lesion segmentation stage 140 of the AI pipeline 100.
[0162] As mentioned above, the lesion segmentation logic (e.g. Figure 1 The lesion segmentation stage 140 in the embodiment of the present invention uses a watershed technique and a corresponding ML / DL computer model to partition the detection map to generate image element partitions of the medical image (slice) of the input volume. The liver lesion segmentation stage also provides other mechanisms (such as one or more other ML / DL computer models) that identify all contours corresponding to lesions present in the slice of the input volume based on the image element partitions and perform operations to identify which contours correspond to the same lesion in three dimensions. The lesion segmentation stage further provides multiple mechanisms (such as one or more additional ML / DL computer models) that aggregate related lesion contours to generate a three-dimensional partition of the lesion.
[0163] Lesion segmentation uses lesion image elements and in-painting of non-liver tissue represented in the medical image to focus on each lesion individually and perform active contour analysis. In this way, individual lesions can be identified and processed without biasing the analysis due to other lesions in the medical image or due to portions of the image outside the liver. The result of lesion segmentation is a list of lesions with their corresponding outlines or contours in the input volume.
[0164] Figure 8 A block diagram depicting an overview of various aspects of the lesion segmentation process performed by lesion segmentation logic according to one illustrative embodiment is shown. Figure 8 As depicted, lesion segmentation includes mechanisms for slice-wise segmentation of 2D detections (i.e., detection of lesions in 2D slices) (block 810), connecting 2D lesions along the z-axis (block 820), and slice-wise refinement of contours (block 830). Each of these blocks will be described in more detail below with respect to subsequent figures. Figure 8The segmentation process shown in is implemented as a process for identifying all lesions in a given input volume in an analysis and distinguishing lesions that are close to each other in images (slices) of the input volume. For example, two lesions that appear to be pixel-wise merged in one or more images may need to be identified as two different regions or different lesions for the purpose of performing other downstream processing of the detected lesions (such as during lesion classification) and for separately identifying lesions in the output of a lesion list for use in downstream computing system operations (such as providing a medical review application, performing a treatment recommendation operation, performing a decision support operation, etc.).
[0165] As part of partitioning the 2D image slice-wise in block 810, the mechanisms of the illustrative embodiments use existing watershed techniques to partition the detection map from the previous lesion detection stage of the AI pipeline (e.g., by Figure 1 The watershed algorithm requires a seed to be defined in order to perform mask partitioning. The watershed algorithm splits the mask into as many regions as there are seeds, so that each region has exactly one seed located approximately at its center, e.g. Figure 10A and 10C As shown. In automatic segmentation, the seeds in the mask can be obtained as the local maxima of their distance map (distance to the mask contour). However, this method is prone to noise and may result in too many seeds, thus over-splitting the mask. Therefore, we need to edit the partition by reorganizing some areas of the region. Considering the empirical observation that most lesions are bubble-shaped, the guiding principle of region reorganization is to make the resulting new regions roughly circular. For example, for Figure 10C , the mechanism will merge the two regions identified by seeds 1051 and 1061, respectively, thereby producing a new mask partition that includes only two roughly circular regions. Thus, for the detected lesion defined in the detection map 135, as in Figure 9 The lesion shown on the left (described later) can be divided into several bubble-like lesions, as shown in Figure 9 They are to be interpreted as cross sections of the 3D lesion on the slice.
[0166] Watershed segmentation is a region-based method that originates from mathematical morphology. In watershed segmentation, the image is viewed as a topographic landscape with ridges and valleys. The elevation values of the landscape are usually defined by the grayscale value of the corresponding pixel or the magnitude of its gradient, so the two dimensions are treated as a three-dimensional representation. The watershed transform decomposes the image into "catchment basins". For each local minimum, the catchment basin includes all points whose steepest descent path ends at this minimum. The watershed separates the basins from each other. The watershed transform completely decomposes the image and assigns each pixel to a region or watershed.
[0167] Watershed segmentation requires selecting at least one label (called a "seed" point) within each object in the image. The seed point can be selected by the operator. In one embodiment, the seed point is selected by an automated process that takes into account application-specific knowledge of the object. Once the object is labeled, the object can be grown using the morphological watershed transform, which will be described in further detail below. Lesions typically have a "bubble" shape. The illustrative embodiments provide techniques for merging watershed segmented regions based on this assumption.
[0168] Thereafter, in block 820, the mechanism of the illustrative embodiment aggregates the voxel partitions along the z-direction on each slice to produce a three-dimensional output. Therefore, the mechanism must determine whether two groups of image elements (e.g., voxels) in different slices belong to the same lesion (i.e., whether they are aligned in three dimensions). The mechanism calculates measurements between lesions in adjacent slices based on the intersection and union of the lesions, and applies a regression model to determine whether two lesions in adjacent slices are part of the same region. One can view each lesion as a group of voxels, and the mechanism determines the intersection of two lesions as the intersection of the two groups of voxels, and the union of two lesions as the union of the two groups of voxels.
[0169] This results in a three-dimensional partitioning of the lesion; however, the contours may not fit the actual image very well. There may be over-segmented lesions. The illustrative embodiment proposes the use of active contours, which is a traditional framework for dealing with segmentation problems. This algorithm attempts to iteratively edit the contours so that they fit the image data better and better, while ensuring that they maintain certain desirable properties (such as shape smoothness). In block 830, the mechanism of the illustrative embodiment initializes the active contours using the partitions obtained from the first stage 810 and the second stage 820, and focuses on one lesion at a time; otherwise, running active contours or random segmentation methods on close lesions may cause them to merge into one contour again, which is counterproductive because it is equivalent to essentially eliminating the benefits brought by the previous zone levels. The mechanism focuses on one lesion and performs "painting" on the lesion voxels and non-liver tissue near the focused lesion.
[0170] The chaining of these three processing stages allows for processing that is not biased towards other lesions in the image or towards pixels or lesions outside the liver.
[0171] Slice-based 2D detection
[0172] Figure 9 Depicts the results of lesion detection and slice-wise partitioning according to one illustrative embodiment. Figure 9 As seen on the left side of FIG, the lesion area 910 is detected by the previous AI pipeline process described above and can be obtained from the lesion detection logic (e.g., Figure 1 130) in contour and detection maps (e.g., Figure 1 135) in the output is limited. Figure 9 As shown on the right side of , according to one illustrative embodiment, Figure 8 The logic of block 810 in attempts to partition the area into three lesions 911, 912, and 913. The partitioning mechanism of the illustrative embodiment is based on existing watershed techniques, which operate to partition the detection map from the previous lesion detection stage of the AI pipeline. Watershed algorithms are primarily used for image processing for segmentation purposes. The principle behind these known watershed algorithms is that grayscale images can be viewed as topographic surfaces, where high intensities represent peaks and hills, while low intensities represent valleys. The watershed technique begins by filling each isolated valley (local minimum) with water (markers) of a different color. As the water rises, water from different valleys with different colors will begin to merge, depending on the nearby peaks (gradients). To avoid this, barriers are built at the locations where the water merges. The work of filling water and building barriers continues until all peaks are underwater, at which point the created barriers give the segmentation results. Again, watershed techniques are generally known, and therefore, a more detailed description is not provided herein. Any known technique for slice-wise partitioning of 2D images may be used without departing from the spirit and scope of the present invention.
[0173] In the context of lesion segmentation, the empirical observation that most lesions are circular in shape strongly suggests that a partitioning that produces a set of circular regions is likely to be a good partitioning. However, as mentioned earlier, the quality of a watershed-type partitioning depends on the quality of the seeds. In fact, an arbitrary set of seeds need not result in a set of circular regions. For example, Figure 10C The watershed partition resulting from three seeds containing only one roughly circular region is shown. The other two are not circular. However, their union is again roughly circular. This configuration is called an over-split because the slanted split in the figure divides the otherwise circular region into two smaller, non-circular regions. Therefore, it is desirable to have an algorithm that can correct for over-splits. The seed relabeling mechanism does this by merging several over-split regions to form a coarser partition containing only circular regions. For example, the mechanism targets Figure 10C The partition in decides to merge the two regions identified by seeds 1051 and 1061 to form a new region that is more circular.
[0174] The illustrative embodiment merges the regions in the partition into smaller and larger regions that may correspond to physical lesions. Partitioning divides a region into smaller regions, or as described herein, partitioning divides a mask into smaller regions. In terms of contours, partitioning thus produces a set of smaller contours from a large contour (see Figure 9 from left to right).
[0175] The seeds are obtained by extracting local maxima from a distance map, which is computed from the input mask to the partitions. This map measures, for each pixel, its Euclidean distance to the mask contour. Depending on the topology of the input mask, local maxima derived from this distance map can lead to over-segmented partitions by the watershed algorithm. In this case, the watershed is said to be over-divisive and tends to produce regions that are not circular, which may be desirable in some applications, but is not ideal for lesion segmentation. Figure 10C In
[15] , we show a synthetic input mask whose distance map has three local maxima. Thus, the watershed produces a partition containing three regions, only one of which is roughly circular (corresponding to seed 1071). The other two are not. The region with seed 1051 is only semicircular. The seed relabeling mechanism then examines all seed pairs and determines that the two regions corresponding to seeds 1051 and 1061 should be merged together, forming a more perfect bubble. This operation results in a new partition containing only two regions, both of which are roughly circular in shape.
[0176] A local maximum is a point that has the largest distance from the contour compared to its immediate neighbors. The local maximum is a point, and its distance to the contour is known. Therefore, the mechanism of the illustrative embodiment can draw a circle centered at this point. The radius of the circle is the distance. For two local maxima, the mechanism can thus calculate the overlap of their corresponding circles. This is in Figure 10A and Figure 10B Depicted in.
[0177] Seed relabeling determines whether to merge two regions as follows. For two regions whose associated seeds are directly adjacent, a merge will occur; otherwise, the mechanism bases its decision on a hypothesis testing procedure. For example, see Figure 10A , the depicted example describes a case where the distance map produces two different local maxima, which leads to the assumption that each maximum represents the center of a different circular lesion. Note that the distance map also allows the mechanism of the illustrative embodiment to tell how far the maximum is from the contour (boundary). The distance is Figure 10B The dashed line segments connecting the maxima and points on the contours are shown in FIG. Thus, if the assumption holds, one can infer the spatial extent of the two lesions since the lesions are assumed to have a roughly circular or "bubble" shape. This allows the mechanism of the illustrative embodiment to be drawn as Figure 10B, the two complete circles shown in . From this, the mechanism then measures the overlap of the two circles (e.g., using the classic dice metric) and compares it to a predetermined threshold. If the value of the overlap metric is greater than the threshold, the mechanism concludes that the two bubbles overlap too much to be significant and will merge. In other words, the mechanism of this illustrative embodiment then concludes that the two local maxima correspond to two "centers" of the same lesion. However, in traditional watershed, there is no such seed (i.e., maximum) relabeling mechanism. As a result, mask over-splitting often occurs.
[0178] Overlap can be measured in a variety of ways. In one example embodiment, the mechanism uses the dice coefficient. Figure 10B The mechanism can calculate the dice metric for the two complete circles corresponding to the two local maxima shown in . In this way, the mechanism can learn from the training dataset what optimal threshold to apply in practice so that once the dice metric is greater than the threshold, the two local maxima are actually the centers of the same lesion.
[0179] Figure 10C and Figure 10D An example of another lesion mask shape is provided, which is similar to Figure 10A and Figure 10B The difference is that the two parts of the circle merged in Figure 10A China and Belgium Figure 10C Since the distance map can be very sensitive to the mask shape, Figure 10C There are three seeds in the example lesion mask shape. Following the above reasoning, the lesion splitting algorithm will Figure 10C The lesion represented in is split into two separate lesions, but not into three separate lesions as would occur in the watershed technique without seed relabeling.
[0180] exist Figure 10C and Figure 10D In the example, seeds 1051 and 1061 represent Figure 10A and Figure 10B 1061 , the seeds 1071 and 1061 are located at the center of a different bubble, resulting in the following: Figure 10C and Figure 10D Equivalently, this results in a label for seed 1071 that is different from the labels assigned to seeds 1051 and 1061. However, similar to Figure 10A and Figure 10B In the case of , the hypothesis testing procedure of the seed relabeling technique of the illustrative embodiment will determine that seeds 1051 and 1061 correspond to the same lesion.
[0181] Figure 11A is a block diagram illustrating a mechanism for lesion splitting and relabeling according to one illustrative embodiment. Figure 11A As shown, the mechanism can be implemented as a computer model including one or more algorithms, machine learning computer models, etc., which is executed by one or more processors of one or more computing devices and the one or more processors operate on an input volume of one or more medical image data structures, receive a two-dimensional lesion mask 1101, and perform a distance transform (block 1102) to generate a distance map 1111. The distance transform (block 1102) is an operation performed on a binary mask that calculates, for each point in the lesion mask, its shortest distance to the mask contour (boundary). The further you move toward the interior of the lesion mask, the further away from its contour (boundary). Thus, the distance transform identifies the center point of the lesion mask (i.e., those points with a greater distance than other points). In one embodiment, the mechanism optionally performs Gaussian smoothing on the distance map 1111.
[0182] The mechanism then performs local maximum identification (box 1103) to generate seeds 1112. As described above, these local maxima are the points in the distance map 1111 that have the highest distance to the contour or boundary. The mechanism performs a watershed technique (box 1104) based on the seeds 1112 to generate a watershed split lesion mask 1113. As described above, this split lesion mask 1113 can be over-split, resulting in areas that do not conform to the assumed bubble shape of the lesion. Therefore, the mechanism performs seed re-labeling (box 1120) based on the distance map 1111, the seeds 1112, and the split 2D lesion mask 1113 to generate an updated split lesion mask 1121. Seed re-labeling is described below with reference to Figure 11B The resulting updated split lesion mask 1121 will have regions that have been merged to form regions that more accurately conform to the bubble shape assumed for the lesion.
[0183] Figure 11B is a block diagram illustrating a mechanism for seed re-labeling according to one illustrative embodiment. Figure 11BAs shown, the mechanism receives a distance map 1111 and seeds 1112, and the mechanism can be implemented as a computer model including one or more algorithms, a machine learning computer model, etc., executed by one or more processors of one or more computing devices and operating on an input volume of one or more medical image data structures. More specifically, the mechanism considers each pair of seeds (seed A and seed B) in seeds 1112. The mechanism determines whether seed A and seed B are direct neighbors (box 1151). If seed A and seed B are direct neighbors, the mechanism assigns the same label to seed A and seed B (box 1155). In other words, seed A and seed B are grouped to represent a single region.
[0184] As described below, in block 1151, if seed A and seed B are not direct neighbors, the mechanism performs spatial range estimation (1152) based on the distance map 1111 and determines the pairwise affinity of seed A and seed B. According to the illustrative embodiment, the spatial range estimation assumes that the region is a "bubble" shape. Thus, the mechanism assumes that each seed represents a circle with the distance from the distance map as the radius of the circle.
[0185] The mechanism then calculates the overlap metric of the circles represented by seed A and seed B (block 1153). In one example embodiment, the mechanism uses the following dice metric:
[0186]
[0187] where |A| represents the area of the circle represented by seed A, and |B| represents the area of the circle represented by seed B. Similarly, |A∩B| represents the area of the intersection of A and B. In an alternative embodiment, the mechanism may compute the overlap metric as follows:
[0188]
[0189] Where |A| represents the area of the circle represented by seed A, |B| represents the area of the circle represented by seed B, |A∩B| represents the area of the intersection of A and B, and |A∪B| represents the area of the union of A and B.
[0190] The mechanism determines whether the overlap metric is greater than a predetermined threshold (block 1154).If the overlap metric is greater than the threshold in block 1154, the mechanism merges corresponding regions in the split 2D lesion mask 1113 (block 1155).
[0191] If the affinity between two seeds is greater than a threshold, they are assigned the same label. Otherwise, at this level, it is not known whether they should belong to the same group. This decision is left to the label propagation level ( Figure 15 1512 in the figure), the tag propagation stage is the same module used in the z-direction connection, which will be described below.
[0192] In the case where we have more than two seeds, repeat for all seed pairs before label propagation Figure 11B For example, there is a case where the seed pairs (a, b) and (b, c) are determined to belong to the same group, while the seed pair (a, c) fails the test, such as Figure 11B As shown. Label propagation would then necessarily place a, b, and c in the same group (i.e., the regions corresponding to seeds a and c would still be merged). However, if seeds a, b, c, and d were present, and the affinity calculation (performed for a total of six pairs) showed that only (a, b) and (c, d) passed the test, label propagation would result in two groups, one containing (a, b) and the other containing (c, d). Therefore, if a seed pair fails the test, it means that it is not known whether they should be placed in the same group, not that they should belong to different groups.
[0193] For example, in Figure 10C In
[15] , there are 3 seed pairs (1051-1061, 1051-1071, 1061-1071), and the mechanism should determine that seeds 1051 and 1061 should be assigned the same label (belong to the same group). The label propagation step then clusters these 3 seeds into 2 groups, the first group contains only 1071 and the second group has both 1051 and 1061.
[0194] Figure 12 is a flowchart outlining example operations for lesion splitting in accordance with one illustrative embodiment. Figure 12 The operations outlined in Figures 11A-11B The mechanism described is implemented. Figure 12 As shown, the operation begins (step 1200) and the mechanism generates a distance map for a two-dimensional lesion mask (block 1101). As described above, the distance map can be generated by performing a distance transform operation on the two-dimensional lesion mask and optionally performing Gaussian smoothing to remove noise. The mechanism then uses local maximum identification to generate groupings of data points (e.g., local maxima for each group) (step 1202). The mechanism performs lesion splitting based on local maxima to generate regions (step 1203). The mechanism then uses the distance map to relabel seeds based on pairwise similarity (step 1204). The mechanism then merges regions corresponding to seeds with the same label (step 1205). It should be understood that, as described above, due to the seed relabeling performed by the mechanism of the illustrative embodiment, the split lesion mask output in step 1205 does not suffer from the over-splitting problem associated with the watershed technique due to mislabeling associated with data points associated with each lesion shape. Thereafter, the operation ends (step 1206).
[0195] Z-connection of the lesion
[0196] The above-described process of lesion splitting and seed relabeling can be performed for each two-dimensional image or slice of the input volume, thereby generating an appropriately labeled lesion mask for each lesion represented in the corresponding two-dimensional image. However, the input volume represents a three-dimensional representation of the internal anatomy of a biological entity, and lesions that may appear to be associated with the same lesion when considered in three dimensions may actually be associated with different lesions. Thus, in order to correctly identify individual lesions within a biological entity represented in three dimensions of the input volume, the illustrative embodiments provide a mechanism for connecting two-dimensional lesions along the z-axis (i.e., in three dimensions).
[0197] A mechanism for performing the linking of two-dimensional lesions along the z-axis (referred to as z-linking of lesions) includes executing a logistic regression model on the split-lesion output generated by the above-described mechanism to determine three-dimensional z-lesion detection. This mechanism links two lesions in adjacent image slices. When the logistic regression model determines that two lesions represent the same lesion, the two lesions are linked. For example, for any two-dimensional lesions on adjacent image slices (i.e., slices whose z-axis coordinates are sequentially ordered along the z-axis in a collection of sliced three-dimensional tissue), as described below, the mechanism determines whether these two-dimensional lesions belong to the same three-dimensional lesion.
[0198] Figures 13A-13C A process for z-linking of lesions according to one illustrative embodiment is described. Figure 13A Lesion mask input is depicted. Figure 13B The lesion is depicted after slice-wise lesion splitting, which may employ the relabeled improved lesion splitting mechanism of the previously described illustrative embodiment. Figures 13A-13B As shown, slice 1310 has lesions 1311 and 1312, slice 1320 has lesion 1321, and slice 1330 has lesions 1331 and 1332. The z-connectivity of the lesion mechanism (i.e., a logistic regression model) is performed on the split lesion mask for each pair of adjacent slices in the input volume so that each lesion in a given slice is compared to each lesion in the paired adjacent slices. For example, the z-connectivity of the lesion mechanism compares lesion 1311 (lesion A) in slice 1310 with lesion 1321 (lesion B) in slice 1320. For each comparison, the mechanism treats each lesion as a set of voxels and determines the intersection between lesion A (a set of voxels in lesion A) and lesion B (a set of voxels in lesion B) for the size of lesion A and for the size of lesion B. The z-connectivity of the lesion mechanism determines whether lesion A and lesion B are connected based on two overlap ratios using a logistic regression model, as follows:
[0199]
[0200] Where |A| represents the area of the circle represented by seed A, |B| represents the area of the circle represented by seed B, and |A∩B| represents the area of the intersection of the circles represented by seed A and seed B. The mechanism uses these two ratios as input features to train a logistic regression model to determine the probability of connecting lesion A and lesion B. That is, using a machine learning process such as previously described above, a logistic regression model is trained on a volume of training images to generate a prediction about the probability that, in each pairwise combination of slices in each training volume, the lesion in one slice is the same or a different lesion as the lesion represented in the adjacent slice. This prediction is compared to a ground truth indication of whether the lesions are the same or different lesions to generate a loss or error. The operating parameters (e.g., coefficients or weights) of the logistic regression model are then modified to reduce this loss or error until a predetermined number of training epochs have been performed or a predetermined stopping condition is met.
[0201] Logistic regression models are widely used to solve binary classification problems. However, in the context of the illustrative embodiment, this logistic regression model predicts the probability that two cross-sections of a lesion are part of the same lesion. To do this, logistic regression uses the two overlap ratios r0 and r1 as mentioned previously. Specifically, the logistic model learns to linearly combine two features as follows:
[0202]
[0203] Where (C0, C1, b) are the operational parameters learned from the training volume via the machine learning training operation. The symbols r0 and r1 are used to represent the minimum overlap ratio and the maximum overlap ratio, respectively. The state of the operational parameters after the training of the logistic regression model can be expressed as At inference time (i.e., after training of the logistic regression model), when processing a new input volume of an image (slice), a threshold t is set such that if and only if the relation When the predicted probability is higher than the set threshold, the two sections are considered to belong to the same lesion.
[0204] There are two extreme cases. First, when the threshold t is set to 0, the z-connection mechanism of the illustrative embodiment always determines that the lesions are the same lesion (i.e., the sections are connected). Then both the true positive rate and the false positive rate are 1. Second, when the threshold t is set to 1, the z-connection mechanism will not identify any sections of the lesion to be connected. In this case, both the true positive rate and the false positive rate will be 0. Therefore, only when the threshold t is in the interval (0, 1) will the logistic regression model determine whether the lesion sections are associated with the same lesion across adjacent slices. Using an ideal logistic regression model, the true positive rate is equal to 1 (all true connections are identified) and at the same time the false positive rate is 0 (zero false connections are made).
[0205] Thus, once the logistic regression model is trained, new pairs of slices can be evaluated by calculating these ratios for these pairs and feeding them as input features into the trained logistic regression model so as to generate a prediction for each of these pairs, and then, if the predicted probability is equal to or greater than a predetermined threshold probability, then lesions A and B are considered to be associated with the same lesion in three dimensions. Appropriate relabeling of lesions across slices can then be performed so as to appropriately associate lesions in a two-dimensional slice with the same lesion representations in other adjacent slices, and thereby identify three-dimensional lesions within the input volume.
[0206] There is a rationale behind using two ratio input features for training a logistic regression model. For example, if lesions A and B are sufficiently different in size, then they cannot be part of the same lesion. Furthermore, if lesions A and B do not intersect (e.g., lesion 1312 in slice 1310 and lesion 1321 in slice 1320), then features r0 and r1 will have a value of zero. As described above, given two features r0 and r1, the logistic regression model performs regression and outputs a probability value between 0 and 1 representing the likelihood that lesions A and B are part of the same lesion.
[0207] Figure 13C Depicts cross-sectional connections between slices according to one illustrative embodiment. Figure 13C As shown, the mechanism determines that lesion 1311 in slice 1310 and lesion 1321 in slice 1321 are part of the same lesion by executing the trained logistic regression model of the illustrative embodiment, which predicts lesion commonality based on the overlap ratio as discussed above. The mechanism also determines in a similar manner that lesion 1321 in slice 1320 and lesion 1331 in slice 1330 are part of the same lesion. Thus, the mechanism propagates intersecting lesions along the z-axis and performs z-axis connection of the lesions.
[0208] Based on the pairwise evaluation of the slices in the input volume with respect to identifying z-connectivity of lesions across the two-dimensional slices, and the determination of whether the lesions are connected along the z-axis by a trained logistic regression model, lesion relabeling can be performed to ensure that the same lesion label is applied to each lesion mask present in each slice of the input volume (e.g., all lesion masks across a set of slices in the input volume), where lesion masks determined by the logistic regression model to be associated with the same lesion A, can be relabeled to specify that they are part of the same lesion A. This can be performed for each lesion cross-section in each slice of the input volume, thereby generating a three-dimensional association of lesion masks for one or more lesions present in the input volume. This information can then be used to represent or otherwise process the lesions in three dimensions (such as in later downstream computing system operations), since all cross-sections associated with the same lesion are correctly labeled in the input volume.
[0209] Figure 14A and Figure 14B Results of a trained logistic regression model are illustrated according to one illustrative embodiment. Figure 14A A receiver operating characteristic (ROC) curve illustrating the maximum overlap ratio (r0) + minimum overlap ratio (r1) metric and the maximum overlap ratio metric. An ROC curve is a graphical plot that illustrates the diagnostic power of a binary classifier system as its discrimination threshold is varied. The ROC curve is created by plotting the true positive rate (TPR) against the false positive rate (FPR) at various threshold settings. Figure 14B Precision-recall curves illustrating the maximum overlap ratio + minimum overlap ratio metric and the maximum overlap ratio metric. A precision-recall curve is a plot of precision (y-axis) and recall (x-axis) for different thresholds, much like an ROC curve, where precision is the fraction of relevant instances among the retrieved instances, and recall (or sensitivity) is the fraction of relevant instances actually retrieved. As shown in these plots, the two-feature logistic model outperforms its single-feature counterpart. The two features thus bring valuable information to this prediction task.
[0210] observe Figure 14A From the maximum overlap ratio (r0) + minimum overlap ratio (r1) metric curve in , we can see that with an appropriate threshold t, the trained logistic regression model can produce a true positive rate of approximately 95% at the cost of a false positive rate of approximately 3%. Figure 14B , the depicted plots evaluate the trained logistic regression model in terms of precision and recall and show that both measures can achieve very good results with an appropriate choice of the threshold t.
[0211] Figure 15is a flow chart outlining an example operation of a mechanism for connecting two-dimensional lesions along the z-axis according to one illustrative embodiment. Figure 15 As shown, the operation begins (step 1500), and the mechanism selects a first image X from an input volume (step 1501) and selects a first lesion A in image X (step 1502). In some illustrative embodiments, the previously described splitting and relabeling mechanism may be used to process images or slices in the input volume, however, this is not required. Rather, the mechanism of the illustrative embodiments involving z-linking of lesions can be performed with virtually any input volume in which a lesion mask has been identified.
[0212] The z-connectivity mechanism of the illustrative embodiment then selects the first lesion B in the adjacent image Y (step 1503). The mechanism then determines the intersection between lesion A and lesion B for lesion A, and the intersection between lesion A and lesion B for lesion B (step 1504). The mechanism determines whether lesion A and lesion B are the same lesion based on these two intersection values by applying a trained logistic regression model to the r0 and r1 features for the intersection of lesions A and B to generate a prediction or probability that lesion A and lesion B are the same lesion, and then comparing this probability to a threshold probability (step 1505). Based on the results of this determination, cross-sections of the lesions in the image can be labeled or relabeled to indicate whether they are part of the same lesion.
[0213] The mechanism determines whether lesion B in image Y is the last lesion in image Y (step 1506). If lesion B is not the last lesion, the mechanism considers the next lesion B in the adjacent image Y (step 1507), and the operation returns to step 1504 to determine the intersection between lesion A and the new lesion B.
[0214] If, at step 1506, lesion B is the last lesion in the adjacent slice or image Y, the mechanism determines whether lesion A is the last lesion in image X (step 1508). If lesion A is not the last lesion in image X, the mechanism considers the next lesion A in image X (step 1509), and the operation returns to step 1502 to consider the first lesion B in the adjacent image Y.
[0215] If lesion A is the last lesion in image X at step 1508, the mechanism determines whether image X is the last image to be considered (step 1510). If image X is not the last image, the mechanism considers the next image X (step 1511), and the operation returns to step 1502 to consider the first lesion A in the new image X.
[0216] If image X is the last image to be considered at step 1510, the mechanism propagates intersecting lesions between the images along the z-axis, where propagation means that the labels associated with the same lesions determined by the above process are set to the same value to indicate that they are part of the same lesion (step 1512). This is performed for each individual lesion identified in the input volume so that the sections in each image associated with the same lesion are appropriately labeled, and thus, a three-dimensional representation of each lesion is generated by connecting the sections in the z-direction. Thereafter, the operation ends (step 1513).
[0217] Contour refinement
[0218] The above process produces accurate results in terms of the number and relative positions of lesions, as well as in terms of connecting lesions across two-dimensional space (within an image or slice) and three-dimensional space (across images or slices in the input volume). However, lesion outlines (boundaries) are not always well-defined and need improvement. The illustrative embodiments provide a mechanism for improving the accuracy of lesion outlines. This additional mechanism can be employed in conjunction with the above mechanism as part of lesion segmentation, or can be employed in other illustrative embodiments that do not require the specific lesion detection, lesion splitting and relabeling, and / or z-connection mechanisms described above.
[0219] Existing contouring algorithms work well only when there is a lesion in the middle of an anatomical structure with no surrounding lesions, but perform poorly when there are different situations that lead to a "leakage" problem, in which two or more close lesions have initially different contours that merge into a single fully enclosing contour, thereby completely eliminating the benefits brought by the earlier 2D lesion mask splitting. In some cases, when the lesion is near an anatomical structure boundary (e.g., a liver boundary), the contouring algorithm distinguishes between pixels of the anatomical structure relative to pixels of other anatomical structures (e.g., organs) in the image, rather than distinguishing one lesion from another, because the contouring algorithm is best able to distinguish pixels of these anatomical structures.
[0220] The mechanisms of the illustrative embodiments inpaint regions of no interest in an image or slice. Figure 16 An example of contouring of two lesions in the same image is shown in accordance with an illustrative embodiment. Figure 16 On the left side of FIG, an active contour algorithm is used to determine the contours of two lesions 1611 and 1612. An active contour algorithm is a type of algorithm that iteratively evolves contours to better fit the image content.
[0221] According to this illustrative embodiment, the mechanism inpaints non-liver tissue within contour 1612 and in the vicinity of contour 1611 but not within contour 1611, where inpainting means that the pixel values of contour 1612 and pixels within contour 1612 and healthy tissue (non-lesioned tissue) near contour 1611 are set to a specified value so that they all have the same value. For example, the value may be the average tissue value in an area identified as not associated with a lesion (i.e., healthy tissue of the anatomical structure, such as the liver).
[0222] The repair can be performed relative to the selected lesion outline 1611 so that the repair is applied to healthy tissue and other lesions in the image (e.g., lesion 1612). In this way, when re-evaluating the outline 1611, the outlines and pixels associated with the selected lesion (e.g., 1611) are considered separately from the rest of the image. The outline 1611 can then be re-evaluated, and it can be determined whether the re-evaluation of the outline 1611 results in an improved definition of the outline 1611. That is, an initial determination of the contrast and variance between pixels associated with the selected lesion outline 1611 and pixels near the selected lesion outline 1611 can be generated. After calculating this contrast and variance before repairing, repair can be performed on the selected lesion 1611 so that pixels associated with other lesion outlines (e.g., 1612) and areas of anatomical structure in the image representing healthy tissue are repaired using the average pixel intensity value of healthy tissue.
[0223] The variance of a set of values is determined as follows. Consider a voxel set, comprising, for example, n voxels. First, calculate the arithmetic mean by summing their intensity values and then dividing the resulting sum by n. This resulting quantity is denoted by A. Second, individually square these voxel values and then calculate the arithmetic mean. The result is denoted by B. The variance is then defined as B – A*A, i.e., the difference between B and the square of A.
[0224] Thus, the variance of a set of n values {x1, ..., xn} is defined as follows:
[0225]
[0226] The variance between voxels inside and outside a given contour is calculated. Voxels inside a contour are those voxels that are enclosed by the contour, and voxels outside are those voxels that are outside the contour but remain within a predetermined distance from the contour.
[0227] This mechanism uses the active contour algorithm described above to recalculate the contour 1611 of the selected lesion after inpainting, and recalculate the contrast and / or variance of the new contour 1611 to determine whether these values have improved (higher contrast values or lower variance values inside and / or outside the lesion). If the contrast and variance have improved, the newly calculated contour 1611 is retained as the contour of the corresponding lesion. The process can then be performed on lesion 1612, where lesion 1612 is treated as the selected lesion by subsequently inpainting pixels associated with healthy tissue near lesion 1611 and contour 1612. In this way, each lesion is evaluated separately to generate the contour of the lesion, thereby preventing lesions from leaking into each other.
[0228] The mechanism used to calculate the outline of the lesion after repair can be based on the Chan-Vese segmentation algorithm, which is designed to segment objects without well-defined boundaries. The algorithm is based on iteratively evolving level sets to minimize energy, with the level set being defined by the sum of the difference intensities from the mean values outside the segmented region, the sum of the differences from the mean values inside the segmented region, and a weighted term depending on the length of the segmented region's boundary. Initialization is done using a partitioned detection map (solving a local energy minimum problem).
[0229] Once the mechanism has the segmentation, the mechanism initializes the contour using the previous estimate and determines whether the new contour is better (e.g., improves the contrast and variance of the contour). If the original contour is better, the original contour is kept. If the new contour is better (e.g., improves the contrast and variance of the contour), the mechanism uses the new contour. In some illustrative embodiments, the mechanism determines which contour is better based on homogeneous areas and calculated variance. If the variance is reduced both inside and outside the contour, the mechanism uses the new contour; otherwise, the mechanism uses the old contour. In another illustrative embodiment, the mechanism determines whether the contrast (the average value inside the contour to the average value near the contour) is improved. Without departing from the spirit and scope of the illustrative embodiments, other techniques using different measurements can be used to select between the old contour and the new contour.
[0230] Figure 17 is a flow chart outlining an example operation of a mechanism for slice-wise contour refinement according to an illustrative embodiment. Figure 17As shown in , for a given contour in an image segmented to show a lesion (such as in a liver), the operation begins (step 1700) and the mechanism determines a first contrast and variance of the initial contour (step 1701). The mechanism inpaints lesion pixels (or three-dimensional voxels) near the lesion (step 1702). The mechanism then determines a contour around the lesion (step 1703). The mechanism then determines a second contrast and variance of the new contour (step 1704). The mechanism determines whether the second contrast and variance represent an improvement over the first contrast and variance (step 1705). If the second contrast and variance represent an improvement, the mechanism uses the updated contour to represent the lesion (step 1706). Thereafter, the operation ends (step 1708).
[0231] If the second contrast and variance do not indicate an improvement in step 1705, the mechanism reverts to the initial contour (step 1707). Thereafter, the operation ends (step 1708). This process may be repeated for each lesion identified in the input slice and / or input volume in order to recalculate the contour and improve the contour associated with each lesion present in the image / input volume.
[0232] False positive removal
[0233] After performing lesion segmentation to generate a list of lesions and their contours, the AI pipeline 100 performs a false positive stage of processing 150 to remove erroneously indicated lesions from the lesion list. This false positive stage 150 can take many forms to reduce the number of erroneously identified lesions in the lesion list, for example, the liver / lesion detection logic 130 outputs Figure 1 , which are then merged by the segmentation and relabeling performed in the lesion segmentation logic 140. The following description will set forth a novel false positive removal mechanism that can be used to perform such false positive removal, but such specific false positive removal is not required. Furthermore, the false positive removal mechanism described below can be used separately from the other mechanisms described above and can be applied to any list of objects identified in an image, with the illustrative embodiments specifically utilizing such false positive removal of lesions in medical images. That is, the false positive removal mechanism described in this section can be implemented separately and differently from the other mechanisms described above.
[0234] For purposes of illustration, it will be assumed that the false positive removal mechanism is implemented as part of the AI pipeline 100 and as part of the false positive removal logic 150 of the AI pipeline 100. Thus, in the false positive stage 150, the false positive removal mechanism described in this section operates on the lesion list generated by the liver / lesion detection logic and the segmentation and re-labeling of the lesions, taking into account the three-dimensional nature of the input volume with z-connectivity and contour refinement of the lesions as described above. Figure 1This list 148 is input to a false positive removal logic stage 150, which processes the list 148 in a manner described below and outputs a filtered or modified list of lesions (in which lesions incorrectly identified in the modified list of lesions are minimized) to a lesion classification stage 160. The lesion classification stage then classifies the various lesions indicated in the modified list of lesions.
[0235] That is, capturing all lesions in the previous stages of the AI pipeline 100 can result in an increased sensitivity setting that causes the AI pipeline 100 to misidentify pixels that do not actually represent a lesion as part of a lesion. Consequently, there may be false positives that should be removed. The false positive stage 150 includes logic that operates on a list of lesions and their contours to remove false positives. It should be understood that such false positive removal must also be balanced against the risk that, during an exam (where the set of input volume levels is the opposite of the lesion levels), false positive removal (if performed improperly) may result in lesions not being detected. This can be problematic because the physician and patient may not be aware of lesions that require treatment. It should be understood that an exam can theoretically include several image volumes from the same patient. However, because in some illustrative embodiments of the AI pipeline implementing a single phase detection, only one volume of images is processed, it is assumed that the processing is performed on a single volume. For clarity, "patient level" will be used hereafter instead of "exam level" as this is the focus of the illustrative embodiments (whether the patient has a lesion). It should be understood that in other illustrative embodiments, the operations described herein can be extended to the exam level, where multiple image volumes from the same patient can be evaluated.
[0236] For these illustrative embodiments, given the outputs of the previous stages of the AI pipeline 100 (slices, masks, lesions, lesion and anatomical structure outlines, etc.) as inputs 148 to the false positive removal stage 150, the false positive removal stage 150 operates at a high specificity operating point at the patient level (input volume level) so as to allow only a few patient-level false positives (normal patients / volumes with at least one lesion detected). This point can be retrieved from analysis of the patient receiver operating characteristic (ROC) analysis (patient-level sensitivity versus patient-level specificity). For those cases where a high specificity operating point (herein referred to as the patient-level operating point OP) is used, the false positive removal stage 150 is operated at a high specificity operating point at the patient level (input volume level) so as to allow only a few patient-level false positives (normal patients / volumes with at least one lesion detected). This point can be retrieved from analysis of the patient receiver operating characteristic (ROC) analysis (patient-level sensitivity versus patient-level specificity). PATIENT ) produces those volumes of at least some lesions at the lesion level (referred to herein as the lesion level operating point OP LESION ) is used at a more sensitive operating point. The lesion-level operating point OP can be identified from analysis of the lesion-level ROC curve (lesion sensitivity versus lesion specificity) LESION , in order to maximize the number of retained lesions.
[0237] These two operating points (OP PATIENTand OP LESION ) can be implemented in one or more trained ML / DL computer models. One or more trained ML / DL computer models are trained to classify the input volume and / or its lesion list (the result of the segmentation logic) as whether the identified lesion is a true lesion or a false lesion (i.e., a true positive or a false positive). The one or more trained ML / DL computer models can be implemented as a binary classifier, wherein the output indicates for each lesion whether it is a true positive or a false positive. The output set of the binary classification for all lesions in the input lesion list can be used to filter the lesion list to remove false positives. In an illustrative embodiment, the one or more trained ML / DL computer models first implement a patient-level operating point to determine whether the result of the classification indicates that any lesion in the lesion list is a true positive while filtering out false positives. If any true positives remain in the first filtered list of lesions after patient-level (input volume-level) filtering, the lesion-level operating point is used to filter out the remaining false positives (if any). Thus, a filtered lesion list is generated that minimizes false positives.
[0238] The implementation of the operating point can be with respect to a single trained ML / DL computer model or multiple trained ML / DL computer models. For example, using a single trained ML / DL computer model, the operating point can be a setting of an operating parameter of the ML / DL computer model that can be dynamically switched. For example, the input to the ML / DL computer model can be processed using the patient-level operating point to generate a result indicating whether the list of lesions includes true positives after classifying each lesion. If so, the operating point of the ML / DL computer model can be switched to the lesion-level operating point and the input can be processed again, wherein each false positive passed through the ML / DL computer model is removed from the final list of lesions output by the false positive removal stage. Alternatively, in some illustrative embodiments, two separate ML / DL computer models can be trained, one for the patient-level operating point and one for the lesion-level operating point, such that the result of the first ML / DL computer model indicating at least one true positive causes the input to be processed by the second ML / DL computer model and the false positives identified by both models are removed from the final list of lesions output by the false positive removal stage of the AI pipeline.
[0239] The training of (multiple) ML / DL computer models may involve a machine learning training operation in which the ML / DL computer model processes a training input comprising an image volume and a corresponding list of lesions, wherein the list of lesions comprises a lesion mask or outline to generate, for each lesion in the image, a classification as to whether it is a true positive or a false positive. The training input is further associated with ground truth information indicating whether the image comprises a lesion, which ground truth information may then be used to evaluate an output generated by the ML / DL computer model to determine a loss or error, and then modify operating parameters of the ML / DL computer model to reduce the determined loss / error. In this way, the ML / DL computer model learns input features that are representative of true / false positive lesion detections. This machine learning may be performed with respect to each of the operating points (i.e., OP PATIENT and OP LESION ), such that the operating parameters of the ML / DL computer model are learned taking into account patient-level sensitivity / specificity and / or lesion-level sensitivity / specificity.
[0240] When classifying lesions as to whether they are true positive or false positive, an input volume (representing a patient at the "patient level") is considered positive if it contains at least one lesion. If the input volume does not contain a lesion, the input volume is considered negative. With this in mind, a true positive is defined as a positive input volume (i.e., an input volume with at least one finding that is classified as a lesion that is actually a lesion). A true negative is defined as a negative input volume (i.e., an input volume without a lesion and without a finding classified as a lesion). A false positive is defined as a negative input volume without a lesion, however, the input indicates a lesion in the finding (i.e., the AI pipeline lists lesions when no lesion is present). A false negative is defined as a positive input volume with a lesion, but the AI pipeline does not indicate a lesion in the finding. The trained ML / DL computer model classifies the lesions in the input as to whether they are true positive or false positive. False positives are filtered out from the output generated by false positive removal. Detection of false positives is performed at different sensitivity / specificity levels at the patient level and the lesion level (i.e., two different operating points).
[0241] Two different operating points at the patient level and at the lesion level can be determined based on ROC curve analysis. The ROC curve can be calculated using ML / DL computer model validation data, which consists of several input volumes (e.g., several input volumes corresponding to different patient examinations) that may contain some lesions (between 0 and K lesions per examination). The input to the trained ML / DL computer model, or "classifier", is the findings previously detected in the input that are either actual lesions or false positives (e.g., the output of the lesion detection and segmentation stage of the AI pipeline). The first operating point (i.e., the patient level operating point OP PATIENT ) is defined as retaining at least X% of lesions identified as true positives, meaning that almost all true positives are retained while some false positives are removed. The value of X can be set based on analysis of the ROC curve and can be any suitable value for a particular implementation. In one illustrative embodiment, the value of X is set to 98%, so that almost all true positives are retained while some false positives are removed.
[0242] Define the second operating point (i.e., lesion level operating point OP LESION ), making the lesion sensitivity higher than that for the first operating point (ie, the patient level operating point OP PATIENT ) and achieve a specificity higher than Y%, where Y depends on the actual performance of the trained ML / DL computer model. In one illustrative embodiment, Y is set to 30%. Examples of ROC curves for patient-level and lesion-level operating point determination are given in Figure 18A As shown in Figure 18A As shown, the lesion-level operating point is selected along the lesion-level ROC curve so that the lesion sensitivity is higher than the lesion sensitivity of the patient-level operating point.
[0243] Figure 18B is an example flow chart of operations for performing false positive removal based on patient and lesion level operating points, according to one illustrative embodiment. Figure 18BAs shown, the results of the segmentation level logic of the AI pipeline are input 1810 to a first trained ML / DL computer model 1820 that implements a first operating point. The input 1810 includes an input volume (or image volume (VOI)) and a lesion list, the lesion list including lesion masks or contour data specifying pixels or voxels corresponding to each lesion identified in the image data of the image volume and labels associated with these pixels specifying that they correspond to lesions in the three-dimensional space of the input volume (i.e., the output of the segmentation, z-connection, and contour refinement previously described). The input can be represented as a set S. The first trained ML / DL computer model 1820 implements a patient-level operating point in its training so that features extracted from the input are classified with X% (e.g., 98%) of true positives retained in the resulting filtered list of lesions generated by the classification of the trained ML / DL computer model 1820, and some false positives are removed in the resulting list. The resulting list includes a subset S containing true positive lesions classified by the first ML / DL computer model 1820 + and a subset S − containing false positive lesions classified by the first ML / DL computer model 1820 .
[0244] The false positive removal logic further includes true positive evaluation logic 1830, which determines whether the subset of true positives output by the first ML / DL computer model 1820 is empty. That is, the true positive evaluation logic 1830 determines whether no elements from S are classified as true lesions by the first ML / DL computer model 1820. If the subset of true positives is empty, the true positive evaluation logic 1830 makes the true positive subset S + is output as a filtered lesion list 1835 (ie, no lesions will be identified in the output sent to the lesion classification stage of the AI pipeline). + If it is not empty, then a second ML / DL computer model 1840 is executed on the input S, wherein the second ML / DL computer model 1840 realizes the second operating point (i.e., the lesion level operating point OP) in its training. LESION ). It should be understood that although two ML / DL computer models 1820 and 1840 are shown for ease of explanation, as described above, these two operating points can be implemented in different sets of training operating parameters for configuring the same ML / DL computer model, such that the second ML / DL computer model can be a process using the same ML / DL computer model as 1820 but with different operating parameters corresponding to the second operating point.
[0245] The second ML / DL computer model 1840 processes the inputs with the trained operating parameters corresponding to the second operating point to again generate a classification of lesions as to whether they are true positives or false positives. The result is a subset S' containing the predicted lesions (true positives)+ and the subset S' containing the predicted false positives - The filtered lesion list 1845 is then output as a subset S' + , thereby effectively eliminating - False positives specified in .
[0246] Figure 18A and Figure 18B The example embodiments shown in FIG are described in terms of patient-level and lesion-level operating points. It should be appreciated that the mechanism for false positive removal can be implemented using operating points at various levels. For example, similar operations can be performed for image volume-level and voxel-level operating points in a "voxel-wise" false positive removal operation. Figure 18C is an example flow diagram of operations for performing voxel-wise false positive removal based on input volume-level and voxel-level operating points in accordance with one illustrative embodiment. Figure 18C The operation in is similar to Figure 18B The operation is similar to the operation of , but is performed with respect to voxels in the input set S. With voxel-wise false positive removal, the first operating point can again be the patient-level or input volume-level operating point, while the second operating point can be at the voxel-level operating point OP VOXEL In this case, true positives and false positives are evaluated at the voxel level, such that if any voxel is indicated to be associated with a lesion and it is actually associated with the lesion, it is a true positive, but if a voxel is indicated to be associated with a lesion but it is not actually associated with the lesion, it is considered a false positive. An appropriate setting of the operating point can be generated again based on the corresponding ROC curve so that a similar balance between sensitivity and specificity is achieved as described above.
[0247] It should also be understood that while the above illustrative embodiments of the false positive removal mechanism assume a single input volume from a patient exam, the illustrative embodiments can be applied to any grouping of one or more images (slices). For example, false positive removal can be applied to a single slice, a group of slices smaller than the input volume, or even multiple input volumes from the same exam.
[0248] Figure 19 is a flow chart outlining an example operation of the false positive removal logic of an AI pipeline according to one illustrative embodiment. Figure 19As shown, the operation begins (step 1900) by receiving an input S from a previous stage of the AI pipeline, where the input may include, for example, an input image volume and a corresponding list of lesions including masks, contours, etc. (step 1910). The input is processed by a first trained ML / DL computer model that is trained to achieve a first operating point (e.g., a relatively more specific and less sensitive patient-level operating point) to generate a first set of classifications for lesions including a true positive subset and a false positive subset (step 1920). A determination is made as to whether the true positive subset is empty (step 1930). If the true positive subset is empty, the operation outputs the true positive subset as a filtered list of lesions (step 1940) and the operation terminates. If the true positive subset is not empty, the input S is processed by a second ML / DL computer model that is trained to achieve a second operating point that is relatively more sensitive and less specific than the first operating point (step 1950). As described above, in some demonstrative embodiments, the first and second ML / DL computer models can be the same model, but configured with different operating parameters corresponding to different training to achieve different operating points. The result of processing the second ML / DL computer model is a second set of classifications of lesions including a second subset of true positives and a second subset of false positives. The second subset of true positives is then output as a filtered list of lesions (step 1960) and the operation terminates.
[0249] Example computer system environment
[0250] The illustrative embodiments can be utilized in many different types of data processing environments. To provide a context for describing the specified elements and functions of the illustrative embodiments, the following is provided. Figure 20 and Figure 21 As an example environment in which aspects of the illustrative embodiments may be implemented. It should be understood that Figure 20 and Figure 21 This is merely an example and is not intended to assert or imply any limitation with respect to the environments in which various aspects or embodiments of the invention may be implemented. Many modifications to the depicted environments may be made without departing from the spirit and scope of the invention.
[0251] Figure 20A schematic diagram depicts an illustrative embodiment of a cognitive system 2000 implementing a request processing pipeline 2008, which in some embodiments can be a question answering (QA) pipeline, a treatment recommendation pipeline, a medical imaging enhancement pipeline, or any other artificial intelligence (AI) or cognitive computing-based pipeline that processes requests using complex AI mechanisms that approximate humans through processes related to the generated results, but through different computer-specified processes. For the purposes of this description, it will be assumed that the request processing pipeline 2008 is implemented as a QA pipeline that operates on structured and / or unstructured requests in the form of input questions. An example of a question processing operation that can be used in conjunction with the principles described herein is described in U.S. Patent Application Publication No. 2011 / 0125734, the entire contents of which are incorporated herein by reference.
[0252] The cognitive system 2000 is implemented on one or more computing devices 2004A-D (including one or more processors and one or more memories, and potentially any other computing device elements known in the art, including buses, storage devices, communication interfaces, etc.) connected to a computer network 2002. For illustrative purposes only, Figure 20 Cognitive system 2000 is depicted as being implemented only on computing device 2004A, but as described above, cognitive system 2000 can be distributed across multiple computing devices, such as multiple computing devices 2004A-D. Network 2002 includes multiple computing devices 2004A-D, which can operate as server computing devices, and 2010-2012, which can operate as client computing devices, communicating with each other and other devices or components via one or more wired and / or wireless data communication links, wherein each communication link includes one or more of wires, routers, switches, transmitters, receivers, etc. In some illustrative embodiments, cognitive system 2000 and network 2002 implement their question processing and answer generation (QA) functionality via respective computing devices 2010-2012 of one or more cognitive system users. In other embodiments, cognitive system 2000 and network 2002 may provide other types of cognitive operations, including but not limited to request processing and cognitive response generation, which may take many different forms depending on the desired implementation (e.g., cognitive information retrieval, training / instruction of a user, cognitive evaluation of data, etc.). Other embodiments of cognitive system 2000 may be used with components, systems, subsystems, and / or devices other than those depicted herein.
[0253] The cognitive system 2000 is configured to implement a request processing pipeline 2008 that receives input from various sources. Requests may be posed in the form of natural language questions, natural language requests for information, natural language requests for performance of cognitive operations, and the like. For example, the cognitive system 2000 receives input from the network 2002, a corpus or corpora of electronic documents 2006, cognitive system users and / or other data, and other possible sources of input. In one embodiment, some or all of the input to the cognitive system 2000 is routed through the network 2002. Various computing devices 2004A-D on the network 2002 include access points for content creators and cognitive system users. Some of the computing devices 2004A-D include a corpus or corpora for storing data 2006 (which are shown for illustrative purposes only). Figure 20 The corpus of data 2006 or portions of the corpora of data 2006 may also be provided on one or more other network attached storage devices, in one or more databases, or on a computer. Figure 20 In various embodiments, network 2002 includes local network connections and remote connections, so that cognitive system 2000 can operate in environments of any size, including local and global (e.g., the Internet).
[0254] In one embodiment, a content creator creates content in a document from a corpus or corpora of data 2006 to be used as part of the corpus of data for cognitive system 2000. A document includes any file, text, article, or data source for use in cognitive system 2000. A cognitive system user accesses cognitive system 2000 via a network connection or an internet connection to network 2002 and inputs questions / requests to be answered / processed based on content from the corpus or corpora of data 2006 into cognitive system 2000. In one embodiment, the questions / requests are formulated using natural language. Cognitive system 2000 parses and interprets the questions / requests via pipeline 2008 and provides a response to the cognitive system user (e.g., cognitive system user 2010), including one or more answers to the posed questions, responses to the requests, results of processing the requests, etc. In some embodiments, cognitive system 2000 provides responses to the user in a ranked list of candidate answers / responses, while in other illustrative embodiments, cognitive system 2000 provides a single final answer / response or a combination of the final answer / response and a ranked list of other candidate answers / responses.
[0255] Cognitive system 2000 implements pipeline 2008, which includes multiple stages for processing an input question / request based on information obtained from the corpus or corpora of data 2006. Pipeline 2008 generates an answer / response to the input question or request based on the processing of the input question / request and the corpus or corpora of data 2006.
[0256] In some demonstrative embodiments, cognitive system 2000 may be IBM Watson, available from International Business Machines, Inc. of Armonk, New York. TM Cognitive systems, augmented with the mechanisms of the illustrative embodiments described below. As previously outlined, IBM Watson TM The cognitive system pipeline receives an input question or request, and IBM Watson TM The pipeline of the cognitive system parses the input question or request to extract key features of the question / request, which are in turn used to formulate a query that is applied to the corpus or corpora of data 2006. Based on applying the query to the corpus or corpora of data 2006, a set of hypotheses or candidate answers / responses to the input question / request are generated by looking across the corpus or corpora of data 2006 for portions of the corpus or corpora of data 2006 (hereinafter referred to as corpus 2006) that have some potential to contain valuable responses to the input question / response (hereinafter assumed to be the input question). IBM Watson TM The cognitive system's pipeline 2008 then uses various reasoning algorithms to perform a deep analysis of the language of the input question and the language used in each of the sections of the corpus 2006 discovered during application of the query.
[0257] The scores obtained from the different inference algorithms are then weighted against a statistical model that summarizes IBM Watson TM The pipeline 2008 of the cognitive system 2000 (in this example) generates confidence levels about the evidence that potential candidate answers are inferred from the question. This process is repeated for each candidate answer to generate a ranked list of candidate answers, which can then be presented to the user who submitted the input question (e.g., the user of the client computing device 2010), or a final answer can be selected from the ranked list and presented to the user. About IBM Watson TM More information about the pipeline 2008 of the cognitive system 2000 can be obtained, for example, from the IBM corporate website, IBM Redbooks, etc. For example, regarding IBM Watson TMInformation on the pipeline of cognitive systems can be found in “Watson and Healthcare” by Yuan et al., IBM developerWorks, 2011, and “The Era of Cognitive Systems: An Inside Look at IBM Watson and How it Works” by Rob High, IBM Redbooks, 2012.
[0258] As mentioned above, while the input from the client device to the cognitive system 2000 can be posed in the form of natural language questions, the illustrative embodiments are not limited thereto. Rather, the input questions can actually be formatted or structured to be able to be analyzed using structured and / or unstructured input analysis (including but not limited to IBM Watson TM The cognitive system 2000 can be configured to process any suitable type of request that is parsed and analyzed using the natural language parsing and analysis mechanisms of the cognitive system (e.g., the natural language parsing and analysis mechanisms of the cognitive system, such as the illustrative embodiment) to determine the basis for performing cognitive analysis and provide the results of the cognitive analysis. For example, a physician, patient, etc. can issue a request for a specific medical imaging-based operation (e.g., "identify liver lesions present in patient ABC" or "provide treatment recommendations for patient" or "identify changes in liver lesions in patient ABC") to the cognitive system 2000 via their client computing device 2010. According to the illustrative embodiment, such requests can be specifically directed to cognitive computer operations that employ the lesion detection and classification mechanisms of the illustrative embodiment to provide a list of lesions, outlines of lesions, classifications of lesions, and outlines of anatomical structures of interest, and the cognitive system 2000 operates accordingly to provide cognitive computing output. For example, the request processing pipeline 2008 can process a request such as "identify liver lesions present in patient ABC" to parse the request and thereby identify the anatomical structure of interest as "liver," the specific input volume being the medical imaging volume of patient "ABC," and the "lesion" in the anatomical structure to be identified. Based on this parsing, a specific medical imaging volume corresponding to patient "ABC" can be retrieved from the corpus 2006 and input into the lesion detection and classification AI pipeline 2020, which operates on this input volume as previously described to identify a list of liver lesions, which is output to the cognitive computing system 2000 for further evaluation by the request processing pipeline 2008, for generating medical imaging viewer application output, etc.
[0259] like Figure 20As shown, one or more of these computing devices (e.g., server 2004) can be specifically configured to implement the lesion detection and classification AI pipeline 2020 (e.g., as Figure 1 ). Configuration of the computing device may include providing dedicated hardware, firmware, etc. to facilitate the execution of the operations and generation of outputs described herein with respect to the illustrative embodiments. Configuration of the computing device may also or alternatively include providing a software application stored in one or more storage devices and loaded into a memory of the computing device (such as server 2004) for causing one or more hardware processors of the computing device to execute the software application, the software application configuring the processor to perform the operations and generate outputs described herein with respect to the illustrative embodiments. Furthermore, any combination of dedicated hardware, firmware, software applications executed on hardware, etc. may be used without departing from the spirit and scope of the illustrative embodiments.
[0260] It should be understood that once a computing device is configured in one of these ways, the computing device becomes a special-purpose computing device that is specifically configured to implement the mechanisms of the illustrative embodiments and is not a general-purpose computing device. Furthermore, as described herein, implementation of the mechanisms of the illustrative embodiments improves the functionality of the computing device and provides useful and specific results that facilitate automatic lesion detection in anatomical structures of interest and classification of such lesions, which reduces errors and improves efficiency relative to manual processes.
[0261] As described above, the mechanisms of the illustrative embodiments utilize specially configured computing devices or data processing systems to perform operations for performing anatomical structure identification, lesion detection, and classification. These computing devices or data processing systems may include various hardware elements that are specifically configured through hardware configuration, software configuration, or a combination of hardware and software configuration to implement one or more of the systems / subsystems described herein. Figure 21 is a block diagram of only one example data processing system in which various aspects of the illustrative embodiments may be implemented. Data processing system 2100 is a computer (such as Figure 20 An example of a server 2004 in which computer usable code or instructions implementing the various processes and aspects of the illustrative embodiments of the present invention can be located and / or executed to achieve the operations, outputs and external effects of the illustrative embodiments described herein.
[0262] In the depicted example, data processing system 2100 employs a hub architecture including north bridge and memory controller hub (NB / MCH) 2102 and south bridge and input / output (I / O) controller hub (SB / ICH) 2104. Processing unit 2106, main memory 2108, and graphics processor 2110 are connected to NB / MCH 2102. Graphics processor 2110 may be connected to NB / MCH 2102 via an accelerated graphics port (AGP).
[0263] In the depicted example, a local area network (LAN) adapter 2112 is connected to the SB / ICH 2104. An audio adapter 2116, a keyboard and mouse adapter 2120, a modem 2122, a read-only memory (ROM) 2124, a hard disk drive (HDD) 2126, a CD-ROM drive 2130, a universal serial bus (USB) port and other communication ports 2132, and PCI / PCIe devices 2134 are connected to the SB / ICH 2104 via bus 2138 and bus 2140. PCI / PCIe devices may include, for example, Ethernet adapters, add-in cards, and PC Cards for notebook computers. PCI uses a card bus controller, while PCIe does not. ROM 2124 may be, for example, a flash memory basic input / output system (BIOS).
[0264] HDD 2126 and CD-ROM drive 2130 are connected to SB / ICH 2104 via bus 2140. HDD 2126 and CD-ROM drive 2130 may use, for example, an integrated drive electronics (IDE) or serial advanced technology attachment (SATA) interface. Super I / O (SIO) device 2136 may be connected to SB / ICH 2104.
[0265] The operating system runs on the processing unit 2106. The operating system coordinates and provides Figure 21 The control of each component in the data processing system 2100 in the client. As a client, the operating system can be such as Commercially available operating systems. Object-oriented programming systems (such as Java TM Programming system) can be run in conjunction with the operating system and provide Java programs executed on the data processing system 200 TM A call made by a program or application to the operating system.
[0266] As a server, data processing system 2100 may be, for example, a server running an advanced interactive execution Operating system or IBM eServer operating system TM System Computer systems, based on PowerTM Data processing system 2100 may be a symmetric multi-processor (SMP) system including a plurality of processors in processing unit 2106. Alternatively, a single processor system may be employed.
[0267] Instructions for the operating system, object-oriented programming system, and applications or programs are located on storage devices, such as HDD 2126, and may be loaded into main memory 2108 for execution by processing unit 2106. Processes for the illustrative embodiments of the present invention may be performed by processing unit 2106 using computer-usable program code, which may be located in memory, such as main memory 2108, ROM 2124, or in one or more peripheral devices 2126 and 2130, for example.
[0268] Bus systems (such as Figure 21 The bus 2138 or bus 2140 shown in FIG may include one or more buses. Of course, the bus system may be implemented using any type of communication structure or architecture that provides for data transfer between different components or devices attached to the structure or architecture. Communication units (such as Figure 21 The modem 2122 or network adapter 2112 may include one or more devices for transmitting and receiving data. The memory may be, for example, the main memory 2108, the ROM 2124 or a memory such as a memory card. Figure 21 The cache found in the NB / MCH2102.
[0269] As described above, in some illustrative embodiments, the mechanisms of the illustrative embodiments may be implemented as application-specific hardware, firmware, etc., application software stored in a storage device (such as HDD 2126) and loaded into a memory (such as main memory 2108) for execution by one or more hardware processors (such as processing unit 2106), etc. In this way, Figure 21 The computing device shown in becomes specially configured for implementing the mechanisms of the illustrative embodiments and is specially configured to perform the operations and generate the outputs described herein with respect to the lesion detection and classification artificial intelligence pipeline.
[0270] Those skilled in the art will appreciate that Figure 20 and Figure 21 The hardware in may vary depending on the implementation. Figure 20 and Figure 21 In addition to or in place of the hardware depicted Figure 20 and Figure 21In addition to the hardware depicted in the drawings, other internal hardware or peripheral devices (such as flash memory, equivalent non-volatile memory, or optical disk drives, etc.) may be used. In addition, the processes of the illustrative embodiments may be applied to multi-processor data processing systems other than the previously mentioned SMP systems without departing from the spirit and scope of the present invention.
[0271] Furthermore, data processing system 2100 may take the form of any of a number of different data processing systems, including a client computing device, a server computing device, a tablet computer, a laptop computer, a telephone or other communication device, a personal digital assistant (PDA), and the like. In some illustrative examples, data processing system 2100 may be a portable computing device configured with flash memory to provide non-volatile memory for storing, for example, operating system files and / or user-generated data. Essentially, data processing system 2100 may be any known or later developed data processing system without architectural limitation.
[0272] As described above, it should be understood that the illustrative embodiments can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment comprising both hardware and software elements. In one exemplary embodiment, the mechanisms of the illustrative embodiments are implemented in software or program code, including but not limited to firmware, resident software, microcode, etc.
[0273] The data processing system that is suitable for storing and / or executing program code will include at least one processor, which is directly or indirectly coupled to the memory element by a communication bus (such as, for example, a system bus). The memory element can include a local memory, a mass storage device, and a temporary storage of at least some program codes that is adopted during the actual execution of the program code in order to reduce the number of times that code must be retrieved from the mass storage device during execution. The memory can be of various types, including but not limited to ROM, PROM, EPROM, EEPROM, DRAM, SRAM, flash memory, solid-state memory, etc.
[0274] Input / output or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) can be coupled to the system directly or through intervening wired or wireless I / O interfaces and / or controllers, etc. I / O devices can take many different forms besides traditional keyboards, displays, pointing devices, etc., such as, for example, communication devices coupled via wired or wireless connections, including but not limited to smartphones, tablet computers, touch screen devices, voice recognition devices, etc. Any known or later developed I / O device is intended to be within the scope of the illustrative embodiments.
[0275] A network adapter may also be coupled to a system so that the data processing system can become coupled to other data processing systems or remote printers or storage devices through an intervening private or public network. Modems, cable modems, and Ethernet cards are just some of the currently available types of network adapters for wired communications. Network adapters based on wireless communications may also be utilized, including but not limited to 802.11a / b / g / n wireless communication adapters, Bluetooth wireless adapters, and the like. Any known or later developed network adapter is intended to be within the spirit and scope of the present invention.
[0276] The description of the present invention is presented for the purpose of illustration and description and is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The embodiments are selected and described in order to best explain the principles of the invention, practical applications, and to enable those of ordinary skill in the art to understand the invention for various embodiments with various modifications suitable for the specific use under consideration. The terms used herein are selected to best explain the principles of the embodiments, practical applications, or technical improvements over technologies found in the marketplace, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method in a data processing system comprising at least one processor and at least one memory, the at least one memory comprising instructions executed by the at least one processor to implement a lesion detection and classification AI pipeline comprising a plurality of trained machine learning computer models, wherein the AI pipeline performs a method comprising: processing an input volume of a medical image by one or more first machine learning computer models of the AI pipeline to determine whether the input volume depicts a predetermined amount of an anatomical structure of interest; determining, by the logic of the AI pipeline, whether the output of the one or more first machine learning computer models satisfies one or more predetermined criteria, wherein the one or more predetermined criteria include a predetermined amount of anatomical structure of interest depicted in the input volume, and wherein, In response to determining that the outputs of the one or more first machine learning computer models satisfy the one or more predetermined criteria, performing a lesion treatment operation, the lesion treatment operation comprising: processing the input volume by one or more second machine learning computer models of the AI pipeline to detect lesions corresponding to the anatomical structure of interest present in the medical image of the input volume; processing the detected lesions by one or more third machine learning computer models of the AI pipeline to perform lesion segmentation, and combining lesion contours of different medical images associated with the same lesion in the input medical image to generate a lesion list and lesion contours; processing the list of lesions and contours associated with the lesions in the list of lesions by one or more fourth machine learning computer models of the AI pipeline to classify the lesions in the list of lesions into one or more predetermined lesion classes corresponding to the anatomical structure of interest; and The lesion list and the classification associated with the lesions are output by the AI pipeline for processing by a downstream computing system.
2. The method according to claim 1, wherein The one or more first machine learning computer models of the AI pipeline include a body part detection machine learning model that performs processing on the input volume to detect a portion of a biological entity's body represented in the input volume and determines whether the portion of the biological entity's body corresponds to a body part in which the anatomical structure of interest is present.
3. The method according to claim 2, wherein: The body part is the human abdomen, and the anatomical structure of interest is the human liver.
4. The method according to claim 1, wherein The one or more first machine learning computer models of the AI pipeline include a phase classification machine learning model and an anatomical structure detection machine learning model that operate on the input volume, wherein the phase classification machine learning model determines the phase of medical imaging represented in each of the medical images of the input volume, and wherein the anatomical structure detection machine learning model determines the amount of the anatomical structure of interest represented in the input volume.
5. The method according to claim 4, wherein: The one or more predetermined criteria further include an input volume having a single phase medical imaging present in the input volume, and wherein a minimum threshold amount of anatomical structure of interest present in the input volume is determined based on the results of the operation of the phase classification machine learning model on the input volume and based on the results of the operation of the anatomical structure detection machine learning model on the input volume.
6. The method of claim 4, wherein: In response to a result of the operation of the phase classification machine learning model indicating that there is more than one phase of medical imaging in the input volume, a subset of medical images in the input volume corresponding to the phases of interest is selected, and the lesion processing operation is performed only on the selected subset of medical images in the input volume.
7. The method according to claim 6, wherein: The phase of interest is one of a pre-contrast phase, an arterial contrast phase, a portal contrast phase, or a delayed phase.
8. The method of claim 1, wherein: Determining whether the output of the one or more first machine learning computer models satisfies one or more predetermined criteria further comprises: An axial scoring operation is performed on the medical images in the input volume by the axial scoring logic of the AI pipeline, wherein the axial scoring operation includes scoring the images in the input volume according to a predefined scoring algorithm to infer axial scores of the lowest slice MISV in the volume and the highest slice MSSV in the volume based on scores associated with the images, and identifying scores of anatomical structures of interest present in the input volume based on the axial scores of MISV and MSSV.
9. The method of claim 1, wherein: The one or more second machine learning computer models of the AI pipeline comprise an ensemble of machine learning computer models, wherein each machine learning computer model in the ensemble of machine learning computer models is trained by a machine learning process to process an input volume differently than other machine learning computer models in the ensemble of machine learning computer models to generate corresponding lesion detection predictions.
10. The method of claim 9, wherein: The integration of the machine learning computer model includes a mask generation machine learning model, an input volume processing machine learning model, and a masked input volume processing machine learning model, wherein the mask generation machine learning model is trained by a machine learning process to generate a mask corresponding to the anatomical structure of interest, and the mask is applied to the input volume to generate a masked input volume, wherein the input volume processing machine learning model is trained by a machine learning process to generate a first lesion prediction output based on the input volume, and the masked input volume processing machine learning model is trained by a machine learning process to generate a second lesion prediction output based on the masked input volume.
11. The method of claim 9, wherein: training a first machine learning model of the ensemble using a first loss function that penalizes errors in false negative classification of lesions, and A second machine learning model in the ensemble, different from the first machine learning model, is trained using a second loss function that penalizes errors in false-positive classification of lesions.
12. The method of claim 11, wherein: The integrated logic applies a third loss function that compares a first lesion detection output of the first machine learning model with a second lesion detection output of the second machine learning model and operates to reconcile the first lesion detection output with the second lesion detection output.
13. The method of claim 1, wherein: The one or more third machine learning computer models of the AI pipeline include: a lesion segmentation machine learning model, a z-connectivity machine learning model, and a contour refinement machine learning model, wherein the lesion segmentation machine learning model segments the image of the input volume into contours corresponding to lesions to generate lesion segmentations, the z-connectivity machine learning model combines a subset of lesion segmentations based on the lesion segmentations generated by the lesion segmentation machine learning model, and the contour refinement machine learning model separates lesions in the input volume and refines the contours of lesions.
14. The method of claim 1, further comprising one or more false positive removal machine learning computer models operating on the list of lesions to remove false positive detections of lesions from the list of lesions.
15. The method of claim 14, wherein: The one or more false positive removal machine learning computer models include one or more false positive removal machine learning computer models trained by a machine learning process based on two different operating points, wherein a first operating point of the two different operating points corresponds to a patient-level operating point and a second operating point of the two different operating points corresponds to a lesion-level operating point.
16. The method of claim 1, wherein: The one or more fourth machine learning computer models classify each lesion in the list of lesions into one of a plurality of predetermined categories, wherein the one or more predetermined categories include lesion type or at least one of a benign category, a malignant category, or an indeterminate category.
17. The method of claim 1, further comprising: The list of lesions and the output of the classifications associated with the lesions are processed by a downstream computing system to generate a visual representation of the lesions in the list of lesions in a graphical user interface.
18. The method of claim 1, wherein: The input volume includes an image corresponding to at least one of a magnetic resonance imaging modality or a computed tomography modality.
19. A computer program product comprising a computer-readable storage medium having a computer-readable program stored therein, wherein The computer readable program, when executed on a computing device, causes the computing device to perform the steps of the method according to any one of claims 1 to 18.
20. A lesion detection device, comprising: processor; as well as A memory coupled to the processor, wherein the memory comprises instructions which, when executed by the processor, cause the processor to perform the steps of the method according to any one of claims 1 to 18.
21. A computer system comprising means for performing the steps of the method according to any one of claims 1 to 18.
22. A method in a data processing system comprising at least one processor and at least one memory, the at least one memory comprising instructions for execution by the at least one processor to implement a lesion detection integrated machine learning model architecture comprising a plurality of trained machine learning computer models, wherein: The lesion detection integrated machine learning model architecture performs the method including the following: processing a medical image input of at least one medical image by a first decoder of a lesion detection machine learning computer model to generate a first lesion map prediction output identifying graphical elements corresponding to lesions in the at least one medical image; processing the medical image input by a second decoder of the lesion detection machine learning computer model to generate a second lesion map prediction output identifying graphical elements corresponding to lesions in the at least one medical image; combining, by combination logic of the lesion detection machine learning computer model, the first lesion map prediction output and the second lesion map prediction output to generate a combined lesion map prediction output; generating, by final lesion map output logic of the lesion detection integrated machine learning model architecture, a final lesion prediction output based on the combined lesion map prediction output; and The final lesion prediction output is output by the final lesion map output logic for further downstream computational operations, wherein the first decoder is trained using a first loss function configured to balance the training of the second decoder, which is trained using a second loss function different from the first loss function.
23. The method of claim 22, further comprising: training the first decoder using the first loss function with machine learning logic implementing a first machine learning process, wherein the first loss function penalizes false negative lesion detections; training the second decoder using the second loss function with machine learning logic implementing a second machine learning process, wherein the second loss function penalizes false positive lesion detections; and The combination of the first decoder and the second decoder is trained by applying a third loss function to the first lesion map prediction output and the second lesion map prediction output by the logic of the lesion detection integrated machine learning model architecture to make the first lesion map prediction output and the second lesion map prediction output consistent with each other.
24. The method of claim 22, further comprising: processing one or more received medical images by a mask-generating machine learning computer model to generate a mask corresponding to an anatomical structure of interest present in said input; and The mask generated by the mask generating machine learning computer model is applied to the one or more received medical images to generate an input of at least one medical image, such that the at least one medical image includes a masked portion of the received medical image corresponding to the anatomical structure of interest.
25. The method of claim 24, wherein: The one or more received medical images include a subset of the medical images of the input volume of medical images.
26. The method of claim 24, wherein: The anatomical structure of interest is the human liver.
27. The method of claim 24, wherein: Generating the final lesion prediction output based on the combined lesion map prediction output further comprises: The one or more received medical images are processed by one or more decoders of an unmasked input processing machine learning computer model comprising one or more decoders and one or more encoders to generate an unmasked lesion map prediction output, wherein generating the final lesion prediction output based on the combined lesion map prediction output further comprises generating the final lesion prediction output by combining the combined lesion map prediction output and the unmasked lesion map prediction output.
28. The method of claim 27, wherein: The one or more encoders include three encoders, wherein each encoder is a convolutional neural network trained to detect lesions in the anatomical structure of interest, wherein the three encoders share a same set of operating parameters optimized by a machine learning process, and wherein the training of the encoders implements two loss functions, the two loss functions comprising a first adaptive loss function configured to penalize false positive errors in lesion detection and a second deeply supervised loss function.
29. The method of claim 28, wherein: The outputs from the three encoders are combined by the combination logic of the unmasked input processing machine learning computer model to generate a combined lesion prediction output of the unmasked input processing machine learning computer model, and the combined lesion prediction output is processed by the decoder of the unmasked input processing machine learning computer model to generate the unmasked lesion map prediction output.
30. The method of claim 27, wherein: Combining the combined lesion map prediction output and the unmasked lesion map prediction output includes generating an average of the combined lesion map prediction output and the unmasked lesion map prediction output.
31. The method of claim 24, wherein: Outputting the final lesion prediction output includes outputting the mask and the final lesion prediction output.
32. A computer program product comprising a computer-readable storage medium having a computer-readable program stored therein, wherein the computer-readable program, when executed on a computing device, causes the computing device to perform the steps of the method according to any one of claims 22 to 31.
33. A lesion detection device comprising: at least one processor; as well as At least one memory coupled to the at least one processor, wherein the at least one memory comprises instructions that, when executed by the at least one processor, cause the at least one processor to perform the steps of the method according to any one of claims 22 to 31.
34. A computer system comprising means for performing the steps of the method according to any one of claims 22 to 31.
Citation Information
Patent Citations
Questions and answers generation
US20110125734A1
Method of determining contrast phase of a computerized tomography image
US20220012927A1
Method and system automatically detecting local lesion in radiographic image
CN105640577A
Computer-aided detection using multiple images from different views of a region of interest to improve detection accuracy
CN109791692A