Method, computer program, and apparatus for a lesion detection artificial intelligence pipeline computing system (lesion detection artificial intelligence pipeline computing system)

An AI pipeline with multiple machine learning models enhances liver lesion detection and classification by minimizing false positives and negatives, addressing inefficiencies in existing automated systems.

JP7910896B2Active Publication Date: 2026-08-25MERATIVE US LP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021176183
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-30
Filing Date
2021-10-28
Publication Date
2026-08-25
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

Existing automated image analysis mechanisms for detecting liver lesions in medical imaging are inefficient and inaccurate, prone to human error, and require manual intervention for accurate lesion identification and classification.

Method used

An AI pipeline comprising multiple trained machine learning computer models processes medical images to detect, segment, and classify lesions, minimizing false positives and negatives by employing various loss functions and models to ensure accurate detection and classification of liver lesions.

Benefits of technology

The AI pipeline provides automated and accurate lesion detection and classification, reducing human error and improving efficiency in identifying liver lesions, enabling reliable identification and classification for healthcare professionals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007910896000006
    Figure 0007910896000006
  • Figure 0007910896000007
    Figure 0007910896000007
  • Figure 0007910896000008
    Figure 0007910896000008
Patent Text Reader

Abstract

To provide a lesion detection and classification artificial intelligence (AI) pipeline comprising a plurality of trained machine learning (ML) computer models.SOLUTION: First ML models process an input volume of medical images (VOI) to determine whether the VOI depicts a predetermined amount of an anatomical structure. The AI pipeline determines whether criteria, such as a predetermined amount of an anatomical structure of interest being depicted in the input volume, are satisfied by output of the first ML models. If satisfied, lesion processing operations are performed including: second ML models processing the VOI to detect lesions which correspond to the anatomical structure of interest; third ML models performing lesion segmentation and combining of lesion contours associated with the same lesion; and fourth ML models processing the listing of lesions to classify the lesions. The AI pipeline outputs the listing of lesions and the classifications.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates in general to improved data processing devices and methods, and more particularly to a mechanism for providing an artificial intelligence pipeline computing system capable of identifying anatomical features of a subject in electronic medical imaging data, such as lesions. [Background technology]

[0002] Liver lesions are groups of abnormal cells within the liver, a biological entity, and are sometimes called lumps or tumors. Non-cancerous or benign liver lesions are common and do not spread to other parts of the body. Such benign liver lesions generally do not cause any health problems. However, some liver lesions result from cancer. Patients with certain medical conditions may be more likely to have cancerous liver lesions than other patients. These medical conditions include, for example, hepatitis B or C, cirrhosis, iron storage (hemoglobinosis), obesity, or exposure to harmful chemicals (e.g., arsenic or aflatoxin).

[0003] Liver lesions are usually only identifiable by the presence of medical imaging tests, such as ultrasound, magnetic resonance imaging (MRI), computed tomography (CT), or positron emission tomography (PET) scans. Such medical imaging tests must be performed by a human subject specialist (SME) in medical imaging, who must use their expertise and human ability to interpret patterns in the images to determine whether the medical imaging test indicates any lesion. If a human SME identifies a possible cancerous lesion, the patient's physician may perform a biopsy to determine whether the lesion is cancerous.

[0004] Abdominal contrast-enhanced (CE)CT is the current standard for evaluating various abnormalities (e.g., lesions) of the liver. These lesions may be evaluated by human SMEs as malignant (hepatocellular carcinoma, cholangiocarcinoma, angiosarcoma, metastases, and other malignant lesions) or benign (hemangiomas, focal nodular hyperplasia, adenomas, cysts or lipomas, granulomas, etc.). Manual evaluation of such images by human SMEs is important in guiding subsequent therapeutic interventions. To adequately evaluate lesions on CE CT, multiple-phase studies have been conducted many times, providing medical images of different stages of palatine tissue enhancement in a healthy liver and comparison with the enhancement of the lesion to determine difference detection. Human SMEs can then determine the diagnosis of the lesion based on these differences. [Prior Art Documents] [Patent Documents] [Patent Document 1] U.S. Patent Application No. 16 / 926,880 [Patent Document 2] U.S. Patent Application Publication No. 2011 / 0125734 [Non-Patent Documents] [Non-Patent Document 1] Abraham et al., "A Novel Focal Tversky Loss function with Improved Attention U-Net for Lesson Segmentation," arXiv:1810.07842[cs], October 2018 [Non-Patent Document 2] "Watson And Heaalcare" by Yuan et al., IBMdeveloperWorks, 2011 [Non-Patent Document 3] "The Era of Cognitive Systems: An Inside Look at IBM Watson and How it Works" by Rob High, IBM RedBooks, 2012 [Overview of the project] [Problems that the invention aims to solve]

[0005] While several automated image analysis mechanisms have been developed, there is still a need to improve such mechanisms to provide more efficient and accurate analysis of medical image data for detecting lesions within imaged anatomical structures (e.g., the liver or other organs). [Means for solving the problem]

[0006] This “Outline of the Invention” is provided to introduce the selection of concepts in a simplified form, which will be further described in the “Modes for Carrying Out the Invention” herein. This “Outline of the Invention” is not intended to identify any important factors or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0007] In one embodiment, a method is provided in a data processing system comprising at least one processor and at least one memory, wherein at least one memory contains instructions for implementing a lesion detection and classification artificial intelligence (AI) pipeline, which is executed by at least one processor and comprises a plurality of trained machine learning computer models. The AI ​​pipeline performs a method which includes processing an input volume of a medical image by one or more first machine learning computer models of the AI ​​pipeline and determining whether the input volume exhibits a predetermined amount of anatomical structures of interest. The method further includes determining, by logic of the AI ​​pipeline, whether one or more default criteria are met by the output of one or more first machine learning computer models. One or more default criteria include that a predetermined amount of anatomical structures of interest are exhibited in the input volume.

[0008] In response to a determination that one or more default criteria are met by the output of one or more first machine learning computer models, this method performs a lesion processing operation. The lesion processing operation includes processing the input volume with one or more second machine learning computer models of the AI ​​pipeline to detect lesions present in the medical images of the input volume that correspond to the anatomical structures of interest. The lesion processing operation further includes processing the detected lesions with one or more third machine learning computer models of the AI ​​pipeline, performing segmentation of the lesions and joining of lesion contours from different medical images within the input medical images associated with the same lesion to generate a list of lesions and their contours. The lesion processing operation also includes processing the list of lesions and the contours associated with the lesions in the list of lesions with one or more fourth machine learning computer models of the AI ​​pipeline, and classifying the lesions in the list of lesions into one or more default classes of lesions corresponding to the anatomical structures of interest.

[0009] This method also includes outputting a list of lesions and their associated classifications via an AI pipeline for processing in downstream computing systems. Using this method, improved lesion detection and classification become possible within images of the target anatomical structure. This method provides an automated computing tool for automatically identifying lesions within anatomical structures, minimizing false positives and identifying the same disease. The contours of lesions are generated by relating the contours of lesions associated with the abnormality. In this way, more accurate detection and classification of lesions becomes possible, surpassing the mechanisms of conventional techniques.

[0010] In some embodiments, one or more first machine learning computer models in the AI ​​pipeline include a body part detection machine learning model that runs on an input volume to detect body parts of biological entities represented within the input volume and determines whether that part of the biological entity's body corresponds to a body part where the anatomical structure of interest resides. In some embodiments, the body part is the human abdomen, and the anatomical structure of interest is the human liver. This operation allows the AI ​​pipeline mechanism to first ensure that the volume of images being processed provides images of a body part (e.g., the abdomen) of a biological entity where the specific anatomical structure of interest (e.g., the liver) resides, so that processing resources are not applied to volumes of images where the anatomical structure does not reside.

[0011] In some embodiments, one or more first machine learning computer models in the AI ​​pipeline include a phase classification machine learning model and an anatomical structure detection machine learning model operating on an input volume, wherein the phase classification machine learning model determines the phase of the medical image represented in each of the medical images of the input volume, and the anatomical structure detection machine learning model determines the amount of anatomical structure of interest represented within the input volume. These operations ensure that all images in the volume represent the same phase, and therefore differences between images due to different phases do not cause errors in lesion detection. Furthermore, the determination of whether a particular amount of anatomical structure of interest is present in the input volume provides another mechanism for ensuring that the images provide useful results by representing at least a minimum amount of anatomical structure of interest, so that the AI ​​pipeline can produce reliable results about whether a patient has a lesion.

[0012] In some embodiments, one or more criteria further include determining that the input volume contains a single phase of medical images present in the input volume, based on the results of a phase classification machine learning model's operation on the input volume, and determining that a minimum threshold amount of the target anatomical structure is present in the input volume, based on the results of an anatomical structure detection machine learning model's operation on the input volume. In some embodiments, in response to the results of a phase classification machine learning model's operation indicating that two or more phases of medical images are present in the input volume, a subset of medical images in the input volume corresponding to the target phase is selected, and lesion processing operations are performed only on the selected subset of medical images in the input volume. These operations also ensure the reliability of the AI ​​pipeline results by eliminating a potential source of error resulting from the presence of multiple phases in the input volume. By providing the ability to select a subset of the volume associated with a single phase, this enables adaptive processing of the input volume without requiring the rejection of the entire volume. In some embodiments, the target phase is one of the following: pre-contrast phase, angiography phase, portal / venography phase, or delayed phase.

[0013] In some embodiments, determining whether one or more default criteria are met by the output of one or more first machine learning computer models further includes performing an axial scoring operation on medical images in an input volume by the axial scoring logic of the AI ​​pipeline, the axial scoring operation including scoring the images in the input volume according to a predefined scoring algorithm, estimating the axial scores of the lowest slice (MISV) and highest slice (MSSV) in the volume from the scores associated with the images, and identifying fragments of the target anatomical structure present in the input volume based on the axial scores of the MISV and MSSV. These operations provide a mechanism for determining the amount of anatomical structure present in the input volume. This scoring provides a measurement mechanism based on the range of slices present in the input volume.

[0014] In some embodiments, one or more second machine learning computer models in the AI ​​pipeline include a population of machine learning computer models, each machine learning computer model within the population is trained by a machine learning process to process the input volume in a different manner than other machine learning computer models in the population to generate corresponding lesion detection predictions. The existence of a population of differently trained machine learning computer models allows the machine learning computer models to compensate for the focus, potential weaknesses, or both of the loss function / error function in the training of the other computer models in the population, thereby improving the overall lesion detection.

[0015] In some embodiments, the collection of machine learning computer models includes a mask generation machine learning model, an input volume processing machine learning model, and a masked input volume processing machine learning model, wherein the mask generation machine learning model is trained by a machine learning process to generate a mask corresponding to the target anatomical structure, and applies this mask to the input volume to generate a masked input volume; the input volume processing machine learning model is trained by a machine learning process to generate a first lesion prediction output based on the input volume; and the masked input volume processing machine learning model is trained by a machine learning process to generate a second lesion prediction output based on the masked input volume. The existence of a machine learning model that generates a mask and operates on a masked input volume allows the processing performed by that machine learning model to focus on the portion of the image corresponding to the target anatomical structure. Combining this machine learning model with the operation of other machine learning models that operate on unmasked inputs enables even better lesion detection.

[0016] In some embodiments, a first machine learning model for a population is trained using a first loss function that penalizes errors in classifying false negatives of lesions, and a second machine learning model for a population is trained using a second loss function different from the first machine learning model, which penalizes errors in classifying false positives of lesions. The existence of multiple machine learning models implementing different loss functions allows each machine learning model to compensate for the weaknesses of the other machine learning models in terms of sensitivity and specificity.

[0017] In some embodiments, the population logic applies a third loss function, which compares the output of the first lesion detection of the first machine learning model with the output of the second lesion detection of the second machine learning model, and operates to make the output of the first lesion detection match the output of the second lesion detection. The third loss function ensures that the outputs of differently trained machine learning models in a population match each other during training, thereby improving the training of different machine learning models even when trained using different loss functions.

[0018] In some embodiments, one or more third machine learning computer models in the AI ​​pipeline include a lesion segmentation machine learning model that segments an image of an input volume into contours corresponding to lesions to generate lesion segments; a z-direction connection machine learning model that joins subsets of lesion segments based on segments generated by the lesion segmentation machine learning model that are determined to be associated with the same lesion in three-dimensional space; and a contour improvement machine learning model that isolates lesions in the input volume and improves the contours of the lesions. These models enable the identification of lesions in three-dimensional space by connecting two-dimensional lesion contours associated with the same lesion and then improving the contours of these lesions. In this way, a more accurate representation of the detected lesions is made possible in three dimensions.

[0019] In some embodiments, one or more false positive removal machine learning computer models are provided that operate on a list of lesions and remove false positive detections of lesions from the list. False positive removal helps to remove those contours and regions of images in the input volume that were detected as lesions but are not actually lesions.

[0020] In some embodiments, one or more false positive removal machine learning computer models include one or more false positive removal machine learning computer models trained by a machine learning process based on two different operating points, the first of which the operating points corresponds to a patient-level operating point, and the second of which the operating points corresponds to a lesion-level operating point. The dual operating point implementation of false positive removal enables more accurate false positive removal by first determining at the patient level whether a patient has a lesion, and if so, performing a more sensitive evaluation of the list of lesions to remove lesions in the list that are not actually lesions.

[0021] In some embodiments, one or more fourth learning computer models classify each lesion in a list of lesions into one of a plurality of default classes, each of which includes at least one of the lesion type, or benign, malignant, or undefined class. Classifying the lesions informs healthcare professionals of which lesions in a patient require treatment. Furthermore, in some embodiments, this lesion classification may be used by a downstream computing system to process the output of the list of lesions and the classifications associated with the lesions, and to generate a visual representation of the lesions in the list of lesions in a graphical user interface. In this way, a visual representation of lesions that require the attention of healthcare professionals to provide treatment to patients is provided to healthcare professionals.

[0022] In some embodiments, the input volume includes an image corresponding to at least one of magnetic resonance imaging or computed tomography. Therefore, the embodiments may be applied to known and future medical imaging technologies in which lesion detection and classification are performed.

[0023] In other embodiments, a computer-readable medium containing a computer-readable program or a computer program product comprising a computer-readable medium is provided. When this computer-readable program is executed on a computing device, it causes the computing device to perform various operations and combinations of operations among those outlined above with respect to embodiments of the method.

[0024] In yet another embodiment, a system / device is provided. This system / device may comprise one or more processors and memory coupled to one or more processors. This memory may contain instructions, which, when executed by one or more processors, cause one or more processors to perform various operations and combinations of operations from among the operations outlined above with respect to embodiments of the method.

[0025] By considering the following detailed description of embodiments of the present invention, these and other features and advantages of the present invention will be explained and will become apparent to those skilled in the art. [Brief explanation of the drawing]

[0026] The present invention, its most commonly used methods of use and other purposes, and its advantages are described in the embodiments. The following detailed explanation of the example will be best understood by reading and referring to it together with the attached diagram.

[0027] [Figure 1] This is an exemplary block diagram of an AI pipeline that implements multiple ML / DL computer models, specifically configured and trained to perform anatomical structure identification and lesion detection in input medical image data, according to one embodiment.

[0028] [Figure 2] This is an exemplary flowchart illustrating the exemplary operation of an AI pipeline according to one embodiment.

[0029] [Figure 3A] This is an exemplary diagram showing an exemplary input volume of a slice (medical image) of a patient's abdomen, according to one embodiment.

[0030] [Figure 3B] This figure shows another representation of the input volume in Figure 3A, including a section of slices represented with corresponding axial scores s'inf and s'sup.

[0031] [Figure 3C] Figure 3A is an illustrative diagram of an input volume in which the volume is divided axially into n completely overlapping sections.

[0032] [Figure 4A] This is an illustrative figure of one embodiment of an ML / DL computer model trained to estimate the values ​​of s'sup and s'inf in sections of an input volume of a medical image, according to one embodiment example. [Figure 4B] This is an illustrative figure of one embodiment of an ML / DL computer model trained to estimate the values ​​of s'sup and s'inf in sections of an input volume of a medical image, according to one embodiment example. [Figure 4C] This is an illustrative figure of one embodiment of an ML / DL computer model trained to estimate the values ​​of s'sup and s'inf in sections of an input volume of a medical image, according to one embodiment example.

[0033] [Figure 5] A flowchart illustrating the exemplary operation of the AI ​​pipeline's liver detection and predetermined amount of anatomical structure determination logic according to one embodiment example.

[0034] [Figure 6]This is an illustrative figure of a group of ML / DL computer models used to perform lesion detection within a target anatomical structure (e.g., the liver) according to one embodiment.

[0035] [Figure 7] A flowchart illustrating the exemplary operation of liver / lesion detection logic in an AI pipeline according to one embodiment.

[0036] [Figure 8] A block diagram showing a segmentation pattern of a lesion according to one embodiment.

[0037] [Figure 9] This figure shows the results of lesion detection and slice-by-slice segmentation according to one embodiment.

[0038] [Figure 10A] This figure shows the positioning of the seed according to one embodiment. [Figure 10B] This figure shows the positioning of the seed according to one embodiment. [Figure 10C] This figure shows the positioning of the seed according to one embodiment. [Figure 10D] This figure shows the positioning of the seed according to one embodiment.

[0039] [Figure 11A] A block diagram showing the mechanism of lesion segmentation according to one embodiment.

[0040] [Figure 11B] This block diagram shows a mechanism for relabeling seeds according to one embodiment.

[0041] [Figure 12] This flowchart shows an overview of exemplary operation of lesion segmentation according to one embodiment.

[0042] [Figure 13A] This figure shows the z-direction connection of the lesion according to one embodiment. [Figure 13B] This figure shows the z-direction connection of the lesion according to one embodiment. [Figure 13C] This figure shows the z-direction connection of the lesion according to one embodiment.

[0043] [Figure 14A] This figure shows the results of a trained model regarding the connectivity of lesions in the z direction, according to one embodiment. [Figure 14B] This figure shows the results of a trained model regarding the connectivity of lesions in the z direction, according to one embodiment.

[0044] [Figure 15] This flowchart shows an illustrative overview of the operation of a mechanism for connecting two-dimensional lesions along the z-axis, according to one embodiment.

[0045] [Figure 16] This figure shows an example of an image containing the contours of two lesions within the same image, according to one embodiment.

[0046] [Figure 17] This flowchart outlines the exemplary operation of a mechanism for improving the contour of each slice, according to one embodiment.

[0047] [Figure 18A] This figure shows an example of an ROC curve for determining operating points at the patient level and the lesion level, according to one embodiment.

[0048] [Figure 18B] This is an exemplary flowchart of the operation for performing false positive removal based on patient-level and lesion-level operating points, according to one embodiment.

[0049] [Figure 18C] This is an illustrative flowchart of the operation for performing per-voxel false positive removal based on input volume level and voxel level operating points, according to one embodiment example.

[0050] [Figure 19] This flowchart outlines the exemplary operation of the false positive removal logic of an AI pipeline according to one embodiment.

[0051] [Figure 20] This is an illustrative diagram of a distributed data processing system in which an embodiment of the example may be implemented.

[0052] [Figure 21] This is an exemplary block diagram of a computing device in which an embodiment of the example may be implemented. [Modes for carrying out the invention]

[0053] In modern pharmacotherapy, the detection of lesions or groups of abnormal cells is primarily a manual process. Because it is a manual process, it is inherently prone to errors, stemming from human limitations in an individual's ability to detect portions of digital medical images that indicate such lesions, especially given the growing demand for individuals to evaluate an increasing number of images in shorter timeframes. While several automated image analysis mechanisms have been developed, there is still a need to improve such mechanisms to provide a more efficient and accurate analysis of medical imaging data for detecting lesions within imaged anatomical structures (e.g., the liver or other organs).

[0054] The examples particularly concern an improved computing tool that provides automated computer-aided artificial intelligence medical image analysis, which is specifically trained via a machine learning / deep learning computer process to perform: detection of anatomical structures; detection of lesions or other biological structures of subject within or associated with such anatomical structures; special segmentation of detected lesions or other biological structures; false positive removal based on the special segmentation; classification of detected lesions or other biological structures; and providing the results of lesion / biological structure detection to a downstream computing system for additional computer operations. The following description of the examples assumes an example mechanism that is specifically trained with respect to liver lesions as the biological structure of subject, but the examples are not limited to such an example. Rather, those skilled in the art will recognize that the machine learning / deep learning-based artificial intelligence mechanism of the examples may be implemented with respect to many other types of biological structures / lesions within or associated with other anatomical structures represented in medical image data, without departing from the spirit and scope of the invention. Furthermore, while the examples of embodiments may be described in relation to medical image data, such as computed tomography (CT) medical image data, the examples of embodiments may be carried out using any digital medical image data from various types of medical imaging techniques, including but not limited to positron emission tomography (PET) and other nuclear medicine imaging, ultrasound, magnetic resonance imaging (MRI), elastography, photoacoustic imaging, echocardiography, magnetic particle imaging, functional near-infrared spectroscopy, elastography, and various X-ray imaging techniques.

[0055] Overall, the embodiment provides an improved artificial intelligence (AI) computer pipeline comprising several particularly configured and trained AI computer tools (e.g., neural networks, cognitive computing systems, or other AI mechanisms trained on a finite set of data to perform a specific task). Each configured and trained AI computer tool is specifically configured / trained to perform a specific type of artificial intelligence processing on a volume of input medical images, represented as one or more collections of data or metadata, or both, defining medical images captured by medical imaging technology. Generally, these AI tools employ machine learning (ML) / deep learning (DL) computer models (or simply ML models) to perform the task, and the ML / DL computer models use different computer processes specific to the computer tools and in particular the ML / DL computer models, learning patterns and relationships between data representing specific outcomes (e.g., image classification or labeling, data values, treatment recommendations, etc.) while emulating human thought processes with respect to the generated outcomes. An ML / DL computer model is essentially a functional set of elements that includes a machine learning algorithm, the configuration settings of the machine learning algorithm, the features of the input data identified by the ML / DL computer model, and the labels (or outputs) generated by the ML / DL computer model. Instances of a particular ML / DL computer model are generated by specifically tuning the functionalities of these elements through a machine learning process. Different ML models may be specifically configured and trained to perform different AI functions on the same or different input data.

[0056] It should be understood that, as artificial intelligence (AI) pipelines implement multiple ML / DL computer models, these ML / DL computer models are trained through ML / DL processes for specific purposes. Therefore, as an overview of the ML / DL computer model training process, it should be understood that machine learning is related to the design and development of techniques that take empirical data (such as medical imaging data) as input and recognize complex patterns within that input data. One pattern common to machine learning techniques is the use of an underlying computer model M, where, given the input data, the parameters of M are optimized to minimize a cost function associated with M. For example, in the context of classification, model M may be a straight line separating the data into two classes (e.g., labels) such that M = a*x + b*y + c, where the cost function is the number of misclassified points. In this case, the learning process operates by tuning parameters a, b, and c to minimize the number of misclassified points. After this optimization (or learning) phase, new data points can be classified using model M. In many cases, M is a statistical model, and given the input data, the cost function is inversely proportional to the possibilities of M. This is merely a simple example to provide a general description of machine learning training, and other types of machine learning using different patterns, cost (or loss) functions, and optimizations may be used in conjunction with the mechanism of the example embodiment without departing from the spirit and scope of the invention.

[0057] For the purpose of anatomical structure detection, lesion detection, or both (a lesion being an "abnormality" in medical imaging data), the learning machine uses ML / DL computers to represent normal structures. Models can be constructed to detect data points in medical images that deviate from this normal structural representation of the ML / DL computer model. Specific ML / DL computer models (e.g., supervised, unsupervised, or semi-supervised models) may be used, for example, to generate anomaly scores and report them to another device, to produce classification outputs indicating one or more classes that classify the input, or to generate probabilities or scores associated with various classes. Examples of machine learning techniques that may be used to construct and analyze such ML / DL computer models include, but are not limited to, nearest neighbor (NN) techniques (e.g., k-NN models, replicator NN models, etc.), statistical techniques (e.g., Bayesian networks, etc.), clustering techniques (e.g., k-means, etc.), neural networks (e.g., reservoir networks, artificial neural networks, etc.), and support vector machines (SVMs).

[0058] The processor-implemented artificial intelligence (AI) pipelines of the embodiments typically include one or both of machine learning (ML) computer models and deep learning (DL) computer models. In some cases, one or more of ML and DL may be used or implemented to achieve a particular outcome. Conventional machine learning includes, or can be used, algorithms such as Bayesian decision, regression, decision trees / forests, support vector machines, or neural networks. Deep learning can be based on deep neural networks and can use multiple layers such as convolutional layers. Such DL using hierarchical networks, etc., can be efficient in its implementation and can achieve improved accuracy compared to conventional ML techniques. Conventional ML can generally be distinguished from DL in that DL models may consume a relatively large amount of processing resources or power resources, or both, although the performance of DL models can outperform that of classical ML models. In relation to the embodiments, references to one or more of ML and DL in this specification should be understood to encompass one or both forms of AI processing.

[0059] In the embodiment, the ML / DL computer model of the AI ​​pipeline is executed after configuration and training via the ML / DL training process to perform complex computer medical image analysis to detect anatomical structures in the input medical image and generate outputs that specifically identify the biological structures of interest (hereinafter assumed to be liver lesions for the purposes of describing the embodiment), their classification, contours that specify the location where these biological structures of interest (e.g., liver lesions) exist in the input medical image (hereinafter assumed to be CT medical image data), and other information that assists human subject matter experts (SMEs) such as radiologists and physicians in understanding the patient's medical condition from the perspective of the captured input medical image. Furthermore, these outputs can be provided to other downstream computer systems to perform additional artificial intelligence operations, such as treatment recommendations and other decision support actions based on classification, contours, etc.

[0060] First, the artificial intelligence (AI) pipeline of the embodiment receives an input volume of computed tomography (CT) medical image data and detects which part of the body of a biological entity is shown in the CT medical image data. The “volume” of medical image data is a three-dimensional representation of the internal anatomical structure of a biological entity, consisting of a stack of two-dimensional slices, which may be individual medical images captured by medical imaging technology. The stack of slices may also be called a “slab,” and these stacks may differ from the slices themselves in that they represent a portion of an anatomical structure that has thickness, and the stack of slices or slab generates a three-dimensional representation of the anatomical structure.

[0061] For the purposes of this explanation, it is assumed that the biological entity is human, but the present invention may operate on medical images of various types of biological entities. For example, in veterinary medicine, the biological entity may be various types of small animals (e.g., pets such as dogs and cats) or large animals (e.g., horses, cattle, or other livestock). In an implementation where the AI ​​pipeline is specifically trained for detecting liver lesions, the AI ​​pipeline determines whether the input CT medical image data represents a scan of the abdomen present in the CT medical image data, and if it does not, the operation of the AI ​​pipeline with respect to the input CT medical image data terminates because it is not directed to the correct location or part of the human body. In accordance with the embodiments, there may be different AI pipelines trained to process input medical images of different parts of the body and different biological structures of different subjects, and input CT medical images may be input to each AI pipeline, or may be routed to an AI pipeline based on the classification of body parts or body regions shown in the input CT medical images. For example, the classification of the input CT medical images as body parts or body regions represented in the input CT medical images may be performed first, and then a corresponding trained AI pipeline for processing the input CT medical images may be selected from several trained AI pipelines of the kind described herein. For the purposes of the following description, a single AI pipeline trained to detect liver lesions will be described, but this extension to a series of AI pipelines or a group of AI pipelines will be apparent to those skilled in the art in consideration of this description.

[0062] Assuming that the input volume of CT medical images includes medical images of the human abdomen (for the purpose of detecting liver lesions), the processing of the input CT medical images is further performed in two initial stages, which may be performed substantially in parallel with each other, sequentially, or both, depending on the desired implementation. The two initial stages include a phase classification stage and an anatomical structure detection stage (for example, a liver detection stage if the AI ​​pipeline is configured to perform liver lesion detection).

[0063] The phase classification stage determines whether the volume of input CT medical images consists of a single imaging phase or multiple imaging phases. In medical imaging, a “phase” is an indication of contrast agent ingestion. For example, in some medical imaging techniques, a phase may be defined in terms of when the contrast agent is introduced into the biological entity, enabling the capture of the medical image, including capturing the pathway of the contrast agent. For example, a phase may include a pre-contrast phase, angiography phase, portal / venography phase, and delayed phase, and the medical image is captured in one or all of these phases. Phases are typically related to the timing after injection and the characteristics of enhancement of structures in the image. Timing information may be considered to “sort” possible phases (e.g., the delayed phase is always acquired after the portal venography phase) and to estimate the possible phases of a particular image. Regarding the use of the properties of enhancing structures within an image, one example of using this type of information to determine the phase is described in the concurrently pending U.S. Patent Application No. 16 / 926,880, filed July 13, 2020, by the same applicant, “Method of Determining Contrast Phase of a Computerized Tomography Image.” Furthermore, timing information may be used in conjunction with other information (such as sampling and reconstruction kernels) to extract the best representation of each phase (a particular acquisition may be reconstructed in multiple ways).

[0064] Based on the timing and / or characteristics of the enhancement, the images in the input volume may be assigned to or classified into corresponding phases. After this classification, it may be determined whether the volume contains images from a single phase (e.g., portal / vein phases are present but arterial phases are not) or from multiple phases (e.g., portal / vein and arterial). If the phase classification indicates that a single phase is present in the volume of input CT medical images, further processing by the AI ​​pipeline is performed as described below. If multiple phases are detected, the volume is not further processed by the AI ​​pipeline. However, in some embodiments, this filtering of volumes based on single / multiple phases accepts only volumes containing images from a single phase and rejects volumes with multiple phases, while in other embodiments, the AI ​​pipeline processing described herein may filter out images from volumes that were not classified as part of the target phase, for example, filtering out images from volumes that were not classified as part of the portal / vein phase while retaining portal / vein phase images in the volume, thereby modifying the input volume to a modified volume containing only a subset of images classified as part of the target phase. Furthermore, as mentioned above, different AI pipelines may be trained for different types of volumes, and in some embodiments, phase classification of images in the input volume may be used to route or distribute the images in the input volume to corresponding AI pipelines that are trained and configured to process images in different phases, thereby allowing the input volume to be further divided into constituent subvolumes that can be routed to corresponding AI pipelines for processing, for example, a first subvolume corresponding to portal / venous phase images being sent to a first AI pipeline for processing, and a second subvolume corresponding to arterial phase images being sent to a second AI pipeline.If the input CT medical image volume contains a single phase, or if the AI ​​pipeline processes images from a single phase input volume or subvolume, the subvolume is filtered and optionally routed to the corresponding AI pipeline, after which the volume (or subvolume) is passed to the next stage of the AI ​​pipeline for further processing.

[0065] The second initial stage is the detection of the target anatomical structure (in this embodiment, the liver), in which a portion of the volume representing the target anatomical structure is identified and passed to the next downstream stage of the AI ​​pipeline. The detection of the target anatomical structure (hereinafter referred to as the liver detection stage according to this embodiment) includes a machine learning (ML) / deep learning (DL) computer model that has been specifically trained and configured to perform computer medical image analysis to identify portions of the input medical image corresponding to the target anatomical structure (e.g., the liver). Such medical image analysis may include training the ML / DL model on labeled training medical image data as input to determine whether the input medical image (training image during training) contains the target anatomical structure (e.g., the liver). Based on the ground truth of the image labels, the operating parameters of the ML / DL model may be adjusted to reduce loss or error in the results produced by the ML / DL model until convergence is reached (i.e., loss is minimized). This process trains the ML / DL model to recognize patterns in medical imaging data indicating the presence of the target anatomical structure (liver in this example). Subsequently, new input medical imaging data, after training, contains patterns indicating the presence of the anatomical structure. To determine whether or not this is the case, an ML / DL model can be run on new input data, and if the probability exceeds a predetermined threshold, it can be determined that the medical image data contains the anatomical structure of interest.

[0066] Therefore, in the liver detection phase, the AI ​​pipeline uses a trained ML / DL computer model to determine whether the volume of the input CT medical image contains images showing a liver. The portion of the volume showing a liver, along with the results of the phase classification phase, is passed to the determination phase of the AI ​​pipeline, which determines whether a single phase of the medical image exists and whether at least a predetermined amount of the target anatomical structure is present in the portion of the volume showing the target anatomical structure (e.g., liver). The determination of whether a predetermined amount of the target anatomical structure is present may be based on a known measurement mechanism that determines a measurement of the structure from the medical image (e.g., calculating the size of the structure from the difference in the pixel location in the image). If the measurement represents at least a predetermined amount or portion of the anatomical structure, the measurement may be compared to a predetermined size (e.g., average size) of the anatomical structure of a similar patient with similar demographics, so that the AI ​​pipeline can perform further processing. In one embodiment, this determination determines, for example, whether at least one-third of the liver is present in the portion of the input CT medical image volume that has been determined to show a liver. In this embodiment, 1 / 3 is used, but any predetermined amount of the structure may be used as determined to be suitable for a particular implementation without departing from the spirit and scope of the present invention.

[0067] In one example embodiment, to determine whether a predefined amount of a target anatomical structure exists within the volume of an input CT medical image, a slice (i.e., the first slice containing the liver (FSL)) corresponding to the medical image within the volume containing the first representation of the target anatomical structure (e.g., the liver) is given a slice score of 0, and an axial score is defined such that the last slice containing the liver (LSL) has a score of 1. Assuming a human biological entity, proceeding from the lowest slice within the volume (MISV) (closest to the lower extremities (e.g., feet)) to the highest slice within the volume (MSSV) (closest to the head), the first and last slices are defined. A pair of slice scores (s sup and s inf ) corresponding to the slice scores of the MSSV slice and the MISV slice, respectively, defines a liver axial score estimate (LAE). The ML / DL computer model is specifically configured and trained to determine the slice scores s sup and s inf of the volume of the input CT medical image, as will be described in more detail later. The mechanism of the example embodiment recognizes these slice scores and, from the above definition, recognizes that the liver extends from 0 to 1, and can determine a fragment of the liver within the field of view of the volume of the input CT medical image.

[0068] In some example embodiments, the slice scores s sup and s inf may be indirectly detected by first dividing the volume of the input CT medical image into sections and then running a configured and trained ML / DL computer model against the slices of the section to estimate the height of each slice to determine the highest slice (closest to the head) and the lowest slice (closest to the feet) s' sup and s' inf of the liver within the section for each section. s sup and s' infAssuming the estimated value of s, and knowing how the section is located relative to the entire volume of the input CT medical image, extrapolation can be performed. sup and s inf An estimate of the height is detected. This method is based on a robust estimate of the height of any slice from the input volume (or a sub-volume associated with the phase of interest). Such an estimate can be obtained, for example, by training a regression model using a deep learning model that performs height estimation from chunks (a set of consecutive slices). For example, a long-shorter-term memory (LSTM) type artificial neural network model is suitable for these tasks and has the ability to encode the ordering of slices containing the biological structures of the liver and abdomen. For each volume, s sup and s inf It should be noted that there are n estimates, where n is the number of sections per volume. In one embodiment, the final estimate is obtained by taking the unweighted average of these n estimates, but in other embodiments, the final estimate may be generated using other functions of the n estimates.

[0069] The volume of input CT medical images sup and s inf The final estimates are determined, and based on these values, fragments of the target anatomical structure (e.g., liver) are calculated. This task is made possible by estimating the height of each slice. From the estimates of the height of the first slice (h1) and the last slice (h2) of the liver in the input volume, assuming that the actual heights of the first and last slices of the liver (whether they are included in the input volume or not) are H1 and H2, the visible portion of the liver in the input volume can be expressed as (min(h1, H1)-max(h2, H2)) / (H1-H2). These calculated fragments may then be compared to a predetermined threshold to determine whether a predetermined minimum amount of the target anatomical structure is present in the volume of the input CT medical image (e.g., at least one-third of the liver is present in the volume of the input CT medical image).

[0070] If this determination concludes that there are multiple phases, or that a predetermined amount of the target anatomical structure is not present in the portion of the input CT medical image volume showing the anatomical structure, or both, then further processing of the volume may be interrupted. If this determination concludes that the input CT medical image volume contains a single phase and at least a predetermined amount of the target anatomical structure (e.g., one-third of the liver is visible in the image), then the portion of the input CT medical image volume showing the anatomical structure is transferred to the next stage of the AI ​​pipeline for processing.

[0071] In the next stage of the AI ​​pipeline, the AI ​​pipeline performs lesion detection on the portion of the input CT medical image volume representing the target anatomical structure (e.g., the liver). This liver and lesion detection stage of the AI ​​pipeline uses a population of ML / DL computer models to detect the liver and lesions within the liver represented within the volume of the input CT medical image. The population of ML / DL computer models performs liver and lesion detection using differently trained ML / DL computer models, and the ML / DL computer models are trained using a loss function to balance false positives and false negatives in lesion detection. Furthermore, the population of ML / DL computer models is configured such that a third loss function forces the outputs of the ML / DL computer models to match each other.

[0072] Assuming that liver detection and lesion detection are performed at this stage of the AI ​​pipeline, a first ML / DL computer model is run on a volume of input CT medical images to detect the presence of a liver. This ML / DL computer model may be the same as the ML / DL computer model employed in a previous stage of the AI ​​pipeline for the detection of the target anatomical structure, and therefore, previously obtained results may be utilized. Multiple (two or more) other ML / DL computer models are configured and trained to perform lesion detection on portions of medical images showing a liver. The first ML / DL computer model is configured using two loss functions. The first loss function penalizes errors in false negatives (i.e., classifications that incorrectly indicate the absence of a lesion (normal anatomical structure)). The second loss function penalizes errors in false positive results (i.e., classifications that incorrectly indicate the presence of a lesion (abnormal anatomical structure)). The second ML / DL model is trained to detect lesions using an adaptive loss function that penalizes false-positive errors in liver slices containing normal tissue and false-negative errors in liver slices containing lesions. The detections output from the two ML / DL models are averaged to produce the final lesion detection.

[0073] The result of the liver / lesion detection stage of the AI ​​pipeline includes a detection map (e.g., a per-voxel map of the detected liver lesions within the volume of the input CT medical image) that identifies one or more contours (contour lines) of the liver and portions of the medical image data elements corresponding to the detected lesions. The image map is then input to the lesion segmentation stage of the AI ​​pipeline. The lesion segmentation stage divides the detection map using watershed techniques, as will be described in more detail later, to generate a division of the image elements (e.g., voxels) of the input CT medical image. Based on this division, the liver lesion segmentation stage identifies all contours corresponding to lesions present in slices of the volume of the input CT medical image and performs actions to identify which contours in three dimensions correspond to the same lesion. Lesion segmentation aggregates the interrelated lesion contours to generate a three-dimensional segmentation of the lesions. Lesion segmentation uses inpainting of lesion image elements (e.g., voxels) and non-liver tissues represented within the medical image to focus on each lesion individually and perform active contour analysis. In this way, individual lesions can be identified and processed without biasing the analysis due to other lesions in the medical image or due to the extravasation of the liver in the image.

[0074] The results of lesion segmentation are a list of lesions containing corresponding contours or outlines within the volume of the input CT medical image. These outputs may include detections that are not actual lesions. To minimize the impact of these false positives, these outputs are fed to the next stage of the AI ​​pipeline, which targets false positive removal using a trained false positive removal model. The false positive removal model in the AI ​​pipeline acts as a classifier to identify which outputs are actual lesions and which detections are false positives. The input consists of volume of images (VOIs) centered on detected detections, associated with masks obtained from the improved lesion segmentation. The false positive removal model is trained using the data that is the result of the detection / segmentation stage. Objects that are lesions from ground truth detected by the detection algorithm are used to represent the lesion class during training, and detections that do not match any lesion from ground truth are used to represent the non-lesion (false positive) class.

[0075] To further improve overall performance, a dual operating point strategy was employed in the lesion detection model and the false-positive model. The idea is that the output of the AI ​​pipeline can be interpreted at different levels. First, the output of the AI ​​pipeline can be used to tell whether the examination volume (i.e., the input volume or image volume (VOI)) contains a lesion. Second, the output of the AI ​​pipeline can be used regardless of whether the lesion is contained in the same patient / examination / volume. The goal is to maximize the detection of lesions. For clarity, measurements performed for testing are referred to herein as "patient level," and measurements performed in relation to lesions are referred to herein as "lesion level." Maximizing sensitivity at the "lesion level" reduces specificity at the "patient level" (one detection is sufficient to say that a patient has a lesion). This ultimately requires choosing between insufficient specificity at the patient level or low sensitivity at the lesion level, which may be suboptimal for clinical use.

[0076] With this in mind, the embodiment uses a dual operating point method for both lesion detection and false-positive removal. The principle is to first perform processing using a first operating point that provides reasonable performance at the patient level. Then, for patients who have at least one detected lesion detected from the first run, a second operating point is used to reinterpret / process the detected lesion. This second operating point is selected to be more sensitive. This second operating point has lower specificity than the first operating point, but this loss of specificity is included at the patient level because all patients for whom no lesions were detected using the first operating point are left as they are, regardless of whether the second operating point detects additional lesions. Therefore, the specificity at the patient level is determined by the first operating point alone. The sensitivity at the patient level lies between the sensitivity of the first operating point selected alone and the sensitivity of the second operating point (some false-negative cases from the first operating point may be turned into true positives by the second operating point). On the lesion side, the sensitivity at the actual lesion level is improved compared to the first operating point alone. When processing is performed using only the first operating point, false positives do not occur, resulting in better lesion specificity than when a second operating point is selected alone and has lower specificity.

[0077] While the embodiment assumes the use of a specific configuration and dual operating point method, it should be understood that the dual operating point method may be used in other configurations and for other purposes if there is interest in measuring performance at the group level (in the embodiment, this group level is the "patient level") and element level (in the embodiment, this element level is the "lesion level"). In the embodiment, the dual operating point method is applied to both lesion detection and false positive removal, but it should be understood that the dual operating point method may be extended beyond these stages of the AI ​​pipeline. For example, lesion detection may be performed at the voxel level (element) and volume level (group) rather than at the patient level and lesion level. In another example, the voxel level or lesion level may be used at the element level, and a slab (set of slices) may be used as the group level. In yet another example, all volumes of the examination may be used as the group level instead of a single volume. It should be understood that this method may be applied to two-dimensional images (e.g., 2D X-ray images such as chest or mammography) to analyze images rather than three-dimensional volumes. Specificity, such as the average number of false positives per patient / group, may be used in selecting the operating point. Furthermore, although the embodiment is described as being applicable to lesion detection and classification, the dual operating point-based method may be applied to other structures (such as clips, stents, and implants) and beyond medical imaging.

[0078] The results of detection and false-positive removal based on dual operating points lead to the identification of a final filtered list of lesions, which are further processed by the lesion classification stage of the AI ​​pipeline. In the lesion classification stage of the AI ​​pipeline, a configured and trained ML / DL computer model is run on a list of lesions and corresponding contour data, thereby classifying the lesions into one of several predefined lesion classifications. For example, each lesion and its attributes (e.g., contour data) in the final filtered list of lesions may be input into the trained ML / DL computer model, which then operates on this data to classify the lesions as a specific type of lesion. This classification may be performed using a classifier previously trained on ground truth data (e.g., a trained neural network computer model) in combination with the results of the previous processing steps in the AI ​​pipeline. This classification task can be more or less complex, and may involve providing labels such as benign, malignant, or indeterminate, or, in another example, providing the actual type of lesion, such as cyst, metastasis, or hemangioma. The classifier can be, for example, a computer model based on a neural network (e.g., SVM, decision tree, etc.) or a deep learning computer model. The actual input to this classifier is the lesion-centered portion, which in some embodiments may be augmented using a mask or contour of the lesion.

[0079] Following the classification of lesions by the lesion classification stage of the AI ​​pipeline, the AI ​​pipeline outputs a list of lesions and their classifications, along with any contour attributes of the lesions. Furthermore, the AI ​​pipeline may output liver contour information for the liver. The information generated by this AI pipeline may be provided to further downstream computing systems for further processing and generation of representations of the target anatomical structure and any detected lesions present within that anatomical structure. For example, a graphic representation of the volume of the input CT medical image may be generated in a medical image viewer or other computer application, and the anatomical structure and detected lesions may be overlaid on the graphic representation or otherwise highlighted using the contour information generated by the AI ​​pipeline. In other embodiments, downstream processing of the information generated by the AI ​​pipeline may include diagnostic decision support actions and automated medical image report generation based on the detected list of lesions, classifications, and contours. In other embodiments, different treatment recommendations may be generated based on the lesion classification for physician review and consideration.

[0080] In some embodiments, a list of lesions, their classifications, and contours may be stored in a patient-associated historical data structure that matches a volume of input CT medical images, allowing for the storage and evaluation of multiple runs of an AI pipeline over different volumes of input CT medical images associated with that patient over time. For example, differences between a list of lesions, their associated classifications or both, and contours may be determined to assess the progression of a patient's disease or medical condition and to present such information to healthcare professionals to assist in patient treatment.

[0081] Other downstream computing systems and processing of specific anatomical structure and lesion detection information generated by the AI ​​mechanism of the embodiment may be implemented without departing from the spirit and scope of the present invention. For example, the output of the AI ​​pipeline may be used by another downstream computing system to process the anatomical structure and lesion information in the output of the AI ​​pipeline to identify discrepancies with other sources of information (e.g., radiology reports) in order to make clinical staff aware of detection results that might otherwise be overlooked.

[0082] Therefore, the embodiment provides a mechanism for realizing an automated AI pipeline that includes multiple configured and trained ML / DL computer models that implement various artificial intelligence operations at various stages of the AI ​​pipeline to perform the following: identify anatomical structures and lesions associated with these anatomical structures within a volume of input medical images; determine contours associated with such anatomical structures and lesions; determine the classification of such lesions; and generate a list of such lesions, as well as contours of lesions and anatomical structures, for further downstream computer processing of the AI-generated information from the AI ​​pipeline. The operation of the AI ​​pipeline is automated so that there is no human intervention at any stage of the AI ​​pipeline; instead, specially configured and trained ML / DL computer models, trained via machine learning / deep learning computer processes, are employed to perform specific AI analyses at various stages. Human intervention may be present only before input of the volume of input medical images (e.g., during acquisition of a patient's medical image) and after output of the AI ​​pipeline (e.g., when displaying an augmented medical image presented via a computer image display application based on the output of a list of lesions and contours generated by the AI ​​pipeline). Therefore, because AI pipelines are particularly related to improved automated computer tools implemented as artificial intelligence using specific machine learning / deep learning processes that exist only within a computer environment, they perform actions that cannot be performed by humans as intelligent processes and do not organize any human activities.

[0083] Before continuing the description of the embodiments and various aspects of the operation of the improved computer performed by the embodiments, it should first be understood that throughout this description, the term “mechanism” is used to refer to elements of the invention that perform various operations, functions, etc. When the term “mechanism” is used herein, it may be an implementation of the function or aspect of the embodiments in the form of an apparatus, procedure, or computer program product. In the case of a procedure, the procedure is implemented by one or more devices, apparatus, computers, data processing systems, etc. In the case of a computer program product, logic represented by computer code or instructions embodied within or on the computer program product is executed by one or more hardware devices in order to implement a function associated with a particular “mechanism” or to perform an operation associated with a particular “mechanism”. Accordingly, the mechanisms described herein may be implemented as special hardware, software that runs on the hardware and thereby configures the hardware to implement special functions of the present invention that the hardware cannot perform in any other way, software instructions stored on a medium that make the instructions easily executable by the hardware and thereby configure the hardware to specifically perform the functions shown herein and the operations of the particular computer described herein, or any combination thereof.

[0084] In this description and claims, the terms “one,” “at least one of,” and “one or more of” may be used with respect to certain features and elements of the embodiments. It should be understood that these terms and phrases are intended to state that at least one of certain features or elements present in a particular embodiment is present, but there may be two or more. That is, these terms / phrases are not intended to limit the description or claims to a single feature / element present, nor are they intended to require the presence of multiple such features / elements. On the contrary, these terms / phrases require only at least a single feature / element, and there may be multiple such features / elements within the description and claims.

[0085] Furthermore, the use of the term "engine" is intended to describe the embodiments and features of the present invention. When used herein in relation to an engine, it should be understood that it is not intended to restrict any particular implementation to perform, execute, or both, an action, step, process, etc., that is caused by, performed by, or both an engine. An engine may be, but is not limited to, any use of a general-purpose processor or a specialized processor or both, combined with appropriate software loaded into or stored in machine-readable memory and executed by the processor, software, hardware, or firmware, or any combination thereof, that performs a specified function. Furthermore, all names associated with a particular engine are for convenience of reference unless otherwise specified and are not intended to restrict any particular implementation. Furthermore, any function caused by one engine may be incorporated into, combined with, or both combined with, a function of the same or different type of another engine, or distributed across one or more engines in various configurations and performed in the same way by multiple engines.

[0086] In addition, it should be understood that the following description uses several different examples of various elements of the embodiment to further illustrate exemplary implementations of the embodiment and to aid in understanding the mechanism of the embodiment. These examples are intended to be non-limiting and do not cover all possible implementations of the mechanism of the embodiment. In consideration of this description, it will be apparent to those skilled in the art that there are many other alternative implementations of these various elements that can be used in addition to or in place of the examples provided herein without departing from the spirit and scope of the invention.

[0087] The present invention may be a system, method, or computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium (or media) containing computer-readable program instructions for causing a processor to perform an aspect of the present invention.

[0088] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction-executing device. A computer-readable storage medium may be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of further specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random-access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random-access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy(R) disks, mechanically encoded devices such as punch cards or raised structures in grooves on which instructions are recorded, and any suitable combination thereof. When used herein, computer-readable storage media should not be interpreted as themselves being radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmitting media (e.g., light pulses passing through optical fiber cables), or transient signals such as electrical signals transmitted through wires.

[0089] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing device / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). This network may include copper transmission cables, optical transmission fibers, wireless transmitters, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing device / processing device receives computer-readable program instructions from the network and transfers those computer-readable program instructions for storage on a computer-readable storage medium within each computing device / processing device.

[0090] The computer-readable program instructions for performing the operation of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java®, Smalltalk®, and C++, and conventional procedural programming languages ​​such as the C programming language or similar programming languages. The computer-readable program instructions may be executed as a whole on the user's computer, partially as a standalone software package on the user's computer, partially on the user's computer and on a remote computer, or entirely on a remote computer or a server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or the connection may be to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, to carry out aspects of the present invention, an electronic circuit including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions to customize the electronic circuit by utilizing state information of computer-readable program instructions.

[0091] Aspects of the present invention will be described herein by reference to flowcharts or block diagrams, or both, of methods, apparatuses (systems), and computer program products, according to embodiments of the present invention. It will be understood that each block in a flowchart or block diagram, or both, and any combination of blocks contained in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.

[0092] These computer-readable program instructions may be provided to a computer or other programmable data processing device processor to create a machine, so that instructions executed via the processor of the computer or other programmable data processing device may create means to perform functions / operations specified in one or more blocks of a flowchart or block diagram or both. These computer-readable program instructions may be stored on a computer-readable storage medium containing instructions that include products containing instructions to perform modes of functions / operations specified in one or more blocks of a flowchart or block diagram or both, and may be used to instruct a computer, a programmable data processing device, or other device, or a combination thereof, to function in a particular manner.

[0093] Computer-readable program instructions may be read into a computer, other programmable data processing device, or other device to generate a computer implementation process in which instructions executed on a computer, other programmable device, or other device perform functions / operations specified in one or more blocks of a flowchart or block diagram, or both, and cause a series of operable steps to be executed on the computer, other programmable device, or other device.

[0094] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a defined logical function. In some alternative implementations, the functions shown in the blocks may occur in an order different from the order shown in the figures. For example, two consecutively shown blocks may actually be executed substantially simultaneously, depending on the functions they contain, or in some cases, they may be executed in reverse order. Note also that each block in the block diagram or flowchart diagram, or both, and any combination of blocks contained in the block diagram or flowchart diagram, or both, may be implemented by a dedicated hardware-based system that performs a defined function or operation, or a combination of dedicated hardware and computer instructions.

[0095] [Overview of AI pipeline for lesion detection and classification]

[0096] Figure 1 is an exemplary block diagram of a lesion detection and classification artificial intelligence (AI) pipeline (hereinafter simply referred to as the “AI pipeline”) that implements multiple ML / DL computer models specifically configured and trained to perform identification of anatomical structures and lesion detection in input medical image data, according to one embodiment example. For illustrative purposes only, the shown AI pipeline is described in particular as being directed toward liver detection and liver lesion detection in medical image data. As stated above, the embodiment example is not limited to such an embodiment example and may be applied to anatomical structures of any subject and lesions associated with such anatomical structures of any subject, which may be represented by image elements of medical image data captured by medical imaging technology and the corresponding computing system. For example, the mechanism of the embodiment example may be applied to the detection, contour identification, classification, etc., of other anatomical structures such as the lungs and heart, and lesions associated with the lungs, heart, or other anatomical structures of any subject.

[0097] Furthermore, it should be understood that the following description outlines the AI ​​pipeline from the level shown in Figure 1, and that subsequent sections of this description provide additional details regarding the individual stages of the AI ​​pipeline. Each stage of the AI ​​pipeline is implemented as a configured and trained ML / DL computer model, such as a neural network of a deep learning neural network, as represented by symbol 103 in various stages of the AI ​​pipeline 100 in some embodiments. These different ML / DL computer models are specifically configured and trained to perform certain AI operations described herein, such as body part identification, liver detection, phase classification, minimum liver volume detection, liver / lesion detection, lesion segmentation, false positive removal, and lesion classification. These additional sections of the following description describe the AI ​​operations of the various stages. While specific embodiments for implementing various stages of an AI pipeline that offer new technologies, mechanisms, and methods are described, it should be understood that other equivalent technologies, mechanisms, or methods may be used in relation to the AI ​​pipeline as a whole, without departing from the spirit and scope of the examples. These other equivalent technologies, mechanisms, or methods will be obvious to those skilled in the art in light of this description and are intended to be within the spirit and scope of the invention.

[0098] As shown in Figure 1, the artificial intelligence (AI) pipeline 100 receives as input an input medical image volume 105, which is a volume of input computed tomography (CT) medical images represented as one or more data structures in the shown example, according to one embodiment example, and this volume is then automatically processed by various stages of the AI ​​pipeline 100 to finally generate an output 170 which includes classification and contour information, as well as contour information about the anatomical structure of the subject (e.g., the liver in the shown example) and a list of lesions. The input medical image volume 105 may be captured by a medical imaging technique 102 that uses one of many commonly known or later developed medical imaging techniques and equipment that draws images of the internal anatomical structures of a biological entity (i.e., a patient) as one or more medical image data structures. In some embodiments, the input medical image volume 105 includes two-dimensional slices (individual medical images) of parts of the patient's anatomical structure in parts of the patient's body, these slices are then joined together to generate slabs (a combination of axial slices to provide a collection of medical images having thickness along the axis), and these slabs are then joined together to generate a three-dimensional representation (i.e., volume) of the anatomical structure of the body part.

[0099] In the first stage logic 110 of the AI ​​pipeline 100, the AI ​​pipeline 100 determines (112) the part of the patient's body corresponding to the input volume 105 of CT medical image data, and the body part determination logic 114 determines whether this part of the patient's body represents the part of the patient's body corresponding to the anatomical structure of interest (e.g., an abdominal scan, rather than a cranial scan, lower body scan, etc.). This evaluation acts as an initial filter for the use of the AI ​​pipeline 100 only on the input volume 105 of CT medical image data (hereinafter referred to as "input volume" 105), and the AI ​​pipeline 100 is specifically configured and trained to perform anatomical structure identification and contouring, as well as lesion identification, contouring, and classification on the input volume 105 of CT medical image data. This detection of body parts represented within the input volume 105 may refer to metadata associated with the input volume 105, which may include fields specifying the scanned region of the patient's body, which may be specified by the source medical imaging technology computing system 102 when performing the medical image scan. Alternatively, the logic 110 of the first stage of the AI ​​pipeline 100 may implement an ML / DL computer model specifically configured and trained for body part detection 112 that performs medical image classification with respect to specific parts of a patient's body, which performs computerized pattern analysis on medical image data of an input volume 105 and predicts the classification of the medical image data with respect to one or more predetermined parts of the patient's body. In some embodiments, this evaluation may be binary (e.g., whether or not it is a medical image volume of the abdomen) or may be a more complex multi-class evaluation that specifically identifies probabilities or scores with respect to the classification of multiple different body parts (e.g., abdomen, skull, lower limbs, etc.).

[0100] If the body part determination logic 114 of the first stage logic 110 of the AI ​​pipeline 100 determines that the input volume 105 does not represent a part of the patient's body where the target anatomical structure can be detected (e.g., the abdomen where the liver can be detected), processing of the AI ​​pipeline 100 may be interrupted (a case of rejection). If the body part determination logic 114 of the first stage logic 110 of the AI ​​pipeline 100 determines that the input volume 105 represents a part of the patient's body where the target anatomical structure can be detected, processing of the input volume 105 by the AI ​​pipeline 100 is further performed, as will be described later. It should be understood that in some embodiments, there may be multiple different instances of the AI ​​pipeline 100, each configured and trained to process input volumes 105 corresponding to different anatomical structures that may be located in different parts of the patient's body. Therefore, the first stage logic 110 may be provided outside of the AI ​​pipeline 100 and may act as routing logic for routing input volumes 105 to corresponding AI pipelines 100 that are specifically configured and trained to process input volumes 105 of a particular classification, for example, one instance of an AI pipeline for liver and liver lesion detection / classification, another instance of an AI pipeline for lung and lung lesion detection / classification, a third instance of an AI pipeline for heart and cardiac lesion detection / classification, and so on. Therefore, the first stage logic 110 may include routing logic that stores a mapping of which instance of an AI pipeline 100 corresponds to the anatomical structure of different body parts / objects, and can automatically route the input volumes 105 to corresponding instances of an AI pipeline 100 that are specifically configured and trained to process input volumes 105 corresponding to the detected body parts, based on the detection of body parts represented in the input volumes 105.

[0101] Assuming that input volume 105 is detected to represent a part of the patient's body where the target anatomical structure exists (for example, an abdominal scan is present in input volume 105 for the purpose of liver lesion detection), the AI ​​pipeline 100 further processes input volume 105 in a second stage logic 120. This second stage logic 120 includes two initial substages 122 and 124, which may be executed substantially in parallel with each other, sequentially, or both, depending on the desired implementation (Figure 1 shows parallel execution as an example). The two initial substages 122 and 124 include a phase classification substage 122 and an anatomical structure detection substage 124 (for example, a liver detection substage 124 if the AI ​​pipeline 100 is configured to perform liver lesion detection).

[0102] The phase classification substage 122 determines whether the input volume 105 contains a single imaging phase (e.g., pre-contrast phase, angiography phase, portal / venography phase, delayed phase, etc.). Again, the phase classification substage 122 may be implemented as logic that evaluates metadata associated with the input volume 105, which may contain fields specifying the phase of the medical imaging study corresponding to the medical image, and which may be generated by the medical imaging technology computing system 102 when performing medical image acquisition. Alternatively, an example embodiment may implement a configured and trained ML / DL computer model, specifically trained to detect patterns of medical images representing different phases of a medical imaging study, thereby classifying the medical images in the input volume 105 according to which phase the medical image corresponds to. The output of the phase classification substage 122 may be a binary value indicating whether the input volume 105 contains one or more phases, or a classification of each of the phases represented within the input volume 105, which can then be used to determine whether a single or multiple phases are represented.

[0103] If the phase classification indicates that a single phase is present in the input volume 105, the AI ​​pipeline 100 will perform further processing through downstream stages 130-170, as will be explained later. If multiple phases are detected, the input volume 105 may not be further processed by the AI ​​pipeline 100, or it may be filtered into subvolumes, split into subvolumes, or both, as described above, each subvolume containing an image of a single phase, such that only one subvolume corresponding to the target phase is processed by the AI ​​pipeline 100, or the subvolume is routed to a corresponding AI pipeline trained to process input volumes of images corresponding to a particular phase classification, or both. It should be understood that an input volume may be rejected for several reasons (e.g., the liver is not present in the image, it is not a single-phase input volume, there is not enough liver in the image, etc.). Depending on the actual root cause of the rejection, the reason for the rejection may be communicated to the user via a user interface or the like. For example, the output of the AI ​​pipeline 100 in response to a rejection may indicate the reason for the rejection and may be used by a downstream computing system (e.g., a viewer or additional automated processing system) to communicate the reason for the rejection via the output. For example, if no liver is detected in the input volume, the input volume may be silently ignored without communicating the rejection to the user. On the other hand, if the input volume contains a liver but includes input volumes of multiple phases, the rejection may be communicated to the user (e.g., a radiologist) by clearly stating in the user interface generated by the downstream viewer computing system that the input volume was not processed by the AI ​​pipeline 100, so as not to be mistaken for an input volume that contains no detection results, for example, because the input volume contains images of two or more phases.

[0104] A second initial substage 124 is a detection substage for detecting the target anatomical structure (in this embodiment, the liver) in portions of the input volume 105. That is, slices, slabs, etc., within the input volume 105 that specifically indicate the target anatomical structure (liver) are identified and evaluated to determine whether a predetermined minimum amount of the target anatomical structure (liver) is present in these slices, slabs, or as a whole in the input volume. As previously stated, the detection substage 124 includes an ML / DL computer model 125 that is specifically trained and configured to perform computer medical image analysis to identify portions of the input medical image corresponding to the target anatomical structure (e.g., the human liver).

[0105] Therefore, in the liver detection sub-stage 124, the AI ​​pipeline 100 uses the trained ML / DL computer model 125 to determine whether the volume of the input CT medical image contains an image showing a liver. The portion of the volume showing a liver, along with the result of the phase classification sub-stage 122, is passed to the determination sub-stage 126 of the AI ​​pipeline 100, which includes single-phase determination logic 127 and structural minimum volume determination logic 128. The phase determination logic 127 determines whether a single phase exists in the medical image, and the structural minimum volume determination logic 128 determines whether at least a predetermined amount of the anatomical structure of the subject is present in the volume showing the anatomical structure of the subject (e.g., liver). The determination of whether a portion is present is made. As mentioned above, the determination of whether a predetermined amount of the target anatomical structure is present may be made based on a known measurement mechanism, for example, by calculating the size of the structure from the difference in pixel positions in the image, determining the measurement value of the structure from the medical image, comparing these measurements to one or more predetermined thresholds, and determining whether the minimum amount of the target anatomical structure (e.g., liver) is present in the input volume 105 (for example, whether 1 / 3 of the liver is present in the portion of the input volume 105 that has been determined to represent, for example, the liver).

[0106] In one embodiment, the axial scoring mechanism described above may be used to determine whether a predetermined amount of the target anatomical structure (liver) is present in the input volume 105, and the portion of the anatomical structure present in the input volume 105 may be evaluated. As described above, the ML / DL computer model assigns slice scores s to the slice scores of the MSSV slices and MISV slices of the input volume 105, respectively. sup and s inf It may be configured and trained to estimate slice scores s. sup and s inf First, the input volume 105 is divided into sections, and then, for each section, the configured and trained ML / DL computer model is run on the slices of the section to obtain the slice scores s' for the first and last slices within the section. sup and s' inf It may be detected indirectly by estimating s'. sup and s' inf Assuming the estimated value of s, and knowing how the section is located relative to the entire volume of the input CT medical image, extrapolation can be performed. sup and s inf An estimated value is detected. For every 105 input volume, s sup and s inf It should be noted that there are n estimates, where n is the number of sections per volume. In one embodiment, the final estimate is obtained by taking the unweighted average of these n estimates, but in other embodiments, the final estimate may be generated using other functions of the n estimates.

[0107] The volume of input CT medical images sup and s infThe final estimates are determined, and based on these values, fragments of the target anatomical structure (e.g., liver) are calculated. These calculated fragments may then be compared to a predetermined threshold to determine whether a predetermined minimum amount of the target anatomical structure is present in the volume of the input CT medical image (e.g., at least one-third of the liver is present in the volume of the input CT medical image).

[0108] If the determination logics 127 and 128 indicate that multiple phases exist, or that a predetermined amount of the target anatomical structure is not present in the portion of the input volume 105 showing the liver, or both, further processing of the input volume 105 by the AI ​​pipeline 100 for steps 130-170 may be interrupted (i.e., the input volume 105 is rejected). If the determination logics 127 and 128 conclude that the input volume 105 contains a single-phase image and shows at least a predetermined amount of the liver, the portion of the input volume 105 showing the anatomical structure is transferred to the next step 130 of the AI ​​pipeline 100 for processing. In this embodiment, the lower portion of the input volume containing the liver is transferred for further processing, but in other embodiments, background around the liver may be provided, which can be done by adding a predetermined amount of margin above and below the selected liver region. Depending on the amount of background required by subsequent processing operations, the margin may be increased to cover the entire range of the original input volume.

[0109] In the next stage 130 of AI pipeline 100, AI pipeline 100 performs lesion detection on the portion of input volume 105 that represents the target anatomical structure (e.g., liver). This liver / lesion detection stage 130 of AI pipeline 100 uses a population of ML / DL computer models 132-136 to detect the liver and lesions within the liver represented in input volume 105. The population of ML / DL computer models 132-136 performs liver and lesion detection using differently trained ML / DL computer models 132-136, which are trained with a loss function to balance false positives and false negatives in lesion detection. Furthermore, the population of ML / DL computer models 132-136 is configured such that a third loss function forces the outputs of ML / DL computer models 132-136 to match each other.

[0110] In one embodiment, a configured and trained ML / DL computer model 132 is run on input volume 105 to detect the presence of a liver. This ML / DL computer model 132 may be the same as the ML / DL computer model 125 employed in the previous AI pipeline stage 120, and therefore, previously obtained results may be utilized. Multiple (two or more) other ML / DL computer models 134-136 are configured and trained to perform lesion detection on the portion of the medical image of input volume 105 that shows a liver. The first ML / DL computer model 134 is configured and trained to operate directly on input volume 105 and generate lesion predictions. The second ML / DL computer model 136 is constructed using two different decoders that implement two different loss functions: one loss function penalizes errors in false negatives (i.e., classifications that incorrectly indicate the absence of a lesion (normal anatomical structure)), and the other loss function penalizes errors in false positive results (i.e., classifications that incorrectly indicate the presence of a lesion (abnormal anatomical structure)). The first decoder of ML / DL computer model 136 is trained to identify patterns representing a relatively large number of different lesions, at the cost of a large number of false positives. The second decoder of ML / DL computer model 136 is trained to be less sensitive to lesion detection, but detected lesions are very likely to be accurately detected. A third loss function for the population of ML / DL computer models compares the results of the decoders of ML / DL computer model 136 as a whole with each other and forces those results to match each other. The lesion prediction results of the first and second ML / DL computer models 134 and 136 are combined to generate a final lesion prediction for the population, while another ML / DL computer model 132, which generates a prediction of the liver mask, provides an output representing the liver and its contours. Exemplary architectures of these ML / DL computer models 132-136 will be described in more detail later with respect to Figure 6.

[0111] The result of the liver / lesion detection stage 130 of the AI ​​pipeline 100 includes a detection map (e.g., a per-voxel map of the liver lesion detected in the input volume 105) that identifies one or more contours (contour lines) of the liver and portions of medical image data elements corresponding to the detected lesion 135. The detection map is then input to the lesion segmentation stage 140 of the AI ​​pipeline 100. The lesion segmentation stage 140 divides the detection map using watershed techniques and corresponding ML / DL computer models 142, as will be described in more detail later, to generate a division of image elements (e.g., voxels) of the medical images (slice) in the input volume 105. Based on this division, the liver lesion segmentation stage 140 provides other mechanisms, such as an ML / DL computer model 144, which perform actions to identify all contours corresponding to lesions present in the slices of the input volume 105 and identify which contours in three dimensions correspond to the same lesion. The lesion segmentation stage 140 further provides mechanisms, such as an ML / DL computer model 146, which aggregates the interrelated lesion contours to generate a three-dimensional division of the lesion. Lesion segmentation uses inpainting of lesion image elements (e.g., voxels) and non-liver tissues represented within the medical image to focus individually on each lesion and perform active contour analysis. In this way, individual lesions can be identified and processed without biasing the analysis due to other lesions in the medical image or due to parts of the image outside the liver.

[0112] The result of lesion segmentation 140 is a list 148 of lesions containing corresponding contours or contours within the input volume 105. These outputs 148 are provided to the false-positive removal stage 150 of the AI ​​pipeline 100. The false-positive removal stage 150 reduces the detection of false-positive lesions in the list of lesions generated by the lesion segmentation stage 140 of the AI ​​pipeline 100 using a configured and trained ML / DL computer model that employs a dual operating point strategy. The first operating point is selected to be sensitive to false positives by configuring the ML / DL computer model of the false-positive removal stage 150 to remove as many lesions as possible. After the highly sensitive removal of false positives, a determination is made as to whether fewer than a predetermined number of lesions remain in the list. If so, the other lesions in the removed list are re-examined using a second operating point that is relatively less sensitive to false positives. The results of both methods identify the final filtered list of lesions that are further processed by the lesion classification stage of the AI ​​pipeline.

[0113] After removing false positives from the list of lesions and their contours generated by the lesion segmentation stage 140, the resulting filtered list of lesions 155 is provided as input to the lesion classification stage 160 of the AI ​​pipeline 100, where a configured and trained ML / DL computer model is run on the list of lesions and their corresponding contour data, thereby classifying the lesions into one of several default lesion classifications. For example, each lesion and its attributes (e.g., contour data) in the final filtered list of lesions may be input to the trained ML / DL computer model in the lesion classification stage 160, where the trained ML / DL computer model then operates on this data and classifies the lesions as a specific default type or class of lesion.

[0114] Following the classification of lesions by the lesion classification stage 160 of the AI ​​pipeline 100, the AI ​​pipeline 100 generates an output 170 containing a completed list of lesions and their classifications, along with any contour attributes of the lesions. Furthermore, the output 170 of the AI ​​pipeline 100 may also include liver contour information of the liver obtained from the liver / lesion detection stage 130. The output generated by this AI pipeline 100 may be provided to a further downstream computing system 180 for further processing and generation of representations of the target anatomical structure and any detected lesions present in the anatomical structure. For example, a graphic representation of the input volume may be generated in a medical image viewer or other computer application of the downstream computing system 180, and the anatomical structure and detected lesions may be superimposed on the graphic representation or otherwise highlighted using the contour information generated by the AI ​​pipeline. In some embodiments, downstream processing by a downstream computing system 180 may include diagnostic decision support actions and automated medical image report generation based on a detected list of lesions, classifications, and contours. In other embodiments, different treatment recommendations may be generated based on the classification of lesions for physician review and consideration. In some embodiments, lists of lesions, their classifications, and contours may be stored in a historical data structure of the downstream computing system 180 associated with a patient identifier, allowing multiple runs of the AI ​​pipeline 100 for different input volumes 105 associated with the same patient to be evaluated over time. For example, differences between the list of lesions and / or their associated classifications and contours may be determined to assess the progression of the patient's disease or medical condition and to present such information to healthcare professionals to assist in the treatment of the patient. Other downstream computing systems 180 and processing of specific anatomical structures and lesion detection information generated by the AI ​​pipeline 100 in the embodiments may be implemented without departing from the spirit and scope of the invention.

[0115] Figure 2 is an exemplary flowchart illustrating an exemplary operation of an AI pipeline according to one embodiment. The operation outlined in Figure 2 may be implemented by various stages of logic involving configured and trained ML / DL computer models, as described above using specific embodiments shown in Figure 1, which will be described later in separate sections of this description. It should be understood that this operation is particularly intended for automated artificial intelligence pipelines implemented in one or more data processing systems, which include one or more computing devices specifically configured to implement the mechanism of an automated computer tool. The operation outlined in Figures 1 and 2 involves no human intervention other than the creation of volumes of medical images and the use of outputs from downstream computing systems. The present invention particularly provides automated artificial intelligence computing mechanisms improved to perform the described operations, which avoid human interaction and reduce potential errors resulting from previous manual processes by providing new and improved processes that enable the implementation of the improved artificial intelligence computing mechanisms of the present invention in automated computing tools, which are particularly different from any previous manual processes.

[0116] As shown in Figure 2, the operation begins by receiving a medical image input volume from a medical imaging technology computing system (e.g., a computing system that provides computed tomography (CT) medical images) (step 210). The AI ​​pipeline operates on the received input volume to perform body part detection (step 212) and enable it to make a determination as to whether the body part of interest is present in the received input volume (step 214). If the body part of interest (e.g., the abdomen in the case of liver lesion detection and classification) is not present in the input volume, the operation ends. If the body part of interest is present in the input volume, phase classification and minimal anatomical structure assessment are performed either sequentially or in parallel.

[0117] Specifically, as shown in Figure 2, a phase classification is performed on the input volume (step 216) to determine whether the input volume contains a single phase of a medical image (e.g., pre-contrast imaging, partially-contrast imaging, delayed phase, etc.) or a medical image (slice) of multiple phases. Next, a determination is made as to whether the phase classification indicates a single phase or multiple phases (step 218). If the input volume contains a medical image covering multiple phases, the operation ends; otherwise, if the input volume contains a medical image covering a single phase, the operation proceeds to step 220.

[0118] In step 220, detection of the target anatomical structure (e.g., the liver in the example shown) is performed to determine whether a minimum amount of the anatomical structure is present in the input volume, so that subsequent stages of the AI ​​pipeline operation can be performed accurately. A determination is made as to whether a minimum amount of the anatomical structure is present (e.g., at least one-third of the liver is represented in the input volume) (step 222). If no minimum amount is present, the operation ends; otherwise, the operation proceeds to step 224.

[0119] In step 224, liver / lesion detection is performed to generate lesion contours and detection maps. These contours and detection maps are provided to the lesion segmentation logic, which performs lesion segmentation (e.g., liver lesion segmentation in the example shown in the figure) based on these contours and detection maps (step 226). Lesion segmentation results in the generation of a list of lesions and their contours, as well as detection and contour information for anatomical structures (e.g., liver) (step 228). Based on this list of lesions and their contours, a false-positive removal operation is performed on the lesions in the list to remove false positives and generate a filtered list of lesions and their contours (step 230).

[0120] A filtered list of lesions and their contours is provided to a lesion classification logic, which performs lesion classification and generates a completed list of lesions, their contours, and lesion classifications (step 232). This completed list, along with liver contour information, is provided to a downstream computing system (step 234), which may act on this information to generate medical image displays for a medical image viewer application, generate treatment recommendations based on the classification of detected lesions, and evaluate the past progression of lesions over time in the same patient based on a comparison of completed lists of lesions generated by the AI ​​pipeline at different points in time.

[0121] Therefore, as outlined above, the embodiment provides an automated artificial intelligence mechanism and ML / DL computer model that operates on an input volume of medical images and generates a list of lesions, their contours, and classifications while minimizing false positives. The embodiment provides an automated artificial intelligence computer tool that specifically identifies, within a particular set of voxels in an image of an input volume, which of a plurality of voxels corresponds to a portion of the anatomical structure of interest (e.g., liver), and which of those voxels corresponds to a lesion within the anatomical structure of interest (e.g., liver lesion). The embodiment offers a clear improvement over previous methods, both manual and automated, in that the embodiment can be integrated into a fully automated computer tool within a clinician's workflow. In practice, based on the initial stages of the AI ​​pipeline design in the embodiment, which accept only input volumes of a single phase (e.g., abdominal scan) and reject input volumes that do not show the target anatomical structure (e.g., liver) or do not show a predetermined amount of the target anatomical structure (e.g., too little liver), the automated AI pipeline processes only meaningful input volumes, thereby preventing radiologists from wasting valuable manual resources on useless or erroneous results when reviewing input volumes other than the target anatomical structure (e.g., other than liver). In addition to preventing radiologists from receiving large amounts of useless information, the automated AI pipeline in the embodiment also ensures the smooth integration of information technology by avoiding congestion in the AI ​​pipeline and downstream computing systems (such as networks, archives, and verification computing systems) with data associated with data that does not correspond to the target anatomical structure or data associated with data that does not provide a sufficient amount of the target anatomical structure.Furthermore, as described above, the automated AI pipelines of the embodiments enable accurate detection, measurement, and characterization of lesions in a fully automated manner, which is technically made possible by the structure of one or more automated AI pipelines of the embodiments and the corresponding automated ML / DL computer model-based components.

[0122] [ML / DL computer model for detecting the minimum amount of anatomical structure present in an input volume]

[0123] As mentioned above, as part of processing the input volume 105, it is important to ensure that the input volume 105 represents a single phase of a medical image and that at least a minimum amount of the anatomical structure of the subject is represented within the input volume 105. To determine whether a minimum amount of the anatomical structure of the subject is present in the input volume 105, in one embodiment, the decision logic 128 implements a specially configured and trained ML / DL computer model that estimates slice scores to determine the portion of the anatomical structure (e.g., liver) present in the input volume 105. Below, an embodiment of this configured and trained ML / DL computer model is described based on a defined axial scoring technique.

[0124] Figure 3A is an exemplary diagram showing an exemplary input volume (medical image) of the abdomen of a human patient according to one embodiment. The depiction in Figure 3A shows a two-dimensional representation of a three-dimensional volume. Slices are horizontal lines within the two-dimensional representation shown in Figure 3A, but are represented as planes extending within or outside the page, or both, representing flattened two-dimensional slices of the human body, and the stacking of these planes constitutes a three-dimensional image.

[0125] As shown in Figure 3A, the embodiment defines axial scores for slices ranging from 0 to 1. The axial score is defined such that the slice corresponding to the first slice containing the liver (FSL) has a slice score of 0, and the slice corresponding to the last slice containing the liver (LSL) has a slice score of 1. In the example shown in the figure, the first and last slices are defined in relation to the lowest (MISV) and highest (MSSV) slices in the volume, respectively, where lower and upper are determined along a specific axis of the volume (e.g., the y-axis in the example shown in Figure 3A). Thus, in this example, the MSSV is the slice with the highest y-axis value, and the MISV is the slice with the lowest y-axis value. For example, the MISV may be closest to the lower extremities of the biological entity (e.g., the subject's feet), and the MSSV may be closest to the upper extremities of the biological entity (e.g., the subject's head). The FSL is the slice showing the anatomical structure of the subject (e.g., the liver) that is relatively closest to the MISV. The LSL is the slice showing the anatomical structure of the subject that is relatively closest to the MSSV. In one embodiment, a trained ML / DL computer model (e.g., a neural network) may assign axial scores by taking a chunk of slices as input and outputting the height (axial score) of the central slice within the chunk. This trained ML / DL computer model is trained using a cost function that minimizes the error (e.g., least squares error) in the actual height. This trained ML / DL computer model is then applied to all chunks covering the input volume. (In some cases, there is some overlap between chunks.)

[0126] Pairs of slice scores corresponding to the slice scores of MSSV slices and MISV slices (s sup and s inf The estimated liver axial score (LAE) is defined by ). The ML / DL computer model of the judgment logic 128 in Figure 1 uses the slice scores s of input volume 105.sup and s inf The mechanism of this embodiment is specifically configured and trained to determine these slice scores and can identify liver fragments within the field of view of the input volume 105.

[0127] In some embodiments, slice scores sup and s inf First, the input volume 105 is divided into sections (for example, sections containing X slices (for example, 20 slices)), and then, for each section, the configured and trained ML / DL computer model is run on the slices of the section to obtain the slice scores s' of the first and last slices within the section. sup and s' inf This can be indirectly detected by estimating the first and last slices, and the "first" and "last" slices may be determined according to the direction of travel along the axes of the 3D volume 105 (for example, the direction of travel from the first slice to the last slice along the y axis, from the slice with the smallest y-axis value to the slice with the largest y-axis value). sup and s' inf Assuming the estimated values, and knowing how the sections are located within the entire volume 105, extrapolation can be performed to s sup and s inf An estimated value is detected for each volume. sup and s inf It should be noted that there are n estimates, where n is the number of sections per volume. In one embodiment, the final estimate is obtained by taking the unweighted average of these n estimates, but in other embodiments, the final estimate may be generated using other functions of the n estimates.

[0128] For example, Figure 3B shows the corresponding axial score s' inf and s' supFigure 3B shows another representation of the input volume, including sections of slices represented together. As shown in Figure 3B, the sections are defined in this example as 20 slices separated by 5 mm. For each section of 20 slices of the volume, the slice score s' is calculated by the ML / DL computer model. sup and s' inf These values ​​are estimated and are based on a specific range (e.g., a range of 0 to 1, a range of -0.5 to 1.2, or any other desirable default range suitable for a particular implementation). sup and s' inf From extrapolating the values, s sup and s inf This is obtained. In this example, assuming the default range of -0.5 to 1.2, the application of the ML / DL computer model and extrapolation results in s sup It is estimated that s is approximately 1.2, inf If it is estimated to be -0.5, these values ​​indicate that the entire liver is included in the volume. Similarly, s sup It is estimated that s is 1.2, inf If it is estimated to be 0.5, these values ​​indicate that the volume contains approximately the top 50% of the axial range of the liver, for example the coverage is (1.2-0.5) / (1.2-(-0.5))=0.41. As an additional example, the liver starts at -2.0 and ends at 0.8 (i.e., s sup It is estimated that s is 0.8, inf In another embodiment (where is estimated to be 2.0), the upper limit of the liver is lower than 1.2, so the liver is cut off at the top, and the lower limit is lower than -0.5, so the lower part of the liver is completely covered. This indicates that approximately 80% of the lower axial range of the liver is included in the volume, i.e., the coverage is (0.8 - max(-2, -0.5)) / (1.2 - (-0.5)) = 0.76.

[0129] Figure 3C is an illustrative diagram of the input volume of Figure 3A, in which the volume is axially divided into n completely overlapping sections. In the example shown in the figure, there are seven sections indicated by arrows. It should be noted that in this example, the last two sections (arrows at the top of the figure) are almost completely identical. As previously mentioned, the ML / DL computer model allows for s' for each of these sections. sup and s' inf The values ​​of MSSV slices and MISV slices are estimated. sup and s inf Used to extrapolate the value of s sup and s inf The value can then be used to determine the amount of anatomical structure of the object present in input volume 105.

[0130] Therefore, MSSV and MISV sup and s inf The value is obtained by first dividing the input volume 105 into sections, and then for each section, the slice score of the first and last slice within the section is s' sup and s' inf It is detected indirectly by estimating . Based on these estimates, we know how the section is located relative to the entire input volume 105, so s sup and s inf The value of is estimated by extrapolation. sup and s inf There are n estimates, where n is the number of sections per volume. The final estimate may be obtained by any suitable combinatorial function that evaluates the n estimates, such as an unweighted average of the n estimates or any other suitable combinatorial function.

[0131] Figures 4A-4C show the section of the medical image input volume according to one embodiment example. sup and s' infFigures 4A–4C illustrate an exemplary diagram of one embodiment of an ML / DL computer model trained and configured to estimate the value of . The ML / DL computer models in Figures 4A–4C are merely examples of ML / DL computer model architectures, and many modifications to this architecture may be made without departing from the spirit and scope of the invention, such as changing the tensor size of the input slices of the input volume, changing the number of nodes in the layers of the ML / DL computer model, or changing the number of layers. Those skilled in the art will recognize, in consideration of this description, how to modify the exemplary ML / DL computer models to a desired implementation.

[0132] As shown in Figure 4A, a series of 20 slices representing section 410 or "slab" of input volume 105 are provided as input to processing blocks (PB) 420-430. In the example embodiment shown, PB420-430 are blocks of logic that combine convolutional and LSTM layers, as shown in Figures 4B and 4C. Features are extracted from the convolutional layers of PB420,430, and these features are then fed as input to the LSTM layers of PB420,430. This is a kind of clever / lightweight modeling of the fact that slices have a specific order within an anatomical region or anatomical structure of the subject (e.g., abdomen / liver) and are driven by the biological structure (e.g., the relative positions of the liver, kidney, heart, etc., in addition to the biological structure of the liver itself). In the example shown in the figure, first, the tensor size of the 20 slices 410 is 128x128 in this example. The first processing block 420, in this embodiment, reduces the size of the tensor to one-eighth to generate sections of 20 slices (it should be understood that the number of slices in a section is implementation-specific and can be changed without departing from the spirit and scope of the invention), and these slices have dimensions of 16x16x32, where 32 is the number of filters. The second processing block 430 converts the slices of the input section into sections of 20 slices, each containing a slice with dimensions of 2x2x64, where 64 is the number of filters. The subsequent neural network 440, configured using flattening layers, high-density layers, and linear layers, estimates the input section 410 of the input volume 105 s' sup and s' inf It is configured and trained to generate the following. Figure 4B shows the composition of a processing block (PB) with respect to convolutional and LSTM layers according to one embodiment, and Figure 4C shows each of these convolutional and LSTM layers of each PB as an example.

[0133] As an example, in the ML / DL computer model architecture shown in Figures 4A-4C, during training of this ML / DL computer model, in one embodiment, medical image data (e.g., medical digital image and communication (DICOM) data) is input volume (e.g., containing values ​​in Hounsfield units (HU)). i It is assembled into a three-dimensional array with a size of ×512×512 (32-bit floating point) (the value in Houndfield units is a normalized physical value that indicates the X-ray attenuation of the material presented at a particular location (e.g., voxel)). i is the number of slices in the i-th volume, where i is in the range of 0 to N-1, and N is the total number of volumes. Each input volume is processed by the body site detector, and as previously described, an approximate region corresponding to the abdomen (in the case of liver detection) is extracted. The abdomen is defined, for example, as a continuous region between axial scores -30 and 23 from the body site detector. Slices outside this continuous region are rejected, and ground truth may be defined as the locations of appropriately adjusted FSL and LSL. For example, assuming the range of the input volume is a to b, the input volume is rejected if there is no overlap between [a:b] and [-30, 23]. In other words, if b > 23 or a < -30.

[0134] The input section 410 or "slab" is sliced ​​again into predetermined slice separations (e.g., 5 mm). The input section 410 is reshaped to 128x128 in x, y dimensions, resulting in shape M. i We obtain N sections of x128x128, resulting in 410. This is called downsampling of the data in the input volume. Because ordering the slices in the input volume relies on coarse information (e.g., organ size), the AI ​​pipeline still works well with the downsampled data, and both the processing time and training time of the AI ​​pipeline are improved because the size of the downsampled data is reduced.

[0135] Input section 410 containing fewer than the default number of slices (e.g., 20) or having a pixel size less than the default pixel size (e.g., 55 mm) is rejected, and N' M i A x128x128 section is obtained. The values ​​within the section are clipped and normalized using a linear transformation from the acquired range (e.g., -1024, 2048) to the range (0, 1). At this point, N' M values ​​have been processed. i The x128x128 section constitutes the training set as described above, and the neural network 440 uses this training set as input to the s' section. sup and s' inf It is trained to generate estimates of [the value].

[0136] Regarding performing inference using the trained neural network 440, the process involves body part detection, selection of slices corresponding to the target body part, reslicing, and reshaping. The above operations for processing input volume 105 by completing, rejecting specific sections that do not meet the default requirements, and generating clipped and normalized sections are performed again for new sections of input volume 105. After generating clipped and normalized sections, input volume 105 is divided into R-ceil(M-10) / 10 subvolumes or sections containing 20 slices, thereby generating a division of slices containing overlapping chunks. For example, if there is a volume with N'=31 slices (slices numbered 0-30), three sections or subvolumes are defined containing overlapping slices 0-19, 10-29, and 11-30. Sections or subvolumes typically have at least about 50% overlap.

[0137] Therefore, an ML / DL computer model is provided, assuming a defined axial score range of 0 to 1, and a predetermined number of slices (medical images) are provided for the sections of the volume. sup and s' infBased on the estimated value of the input volume s sup and s inf The system is configured and trained to estimate the values ​​of . From these estimates, a determination can be made as to whether the input volume contains medical slices that together constitute at least a predetermined amount of the target anatomical structure (e.g., liver). This determination may be part of the determination logic 128 of the AI ​​pipeline 100 to determine whether a sufficient representation of the anatomical structure is present in the input volume 105 in order to enable accurate liver / lesion detection, lesion segmentation, etc., at further downstream stages of the AI ​​pipeline 100, as described above.

[0138] Figure 5 is a flowchart illustrating the exemplary operation of the AI ​​pipeline's liver detection and predetermined amount of anatomical structure determination logic according to one embodiment. As shown in Figure 5, the AI ​​pipeline's liver detection operation begins with receiving an input volume (step 510), and dividing the input volume into multiple overlapping sections of a predetermined number of slices per section (step 520). The slices for each section are fed into a trained ML / DL computer model that estimates the axial scores of the first and last slices within each section (step 530). The axial scores of the first and last slices are used to extrapolate to the scores of the lowest-level slice (MISV) and highest-level slice (MSSV) in the input volume (step 540). This yields multiple estimates of the axial scores of the MISV and MSSV, which are then combined by a function of the individual estimates to generate an estimate of the axial scores of the MISV and MSSV in the input volume (e.g., a weighted average) (step 550). Based on the estimated axial scores of MISV and MSSV, the axial scores are compared to a criterion for determining whether a predetermined amount of the target anatomical structure (e.g., liver) is present in the input volume (step 560). The operation then terminates.

[0139] [Liver & Lesion Detection]

[0140] As described above, assuming that input volume 105 is determined to represent a single phase and that input volume 105 is determined to represent a predetermined amount of target anatomical structure in a slice of input volume 105, liver / lesion detection is performed on the portion of input volume 105 containing the target anatomical structure. In one embodiment, the liver / lesion detection logic of stage 130 of AI pipeline 100 employs a configured and trained ML / DL computer model (which, in some embodiments, may be the same ML / DL computer model 125 used in stage 120 for liver detection) that operates to detect the target anatomical structure (e.g., liver) in a slice of input volume 105. The liver / lesion detection logic of stage 130 of AI pipeline 100 also includes a group of several other ML / DL computer models that are configured and trained to detect lesions in images of the target anatomical structure (liver).

[0141] Figure 6 is an illustrative diagram of a population of ML / DL computer models used to perform lesion detection in a target anatomical structure (e.g., liver) according to one embodiment. The population of ML / DL computer models 600 includes a first ML / DL computer model 610 for detecting a target anatomical structure, e.g., liver, and generating a corresponding mask. The population of ML / DL computer models 600 also includes a second ML / DL computer model 620, which is configured and trained to process a liver mask input and generate lesion predictions using two computing loss functions implemented in two decoders of the second ML / DL computer model 620. One loss function is configured to penalty false positive errors (resulting in low sensitivity but high accuracy), and the other is configured to penalty false negative errors (resulting in high sensitivity but low accuracy). In Figure 6, an additional loss function called consistency loss 627 is employed in the second ML / DL computer model 620 to make the outputs generated by the two computing decoders similar (consistent) to each other. The collection of ML / DL computer models further includes a third ML / DL computer model 630 that is trained and configured to directly process the input volume 105 and generate lesion predictions.

[0142] As shown in Figure 6 and described above, the population 600 specifically includes a first configured and trained ML / DL computer model 610, which is trained to identify the anatomical structures of a target in an input medical image. In some embodiments, this first ML / DL computer model 610 includes a U-Net neural network model 612, which is trained to perform image analysis and detect the liver in a medical image. However, it should be understood that the embodiments are not limited to this particular neural network model capable of performing segmentation, and any ML / DL computer model may be used without departing from the spirit and scope of the invention. U-Net is a convolutional neural network developed for medical bioimage segmentation at the Faculty of Computer Science, University of Freiburg, Germany. The U-Net neural network is based on a fully convolutional network whose architecture has been modified and extended to work with fewer training images and to result in more accurate segmentation. U-Net is well known in the art and will not be described in further detail herein.

[0143] As shown in Figure 6, in one embodiment, the first ML / DL computer model 610 can be trained to process a predetermined number of slices when it is determined, for example, through an empirical process, that three slices yield good results, once it is determined that the number is appropriate for a desired implementation. For example, other implementations may use different slice dimensions without departing from the idea and scope of the embodiment, but in one embodiment, the slices of this input volume were 512x512 pixel medical images. U-Net generates anatomical segmentation in the input slices, resulting in one or more segments corresponding to the anatomical structure of interest, for example, the liver. As part of this segmentation, the first ML / DL computer model 610 generates a segment corresponding to a liver mask 614. This liver mask 614 is provided as input to at least one of the other ML / DL computer models 620 in population 600 in order to focus on processing by the ML / DL computer model 620 only on the portion of the input slice of the input volume 105 corresponding to the liver. By preprocessing the input to the ML / DL computer model with the liver mask 614, this processing by the ML / DL computer model can be focused on the portion of the input slice corresponding to the anatomical structure of the subject, and not on the "noise" in the input image. Other ML / DL computer models, such as ML / DL computer model 630, receive the input volume 105 directly without liver masking, using the liver mask 614 generated by the first ML / DL computer model 610.

[0144] In the embodiment of the shown population 600, the third ML / DL computer model 630 consists of encoder sections 634-636 and a decoder section 638. The ML / DL computer model 630 receives nine slice slabs from input volume 105, which are then divided into groups 631-633 of three slices each, in this case each group 631-633 being input to the corresponding encoder networks 634-636. Each encoder 634-636 is a fully connected, headless convolutional neural network (CNN:) such as DenseNet-121 (D121), which is pre-trained to recognize different types of objects (e.g., lesions) in the input slices and output a classification output indicating the detected type of object in the input slice, for example, as an output classification vector. CNN634-636 can, for example, work on three channels of an input slice, and the resulting output features of CNN634-636 are provided to concatenated NHWC logic 637, where NHWC refers to the number of images in a batch (N), image height (H), image width (W), and number of channels in the image (C). The original DenseNet network architecture includes many convolutional layers and skip-true connections that de-resolution the three full-resolution slice inputs into many feature channels of lower resolution. From there, a fully concatenated head aggregates all the features and maps them to multiple classes in the final output of DenseNet. Since the DenseNet network is used as an encoder in the shown architecture, this head is removed, and only the de-resolution features are retained. As a result, the concatenated NHWC logic 637 concatenates all feature channels so that they are forwarded to a decoder stage 638, which is responsible for high-resolution processing of the image until a desired output probability map resolution (e.g., 512x512) is reached.

[0145] Encoders 634-636 share the same parameters optimized throughout the training process, such as weights, sampling for lesion types during training, weights for loss, and scaling type. Training the ML / DL computer model 630 uses two different loss functions. The primary loss function is an adaptive loss specifically configured to penalize false positive errors in slices without lesions in the ground truth and false negative errors in slices with lesions in the ground truth. The loss function is a revised Tversky loss as follows: For each output slice: TP=sum(prediction*target)FP=sum(1-target)*prediction)FN=sum((1-prediction)*target)LOSS= 1 - ((TP+1) / TP+1+alpha*FN+beta*FP)) where “prediction” is the output probability of the ML / DL computer model 630, and “target” is the ground truth lesion mask. The output probability value is in the range of 0 to 1. Target is a value of either 0 or 1 for each pixel in the slice. In slices where there are no lesions, the term “alpha” is small (e.g., zero) and “beta” is large (e.g., 10). In slices where there are lesions, the term “alpha” is large (e.g., 10) and “beta” is small (e.g., 1).

[0146] The second loss function 639 is a function connected to the outputs of encoders 634-636. Since the input to this loss comes from the middle of the ML / DL computer model 630, it is also called "deep monitoring" 639. Deep monitoring is known to compel the encoder neural networks 634-636 to learn better representations of the input data during training. In one embodiment, this second loss is the simple mean squared error in predicting whether a slice has a lesion or not. Therefore, a mapping network is used to map the output features of encoders 634-636 to nine values ​​between 0 and 1, corresponding to the probability of having a lesion in each of the nine slice inputs. Decoder 638 generates an output specifying a probability map for lesions detected in the input image.

[0147] The second ML / DL computer model 620 receives from the input volume a pre-processed input of three slices, which have been pre-processed by the liver mask 614 generated by the first ML / DL computer model 610 to identify portions of the three slices corresponding to the liver mask 614. The resulting pre-processed input slices (192x192x3 in the illustrated embodiment shown) are provided to the second ML / DL computer model 620, which has a DenseNet-169 (D169) encoder 621 connected to two decoders (representing 2D DEC, where the decoders consist of two-dimensional neural network layers). The D169 encoder 621 is a neural network feature extractor widely used in computer vision applications. It consists of a series of convolutional layers, where features extracted from each layer are connected in a feed-forward manner to any other layer. The features extracted by encoder 621 are transferred to two independent decoders 622 and 623, each consisting of a two-dimensional convolutional layer and a high-resolution layer (referred to as 2D DEC in Figure 6). Each decoder 622 and 623 is trained to detect lesions (e.g., liver lesions) in the input slice. Both decoders 622 and 623 are trained to perform the same task, namely lesion detection, but the key difference in their training is that the two decoders 622 and 623 each utilize different loss functions to drive detection training in two competing directions, as previously and further described. The final detection map of the second ML / DL model 620 is combined with the final detection map of the third ML / DL model 630 using an averaging operation 640. This procedure is applied across the entire input slab of the input volume 105 to generate the final detection map (e.g., liver lesions).

[0148] As mentioned above, a second ML / DL computer model 620, which attempts to achieve the opposing detection operating point performance, is trained using two different loss functions. That is, one encoder 622 uses one loss function for training, which penalizes the error in false-negative lesion detection, thereby producing high-sensitivity detection with relatively low accuracy, while the other encoder 623 uses another loss function for training, which penalizes the error in false-positive lesion detection, resulting in low-sensitivity but high-accuracy detection. An example of these loss functions is the Focal Tversky Loss, whose parameters are tuned to high or low penalties for false positives and false negatives according to one embodiment (see Abraham et al., "A Novel Focal Tversky Loss function with Improved Attention U-Net for Lesion Segmentation," arXiv:1810.07842[cs], October 2018). A third loss function, the consistency loss 627, is used to enforce consistency between the predictive detections of each decoder 622, 623. The consistency loss logic 627 compares the outputs 624, 625 of the two encoders 622, 623 with each other and enforces that these outputs are similar to each other. This loss could be, for example, the mean squared error loss between two predictive detections, a structural similarity loss, or any other loss that enforces consistency / similarity between the compared predictive detections.

[0149] At runtime, these opposing operating point encoders 622, 623 are used to generate two lesion outputs 624, 625, which are then fed into a slice averaging (SLC AVG) logic 623 that generates an average of the lesion outputs. The average of these lesion outputs is then sampled again to produce an output that is dimensionally balanced with the output of a third ML / DL computer model 630 for comparison (note that this process involves reversing the liver masking operation, thereby calculating the lesion outputs at the original 512x512x3 resolution).

[0150] At runtime, the Slice Averaging (SLC AVG) logic 626 operates to generate the final detection map of the ML / DL model 620 for the lesion prediction outputs 624 and 625 of the encoders 622 and 623. While consistency loss 627 was applied during training to drive each decoder 622 and 623 to learn consistent detections, it should be understood that at runtime, when this consistency loss is no longer used, the ML / DL model 620 instead outputs two detection maps that need to be aggregated by the SLC AVG module 626. The results of the SLC AVG logic 626 are sampled again to produce an output with dimensions that balance with the input slab (512x512x3). All detections of the ML / DL model 620 generated for each slab of the input volume 105 are combined with detections of the ML / DL model 630 generated via the Volume Averaging (VOL AVG) logic 640. This logic calculates the average value of the two detection masks at the voxel level. This result is the Final Leision mask 650 corresponding to the lesion detected at input volume 105.

[0151] Thus, after training the ML / DL computer models 620 and 630, when presented with a new slice of a new input volume 105, the first ML / DL computer model 610 generates a liver mask 614 to preprocess the input to the second ML / DL computer model 620, and the two ML / DL computer models 620 and 630 process the input slice to generate lesion predictions that are averaged over that volume by a volume averaging logic 640. The result is a final lesion output 650 with a liver mask output 660 based on the operation of the first ML / DL computer model 610. These outputs may be provided as outputs of the liver / lesion detection logic stage 130 of the AI ​​pipeline 100, which are provided to the lesion segmentation logic stage 140 of the AI ​​pipeline 100, as previously described above and as will be described in more detail below. Thus, the mechanism of the embodiment provides a population 600 method for anatomical structure recognition and lesion detection in an input volume 105 of medical images (slices).

[0152] The population architecture shown in Figure 6 achieves improved performance that surpasses the use of a single ML / DL computer model. Specifically, it is observed that the improved detection specificity at the same sensitivity level as a single ML / DL computer model is achieved through the combined detection outputs of multiple ML / DL computer models in the population. In other words, when the detection outputs of ML / DL models 620 and 630 are averaged out due to errors (false positives) committed by the models at various locations, signals from true positive lesions become dominant while signals from false positives are reduced, leading to improved performance.

[0153] Figure 7 is a flowchart outlining the exemplary operation of liver / lesion detection logic in an AI pipeline according to one embodiment. As shown in Figure 7, the operation begins by receiving an input volume (step 710) and performing anatomical structure detection, e.g., liver detection (step 720), using a first trained ML / DL computer model, such as a U-Net computer model, which is configured and trained to identify anatomical structures (e.g., liver). The result of the anatomical structure detection, which identifies masks for anatomical structures (e.g., liver masks), is segmentation of the input volume (step 730). This input volume is also processed via a first trained ML / DL computer model of a population, which is specifically configured and trained to perform lesion detection (step 740). The first trained ML / DL computer model generates a first set of lesion detection prediction outputs based on its processing of the input volume (step 750).

[0154] The second trained ML / DL computer model of the population receives a masked input generated by applying a generated anatomical structure mask to the input volume, thereby identifying the portion of the medical image corresponding to the anatomical structure of interest in the input volume (step 760). The second trained ML / DL computer model processes the masked input through two different decoders having two different, competing loss functions, e.g., one that penalizes error in false-positive lesion detection and another that penalizes error in false-negative lesion detection (step 770). The result is two sets of lesion prediction outputs, which are then combined through coupling logic to produce lesion prediction outputs for the second DL / ML computer model (step 780). If necessary, the second lesion prediction outputs are sampled again and combined with the first lesion prediction output generated by the first ML / DL computer model of the population to produce a final lesion prediction output (step 790). Thus, the final lesion prediction output is output with the anatomical structure mask (step 795), and the operation ends.

[0155] [Lesion Segmentation]

[0156] As mentioned above, the lesion prediction output is generated through the operation of various ML / DL computer models and through logical stages of the AI ​​pipeline, including body part detection, target body part determination, phase classification, target anatomical structure identification, and anatomical structure / lesion detection. For example, in the AI ​​pipeline 100 shown in Figure 1, the result of the liver / lesion detection stage 130 of the AI ​​pipeline 100 includes a detection map that identifies one or more contours (contour lines) of the liver, and furthermore, a portion of the medical image data elements corresponding to the detected lesion 135, for example, a voxel-by-voxel map of the liver lesion detected in the input volume 105. Thus, the detection map is used in the AI ​​pipeline 10 The lesion is entered into segmentation stage 140, which is 0.

[0157] As mentioned above, in the lesion segmentation logic, for example in lesion segmentation stage 140 in Figure 1, the detection map is segmented to generate image element segmentation of the medical image (slice) of the input volume using watershed techniques and corresponding ML / DL computer models. The liver lesion segmentation stage also provides other mechanisms, such as one or more other ML / DL computer models, which identify all contours corresponding to lesions in the slices of the input volume based on image element segmentation and perform the operation of identifying which contours correspond to the same lesion in three dimensions. This lesion segmentation stage further provides mechanisms, such as one or more further ML / DL computer models, which aggregate the interrelated lesion contours to generate a three-dimensional lesion segmentation.

[0158] In lesion segmentation, inpainting of lesion image elements and non-liver tissue represented in the medical image are used to focus on each lesion individually and perform active contour analysis. In this way, individual lesions can be identified and processed without biasing the analysis due to other lesions in the medical image or due to parts of the image outside the liver. The result of this lesion segmentation is a list of lesions by their corresponding contour lines or contours in the input volume.

[0159] Figure 8 shows a block diagram illustrating an overview of the lesion segmentation process performed by lesion segmentation logic according to one embodiment. As shown in Figure 8, lesion segmentation involves a mechanism for 2D detection, i.e., dividing the detection of lesions in a 2D slice slice by slice (block 810), connecting the 2D lesions along the z-axis (block 820), and refining the contour slice by slice (block 830). Each of these blocks will be described in more detail below with reference to subsequent drawings. The segmentation process shown in Figure 8 is performed as a process for identifying all lesions in a given input volume under analysis and for distinguishing lesions that are close to another lesion in the image (slice) of the input volume. For example, two lesions that appear to merge in pixel time in one or more images may need to be identified as two different regions, or distinguishable lesions, for the purpose of other downstream processing of the detected lesions, such as lesion classification and distinguishing them as separate lesions in the output of a list of lesions for downstream computing system operation, such as providing a medical viewing application, performing treatment recommendation actions, and performing decision support actions.

[0160] As part of the segmentation of 2D images per slice in block 810, the mechanism of the embodiment uses existing watershed techniques to segment detection maps from previous lesion detection stages of the AI ​​pipeline, for example, detection map 135 generated by the liver / lesion detection logic 130 of AI pipeline 100 in Figure 1. The watershed algorithm requires seed definition to perform mask segmentation. This results in dividing the mask into as many regions as possible, with seeds such that there is approximately one seed exactly in the center of each region, as shown in Figures 10A and 10C. In automated segmentation, seeds within a mask can be obtained as the maximal values ​​of its distance map (distance to the mask contour). However, such methods are prone to noise and can lead to too many seeds, thereby over-segmenting the mask. Therefore, the segmentation needs to be modified by regrouping some of the regions. Given the empirical observation that most lesions are bubble-shaped, the guiding principle for regrouping regions is to result in new, nearly circular regions. For example, in the case of the mask shown in Figure 10C, the mechanism merges two regions identified by seeds 1051 and 1061, respectively, resulting in a new mask segmentation consisting of only two nearly circular regions. Thus, in the case of detected lesions defined in detection map 135, such as the lesion shown on the left side of Figure 9, which will be described later, they can be segmented into several bubble-shaped lesions, as shown on the right side of Figure 9. These will be considered cross-sections of the 3D lesion on the slice.

[0161] Watershed segmentation is a region-based method whose origins lie in mathematical morphology. In watershed segmentation, an image is viewed as a local landscape with ridges and valleys. The elevation values ​​of this landscape are typically defined by the grayscale values ​​of each pixel or the magnitude of their gradients, thereby viewing the two-dimensional representation as a three-dimensional representation. In watershed transformation, the image is decomposed into "catchment basins." For each local minimum, the catchment basin contains all the points where its steepest descent ends at its minimum. Watersheds divide the basins from one another. In watershed transformation, the image is completely decomposed, and each pixel is assigned to either a region or a watershed.

[0162] Watershed segmentation requires the selection of at least one marker, also known as a “seed” point, within each object in the image. The seed point may be selected by an operator. In one embodiment, the seed point is selected by an automated procedure that takes into account the object’s specific use-specific knowledge. Once the objects are marked, they can be grown using morphological watershed transformation, as described in more detail below. Lesions are typically “bubble” in shape. This embodiment provides a technique for fusing watershed-segmented regions based on this assumption.

[0163] Next, in block 820, the mechanism of the embodiment aggregates the voxel divisions for each slice along the z-direction so as to generate a three-dimensional output. Therefore, the mechanism needs to determine whether two sets of image elements in different slices, for example, voxels, belong to the same lesion, i.e., whether they are aligned in three dimensions. Based on the intersection and union of the lesions, the mechanism calculates measurements between lesions in adjacent slices and applies a regression model to determine whether two lesions in adjacent slices are part of the same region. Each lesion can be viewed as a set of voxels, and the mechanism determines the intersection of two lesions as the intersection of two sets of voxels, and the union of two lesions as the union of two sets of voxels.

[0164] This results in three-dimensional segmentation of the lesion, but the contours may not match the actual image well. There may be over-segmented lesions. In the embodiment, we propose using active contour trimming, a conventional framework, to address the segmentation problem. Such an algorithm requires iteratively trimming the contours to gradually match them well to the image data, ensuring that it maintains several desirable properties, such as shape smoothness, during this time. In block 830, the mechanism of the embodiment starts the segmentation with active contours obtained from the first stage 810 and the second stage 820, focusing on one lesion at a time, otherwise active contouring or random segmentation methods in operation on similar lesions could result in them merging again into a single contour, which would be unproductive as it would effectively negate the benefits obtained by the previous segmentation stages. This mechanism focuses on one lesion and performs "inpainting" on lesion voxels and non-liver tissue near the lesion under focus.

[0165] This chain of three processing steps allows for processing that eliminates bias caused by other lesions in the image, or by pixels outside the liver, i.e., lesions.

[0166] [Splitting 2D detection per slice]

[0167] Figure 9 shows the results of lesion detection and slice-by-slice segmentation according to one embodiment. As seen on the left side of Figure 9, a lesion region 910 is detected through the AI ​​pipeline process described above up to that point and can be defined in the output of a contour and detection map, e.g., 135 in Figure 1, from lesion detection logic, e.g., 130 in Figure 1. According to one embodiment, the logic of block 810 in Figure 8 attempts to segment this region into three lesions 911, 912, and 913, as shown on the right side of Figure 9. The segmentation mechanism in this embodiment is based on existing watershed techniques that work to segment detection maps from the preceding lesion detection stage of the AI ​​pipeline. Watershed algorithms are used primarily in image processing for segmentation purposes. The underlying principle behind these known watershed algorithms is that a grayscale image can be viewed as a geographical surface where high intensity means peaks and hills, while low intensity means valleys. Watershed techniques begin by filling each isolated valley (local minimum) with water of a different color (label). As the water level rises according to the nearby peaks (gradients), water from valleys of different colors begins to merge. To avoid this, partitions are built where the waters merge. The process of filling with water and building partitions continues until all peaks are submerged at the point when the created partitions give the segmented result. Again, since watershed techniques are generally known, they will not be described in detail herein. Any known technique can be used to divide a 2D image into slices without departing from the spirit and scope of the present invention.

[0168] In the context of lesion segmentation, the empirical observation that most lesions are circular strongly suggests that a segmentation resulting in a set of round regions is likely a good segmentation. However, as previously mentioned, the quality of watershed segmentation depends on the quality of the seeds. In fact, any set of seeds does not necessarily lead to a set of round regions. For example, Figure 10C shows a watershed segmentation induced by three seeds containing only one nearly circular region. The other two regions are not circular. However, their union is also nearly circular. Such a configuration is called over-segmentation, as the diagonal segmentation in the figure divides the other circular region into two smaller non-circular regions. Therefore, it is desirable to have an algorithm that can correct for over-segmentation. The seed relabeling mechanism does this by fusing some over-segmented regions to form a coarser segmentation containing only round regions. For example, this mechanism determines that for the segmentation in Figure 10C, a new, more circular region is created by fusing the two regions identified by seeds 1051 and 1061.

[0169] In this embodiment, regions are merged during the division process to create a larger, more rounded region that can correspond to a physical lesion. This division divides an area into smaller regions, or, as described herein, divides a mask into smaller regions. In terms of contours, the division generates a set of smaller contours from a larger contour (see left and right in Figure 9).

[0170] The seed is obtained by extracting maxima from a distance map calculated from the input mask to be segmented. The map measures the Euclidean distance to the mask contour for each pixel. Depending on the phase of the input mask, the maxima derived from this distance map may lead to over-fragmentation segmentation by the watershed algorithm. In this case, the watershed is called over-segmentation and tends to produce non-circular regions, which may be desirable in some applications but are not ideal for lesion segmentation. Figure 10C shows a composite input mask with three maxima in its distance map. This results in a segmentation containing three regions, only one of which (corresponding to seed 1071) is nearly circular. The other two regions are not circular. Seed 1 The region containing 051 is only semicircular. This allows the seed relabeling mechanism to verify all seed pairs and determine that the two regions corresponding to seeds 1051 and 1061 should merge so that they together form a more complete bubble. This behavior leads to a new division containing only two regions, both of which are nearly circular in shape.

[0171] A local maximum is a point where the distance to the contour is the longest compared to its directly adjacent points. The local maximum is a point, and its distance to the contour is known. As a result, the mechanism in the embodiment can draw a circle centered at this point. The radius of the circle is this distance. Thus, for two local maximums, the mechanism can calculate the overlap of their respective circles. This is shown in Figures 10A and 10B.

[0172] Seed relabeling determines whether two regions should be merged as follows: If the associated seeds are directly adjacent to each other, merging occurs; otherwise, the mechanism bases its decision-making on a hypothesis-testing procedure. For example, referring to Figure 10A, the example illustrates a situation where a distance map may result in two distinguishable maxima, leading to the assumption that each maximum value corresponds to the center of a distinguishable circular lesion. It should be noted that the distance map also allows the mechanism of the embodiment to say how far the maxima are from the contour (boundary). This distance is represented in Figure 10B by dotted line segments connecting the maximum value and the point on the contour. Thus, assuming the assumption is maintained, the spatial extent of these two lesions can be inferred from the assumption that the lesions are approximately round or "bubble" shaped. This allows the mechanism of the embodiment to draw two perfect circles as shown in Figure 10B. From this, the mechanism measures the overlap of the two circles (e.g., by a classical dice index) and compares it to a predetermined threshold. If the value of the overlap index is greater than this threshold, the mechanism concludes that the two bubbles are excessively overlapping and indistinguishable, and fusion occurs. In other words, the mechanism in the embodiment concludes that the two maxima correspond to two "centers" of the same lesion. However, in conventional watersheds, no such seed (i.e., maximum value) relabeling mechanism exists. As a result, mask over-fractionation occurs frequently.

[0173] This overlap can be measured in several ways. In one embodiment, the mechanism uses a Dice coefficient. For two perfect circles corresponding to two maxima, as shown in Figure 10B, the mechanism can calculate the Dice index for these two circles. In this way, the mechanism can learn from the training dataset what the optimal threshold should actually be applied to such that the two maxima are actually at the center of the same lesion, upon receiving a Dice index greater than the threshold.

[0174] Figures 10C and 10D provide an example of a different lesion mask shape, distinct from the lesion mask shapes in Figures 10A and 10B, where two partially fused circles are more similar to each other in Figure 10A than in Figure 10C. Due to distance maps that can be highly sensitive to mask shapes, the example lesion mask shape in Figure 10C has three seeds. Following the reasoning above, a lesion partitioning algorithm would partition the lesion represented in Figure 10C into two separate lesions, rather than into three separate lesions, as might occur in watershed techniques without seed relabeling.

[0175] In Figures 10C and 10D, seeds 1051 and 1061 represent a more extreme case than those shown in Figures 10A and 10B. Without using the seed relabeling technique of the embodiment, a splitting would occur to separate them (represented by diagonal solid lines). However, with the seed relabeling mechanism of the embodiment, this undesirable outcome can be virtually avoided. Conversely, since seed 1071 is sufficiently far from seeds 1051 and 1061, the same hypothesis verification procedure described above helps to accept the assumption that seed 1071 corresponds to the center of a distinguishable bubble, leading to a vertical splitting as shown in Figures 10C and 10D. Similarly, this would result in a shift from the labels assigned to seeds 1051 and 1061 to a different label for seed 1071. However, as with the situation in Figures 10A and 10B, the hypothesis verification procedure of the seed relabeling technique of the embodiment would determine that seeds 1051 and 1061 correspond to the same lesion.

[0176] Figure 11A is a block diagram illustrating a mechanism for lesion segmentation and relabeling according to one embodiment. As shown in Figure 11A, the mechanism, which may be implemented as a computer model including one or more algorithms, machine learning computer models, etc., and which operates on one or more input volumes of medical image data structures, receives two 2D lesion masks 1101, performs a distance transformation (block 1102), and generates a distance map 1111. This distance transformation (block 1102) is an operation performed on a binary mask, which calculates the shortest distance to the mask contour (boundary) for each point within the lesion mask. The further one moves inward into the lesion mask, the further the other moves from its contour (boundary). In this way, the distance transformation identifies the center point of the lesion mask, i.e., a point that is at a longer distance than others. In one embodiment, the mechanism may optionally perform Gaussian smoothing on the distance map 1111.

[0177] Next, the mechanism performs maximal identification (block 1103) to generate seed 1112. As described above, these maximals are the points in the distance map 1111 that are furthest from the contour or boundary. Based on seed 1112, the mechanism performs watershed technique (block 1104) to generate watershed segmented lesion mask 1113. As previously explained, this segmented lesion mask 1113 may be over-segmented, resulting in areas that do not match the assumed bubble shape of the lesion. Therefore, the mechanism performs seed relabeling based on distance map 1111, seed 1112, and segmented 2D lesion mask 1113 (block 1120) to generate updated segmented lesion mask 1121. See Figure 11B for further details on seed relabeling. The resulting updated segmented lesion mask 1121 will have fused regions that better match the assumed bubble shape of the lesion.

[0178] Figure 11B is a block diagram illustrating a seed relabeling mechanism according to one embodiment. As shown in Figure 11B, a mechanism, which may be implemented as a computer model including one or more algorithms, machine learning computer models, etc., executed by one or more processors of one or more computing devices, and which operates on one or more input volumes of medical image data structures, receives a distance map 1111 and seeds 1112. More specifically, the mechanism considers each seed pair (seed A and seed B) in seed 1112. The mechanism determines whether seed A and seed B are directly adjacent seeds (block 1151). If seed A and seed B are directly adjacent seeds, the mechanism assigns the same label to seed A and seed B (block 1155). In other words, seed A and seed B are grouped together so that they correspond to only one region.

[0179] In block 1151, if seed A and seed B are not directly adjacent seeds, the mechanism performs spatial range estimation based on the distance map 1111 (1152) and determines pairwise affinity for seed A and seed B as follows. In this embodiment, the spatial range estimation assumes that a certain region is "bubble" shaped. Accordingly, the mechanism assumes that each seed corresponds to a circle whose distance from the distance map is the radius of the circle.

[0180] The mechanism then calculates the overlap index for the circles corresponding to seeds A and B (block 1153). In one embodiment, the mechanism uses the dice index as follows:

[0181]

number

[0182]

number

[0183] The mechanism determines whether the overlap index is greater than a predetermined threshold (block 1154). If the overlap index is greater than a predetermined threshold in block 1154, the mechanism fuses the corresponding region into a segmented 2D lesion mask 1155 (block 1113).

[0184] If the affinity between two seeds is greater than this threshold, they are assigned the same label. Otherwise, at this stage, it is unknown whether they should belong to the same group or not. This decision is left to the label propagation stage (block 1512 in Figure 15), and for the same module used for z-direction connections, it is described below.

[0185] In situations with more than two seeds, the same process as in Figure 11B is repeated for all seed pairs before label propagation generates seed groups. For example, seed pairs (a, b) and (b, c) are determined to belong to the same group, but seed pair (a, c) fails the check as shown in Figure 11B. In this case, label propagation must place a, b, and c in the same group, meaning the regions corresponding to seeds a and c will still merge. However, if we have seeds a, b, c, and d, and affinity calculations (performed for a total of 6 pairs) show that only (a, b) and (c, d) pass the check, then label propagation could result in two groups, each containing (a, b) and (c, d). Therefore, if a seed pair fails the check, it means we don't know whether they should be in the same group or belong to different groups.

[0186] For example, in Figure 10C, there are three seed pairs (1051-1061, 1051-1071, 1061-1071), and the mechanism should determine that seeds 1051 and 1061 should be assigned the same label (belong to the same group). As a result, in the label propagation step, these three seeds are clustered into two groups: the first group contains only 1071, and the second group contains both 1051 and 1061.

[0187] Figure 12 is a flowchart outlining an exemplary operation of lesion segmentation according to one embodiment. The operation outlined in Figure 12 can be performed by the mechanism described above with respect to Figures 11A-11B. As shown in Figure 12, the operation begins (step 1200), and the mechanism generates a distance map for the two-dimensional lesion mask (block 1101). As described above, this distance map performs a distance transformation operation on the two-dimensional lesion mask and, optionally, Gaussian smoothing to remove noise. This can be generated by the following: The mechanism uses maximal identification to generate data points, for example, groupings of maximals by group (step 1202). The mechanism performs lesion segmentation based on the maximals to generate regions (step 1203). Next, the mechanism uses a distance map to relabel the seeds based on pairwise affinity (step 1204). This allows the mechanism to merge regions corresponding to seeds with the same label (step 1205). It should be understood that, due to the seed relabeling performed by the mechanism of the embodiment, the segmented lesion mask output in step 1205 does not suffer from the over-segmentation problem associated with watershed techniques, which occurs when incorrect labels are associated with data points associated with each of the lesion shapes, as described above. The operation then terminates (step 1206).

[0188] [Z-direction connection of lesions]

[0189] The above processes for lesion segmentation and seed relabeling are performed on each of the 2D images or 2D slices of the input volume, thereby generating lesion masks that are appropriately labeled for each of the lesions represented in the corresponding 2D images. However, the input volume, when considered in 3D, corresponds to a 3D representation of the internal anatomical structure and lesions of a biological entity, which may appear to be associated with the same lesion as may actually be associated with various lesions. Therefore, to enable the correct identification of separate lesions within the biological entity when represented in the 3D input volume, the embodiment provides a mechanism for connecting 2D lesions along the z-axis, i.e., in 3D.

[0190] This mechanism, which connects two-dimensional lesions along the z-axis and is called z-direction lesion connection, includes a logistic regression model that is executed to determine three-dimensional z-direction lesion detection on the segmented lesion output generated by the above mechanism. The mechanism connects two lesions in neighboring image slices. If the logistic regression model determines that the two lesions represent the same lesion, the two lesions are connected. For example, for any two-dimensional lesions on neighboring image slices, i.e., slices with z-axis coordinates that are continuously ordered along the z-axis in a three-dimensionally organized set of slices, the mechanism determines whether these two-dimensional lesions belong to the same three-dimensional lesion, as described below.

[0191] Figures 13A-13C illustrate the process of z-direction lesion connection according to one embodiment. Figure 13A shows the lesion mask input. Figure 13B shows the lesions after per-slice lesion splitting, in which the relabeling in the embodiment described above may employ an improved lesion splitting mechanism. As shown in Figures 13A and 13B, slice 1310 contains lesions 1311 and 1312, slice 1320 contains lesion 1321, and slice 1330 contains lesions 1331 and 1332. The z-direction lesion connection mechanism, i.e., the logistic regression model, is executed by comparing each lesion in a given slice with each lesion in the neighboring slices of the pair, for each per-slice segmented lesion mask in the neighboring slices of the pair in the input volume. For example, the z-direction lesion connection mechanism compares lesion 1311 (lesion A) in slice 1310 with lesion 1321 (lesion B) in slice 1320. For each comparison, the mechanism considers each lesion as a set of voxels and determines the common set between lesion A (the set of voxels in lesion A) and lesion B (the set of voxels in lesion B) relative to the size of lesion A and the size of lesion B. The z-direction lesion connectivity mechanism uses a logistic regression model based on two overlap rates to determine whether lesion A and lesion B are connected, as follows.

number

[0192] Here, |A| represents the area of ​​the circle corresponding to seed A, |B| represents the area of ​​the circle corresponding to seed B, and |A∩B| represents the area of ​​the intersection of the circles corresponding to seed A and seed B. The mechanism uses these two ratios as input features to train a logistic regression model to determine the probability that lesion A and lesion B are connected. That is, using machine learning processes such as those described above, the logistic regression model is trained to generate predictions for the volume of training images, for each pair of unit slice joins per training volume, on the probability that a lesion in one slice is the same lesion or a different lesion represented in neighboring slices. This prediction is compared to a ground truth representation of whether the lesions are the same lesion or different lesions to generate a loss or error. Computational parameters, such as the coefficients or weights of the logistic regression model, are modified to reduce this loss or error until a predetermined number of training epochs are performed or a predetermined stopping condition is met.

[0193] Logistic regression models are widely used to solve binary classification problems. In the context of this embodiment, this logistic regression model predicts the probability that two cross-sections of a lesion are part of the same lesion. For this purpose, logistic regression uses two overlap rates γ0 and γ1, as previously described. Specifically, the logistic regression model learns to linearly combine the two features as follows:

number

[0194] There are two extreme cases. First, when the threshold is set to 0, the z-direction connection mechanism of the embodiment will always determine that the lesions are the same lesion, i.e., the cross-sections are connected. As a result, both the true positive rate and the false positive rate become 1. Second, when the threshold t is set to 1, the z-direction connection mechanism will no longer identify any cross-section of the lesions being connected. In this case, both the true positive rate and the false positive rate become 0. Therefore, the logistic regression model makes a determination regarding whether the lesion cross-sections correspond to the same lesion or do not span neighboring slices only when the threshold t is in the interval (0, 1). Using an ideal logistic regression model, the true positive rate is equal to 1 (all true connections are identified) and the false positive rate is 0 (zero false connections occur).

[0195] Thus, having trained the logistic regression model, new slice pairs can be evaluated by inputting them as input features into the trained logistic regression model so that these ratios for each pair are calculated and predictions are generated for each of these pairs. If the predicted probabilities are greater than or equal to a predetermined threshold probability, then lesions A and B are considered to be associated with the same lesion of three dimensions. Corresponding relabeling of lesions across slices can then be performed to identify three-dimensional lesions within the input volume by associating lesions in two-dimensional slices with the same lesion representations in other neighboring slices.

[0196] There is a relationship that corresponds to the two ratio input features used to train the logistic regression model. For example, if lesions A and B are quite different in size, they are probably not part of the same lesion. Also, if lesions A and B are not crossover, such as lesion 1312 in slice 1310 and lesion 1321 in slice 1320, features γ0 and γ1 will be zero. As can be seen above, the logistic regression model performs regression assuming two feature values ​​γ0 and γ1 and outputs probability values ​​between 0 and 1, which indicate that likelihood lesion A and likelihood lesion B are part of the same lesion.

[0197] Figure 13C shows cross-sectional connections between slices according to one embodiment. As shown in Figure 13C, the mechanism determines that lesion 1311 in slice 1310 and lesion 1321 in slice 1320 are part of the same lesion by running a pre-trained logistic regression model of the embodiment, which predicts lesion commonality based on the overlap rate described above. Similarly, the mechanism also determines that lesion 1321 in slice 1320 and lesion 1331 in slice 1330 are part of the same lesion. In this way, the mechanism propagates common lesions along the z-axis and performs z-axis lesion connections.

[0198] Based on a pairwise evaluation in the input volume for identifying z-direction lesion connections across two-dimensional slices, and a determination by a trained logistic regression model of whether lesions are connected or not along the z-axis, the same label as a given lesion may be applied to each of the lesion masks present in each slice of the input volume, for example, all lesion masks across a set of slices in the input volume, in which case the logistic regression model determines that these lesion masks are associated with the same lesion, and lesion relabeling may be performed to specify that they are part of the same lesion A. This may be done for each lesion cross-section in each slice of the input volume, thereby generating associations of three-dimensional lesion masks for one or more lesions in the input volume. This information can then be used to represent or otherwise process lesions in subsequent downstream computing system operations, etc., in the subsequent three dimensions where all cross-sections associated with the same lesion are correctly labeled in the input volume.

[0199] Figures 14A and 14B show the results of a trained logistic regression model according to one embodiment. Figure 14A shows the receiver operating characteristic (ROC) curve against the maximum overlap rate (γ0) + minimum overlap rate (γ1) and against the maximum overlap rate index. The ROC curve is a graphical plot showing the diagnostic capability of a binary classification system when its discrimination threshold is valid. The ROC curve is produced by plotting the true positive rate (TPR) against the false positive rate (FPR) at various threshold settings. Figure 14B shows the accuracy call curve against the maximum overlap rate + minimum overlap rate index and against the maximum overlap rate index. The accuracy call curve is a plot of accuracy (y-axis) and callback (x-axis) against various thresholds, very similar to the ROC curve, where accuracy is the proportion of relevant instances among the retrieved instances and callback (or sensitivity) is the proportion of the total number of relevant instances actually retrieved. As shown in these figures, the two-feature logistic regression model outperforms its one-feature counterpart. Therefore, both features bring variable information to this prediction task.

[0200] Looking at the maximum overlap rate (γ0) + minimum overlap rate (γ1) index curve in Figure 14A, at the corresponding threshold t, training We can see that the trained logistic regression model can produce a true positive rate of ~=95% at the expense of a false positive rate of approximately 3%. Looking at Figure 14B, the plots shown identify the trained logistic regression model in terms of accuracy and recall, and both measures can yield very good results with the correct selection of threshold t.

[0201] Figure 15 is a flowchart illustrating the exemplary operation of a mechanism for connecting two-dimensional lesions along the z-axis according to one embodiment. As shown in Figure 15, the operation begins (step 1500), the mechanism selects the first image X from the input volume (step 1501), and selects the first lesion A in image X (step 1502). In some embodiments, images or slices in the input volume may be processed using the splitting and relabeling mechanism described above, but this is not required. Conversely, the mechanism of the embodiment targeting z-directional lesion connection can actually be performed for any input volume in which a lesion mask has been identified.

[0202] Next, the z-direction connection mechanism of the embodiment selects the first lesion B in the adjacent image Y (step 1503). The mechanism then determines the intersection of lesion A and lesion B with respect to lesion A, and the intersection of lesion A and lesion B with respect to lesion B (step 1504). The mechanism applies a trained logistic regression model to the γ0 and γ1 features on the common set of lesions A and B to determine, based on the two intersection values, whether lesions A and B belong to the same lesion, generates a prediction or probability that lesions A and B belong to the same lesion, and then compares this probability to a threshold probability (step 1505). Based on the result of this determination, lesion cross-sections in the image may be labeled or relabeled to indicate whether they are part of the same lesion.

[0203] The mechanism determines whether lesion B in image Y is the last lesion in image Y (step 1506). If lesion B is not the last lesion, the mechanism considers the next lesion B to be in the adjacent image Y (step 1507), and the operation returns to step 1504 to determine the intersection of lesion A and the new lesion B.

[0204] In step 1506, if lesion B is the last lesion in an adjacent slice or image Y, the mechanism determines whether lesion A is the last lesion in image X (step 1508). If lesion A is not the last lesion in image X, the mechanism considers the next lesion A to be in image X (step 1509), and the operation returns to step 1502 to consider the first lesion B in the adjacent image Y.

[0205] In step 1508, if lesion A is the last lesion in image X, the mechanism determines whether image X should be considered the last image (step 1510). If image X is not the last image, the mechanism considers the next image X (step 1511), and the operation returns to step 1502 to consider the first lesion A in the new image.

[0206] In step 1510, if image X is to be considered the last image, the mechanism propagates lesions that intersect between images along the z-axis, where propagation means that labels corresponding to the same lesions determined through the above process are set to the same value to indicate that they are part of the same lesion (step 1512). This is performed for each separate lesion identified in the input volume such that the cross-sections in each image corresponding to the same lesion are appropriately labeled, and therefore a three-dimensional representation of each lesion is generated through the z-direction connections of the cross-sections. The operation then terminates (step 1513).

[0207] [Contour refinement]

[0208] The above process yields accurate results regarding the number and relative location of lesions, as well as the interconnectedness of lesions across two-dimensional space (within an image or slice) and three-dimensional space (across images or slices in the input volume). However, lesion contours (boundaries) are not always well defined and require improvement. An example embodiment provides a mechanism for improving lesion contour accuracy. This further mechanism may be employed as part of the lesion segmentation in the above mechanism, or in other examples of embodiments that do not require the above-described specific lesion detection, lesion segmentation and relabeling, or z-direction connectivity mechanism or combination thereof.

[0209] While existing contour algorithms work well when a lesion is located in the center of an anatomical structure with no surrounding lesions, there are different situations that lead to a "leakage" problem when two or more adjacent lesions merge their initially distinguishable contours into a single contour covering all of them, thereby completely negating the benefits provided by early 2D lesion masking. In some cases, when a lesion is near an anatomical structure boundary, such as the liver boundary, contour trimming algorithms do not distinguish one lesion from another, but rather distinguish pixels of the anatomical structure from pixels of other anatomical structures in the image, such as organs, because the contour trimming algorithm can distinguish most of them.

[0210] The mechanism of the embodiment inpaints areas in an image or slice that are not the target area. Figure 16 shows an example of contouring two lesions in the same image according to one embodiment. On the left side of Figure 16, contours 1611 and 1612 are determined for the two lesions using an effective contour algorithm. An effective contour algorithm is a class of algorithms that iteratively evolves the contour to better match the image content.

[0211] According to this embodiment, the mechanism inpaints non-liver tissue within contour 1612 and near contour 1611 but not within contour 1611, but does not inpaint within contour 1611 if this inpainting means that the pixel values ​​for pixels of healthy tissue (non-lesional tissue) within contour 1612 and near contour 1611 are set to a specific value such that they all have the same value. For example, this value could be the average tissue value in an area identified as not associated with a lesion, i.e., an anatomical structure, such as healthy liver tissue.

[0212] This inpainting can be performed on a selected lesion contour 1611 such that the inpainting is applied to healthy tissue and other lesions, for example, lesion 1612 in the image. In this way, the contour and pixels associated with the selected lesion, e.g., 1611, are considered separately from the rest of the image when re-evaluating the numerical values ​​of contour 1611. This allows for a determination of whether contour 1611 can be re-evaluated numerically and whether the re-evaluation of contour 1611 results in an improved definition of contour 1611. That is, an initial determination of the contrast and variance between the pixels associated with the selected lesion contour 1611 and pixels near the selected lesion contour 1611 can be generated. After calculating this contrast and variance before inpainting, inpainting can be performed on the selected lesion 1611 such that pixels associated with other lesion contours, e.g., 1612, and areas of anatomical structures corresponding to healthy tissue in the image are inpainted with the average pixel intensity values ​​of healthy tissue.

[0213] The variance of a set of values ​​is determined as follows. Let's consider a voxel set, which consists of n voxels. First, their intensity values ​​are summed up, and then the arithmetic mean is calculated by dividing the sum by n. This result is denoted as A. Second, these voxel values ​​are individually squared, and then the arithmetic mean is calculated. This result is denoted as B. The variance is then defined as BA*A, that is, the difference between B and squared A.

[0214] Therefore, a set of n values ​​{x1, ..., x n The variance of} is defined as follows:

[0215]

number

[0216] This variance is calculated between voxels inside and outside a given contour. Voxels inside the contour are those enclosed by the contour, while voxels outside the contour refer to voxels that are outside the contour but remain within a predetermined distance from the contour.

[0217] The mechanism involves, after inpainting using the effective contour-reducing algorithm described above, recalculating the contour 1611 of the selected lesion, recalculating the contrast or variance of the new contour 1611, or both, and determining whether these values ​​have improved (higher contrast or lower variance values ​​inside or outside the lesion, or both). If the contrast and variance have improved, the newly calculated contour 1611 is maintained as the contour of the corresponding lesion. This process can then be performed on lesion 1612 as the selected lesion by inpainting pixels associated with the lesion 1611 and healthy tissue near contour 1612. In this way, each lesion is individually evaluated to generate a contour for each lesion, thereby preventing leakage of lesions from each other.

[0218] The mechanism for calculating the lesion contour after inpainting may be based on the Chan-Vase segmentation algorithm, which is designed to segment objects without clearly defined boundaries. This algorithm is based on a set of levels that are iteratively expanded to minimize energy defined by a weighted value corresponding to a term determined by the sum of the difference intensities from the mean outside the segmented region, the sum of the differences from the mean inside the segmented region, and the length of the segmented region's boundary. Initialization is performed using the segmented detection map (solving the energy minimum problem).

[0219] Upon the mechanism's segmentation, it initializes the contours with previous estimates and determines whether the new contours are better, for example, whether the contour's contrast and variance have improved. If the original contours are better, they are retained. If the new contours are better, for example, if the contour's contrast and variance have improved, the mechanism uses the new contours. In some embodiments, the mechanism determines which contours are better based on calculating homogeneous areas and variance. If the variance has decreased both inside and outside the contour, the mechanism uses the new contours; otherwise, the mechanism uses the previous contours. In another embodiment, the mechanism determines whether the contrast (mean value inside the contour versus mean value near the contour) has improved. Without departing from the ideas and scope of the embodiments, other techniques may be used that employ various measures to select between the previous and new contours.

[0220] Figure 17 is a flowchart illustrating the exemplary operation of a mechanism for improving the contour of each slice according to one embodiment. As shown in Figure 17, the operation begins with a given contour in an image segmented to show a lesion, for example in the liver, (step 1700), and the mechanism checks the first contrast and variance against the initial contour (step 1701). The mechanism imprints lesion pixels (or voxels of three dimensions) near the lesion (step 1702). Next, the mechanism checks the contour around the lesion (step 1703). Next, the mechanism checks the second contrast and variance against the new contour (step 1704). The mechanism checks the second contrast and variance against the first contrast The system then determines whether the improvement in contrast and variance is equivalent (step 1705). If the improvement in the second contrast and variance is equivalent, the mechanism uses the updated contour to represent the lesion (step 1706). The operation then ends (step 1708).

[0221] In step 1705, if the second contrast and variance do not result in improvement, the mechanism returns to the initial contour (step 1707). The operation then terminates (step 1708). This process may be repeated for each lesion identified in the input slice or input volume, or both, to recalculate the contour corresponding to each lesion in the image / input volume and improve this contour.

[0222] [False positive removal]

[0223] After performing lesion segmentation to generate a list of lesions and their contours, AI pipeline 100 performs a false positive processing stage 150 to remove misindicated lesions from the list of lesions. This false positive stage 150 can take many forms to reduce the number of lesions misidentified in the contours and map 135 in FIG. 1, which are output by a lesion list, e.g., liver / lesion detection logic 130 and then fused by segmentation and relabeling performed in lesion segmentation logic 140. In the following description, a novel false positive removal mechanism that can be used to perform this false positive removal is disclosed, but this special false positive removal is not necessary. Also, the false positive removal mechanisms described hereinafter can be used separately from the other mechanisms described above and can be applied to any list of objects identified in an image, and embodiments specifically utilize such false positive removal in the case of lesions in medical images. That is, the false positive removal mechanisms described in this section can be implemented separately and distinguishable from the other mechanisms described hereinbefore.

[0224] For the sake of convenience of explanation, it is assumed that the false positive removal mechanism is implemented as part of the AI pipeline 100 and also as part of the false positive removal logic 150 of the AI pipeline 100. Therefore, in the false positive stage 150, the false positive removal mechanisms described in this section operate on the list of lesions obtained from the liver / lesion detection logic and the segmentation and relabeling of the lesions, taking into account the three-dimensional nature of the input volume in the case of the z-direction lesion connection and contour improvement described above. This list 148 in FIG. 1 is input to a false positive removal logic stage 150 that processes the list 148 as described below and outputs a filtered or corrected list of lesions to the lesion classification stage 160, in which misidentified lesions are minimized in the corrected list of lesions. Thereby, the lesion classification stage classifies the various lesions shown in the corrected list of lesions.

[0225] In other words, the capture of all lesions in the AI ​​pipeline 100 up to this point may lead to a setting that increases the sensitivity of the AI ​​pipeline 100, causing it to misidentify pixels that are not actually lesions as being part of a lesion. As a result, there may be false positives that need to be removed. The false positive stage 150 includes logic that operates on a list of lesions and their contours to remove false positives. It should be understood that such false positive removal must also be balanced against the risk that if false positive removal is not performed appropriately during the examination (a set of input volume levels relative to the lesion level), it may result in undetected lesions. This can be challenging as it may lead to physicians and patients being unaware of lesions that require treatment. It should be understood that, theoretically, a single examination may contain multiple volumes of images for the same patient. However, in some embodiments where the single-phase detection implementation AI pipeline has a single-phase detection implementation, only one volume of images is processed, and it is assumed that this processing is performed on a single volume. For clarity, "patient level" will be used from here on instead of "examination level" as this is relevant to this embodiment (patient has a lesion or does not have a lesion). In other embodiments, it should be understood that the operations described herein may extend to examination levels in which multiple volumes of images may be evaluated for the same patient.

[0226] In this embodiment, assuming that the output of the preceding stage of the AI ​​pipeline 100 (such as slices, masks, multiple lesions, lesion and anatomical structure contours) is the input 148 to the false positive removal stage 150, the false positive removal stage 150 is performed at a highly specific operating point at the patient level (input volume level) to allow only a very small number of patient-level false positives (at least one lesion detected in a normal patient / volume). This point can be derived from the analysis of the patient-recipient operating characteristic (ROC) (patient-level sensitivity vs. patient-level specificity) analysis. Hereinafter referred to as the patient-level operating point OP. patientIn a volume that results in at least some lesions when using a very specific operating point, called the lesion-level operating point OP in this specification, a more sensitive operating point lesion is used at the lesion level, called lesion The lesion-level operating point OP can be identified from the analysis of the lesion-level ROC curve (lesion sensitivity vs. lesion specificity) to maximize the number of lesions to be retained.

[0227] Two operating points, namely OP patient and OP lesion can be implemented in one or more trained ML / DL computer models. The one or more trained ML / DL computer models are trained to classify whether the identified lesions are true lesions or false lesions, i.e., true positives or false positives, with respect to the input volume or a list of its lesions (the result of segmentation logic) or both. The one or more trained ML / DL computer models can be implemented as a binary classifier whose output indicates, for each lesion, whether it is a true positive or a false positive. A set of outputs including binary classification for all lesions in the list of input lesions can be used to filter the list of lesions to remove false positives. In one example embodiment, the one or more trained ML / DL computer models first implement a patient-level operating point to determine whether the classification result shows any of the lesions becoming true positives while false positives are filtered out. If any true positives remain in the list of initially filtered lesions after patient-level (input volume level) filtering, the lesion-level operating point is used to filter out any remaining false positives if any. As a result, a list of filtered lesions with minimized false positives is generated.

[0228] The implementation of the operating point can be for one or more trained ML / DL computer models. For example, if only one trained ML / DL computer model is used, the operating point can be a setting of the operating parameters of the ML / DL computer model that can be dynamically switched. For example, the input to the ML / DL computer model may be processed using a patient-level operating point to produce a result indicating whether a list of lesions contains true positives after each classification of the lesions, thereby switching the operating point of the ML / DL computer model between a lesion-level operating point and input processed again by false positives passing through the ML / DL computer model, each of which is removed from the final list of lesions output by a false positive removal step. Alternatively, in some embodiments, two separate ML / DL computer models can be trained, one for patient-level operating points and the other for lesion-level operating points, such that the results of a first ML / DL computer model showing at least one true positive trigger the processing of inputs through a second ML / DL computer model, and both models are identified as false positives that are removed from the final list of lesions output by the false positive removal step of the AI ​​pipeline.

[0229] Training an ML / DL computer model may involve a machine learning training operation in which the ML / DL computer model processes a training input that includes the volume of images and a list of corresponding lesions, in which case the list of lesions includes lesion masks or contours so that each lesion in the image generates a classification as to whether it is true positive or false positive. The training input may also be associated with ground truth information that indicates whether the image contains lesions, then evaluates the output generated by the ML / DL computer model to determine the loss or error, and then can be used to modify the operating parameters of the ML / DL computer model to reduce the determined loss / error. In this way, the ML / DL computer model learns the input features that represent true positive / false positive lesion detection. The machine learning operation is performed such that the operating parameters of the ML / DL computer model are learned taking into account patient-level sensitivity / specificity or lesion-level sensitivity / specificity or both, i.e., the operating point, i.e., OP. patient and OP lesion This can be done for each of them.

[0230] When classifying lesions in terms of whether they are true positives or false positives, an input volume (corresponding to a patient at the "patient level") is considered positive if it contains at least one lesion. An input volume is considered negative if it contains no lesions. With this in mind, a true positive is defined as a positive input volume, i.e., an input volume that has at least one detection result classified as a lesion, which is actually a lesion. A true negative is defined as a negative input volume, i.e., an input volume with no lesions, i.e., no detection results classified as lesions. A false positive is defined as a negative input volume with no lesions, but this input indicates a lesion in the detection results, i.e., the AI ​​pipeline lists that lesion if there is no lesion. A false negative is defined as a positive input volume that has a lesion, but the AI ​​pipeline does not indicate a lesion in the detection results. A trained ML / DL computer model classifies lesions in the input in terms of whether they are true positives or false positives. False positives are filtered out of the output resulting from false positive removal. False positive detection is performed at the patient level and the lesion level, i.e., at two different operating points, and at varying levels of sensitivity / specificity.

[0231] Two distinct operating points, one at the patient level and the other at the lesion level, can be determined based on ROC curve analysis. The ROC curve can be computed on a computer using ML / DL computer model activation data consisting of several input volumes (e.g., several input volumes corresponding to various patient examinations) that may contain several lesions (0 to K lesions per examination). The input to the trained ML / DL computer model, or "classifier," is the detection result previously detected in inputs that are either actual lesions or false positives, e.g., the outputs of the lesion detection and segmentation stages of the AI ​​pipeline. The first operating point, i.e., the patient-level operating point OP, is the first operating point, i.e., the patient-level operating point OP. patientis defined to maintain at least X% of the lesions identified as true positives, meaning that almost all true positives are maintained while removing some false positives. The value of X can be set based on the analysis of the ROC curve and can be any value suitable for a particular implementation. In one example embodiment, the value of X is set to 98% such that almost all true positives are maintained while some false positives are removed. Any value can be used. In one example embodiment, the value of X is set to 98% such that almost all true positives are maintained while some false positives are removed.

[0232] The lesion sensitivity is obtained for a first operating point, i.e., the patient-level operating point OP patient and a second operating point, i.e., the lesion-level operating point OP, is defined such that the lesion sensitivity exceeds the lesion sensitivity obtained for the patient-level operating point and the specificity exceeds Y%, where Y is determined by the actual performance of the trained ML / DL computer model. In one example embodiment, Y is set to 30%. An example of the ROC curve for patient-level and lesion-level operating point determination is shown in FIG. 18A. As shown in FIG. 18A, the lesion-level operating point is selected along the lesion-level ROC curve such that the lesion sensitivity exceeds the lesion sensitivity for the patient-level operating point.

[0233] ​​Figure 18B is an illustrative flowchart of the operation when performing false positive removal based on patient-level operating points and lesion-level operating points according to one embodiment. As shown in Figure 18B, the result of the segmentation stage logic of the AI ​​pipeline is input 1810 to a first trained ML / DL computer model 1820 that implements the first operating point. Input 1810 includes an input volume (i.e., an image volume (VOI)), lesion mask data or contour data that highlights the pixels or voxels corresponding to each of the lesions identified in the image data of the image volume, and a list of lesions including labels associated with pixels that highlight which lesions correspond to in the three-dimensional space of the output of the input volume, i.e., the segmentation, z-direction connection, and contour improvement described above. The input may be denoted as set S. The first trained ML / DL computer model 1820 implements a patient-level operating point in its training to classify features extracted from the input with X%, e.g., 98%, of true positories maintained in the filtered lesion list as a result of the classification of the trained ML / DL computer model 1820, and some of the false positories removed in the resulting list. The resulting list includes subset S containing true positories classified by the first ML / DL computer model 1820. + This includes subset S, which contains false-positive lesions classified by the first ML / DL computer model 1820. - It includes.

[0234] As a false positive removal logic, there may also be a true positive evaluation logic 1830 that determines whether the true positive subset output by the first ML / DL computer model 1820 is empty. That is, the true positive evaluation logic 1830 determines whether there are any elements from S that would be classified as true lesions by the first ML / DL computer model 1820. If it is determined that the true positive subset is empty, the true positive evaluation logic 1830 then evaluates the true positive subset S. +However, this is output as a filtered list of lesions 1835, meaning that there are no lesions identified in the output sent to the lesion classification stage of the AI ​​pipeline. True positive evaluation logic 1830 is true positive subset S + If it is determined that it is not empty, a second ML / DL computer model 1840 is executed upon receiving input S, and this second ML / DL computer model 1840 implements the second operating point during its training, namely the lesion level operating point OP. leision This is implemented. As can be seen above, two ML / DL computer models 1820 and 1840 are shown for explanatory purposes, but it should be understood that these two operating points can be implemented in various sets of trained operating parameters when constructing the same ML / DL computer model that can process input S, where the second ML / DL computer model is the same ML / DL computer model as 1820, but the operating parameters corresponding to the second operating point are different.

[0235] The second ML / DL computer model 1840 processes the input with trained operating parameters corresponding to the second operating point to generate a lesion classification again regarding whether the lesion is true positive or false positive. This result is a subset S' containing the predicted lesions (true positives). + and subset S' including predicted false positives ― Therefore, the filtered list of lesions 1845 is subset S' + It is output as such, and as a result, subset S' ― This effectively eliminates false positives.

[0236] The embodiments shown in Figures 18A and 18B will be described in terms of patient-level operating points and lesion-level operating points. It should be understood that the false positive removal mechanism can be implemented by operating points at various different levels. For example, similar operations are performed on image volume-level operating points and voxel-level operating points in "voxel-level" false positive removal operations. Figure 18C is an exemplary flowchart of the operation when performing voxel-level false positive removal based on input volume-level operating points and voxel-level operating points according to one embodiment. The operation in Figure 18C is similar to the operation in Figure 18B, but this operation may be performed with respect to voxels in the input set S. In voxel-level false positive removal, the first operating point may also be a patient-level operating point or an input volume-level operating point, while the second operating point is a voxel-level operating point OP. voxel This is possible. In this case, true positives and false positives are evaluated at the level such that any voxel is shown as being associated with a lesion, and it is true positive, as is the case when a voxel is actually associated with a lesion. If a voxel is shown as being associated with a lesion, but it is not actually associated with a lesion, it is considered a false positive. The appropriate setting of the operating point can again be generated based on the corresponding curve ROC, which yields a balance between sensitivity and specificity similar to that described above.

[0237] It should be understood that the above embodiment of the false positive removal mechanism assumes only one input volume from a patient examination, and that this embodiment may be applicable to any group of one or more images (slice). For example, false positive removal may be applied to only one slice, a set of slices smaller than the input volume, or even multiple input volumes from the same examination.

[0238] Figure 19 is a flowchart outlining the exemplary operation of the false positive removal logic of an AI pipeline according to one embodiment. As shown in Figure 19, the operation begins (step 1910) with receiving input S from a preceding stage of the AI ​​pipeline, where the input includes, for example, an input volume of images and a list of corresponding lesions including masks, contours, etc. (step 1900). The input is processed by a trained first trained ML / DL computer model that implements a first operating point, e.g., a patient-level operating point that is relatively more specific and less sensitive, to yield a first set of classifications for lesions including true positive subsets and false positive subsets (step 1920). A determination is made as to whether the true positive subset is empty (step 1930). If the true positive subset is empty, the operation ends by outputting the true positive subset as a list of filtered lesions (step 1940). If the true positive subset is not empty, the input S is processed by a trained second ML / DL computer model that implements a second operating point that is relatively more sensitive and less specific than the first operating point (step 1950). As can be seen above, in some embodiments, the first and second ML / DL computer models may be the same model but may be configured with different operating parameters to implement different operating points through different training. The result of processing by the second ML / DL computer model is a second set of classifications for lesions, including a second true positive subset and a second false positive subset. Thus, this second true positive subset is output as a filtered list of lesions (step 1960), and this operation ends.

[0239] [Example of a computer system environment]

[0240] The embodiments can be used in many different types of data processing environments. To provide some context for describing the specific elements and functionality of the embodiments, Figures 20 and 21 are provided hereafter as exemplary environments in which embodiments of the embodiments may be implemented. It should be understood that Figures 20 and 21 do not definitively or implicitly limit any aspects of the present invention or the environments in which embodiments may be implemented. Many modifications can be made to the shown environments without departing from the spirit and scope of the present invention.

[0241] Figure 20 shows a schematic diagram of one embodiment of a cognitive system 2000 that implements a request processing pipeline 2008, which, depending on the embodiment, may be a question-and-answer (QA) pipeline, a treatment recommendation pipeline, a medical image magnification pipeline, or any other artificial intelligence (AI) or cognitive computing-based pipeline that processes requests using a composite artificial intelligence mechanism that approximates a human being through a process, but through various computer-specific processes, on the generated results. For convenience of this explanation, the request processing pipeline 2008 is assumed to be implemented as a QA pipeline that operates in the form of input questions, whether structured or unstructured. One example of a question processing operation that may be used in conjunction with the principles described herein is described in U.S. Patent Application Publication 2011 / 0125734, which is incorporated herein by reference in its entirety.

[0242] The cognitive system 2000 is implemented in one or more computing devices 2004A-D connected to a computer network 2002 (including one or more processors and one or more memories, and optionally any other computing device elements commonly known in the art, such as buses, storage devices, and communication interfaces). For illustrative purposes only, Figure 20 shows the cognitive system 2000 implemented only in computing device 2004A, but as can be seen so far, the cognitive system 2000 may be distributed across multiple computing devices, such as multiple computing devices 2004A-D. The network 2002 includes multiple computing devices 2004A-D that can act as server computing devices and 2010-2012 that can act as client computing devices that communicate with each other and with other devices or components via one or more wired data communication links or wireless communication links, or both, in which case each communication link may be one or more of the following: wires, routers, switches, transmitters, receivers, etc. In some embodiments, the cognitive system 2000 and network 2002 enable question processing and response generation (QA) functionality for one or more cognitive system users via their respective computing devices 2010-2012. In other embodiments, the cognitive system 2000 and network 2002 include, but are not limited to, request processing and cognitive response generation, which can take many different forms depending on the desired implementation, such as cognitive information acquisition, user training / instruction, and cognitive evaluation of data. It can provide other types of cognitive operations. Other embodiments of the cognitive system 2000 may be used with components, systems, subsystems, or devices other than those described herein, or in combination thereof.

[0243] The cognitive system 2000 is configured to implement a request processing pipeline 2008 that receives input from various sources. Requests may be presented in the form of natural language questions and natural language responses for information, or in the form of natural language requests for the performance of cognitive operations. For example, the cognitive system 2000 receives input from network 2002, an electronic document corpus or electronic document corpora 2006, cognitive system users, or other data and other possible input sources or combinations thereof. In one embodiment, some or all of the input to the cognitive system 2000 passes through network 2002. Various computing devices 2004A-D in network 2002 include access points for content creators and cognitive system users. Some of the computing devices 2004A-D include devices for databases that store one or more data corpora or data corpora 2006 (shown as separate entities in Figure 20 for illustrative purposes only). Portions of the data corpus or data corpora 2006 may also be provided to one or more other network-mounted storage devices, one or more databases, or other computing devices not explicitly shown in Figure 20. The network 2002 includes local and remote connections in various embodiments, such that the cognitive system 2000 can operate in environments of any scale, including local and global, such as the Internet.

[0244] In one embodiment, a content creator uses the cognitive system 2000 to create content in the data corpus or data corpora 2006 documents for use as part of the data corpus. Documents include any files, texts, or data sources for use in the cognitive system 2000. A cognitive system user accesses the cognitive system 2000 via a network connection to network 2002 or an internet connection and inputs questions / requests into the cognitive system 2000 that are to be answered / processed based on the content in the data corpus or data corpora 2006. In one embodiment, the questions / requests are formed using natural language. The cognitive system 2000 parses and interprets these questions / requests via pipeline 2008 and provides a response to the cognitive system user, e.g., cognitive system user 2010, including one or more answers to the presented questions, responses to requests, and the results of processing requests. The cognitive system 2000 provides the user with a response from a ranked list of candidate answers / responses. In some embodiments, the cognitive system 2000 provides only one final answer / response, or a combination of the final answer / response and a ranked list of other candidate answers / responses. Other embodiments also exist.

[0245] The cognitive system 2000 implements a pipeline 2008 that includes multiple stages in processing input questions / requests based on information obtained from the data corpus or data corpora 2006. The pipeline 2008 generates answers / responses to the input questions or requests based on the processing of the input questions / requests and the data corpus or data corpora 2006.

[0246] In some example embodiments, cognitive system 2000 can be an IBM Watson (trademark) cognitive system obtained from International Business Machines Corporation, Armonk, New York, augmented by the mechanisms of the example embodiments described hereinafter. As outlined previously, the pipeline of the IBM Watson (trademark) cognitive system receives the key features of an input question or request, then parses the question or request to elicit the key features of the question / request, and then formulates a query that applies to a data corpus or corpora 2006 using these key features. Based on the application of the query to the data corpus or corpora 2006, a set of hypotheses, or candidate answers / responses to the input question / request (hereinafter assumed to be the input question), a portion of the data corpus or corpora 2006 (hereinafter simply referred to as corpus 2006) that has the potential to include a valuable response to the input question is identified across the data corpus or corpora 2006. Thereby, the pipeline 2008 of the IBM Watson (trademark) cognitive system performs a deep analysis on the language of the input question and the language used for each portion of corpus 2006 identified during the application of the query, using a variety of inference algorithms.

[0247] This allows scores from various inference algorithms to be weighted against a statistical model that summarizes the confidence level that the IBM Watson® Cognitive System 2000 pipeline 2008 has regarding the evidence that possible candidate answers are predicted by the question. This process is then repeated for each candidate answer to generate a ranked list of candidate answers that may be presented to the user who issued the input question, e.g., the user of a client computing device 2010, or from which a final answer is selected and presented to the user. Further information about the IBM Watson® Cognitive System 2000 pipeline 2008 can be found, for example, on the IBM Corporation website, IBM Redbooks, etc. For example, information about the IBM Watson® Cognitive System pipeline can be found in "Watson And Heaalcare" by Yuan et al., IBMdeveloperWorks, 2011, and "The Era of Cognitive Systems: An Inside Look at IBM Watson and How it Works" by Rob High, IBM RedBooks, 2012.

[0248] As can be seen above, input from a client device to the cognitive system 2000 may be presented in the form of a natural language question, but the embodiments are not limited thereto. Rather, the input question may actually be formatted or structured as any appropriate type of request that can be parsed and analyzed to perform cognitive analysis using structured input analysis, unstructured input analysis, or both, including but not limited to natural language parsing and the analysis mechanisms of the cognitive system such as IBM Watson™, and to make decisions based on the results of this cognitive analysis. For example, a physician or patient may issue a request to the cognitive system 2000 via their client computing device 2010 when performing a specific medical image-based operation, such as "identify a liver lesion in patient A, B, or C" or "provide treatment recommendations to the patient" or "identify changes in liver lesions in patient A, B, or C." According to an example embodiment, such a request can instruct cognitive computer operations, specifically employing the lesion detection and classification mechanism of the embodiment, to work so that the cognitive system 2000 provides a list of lesions, lesion contours, lesion classifications, and contours of the anatomical structure of interest. For example, a request processing pipeline 2008 can process a request such as "Identify liver lesions in patient ABC" by parsing the request and thereby identifying the anatomical structure as "liver," that is, the "liver" which is a medical imaging volume for patient "ABC" and is also a medical imaging volume in which the "lesions" in the anatomical structure of interest will be identified. Based on this parsing, a specific medical imaging volume corresponding to patient "ABC" can be retrieved from corpus 2006 and input into the lesion detection and classification AI pipeline 2020. The lesion detection and classification AI pipeline 2020 operates on this input volume as described above to identify a list of liver lesions.The list of liver lesions is output to the cognitive computing system 2000 for further evaluation via the request processing pipeline 2008, which is used to generate output for medical image viewer applications, etc.

[0249] As shown in Figure 20, one or more computing devices, for example, Server 2004, may be configured to implement a lesion detection and classification AI pipeline 2020, such as AI pipeline 100 in Figure 1. The configuration of a computing device may include equipment such as application-specific hardware and firmware to facilitate the performance of the operations and generation of outputs described herein with respect to the examples of embodiments. The configuration of a computing device may further or alternatively include software applications stored in one or more storage devices and loaded into the memory of a computing device, such as Server 2004, to cause one or more hardware processors of the computing device to run software applications that constitute the processors in order to perform the operations and generate outputs described herein with respect to the examples of embodiments. Any combination of application-specific hardware, firmware, and software applications running on hardware may be used without departing from the idea and scope of the examples of embodiments.

[0250] Given that the computing device is comprised of one of these methods, it should be understood that the computing device is a specialized computing device configured to implement the mechanism of the embodiment, but not a general-purpose computing device. Furthermore, as described herein, the implementation of the mechanism of the embodiment improves the functionality of the computing device, yielding useful and concrete results that facilitate not only automated lesion detection in the anatomical structure of interest but also the classification of such lesions, thereby reducing errors and increasing efficiency compared to manual processes.

[0251] As can be seen above, the mechanisms of the embodiments utilize a specially configured computing device or data processing system to perform operations in anatomical structure identification, lesion detection, and lesion classification. These computing devices or data processing systems may comprise various specially configured hardware elements, either through hardware configuration, software configuration, or a combination of hardware and software, to implement one or more of the systems / subsystems described herein. Figure 21 is a block diagram of just one example of a data processing system in which an aspect of the embodiments may be implemented. The data processing system 2100 is positioned and / or executed such that computer-readable code or instructions for performing the processes and aspects of the embodiments of the present invention result in the operations, outputs, and external effects of the embodiments as described herein, as shown in Figure 20. This is an example of a computer such as Server 2004.

[0252] In the example description, the data processing system 2100 employs a hub architecture including a northbridge and memory controller hub (NB / MCH) 2102, and a southbridge and input / output (I / O) controller hub (SB / ICH) 2104. The processing unit 2106, main memory 2108, and graphics processor 2110 are connected to the NB / MCH 2102. The graphics processor 2110 may be connected to the NB / MCH 2102 through an accelerated graphics port (AGP).

[0253] In the example shown, the local area network (LAN) adapter 2112 connects to the SB / ICH2104. The audio adapter 2116, keyboard and mouse adapter 2120, modem 2122, read-only memory (ROM) 2124, hard disk drive (HDD) 2126, CD-ROM drive 2130, Universal Serial Bus (USB) port and other communication ports 2132, and PCI / PCIe device 2134 connect to the SB / ICH2140 via buses 2138 and 2104. Examples of PCI / PCIe devices include Ethernet® adapters, add-in cards, and PC cards for notebook computers. PCI uses a card bus controller, while PCIe does not. ROM 2124 could be, for example, a flash basic input / output system (BIOS).

[0254] The HDD2126 and CD-ROM drive 2130 connect to the SB / ICH2104 via bus 2140. The HDD2126 and CD-ROM drive 2130 can use, for example, an integrated drive electronic (IDE) or Serial Advanced Technology Attachment (SATA) interface. A Super I / O (SIO) device 2136 may also be connected to the SB / ICH2104.

[0255] The operating system operates in the processing unit 2106. This operating system coordinates and controls the various components within the data processing system 2100 in Figure 21. As a client, the operating system may be a commercially available operating system such as Microsoft® Windows 10®. An object-oriented programming system such as the Java® programming system can work in conjunction with the operating system, making calls to the operating system from Java® programs or applications running in the data processing system 200.

[0256] As a server, the data processing system 2100 may be, for example, an IBM eServer System p® computer system or a Power® processor-based computer system running the Advanced Interactive Executive (AIX®) operating system or the LINUX® operating system. The data processing system 2100 may be a symmetric multiprocessor (SMP) system with multiple processors in the processing unit 2106. Alternatively, a single-processor system may be employed.

[0257] The operating system, object-oriented programming system, and instructions for an application or program are located on a storage device such as an HDD 2126 and loaded into main memory 2108 for execution by the processing unit 2106. The processes of the embodiments of the present invention may be performed by the processing unit 2106 using computer-readable program code that may be located, for example, in memory such as main memory 2108, ROM 2124, or in, for example, one or more peripheral devices 2126 and 2130.

[0258] A bus system, such as bus 2138 or bus 2140 as shown in Figure 21, may consist of one or more buses. Naturally, a bus system can be implemented using any kind of communication fabric or communication architecture that provides for data transfer between various components or devices attached to the fabric or architecture. Communication units, such as modem 2122 or network adapter 2112 in Figure 21, may include one or more devices used to send and receive data. Memory may be main memory 2108, ROM 2124, or cache, such as the NB / MCH 2102 shown in Figure 21.

[0259] As described above, in some embodiments, the mechanism of the embodiment may be implemented as application-specific hardware, firmware, or application software stored in a storage device such as the HDD 2126 and loaded into memory such as the main memory 2108 for execution by one or more hardware processors such as the processing unit 2106. Thus, the computing device shown in Figure 21 is configured to implement the mechanism of the embodiment and to perform the operations described herein with respect to the lesion detection and classification artificial intelligence pipeline and to generate output.

[0260] Those skilled in the art will understand that the hardware in Figures 20 and 21 may vary depending on the implementation. They will also understand that other internal hardware or peripherals, such as flash memory, equivalent non-volatile memory, or optical disc drives, may be used in conjunction with or instead of the hardware shown in Figures 20 and 21. Furthermore, the processes of this embodiment may be applied to multiprocessor data processing systems other than the SMP systems described herein without departing from the spirit and scope of the invention.

[0261] Furthermore, the data processing system 2100 can take any form of several different data processing systems, including client computing devices, server computing devices, tablet computers, laptop computers, telephones or other communication devices, and personal digital assistants (PDAs). In some exemplary cases, the data processing system 2100 may be a portable computing device configured such that flash memory provides non-volatile memory for storing, for example, operating system files or user-generated data, or both. In essence, the data processing system 2100 can be any known or future-developed data processing system, without any construction constraints.

[0262] As can be seen above, it should be understood that the examples of embodiments may take the form of an entire hardware embodiment, an entire software embodiment, or an embodiment that includes both hardware and software elements. In one example of an embodiment, the mechanism of the embodiment may be implemented in software or program code, including but not limited to firmware, resident software, and microcode.

[0263] A data processing system suitable for storing and / or executing program code would include at least one processor directly or indirectly connected to memory elements via a communication bus, such as a system bus. Memory elements include a large-capacity storage area and local memory, which is used during the actual execution of the program code, and a cache memory that provides temporary storage for at least some program code to reduce the number of times the code must be retrieved from the large-capacity storage area during execution. Memory can be of various types, including but not limited to ROM, PROM, EPROM, EEPROM, DRAM, SRAM, flash memory, and solid-state memory.

[0264] Input / output devices, or I / O devices (including, but not limited to, keyboards, displays, and pointing devices), can be connected to this system directly, or through intermediary wired or wireless I / O interfaces or controllers, or both. I / O devices can take many different forms beyond conventional keyboards, displays, and pointing devices, including, but not limited to, smartphones, tablet computers, touchscreen devices, and voice recognition devices, as well as communication devices connected via wireless or wired connections. Any known or future I / O devices are intended to fall within the scope of these embodiments.

[0265] Network adapters can also be connected to data processing systems to enable them to connect to other data processing systems or remote printers or storage devices via intervening private or public networks. Modems, cable modems, and Ethernet® cards are just a few examples of currently available network adapter types for wired communication. Wireless communication-based network adapters, including but not limited to 802.11a / b / g / n wireless communication adapters and Bluetooth® wireless adapters, may also be used. Any network adapters known or to be developed are intended to fall within the scope of the ideas of this invention.

[0266] While the description of the present invention has been presented for illustrative and explanatory purposes, it is not intended to encompass or limit the invention to the embodiments disclosed herein. Those skilled in the art will see many modifications and variations that do not depart from the scope and spirit of the embodiments described. The embodiments have been selected and described to best illustrate the principles and practical applications of the present invention and to enable those skilled in the art to understand the invention in terms of various embodiments with various modifications suitable for specific intended uses. The terminology used herein has been chosen to best illustrate the principles, practical applications, or technological advancements of embodiments that surpass the art available on the market, or to enable those skilled in the art to understand the embodiments disclosed herein. [Explanation of Symbols]

[0267] 100 AI Pipeline 128 Minimum Structure Determination 130 Liver / Legion Detection Logical Stage 132 ML / DL Computer Model Set 133 ML / DL Computer Model Set 134 ML / DL Computer Model Set 135 ML / DL Computer Model Set, Lesion, Detection Map 136 ML / DL Computer Model Set 140 Lesion Segmentation Logical Stage 148 List, Input 150 False Positive Processing Stage 160 Lesion Classification Stage 600 ML / DL Computer Model Set 610 First ML / DL Computer Model 612 U-Net Neural Network Model 614 Liver Mask 620 Second ML / DL Computer Model 621 DenseNet-169 (D169) Encoder 622 Decoder, Encoder 623 Decoder, Encoder, Slice Averaging (SLC AVG) Logical 624 Output, Lesion Output, Lesion Prediction Output 625 Output, lesion output, lesion prediction output 626 SLC AVG module 627 Loss of consistency, loss of consistency logic 630 Third ML / DL computer model 634 Encoder section, encoder, encoder neural network Work, CNN 635 Encoder section, encoder, encoder neural network, CNN 636 Encoder section, encoder, encoder neural network, CNN 638 Decoder stage, decoder 639 Second loss function, deep monitoring 640 Averaging, volume averaging (VOL ACG) logic 650 Final Leision mask, final lesion output 660 Liver mask output 810 First stage 820 Second stage 910 Lesion area 911 Lesion 912 Lesion 913 Lesion 1051 Seed 1061 Seed 1071 Seed 1101 2D lesion mask 1111 Distance map 1112 Seed 1113 Watershed segmented lesion mask 1121 Updated segmented lesion mask 1310 Slice 1311 Lesion 1312 Lesion 1320 Slice 1321 Lesion 1330 Slice 1331 Lesion 1332 Lesion 1611 Contour 1612 Contour 1810 Input 1820 First trained ML / DL computer model 1830 True positive evaluation logic circuit 1835 Filtered lesion list 1840 Second ML / DL computer model 1845 Filtered lesion list 2000 Cognitive systems 2002 Computer networks 2004A Computing devices, servers 2004B Computing devices 2004C Computing devices 2004D Computing devices 2006 Corpus 2006 Electronic document corpus or electronic document corpora, data corpus or data corpora 2007 Cognitive systems 2008 Request processing pipeline 2010 Cognitive system users 2010-2012 Client computing devices 2020 Lesion detection and classification AI pipeline 2100 Data processing system 2102 Memory controller hub (NB / MCH) 2104 Input / output (I / O) controller hub (SB / ICH) 2106 Processing unit 2108 Main memory 2110 Graphics processor 2112 Local area network (LAN) adapter 2116 Audible adapter 2120Keyboard and mouse adapters 2122 Modem 2124 Read-only memory (ROM) 2126 Hard disk drives (HDDs) and peripherals 2130 CD-ROM drives and peripherals 2132 Universal Serial Bus (USB) ports and other communication ports 2134 PCI / PCIe devices 2136 Super I / O (SIO) devices 2138 Bus 2140 Bus

Claims

1. A method in a data processing system comprising at least one processor and at least one memory, wherein the at least one memory contains instructions executed by the at least one processor for implementing a lesion detection and classification artificial intelligence (AI) pipeline comprising a plurality of trained machine learning computer models, The AI ​​pipeline processes the input volume of medical images using one or more first machine learning computer models and determines whether the input volume contains a predetermined amount of the target anatomical structure. The logic of the AI ​​pipeline determines whether one or more default criteria are met by the output of one or more first machine learning computer models, wherein one or more default criteria include a default amount of anatomical structures of a subject being shown in the input volume, the one or more first machine learning computer models include a phase classification machine learning model operating on the input volume, the phase classification machine learning model determines the phase of the medical image represented in each of the medical images of the input volume, and the one or more default criteria further includes determining, based on the results of the operation of the phase classification machine learning model on the input volume, that a single phase of medical image exists in the input volume. In response to a determination that one or more of the aforementioned default criteria are met by the output of one or more of the aforementioned first machine learning computer models, the pathological processing operation is performed. The method includes, and the lesion treatment operation is, The input volume is processed by one or more second machine learning computer models of the AI ​​pipeline to detect lesions present in the medical images of the input volume that correspond to the anatomical structures of the target, One or more third machine learning computer models of the AI ​​pipeline process the detected lesions, perform segmentation of the lesions, and combine lesion contours from different medical images within the input medical images associated with the same lesion to generate a list of lesions containing the contours associated with the lesions. One or more fourth machine learning computer models of the AI ​​pipeline process the list of lesions and the contours associated with the lesions in the list of lesions, and classify the lesions in the list of lesions into one or more predetermined classes of lesions corresponding to the anatomical structures of the subject, The AI ​​pipeline includes outputting a list of lesions and the classification associated with the lesions for processing by a downstream computing system. Determining whether one or more of the aforementioned default criteria are met by the output of one or more of the aforementioned first machine learning computer models is: The AI ​​pipeline further includes performing an axial scoring operation on the medical images in the input volume using the axial scoring logic of the AI ​​pipeline, The aforementioned axial scoring operation, The process involves scoring the images in the input volume according to a predefined scoring algorithm, and estimating the axial scores of the lowest slice (MISV) and the highest slice (MSSV) in the input volume from the scores associated with the images. A method comprising identifying fragments of the target anatomical structure present in the input volume based on the axial scores of the MISV and the MSSV.

2. The method according to claim 1, wherein the one or more first machine learning computer models of the AI ​​pipeline include a body part detection machine learning model, the body part detection machine learning model is executed on the input volume to detect body parts of a biological entity represented in the input volume, and determines whether the body part of the biological entity corresponds to a body part where the anatomical structure of the subject is located.

3. The method according to claim 2, wherein the body part is the abdomen of a human and the anatomical structure of the object is the liver of a human.

4. The method according to any one of claims 1 to 3, wherein the one or more first machine learning computer models of the AI ​​pipeline further include an anatomical structure detection machine learning model that operates on the input volume, the anatomical structure detection machine learning model determines the amount of the target anatomical structure represented in the input volume.

5. The method according to claim 4, further comprising determining whether a minimum threshold amount of the target anatomical structure is present in the input volume based on the results of the operation of the anatomical structure detection machine learning model on the input volume.

6. The method according to claim 4 or 5, wherein, in response to the result of the operation of the phase classification machine learning model indicating that there are two or more phases of medical images in the input volume, a subset of medical images in the input volume corresponding to the target phase is selected, and the lesion processing operation is performed only on the selected subset of medical images in the input volume.

7. The method according to claim 6, wherein the target phase is one of the pre-contrast phase, angiography phase, portal / venography phase, or delayed phase.

8. The method according to any one of claims 1 to 7, wherein the one or more second machine learning computer models of the AI ​​pipeline comprises a collection of machine learning computer models, and each machine learning computer model in the collection of machine learning computer models is trained by a machine learning process to process the input volume in a manner different from other machine learning computer models in the collection of machine learning computer models to generate predictions for corresponding lesion detection.

9. The method according to claim 8, wherein the group of machine learning computer models includes a mask generation machine learning model, an input volume processing machine learning model, and a masked input volume processing machine learning model, the mask generation machine learning model being trained by a machine learning process to generate a mask corresponding to the anatomical structure of the subject, applying the mask to the input volume to generate a masked input volume, the input volume processing machine learning model being trained by a machine learning process to generate a first lesion prediction output based on the input volume, and the masked input volume processing machine learning model being trained by a machine learning process to generate a second lesion prediction output based on the masked input volume.

10. The first machine learning model of the aforementioned population is trained using a first loss function that penalizes errors in classifying false negatives of lesions. The method according to claim 8 or 9, wherein a second machine learning model of the population is trained using a second loss function different from that of the first machine learning model, which penalizes errors in classifying false positives of lesions.

11. The method according to claim 10, wherein the logic of the group applies a third loss function, the third loss function operates to compare the output of the first lesion detection of the first machine learning model with the output of the second lesion detection of the second machine learning model, and to make the output of the first lesion detection match the output of the second lesion detection.

12. The method according to any one of claims 1 to 11, wherein the one or more third machine learning computer models of the AI ​​pipeline includes a lesion segmentation machine learning model that segments an image of the input volume into contours corresponding to lesions to generate lesion segments; a z-direction connection machine learning model that joins subsets of lesion segments based on the lesion segments generated by the lesion segmentation machine learning model which are determined to be associated with the same lesion in three-dimensional space; and a contour improvement machine learning model that isolates lesions in the input volume and improves the contours of the lesions.

13. The method according to any one of claims 1 to 11, further comprising one or more false positive removal machine learning computer models that operate on the list of lesions to remove false positive detections of lesions from the list of lesions.

14. The one or more false positive removal machine learning computer models include one or more false positive removal machine learning computer models trained by a machine learning process based on two different operating points, The method according to claim 13, wherein the first of the two different operating points corresponds to a patient-level operating point, and the second of the two different operating points corresponds to a lesion-level operating point.

15. The method according to any one of claims 1 to 14, wherein the one or more fourth machine learning computer models classify each of the lesions in the list of lesions into one of a plurality of default classes, the one or more default classes comprising at least one of lesion types, or benign class, malignant class, or undefined class.

16. The method according to any one of claims 1 to 15, further comprising processing the output of the list of lesions and the classification associated with the lesions by a downstream computing system, and generating a visual representation of the lesions in the list of lesions for a graphical user interface.

17. The method according to any one of claims 1 to 16, wherein the input volume includes an image corresponding to at least one of magnetic resonance imaging or computed tomography.

18. A computer program, wherein the computer program causes a computing device to perform a procedure to implement an artificial intelligence (AI) pipeline for lesion detection and classification, which includes a plurality of trained machine learning computer models, and the AI ​​pipeline is The AI ​​pipeline processes the input volume of medical images using one or more first machine learning computer models and determines whether the input volume contains a predetermined amount of the target anatomical structure. The logic of the AI ​​pipeline determines whether one or more default criteria are met by the output of one or more first machine learning computer models, wherein one or more default criteria include that a default amount of anatomical structures of a subject are shown in the input volume, the one or more first machine learning computer models include a phase classification machine learning model that operates on the input volume, the phase classification machine learning model determines the phase of the medical image represented in each of the medical images of the input volume, and the one or more default criteria further includes determining, based on the results of the operation of the phase classification machine learning model on the input volume, that the input volume contains a single phase of medical image. In response to a determination that one or more of the aforementioned default criteria are met by the output of one or more of the aforementioned first machine learning computer models, the pathological processing operation is performed. The lesion processing operation is performed as follows: The input volume is processed by one or more second machine learning computer models of the AI ​​pipeline to detect lesions present in the medical images of the input volume that correspond to the anatomical structures of the target, One or more third machine learning computer models of the AI ​​pipeline process the detected lesions, perform segmentation of the lesions, and combine lesion contours from different medical images within the input medical images associated with the same lesion to generate a list of lesions containing the contours associated with the lesions. One or more fourth machine learning computer models of the AI ​​pipeline process the list of lesions and the contours associated with the lesions in the list of lesions, and classify the lesions in the list of lesions into one or more predetermined classes of lesions corresponding to the anatomical structures of the subject, The AI ​​pipeline includes outputting a list of lesions and the classification associated with the lesions for processing by a downstream computing system. Determining whether one or more of the aforementioned default criteria are met by the output of one or more of the aforementioned first machine learning computer models is: The AI ​​pipeline further includes performing an axial scoring operation on the medical images in the input volume using the axial scoring logic of the AI ​​pipeline, The aforementioned axial scoring operation, The process involves scoring the images in the input volume according to a predefined scoring algorithm, and estimating the axial scores of the lowest slice (MISV) and the highest slice (MSSV) in the input volume from the scores associated with the images. A computer program comprising identifying fragments of the target anatomical structure present in the input volume based on the axial scores of the MISV and the MSSV.

19. Processor and A device comprising a memory coupled to the processor, The memory includes instructions, when executed by the processor, that cause the processor to implement a lesion detection and classification artificial intelligence (AI) pipeline comprising a plurality of trained machine learning computer models, the AI ​​pipeline, The AI ​​pipeline processes the input volume of medical images using one or more first machine learning computer models and determines whether the input volume contains a predetermined amount of the target anatomical structure. The logic of the AI ​​pipeline determines whether one or more default criteria are met by the output of one or more first machine learning computer models, wherein one or more default criteria include that a default amount of anatomical structures of a subject are shown in the input volume, the one or more first machine learning computer models include a phase classification machine learning model that operates on the input volume, the phase classification machine learning model determines the phase of the medical image represented in each of the medical images of the input volume, and the one or more default criteria further includes determining, based on the results of the operation of the phase classification machine learning model on the input volume, that the input volume contains a single phase of medical image. In response to a determination that one or more of the aforementioned default criteria are met by the output of one or more of the aforementioned first machine learning computer models, the pathological processing operation is performed. The lesion processing operation is performed for the purpose of, The input volume is processed by one or more second machine learning computer models of the AI ​​pipeline to detect lesions present in the medical images of the input volume that correspond to the anatomical structures of the target, One or more third machine learning computer models of the AI ​​pipeline process the detected lesions, perform segmentation of the lesions, and combine lesion contours from different medical images within the input medical images associated with the same lesion to generate a list of lesion boundaries containing the contours associated with the lesions. One or more fourth machine learning computer models of the AI ​​pipeline process the list of lesions and the contours associated with the lesions in the list of lesions, and classify the lesions in the list of lesions into one or more predetermined classes of lesions corresponding to the anatomical structures of the subject, The AI ​​pipeline includes outputting a list of lesions and the classification associated with the lesions for processing by a downstream computing system. Determining whether one or more of the aforementioned default criteria are met by the output of one or more of the aforementioned first machine learning computer models is: The AI ​​pipeline further includes performing an axial scoring operation on the medical images in the input volume using the axial scoring logic of the AI ​​pipeline, The aforementioned axial scoring operation, The process involves scoring the images in the input volume according to a predefined scoring algorithm, and estimating the axial scores of the lowest slice (MISV) and the highest slice (MSSV) in the input volume from the scores associated with the images. An apparatus comprising identifying fragments of the target anatomical structure present in the input volume based on the axial scores of the MISV and the MSSV.

Citation Information

Patent Citations

  • Method and system for estimating tumor extent in magnetic resonance imaging

    JP2002526190A

  • Method and apparatus for classifying detected input in medical images

    JP2009525142A

  • Medical image processing device

    JP2020146455A

  • Deep convolutional neural networks for tumor segmentation with positron emission tomography

    WO2020190821A1