Method for tumor assessment

By analyzing the distribution of lymphocytes in tumor images using a deep learning model and combining it with the Cox proportional hazards model, the problem of assessing the prognosis of cancer tumors has been solved, enabling more accurate prognostic prediction and medical decision support.

CN121095142APending Publication Date: 2025-12-09NANTCELL INC +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511074891.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-01-11
Filing Date
2021-01-11
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Assessing the prognosis of cancer tumors is difficult, and existing technologies are unable to effectively handle the multiple factors that affect prognosis and their correlations, resulting in insufficient predictive accuracy.

Method used

Using a deep learning model that combines convolutional neural networks and Gaussian mixture models, the distribution of lymphocytes in tumor images is analyzed, and combined with the Cox proportional hazards model, the tumor category is determined and individual clinical values, such as prognosis, are predicted.

Benefits of technology

It improves the accuracy of tumor prognosis assessment, enabling more precise prediction of individual survival rates and guiding medical decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095142A_ABST
    Figure CN121095142A_ABST
Patent Text Reader

Abstract

A method of tumor assessment operates a device including processing circuitry, and may include executing, by the processing circuitry, instructions that cause the device to: determine a lymphocyte distribution of lymphocytes in a tumor based on an image; applying a classifier to the lymphocyte distribution to classify the tumor, the classifier having been trained to classify the tumor as a category selected from at least two categories respectively associated with the lymphocyte distribution; and determining a clinical value for the individual based on a pre-post of the individual suffering from the tumor of the category into which the classifier classifies the tumor.
Need to check novelty before this filing date? Find Prior Art

Description

This application is a divisional application of Chinese Invention Patent Application No. 202180008820.3, filed on January 11, 2021, entitled “Deep Learning Model for Tumor Assessment”. Cross Reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Application No. 62 / 959,931, filed January 11, 2020, the entire disclosure of which is incorporated by reference. TECHNICAL FIELD

[0002] The present disclosure relates to the field of image analysis using machine learning models, and more particularly to determining clinical variables related to tumors using deep learning models. BACKGROUND

[0003] The background description provided herein is for the purpose of generally presenting the context of the disclosure. The work of the presently named inventors, to the extent the work is described in this background section, as well as aspects of the description that can not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art against the disclosure.

[0004] In the medical field, many scenarios involve analysis of cancerous tumors to assess tumor class based on characteristics such as location, size, shape, and constituent components. The assessment can enable predictions such as tumor behavior and likely aggressiveness, such as probability and rate of growth and / or metastasis. These characteristics of the tumor can in turn determine clinical value (such as prognosis) for the individual, such as likely survival rate, and can guide medical decisions such as selection, type, and / or timing of chemotherapy, surgery, and palliative care. However, determination of prognosis, including survivability, becomes difficult due to the number and variety of relevant factors that can affect the determination, as well as factor interdependencies.

[0005] A variety of diagnostic and prognostic techniques can be used to make tumor class assessments. For example, a data set can include features about individuals with tumors, such as age, physiology, medical history, and / or behavioral habits such as smoking for each individual, which can be correlated with prognosis data, such as typical survival rate, for the individuals. A Cox proportional hazards model can be applied to determine the relevance of features in a set of clinical features based on the collected data set and clinical data, which can support some conclusions about the relevance of individual risk factors for the clinical value (such as prognosis) of the individuals based on the tumors. Thereafter, a clinician can use this information to guide determination or prediction of diagnosis, prognosis, and / or effective care selection for individuals with similar tumors. Also, similar risk factors about these individuals can be collected and processed through a Cox proportional hazards model developed based on the similar tumors to predict the clinical value (such as prognosis) of the individuals based on the similar tumors.

[0006] Other techniques for evaluating an individual's tumor-based clinical value, such as prognosis, can utilize one or more machine learning models. For example, a training dataset of tumor samples with known characteristics, such as data about the tumor, the individual from which the tumor was excised, and / or the individual's clinical value, such as prognosis, can be generated. A machine learning classifier can be trained using the training dataset with labels indicating the class represented by each input, such as whether each tumor represents a high-risk tumor with a poor prognosis or a low-risk tumor with a good prognosis. The training process can yield a trained machine learning classifier that classifies new inputs in a manner consistent with the examples of the training dataset.

[0007] Various different machine learning models can be selected as classifiers for tumors, such as Bayesian classifiers, artificial neural networks, and support vector machines (SVMs). As a first such example, a convolutional neural network (CNN) can process n-dimensional inputs to detect features of tumor images that can appear therein. A convolutional neural network comprising a sequence of neuron convolutional layers can be provided with a feature vector such as an array of pixels of a tumor image. Each convolutional layer can produce a feature map indicating some tumor image features detected at a level of detail that can be processed by the next convolutional layer in the sequence. The feature map produced by a final convolutional layer of the convolutional neural network can be processed by a classifier for tumors that can be trained to indicate whether the feature map is similar to feature maps of objects in images of a training dataset. For example, the CNN can be trained to recognize tumor visual features associated with high-risk prognoses and low-risk prognoses.

[0008] As a second such example, a Gaussian mixture model (GMM) can be generated to classify data about tumors into different clusters of tumors with representative characteristics. For each sample representing a tumor in the training dataset, a set of features can be identified, such as location, size, shape, and composition. The samples of the training dataset can be positioned within a multi-dimensional feature space, where each feature is represented along a dimension axis. Machine learning techniques can be applied to identify clusters of tumors within the feature space with similar prognoses, such as a first cluster representing high-risk tumors and a second cluster representing low-risk tumors, where each cluster is represented as a set of Gaussian probability distributions of the respective features within the feature space. Even though some portions of the clusters overlap (e.g., even though tumors with a particular set of features can be included in either the high-risk cluster or the low-risk cluster), the Gaussian probability distributions of the clusters can enable probabilistic predictions about the likelihood that a tumor belongs to each cluster. In this way, the Gaussian mixture model can enable individual prognosis predictions based on clustering of similar tumor samples in the training dataset. SUMMARY

[0009] Some example embodiments can include a method of operating an apparatus comprising processing circuitry, wherein the method comprises execution of instructions by the processing circuitry that cause the apparatus to: receive an image depicting at least a portion of a tumor; determine a lymphocyte distribution of lymphocytes in the tumor based on the image; apply a classifier to the lymphocyte distribution to classify the tumor, the classifier having been trained to classify tumors into a class selected from at least two classes that are respectively associated with lymphocyte distributions; and determine a clinical value for the individual based on a prognosis dataset corresponding to individuals having tumors classified by the classifier into the class to which the tumor is classified.

[0010] In some example embodiments, the tumor is a pancreatic cancer tumor or a breast cancer tumor.

[0011] In some example embodiments, the apparatus can further comprise a convolutional neural network trained to determine a lymphocyte distribution of lymphocytes in an image region, and the instructions can cause the apparatus to invoke the convolutional neural network to determine a lymphocyte distribution of lymphocytes in a respective region of the image of the tumor. In some example embodiments, the convolutional neural network can be further trained to classify a region of the image into one or more region types selected from a set of region types comprising a tumor region, a lymphocyte region, or an interstitial region. In some example embodiments, determining a lymphocyte distribution of lymphocytes in the tumor can comprise, for a respective lymphocyte region of the image: determining a distance of the lymphocyte region to one or both of a tumor region or an interstitial region; characterizing the lymphocyte region as one of a tumor-infiltrating lymphocyte region, a tumor-adjacent lymphocyte region, an interstitial-infiltrating lymphocyte region, or an interstitial-adjacent lymphocyte region based on the distance, and the classifier can further classify the tumor based on the characterization of the lymphocyte region. In some example methods, determining a lymphocyte distribution of lymphocytes in the tumor can comprise, for a respective interstitial region of the image: determining a distance of the interstitial region to a tumor region; and characterizing the interstitial region as one of a tumor-infiltrating interstitial region or a tumor-adjacent interstitial region based on the distance, and the classifier can further classify the tumor based on the characterization of the interstitial region.

[0012] In some example embodiments, the at least two classes can comprise a high-risk tumor class associated with a first survival probability, and a low-risk tumor class associated with a second survival probability that is longer than the first survival probability.

[0013] In some example embodiments, the classifier can further include a Gaussian mixture model configured to determine, for a respective class, a probability distribution of features of tumors of the class within a feature space. In some example embodiments, the features of the feature space of the Gaussian mixture model can be selected from a feature set including: measurements of tumor regions of the image, measurements of interstitial regions of the image, measurements of lymphocyte regions of the image, measurements of tumor infiltrating lymphocyte regions of the image, measurements of tumor adjacent lymphocyte regions of the image, measurements of interstitial infiltrating lymphocyte regions of the image, measurements of interstitial adjacent lymphocyte regions of the image, measurements of tumor infiltrating interstitial regions of the image, and measurements of tumor adjacent interstitial regions of the image. In some example embodiments, a feature subset can be selected from the feature set based on a relevance of the respective classes to respective features of the subset. In some example embodiments, the relevance of the respective classes to the respective features can be based on one or both of a silhouette score or a consensus index of the feature space. In some example embodiments, the feature subset can consist essentially of measurements of lymphocyte regions of the image, measurements of tumor infiltrating lymphocyte regions of the image, measurements of tumor adjacent lymphocyte regions of the image, and measurements of tumor infiltrating interstitial regions of the image.

[0014] In some example embodiments, the instructions can further cause the apparatus to apply a Cox proportional hazards model to clinical features of the tumor to determine a class of the tumor, and determine a clinical value (such as a prognosis) of the individual based on a prognosis of individuals having tumors classified into the class by the classifier and the class determined by the Cox proportional hazards model. In some example embodiments, the clinical features of the Cox proportional hazards model can be selected from a clinical feature set including: a primary diagnosis of the tumor, a location of the tumor, a treatment of the tumor, a measurement of the tumor, a metastasis status of the tumor, a primary diagnosis of the individual, a previous cancer history of the individual, a gender of the individual, a frequency of a smoking habit of the individual, a duration of a smoking habit of the individual, a drinking history of the individual. In some example embodiments, a clinical feature subset of the clinical features can be selected from the clinical feature set based on a relevance of the respective classes to respective clinical features in the subset for the Cox proportional hazards model. In some example embodiments, the clinical feature subset can consist of the measurement of the tumor and the metastasis status of the tumor.

[0015] In some example embodiments, the instructions can further cause the apparatus to display a visualization of a clinical value (e.g., prognosis) of the individual. In some example embodiments, the visualization is a Kaplan Meier survival ability projection of the tumor. In some example embodiments, the instructions can further cause the apparatus to determine a diagnostic test for the tumor based on the clinical value (e.g., prognosis) of the individual. In some example embodiments, the instructions can further cause the apparatus to determine a treatment for the individual based on the clinical value (e.g., prognosis) of the individual. In some example embodiments, the instructions can further cause the apparatus to determine a schedule of therapeutic agents for treating the tumor based on the clinical value (e.g., prognosis) of the individual.

[0016] In some example embodiments, the at least two classes are a low-risk tumor class and a high-risk tumor class, determining the lymphocyte distribution further comprises applying a convolutional neural network to the image, the convolutional neural network configured to measure the lymphocyte distribution of lymphocytes of different region types of the image, the classifier is a bi-directional Gaussian mixture model configured to determine, for a respective class, a probability distribution of features of tumors of the class within a feature space, the method further comprises applying a Cox proportional hazards model to clinical features of the tumor to determine a class of the tumor, and determining a clinical value (e.g., prognosis) of the individual is further based on the class predicted by the Cox proportional hazards model.

[0017] Some example embodiments can include a system comprising: memory hardware configured to store instructions embodying any of the above-described methods; and processing hardware configured to execute the instructions stored by the memory hardware.

[0018] Some example embodiments can include a system comprising: an image evaluator configured to determine a lymphocyte distribution of lymphocytes in an image; a classifier configured to classify a tumor into a class selected from at least two classes respectively associated with the lymphocyte distribution; and a tumor evaluator configured to determine a clinical value (such as a prognosis and / or survivability) of an individual based on a tumor in the image by invoking the image evaluator with the image to determine the lymphocyte distribution of lymphocytes in the tumor, invoking the classifier to classify the tumor into a class based on the lymphocyte distribution, and outputting the clinical value (such as a prognosis) of the individual based on a prognosis of individuals having tumors of the class into which the classifier classifies the tumor. In some example embodiments, the at least two classes are a low-risk tumor class and a high-risk tumor class, the image evaluator is a convolutional neural network configured to measure the lymphocyte distribution of lymphocytes of different region types of the image, the classifier is a bi-directional Gaussian mixture model configured to determine a probability distribution of features of tumors of the respective class within a feature space for the respective class, the system further comprises a Cox proportional hazards model for clinical features of the tumor to determine the class of the tumor, and the tumor evaluator is further configured to determine the clinical value (such as a prognosis) of the individual based on a prognosis of individuals having tumors of the class into which the classifier classifies the tumor and the class determined by the Cox proportional hazards model.

[0019] Some example embodiments can include a system comprising: an image evaluation apparatus for determining a lymphocyte distribution of lymphocytes in an image; a classification apparatus for classifying a tumor into a class selected from at least two classes respectively associated with the lymphocyte distribution; and a tumor evaluator apparatus for determining a clinical value (such as a prognosis) of an individual based on a tumor in the image by invoking the image evaluation apparatus with the image to determine the lymphocyte distribution of lymphocytes in the tumor, invoking the classifier to classify the tumor into a class based on the lymphocyte distribution, and outputting the clinical value (such as a prognosis) of the individual based on a prognosis of individuals having tumors of the class into which the classifier classifies the tumor.

[0020] Some example embodiments can include an apparatus comprising a memory storing instructions and processing circuitry configured by execution of the instructions stored in the memory to determine a clinical value (such as a prognosis) for an individual based on a tumor in an image by: determining a lymphocyte distribution of lymphocytes in the tumor based on an image of the tumor; applying a classifier to the lymphocyte distribution to classify the tumor, the classifier configured to classify tumors into a class selected from at least two classes respectively associated with lymphocyte distributions; and outputting the clinical value (such as a prognosis) for the individual based on a prognosis of individuals having tumors of the class to which the classifier classifies the tumor. In some example embodiments, the at least two classes are a low-risk tumor class and a high-risk tumor class, determining the lymphocyte distribution further comprises applying a convolutional neural network to the image, the convolutional neural network configured to measure the lymphocyte distribution of lymphocytes of different region types of the image, the classifier comprises a bi-directional Gaussian mixture model configured to determine, for a respective class, a probability distribution of features of tumors of the class within a feature space, and the instructions further cause the processing circuitry to apply a Cox proportional hazards model to clinical features of the tumor to determine the class of the tumor, and determine the clinical value (such as a prognosis) for the individual based on a prognosis of individuals having tumors of the class to which the classifier classifies tumors and the class determined by the Cox proportional hazards model.

[0021] Some example embodiments can include a non-transitory computer-readable medium storing instructions that, when executed by processing circuitry, cause the processing circuitry to determine a clinical value (such as a prognosis) for an individual based on a tumor in an image by: determining a lymphocyte distribution of lymphocytes in the tumor based on an image of the tumor; applying a classifier to the lymphocyte distribution to classify the tumor, the classifier configured to classify tumors into a class selected from at least two classes respectively associated with lymphocyte distributions; and outputting the clinical value (such as a prognosis) for the individual based on a prognosis of individuals having tumors of the class to which the classifier classifies the tumor. In some example embodiments, the at least two classes are a low-risk tumor class and a high-risk tumor class, determining the lymphocyte distribution further comprises applying a convolutional neural network to the image, the convolutional neural network configured to measure the lymphocyte distribution of lymphocytes of different region types of the image, the classifier comprises a bi-directional Gaussian mixture model configured to determine, for a respective class, a probability distribution of features of tumors of the class within a feature space, and the instructions further cause the processing circuitry to apply a Cox proportional hazards model to clinical features of the tumor to determine the class of the tumor, and determine the clinical value (such as a prognosis) for the individual based on a prognosis of individuals having tumors of the class to which the classifier classifies tumors and the class determined by the Cox proportional hazards model. BRIEF DESCRIPTION OF DRAWINGS

[0022] The present disclosure will be more fully understood from the following detailed description taken in conjunction with the accompanying drawings. In the drawings, like reference numerals can be repeated to identify similar and / or identical elements.

[0023] Figure 1 is an illustration of an example convolutional neural network.

[0024] Figure 2A is an illustration of an example image analysis for identifying region types and lymphocyte distribution in a tumor image, according to some example embodiments.

[0025] Figure 2B is an illustration of an example image analysis for classifying lymphocyte distribution in a tumor image, according to some example embodiments.

[0026] Figure 3 is an illustration of a set of masks of a lung tissue sample including lymphocyte distribution of lymphocytes by an example machine learning model, according to some example embodiments.

[0027] Figure 4 is an illustration of an example machine learning model for classifying tumors, according to some example embodiments.

[0028] Figure 5 is an illustration of a representation of a set of images of a pancreatic cancer tissue sample, according to some example embodiments.

[0029] Figure 6A is an illustration of a set of samples arranged in a two-dimensional feature space.

[0030] Figure 6B is an illustration of a Gaussian mixture model configured to classify the set of samples as a set of clusters of probability distributions within the two-dimensional feature space.

[0031] Figure 6C is another illustration of a Gaussian mixture model configured to classify the set of samples as a set of clusters of probability distributions within the two-dimensional feature space.

[0032] Figure 7 is an illustration of selecting a subset of features from a set of features of features within a feature space for a classifier based on a relevance of respective features to respective classes, according to some example embodiments.

[0033] Figure 8 is an illustration of a classification of tumors of different classes based on the subset of features, according to some example embodiments.

[0034] Figure 9 is an illustration of a Kaplan Meier survival ability plot based on image analysis, according to some example embodiments.

[0035] Figure 10 is an illustration of a selection of a subset of clinical features from a set of clinical features of clinical features within a feature space for a Cox proportional hazards model based on a relevance of respective features to respective categories according to some example embodiments.

[0036] Figure 11 is an illustration of a Kaplan Meier survival ability plot based on image analysis and a Cox proportional hazards model according to some example embodiments.

[0037] Figure 12 is an illustration of a result set of a classification of a tumor training dataset and a tumor testing dataset based on image analysis and a Cox proportional hazards model according to some example embodiments.

[0038] Figure 13 is a flowchart of a first example method according to some example embodiments.

[0039] Figure 14 is a flowchart of a second example method according to some example embodiments.

[0040] Figure 15 is a component block diagram of an example apparatus according to some example embodiments.

[0041] Figure 16 is a component block diagram of another example apparatus according to some example embodiments.

[0042] Figure 17 is an illustration of an example computer-readable medium according to some example embodiments.

[0043] Figure 18 is an illustration of an example apparatus in which some example embodiments can be implemented. DETAILED DESCRIPTION A. INTRODUCTION

[0044] The following introduction is intended to provide an overview of some machine learning features related to some example embodiments.

[0045] Figure 1 is an example of a convolutional neural network (CNN) 110 trained to process n-dimensional inputs to detect a plurality of features.

[0046] In the example of Figure 1 a convolutional neural network 110 processes an image 102 as a two-dimensional array of pixels 108 of one or more colors, but some such convolutional neural networks 110 can process other forms of data, such as sound, text, or signals from sensors.

[0047] In the example of Figure 1In the example of FIG. 1, a training dataset 100 is provided as a set of images 102 each associated with a class 106 from a set of classes 104. For example, the training dataset 100 can include a first vehicle image 102-1 associated with a first class 106-1 for vehicle images; a second house image 102-2 associated with a second class 106-2 for house images; and a third cat image 102-3 associated with a third class 106-3 for cat images. The association of an image 102 with a corresponding class 106 is sometimes referred to as a label of the training dataset 100.

[0048] As Figure 1 As further shown, each image 102 can be processed by a convolutional neural network 110 organized as a sequence of convolutional layers 112 each having a set of neurons 114 and one or more convolutional filters 116. In a first convolutional layer 112-1, each neuron 114 can apply a first convolutional filter 116-1 to a certain region of the image and can output an activation indicating whether the pixels in that region correspond to the first convolutional filter 116-1. The set of activations produced by the neurons 114 of the first convolutional layer 112-1, referred to as a feature map 118-1, can be received as input by a second convolutional layer 112-2 in the sequence of convolutional layers 112 of the convolutional neural network 110, and the neurons 114 of the second convolutional layer 112-2 can apply a second convolutional filter 116-2 to the feature map 118-1 produced by the first convolutional layer 112-1 to produce a second feature map 118-2. Similarly, the second feature map 118-2 can be received as input by a third convolutional layer 112-3, and the neurons 114 of the third convolutional layer 112-3 can apply a third convolutional filter 116-3 to the feature map 118-2 produced by the second convolutional layer 112-2 to produce a third feature map 118-3. Such machine learning models including a large number of multiple layers or more complex layer architectures are sometimes referred to as deep learning models.

[0049] As Figure 1Further shown, the third feature map 118-3 produced by the third and final convolutional layer 112-3 can be received by a classification layer 120, such as a“dense” or fully connected layer, which can perform a classification of the third feature map 118-3 to determine a classification of the content of the image 102-1. For example, each neuron 114 of the classification layer 120 can apply a weight to each activation of the third feature map 118-3. Each neuron 114 outputs an activation that is a sum of each activation of the third feature map multiplied by the weight connecting the neuron 114 to the activation. Thus, each neuron 114 outputs an activation that indicates a degree to which the activations included in the third feature map 118-3 match the corresponding weight of the neuron 114. Further, the weight of each neuron 114 is selected based on the activations of the third feature map 118-3 produced by an image 102 of one of the classes 106 in the class set 104. That is, each neuron 114 outputs an activation based on a similarity of the third feature map 118-3 of the currently processed image 102-1 to the third feature map 118-3 produced by the convolutional neural network 110 for an image 102 of one of the classes 106 in the class set 104. A comparison of the outputs of the neurons 114 of the classification layer 120 can allow the convolutional neural network 110 to perform the classification 122 by selecting the class 106-4 for which the probabilities corresponding to the third feature map 118-3 are highest. In this way, the convolutional neural network 110 can perform the classification 122 of the image 102-1 as the class 106 of the most similar image 102 in the plurality of images 102.

[0050] As Figure 1 Further shown, a training process can be applied to train the convolutional neural network 110 to recognize the class set 104 represented by the particular set of images 102 of the training data set 100. During the training process, each image 102 of the training data set 100 can be processed by the convolutional neural network 110, producing a classification 122 of the image 102. If the classification 122 is incorrect, the convolutional neural network 110 can be updated by adjusting the weights of the neurons 114 of the classification layer 120 and the filters 116 of the convolutional layers 112, such that the classification 122 of the convolutional neural network 110 is closer to the correct classification 122 of the image 102 being processed. Repeating the training of the convolutional neural network 110 on the training data set 100, while incrementally adjusting the convolutional neural network 110 to produce the correct classification 122 of each image 102, can result in a convergence of the convolutional neural network 110, where the convolutional neural network 110 correctly classifies the images 102 of the training data set 100 within an acceptable range of error. Examples of convolutional neural network architectures include ResNet and Inception.

[0051] As Figure 1The machine learning models at issue, such as the convolutional neural network 110, can classify inputs such as the image 102 based on arrangements of features relative to one another, such as the number, orientation, and positioning of identifiable features. For example, the convolutional neural network 110 can classify the image 102 by producing a first feature map 118-1 indicating detection of certain geometric shapes such as curves and lines appearing at various locations within the image 102, a second feature map 118-2 indicating that the geometric shapes are arranged to produce certain higher-level features such as a set of curves arranged as a circle or a set of lines arranged as a rectangle, and a third feature map 118-3 indicating that the higher-level features are arranged to produce still higher-level features such as a set of circles arranged as wheels or a set of rectangles arranged as a doorframe. The neurons 114 of the classification layer 120 of the convolutional neural network 110 can determine that the features of the third feature map 118-3, such as two wheels positioned between two doorframes, are arranged in a way that depicts a side of a vehicle such as a car. Similar classifications 122 can be made by other neurons 114 of the classification layer 120 to classify the image 102 as belonging to other classes 106 of the set of classes 104, such as an arrangement of two eyes, two triangular ears, and a nose that depicts a cat, or an arrangement of windows, doorframes, and a roof that depicts a house. In this way, machine learning models such as convolutional neural networks can classify inputs such as the image 102 based on feature arrangements corresponding to similar feature arrangements depicted in inputs such as the training dataset 100. Additional details regarding convolutional neural networks and other machine learning models, including support vector machines, can be found in U.S. Patent Application 62 / 959,931, which is incorporated by reference herein as if fully rewritten. B. Distribution-Based Classification

[0052] In some machine learning scenarios, classification of an input such as an image can occur based on arrangements of identifiable features relative to one another, such as the number, orientation, and / or positioning of identifiable pixel patterns detected in a feature map 118 of a convolutional neural network 110. However, in some other scenarios, classification can not be based on feature arrangements relative to one another Arrangement but rather on features in the input DistributionAs an example, a density of features on an input region can be compared to a recognizable density of features as a characteristic of a class 106. That is, a class 106 in a set of classes 104 can not be recognizable when a set of lower level features corresponds to a recognized arrangement (e.g., number, orientation, and / or positioning) of higher level features relative to each other that corresponds to a class 106. Instead, each class 106 can be recognized as a correspondence of an activation profile of features of an input to some characteristic of the input. Such a profile can not reflect any particular number, orientation, and / or positioning of activations of features of an input, but can instead indicate whether a profile of activations of features corresponds to a profile of activations of features of a respective class 106. In such a scenario, an input (e.g., a training dataset) for each class 106 can be associated with a characteristic profile of activations of features, and a classification of an input can be based on whether a profile of activations of features of the input corresponds to a profile of activations of features in an input for each class 106. Such a distribution-based classification can also occur in various scenarios.

[0053] Figure 2A and Figure 2B Together, examples of several types of image analysis that can be used in some example embodiments to identify region types and lymphocyte distributions in tumor images are shown.

[0054] Figure 2A is an illustration of example image analysis for identifying region types and lymphocyte distributions in tumor images according to some example embodiments. As shown, Figure 2A a dataset can include images 102 of tissue of individuals having a type of tumor, as well as interstitium including connective tissue and tumor stroma. Classification of features of the images 102 can enable determination of regions, e.g., portions of the images 102 having similar features. Further identification of regions of feature maps 118 including certain features of filters 116 can enable determination 200 of region types of respective regions, such as a first region of the images 102 depicting a tumor and a second region of the images 102 depicting interstitium. Thus, each filter 116 of a convolutional neural network 110 can be considered a mask of a region of the images 102 indicative of a particular region type, such as a presence, size, shape, and extent of a tumor or interstitium adjacent to a tumor.

[0055] Further, the image 102 can show the presence of lymphocytes, which can be distributed with respect to the tumor, stroma, and other tissue. Further analysis of the image can enable determination 202 of lymphocyte clusters 204 as areas of continuous regions and / or high lymphocyte concentration, e.g., by counting the number of lymphocytes present within a particular region of the image 102. Thus, the lowest convolutional layers 112 and filters 116 of the convolutional neural network 110 are able to identify features indicative of tumor, stroma, and lymphocytes.

[0056] Figure 2B is an illustration of example image analysis for classifying lymphocyte distribution in a tumor image, in accordance with some example embodiments. In Figure 2B In the example of FIG. 2, the first image analysis 206 can be performed by first dividing the image 102 into a set of regions, and classifying each region of the image 102 as tumor, tumor adjacent, stroma, stroma adjacent, or elsewhere. Thus, each region including a lymphocyte cluster 204 can further characterize the lymphocyte cluster 204 based on the region type, e.g., a first lymphocyte cluster 204-3 appearing within a tumor and a second lymphocyte cluster 204-4 appearing within stroma.

[0057] Second image analysis 208 can be performed to further compare the location of lymphocyte clusters 204 to the location of different region types, to further characterize the lymphocyte clusters 204. For example, first lymphocyte cluster 204-3 can be identified as occurring within a central portion of the tumor region, and / or within a first threshold distance of a location identified as a tumor centroid, and thus can be characterized as a tumor infiltrating lymphocyte (TIL) cluster. Similarly, second lymphocyte cluster 204-4 can be identified as occurring within a central portion of the interstitial region, and thus represents an interstitial infiltrating lymphocyte cluster. However, third cluster 204-5 can be identified as occurring within a peripheral portion of the tumor region, and / or a second threshold distance of the tumor (the second threshold distance being greater than the first threshold distance), and thus can be characterized as a tumor-adjacent lymphocyte cluster. Alternatively or additionally, third cluster 204-5 can be identified as occurring within a peripheral portion of the interstitial region, and / or a second threshold distance of the interstitium (the second threshold distance being greater than the first threshold distance), and thus can be characterized as an interstitial-adjacent lymphocyte cluster. Some example embodiments can classify regions as tumor, interstitial, or lymphocyte; with two labels, such as tumor and lymphocyte, interstitial and lymphocyte, or tumor and interstitial; and / or with three labels, such as tumor, interstitial, and lymphocyte. Some example embodiments can then be configured to identify lymphocyte clusters 204 occurring in each region of image 102, and tabulate these regions to determine a distribution. In this way, in some example embodiments, image analysis of image 102, including feature maps 118 provided by different filters 116 of convolutional neural network 110, can be used to identify and characterize the distribution and / or concentration of lymphocytes in a tumor image.

[0058] Figure 3 is an illustration of a mask set 300 of masks 302 of lung tissue samples including lymphocyte distributions by example machine learning models, in accordance with some example embodiments. As Figure 3As shown, masks 302 of the image 102 can be prepared, each mask 302 indicating regions of the image 102 corresponding to one or more region types. For example, a first mask 302-1 can indicate regions of the image 102 identified as tumor regions. A second mask 302-2 can indicate regions of the image 102 identified as stromal regions. A third mask 302-3 can indicate regions of the image 102 identified as lymphocyte regions. Other masks can also be characterized based on the feature distributions in the feature map 118. For example, a fourth mask 302-4 can indicate regions of the image 102 identified as tumor infiltrating lymphocyte regions. A fifth mask 302-5 can indicate regions of the image 102 identified as tumor adjacent lymphocyte regions. A sixth mask 302-6 can indicate regions of the image 102 identified as stromal infiltrating lymphocyte regions. A seventh mask 302-7 can indicate regions of the image 102 identified as stromal adjacent lymphocyte regions. An eighth mask 302-8 can indicate regions of the image 102 identified as overlapping stromal regions and tumor regions. A ninth mask 302-9 can indicate regions of the image 102 identified as tumor adjacent stromal regions.

[0059] In some example embodiments, the image 102 can also be processed to determine various measurements 304 of respective region types of the image 102. For example, a concentration (as a percentage of the image 102) of each region type can be calculated (e.g., a number of pixels 108 corresponding to each region, optionally taking into account an apparent concentration of features, such as a density or count of lymphocytes in respective regions of the image 102, compared to a total number of pixels of the image 102). In this way, in some example embodiments, based on the feature map 118, various measurements 304 of respective region types of the image 102 can be determined. Figure 2A and Figure 2B The image analysis of the image 102 shown by the distribution analysis can be aggregated into a mask 302 and / or a mask group 300 of quantities.

[0060] Figure 4 is an illustration of an example machine learning model to classify tumors according to some example embodiments. Figure 4 The example machine learning model of includes a first convolutional neural network 112-1 configured to perform region classification 402 of respective regions 400 of a tumor image 102 according to different classes such as tumor regions, stromal regions, lymphocyte regions, tumor infiltrating lymphocyte regions, etc. Based on the region classification 402, a mask group 300 of masks 302 can be generated, for example, a first mask 302-1 indicating regions 400 of the image 102 as tumor regions, a second mask 302-2 indicating regions 400 of the image 102 as stromal regions, and a third mask 302-3 indicating regions 400 of the image 102 as lymphocyte regions. Figure 4An example system of the includes a second convolutional neural network 112-2 configured to determine a density or concentration of various features, such as lymphocyte density range estimates 404 indicative of a count or percentage of lymphocytes in respective regions 400 of the image 102. Based on the lymphocyte density range estimates 404, a lymphocyte density map 406 can be generated that indicates regions 400 of the image 102 having a high density of lymphocytes, such as lymphocyte clusters. Based on the masks 302 in the mask set 300, the region classifications 402, and / or the lymphocyte density map 406 based on the lymphocyte density range estimates 404, the image evaluator 408 can identify: aggregate regions of the image 102, such as tumor regions 410-1, stromal regions 410-2, and lymphocyte regions 410-3; one or more region measurements, such as tumor measurements 304-1, stromal measurements 304-2, and lymphocyte measurements 304-3; and / or one or more regions indicative of a feature distribution, such as tumor-infiltrating lymphocyte regions, tumor-adjacent lymphocyte regions 412-1, tumor and stromal regions, and tumor-adjacent stromal regions 412-2.

[0061] Generally, in some such example embodiments, one or more convolutional neural networks can be trained to determine a lymphocyte distribution of lymphocytes in an image region, e.g., to classify an image region as one or more region types selected from a set of region types including a tumor region, a lymphocyte region, or a stromal region. In some example embodiments, a convolutional neural network can determine a lymphocyte distribution of lymphocytes in a tumor, including, for a respective lymphocyte region of an image: determining a distance of the lymphocyte region to one or both of a tumor region or a stromal region; and characterizing the lymphocyte region as one of a tumor-infiltrating lymphocyte region, a tumor-adjacent lymphocyte region, a stromal-infiltrating lymphocyte region, or a stromal-adjacent lymphocyte region based on the distance. In some example embodiments, a convolutional neural network can determine a lymphocyte distribution of lymphocytes in a tumor by, for a respective stromal region of an image: determining a distance of the stromal region to a tumor region, and characterizing the stromal region as one of a tumor-infiltrating stromal region or a tumor-adjacent stromal region based on the distance, and the classifier further classifies the tumor based on the characterization of the stromal region. Thus, the classifier can further classify the tumor based on the characterization of the lymphocyte region. In some example embodiments, a number of such convolutional neural networks can perform various analyses of an image that can inform a determination of a clinical value for an individual, such as a prognosis. C. Learning parameter determination

[0062] Figure 5 is an illustration of a characterization 500 of a set of images of pancreatic cancer tissue samples in accordance with some example embodiments. In Figure 5 In a graph of the, a set of tumors is characterized by Figure 4The illustrated system characterizes to determine the distribution of detected tumor features, such as tumor region 502-1, interstitial region 502-2, lymphocyte region 502-3, tumor-invading-lymphocyte region 502-4, tumor-adjacent-lymphocyte region 502-5, interstitial-and-tumor-invading-lymphocyte region 502-6, interstitial-adjacent-lymphocyte region 502-7, tumor-and-interstitial region 502-8, and tumor-adjacent-interstitial region 502-9. Each feature can be evaluated in terms of the density or concentration (vertical axis) and percentage (horizontal axis) of each feature in the tumor image 102. The set of tumor images can be further divided into a training image subset, which can be used to train a machine learning classifier, such as neural networks 112-1, 112-2, etc., to determine the features, and a test image subset, which can be used to evaluate the effectiveness of the machine learning classifier in determining features in previously unseen tumor images. In this way, the machine learning classifier can be validated to determine the consistency of the underlying logic when applied to new data. For example, Figure 5 The graph in FIG. 1 1 is developed based on diagnostic hematoxylin and eosin staining (H&E staining) of pathological images of pancreatic cancer patients who received chemotherapy.

[0063] Based on factors that are characteristics of tumors in each class, it can further be desirable to characterize the tumor as one of several classes, such as low-risk tumors and high-risk tumors. For example, tumors of the respective classes can also be associated with different features, such as the concentration and / or percentage of particular types of tumor regions (e.g., tumor-invading-lymphocyte), and differences in such characteristic features can enable distinguishing tumors of one class from tumors of another class. Further, different classes of tumors can be associated with different clinical characteristics, such as responsiveness to various treatment options and prognosis, such as survivability. To determine such clinical characteristics of a particular tumor in an individual, it can be desirable to determine the tumor class of the tumor in order to guide the selection of a diagnostic and / or treatment regimen for the individual.

[0064] However, in many diagnostic scenarios, it can be difficult to correlate features that are characteristic of different classes, such as different tumor classes. As a first such example, features of one class of tumors can differ from features of another class of tumors within a range of probabilities, and the range of probabilities can overlap substantially. For example, the density and percentage of tumor invasion lymphocytes for a high-risk tumor class and a low-risk tumor class can each fit bell curves of probabilities within the tumor classes, and the means of the bell curves can be only slightly offset, such that the probability distributions can overlap. Thus, it can be difficult to determine whether a tumor exhibiting features within the overlap region belongs to the high-risk class or the low-risk class. As a second such example, different features of a tumor can be co variant; for example, high-risk tumors and low-risk tumors can be distinguished based on a combined probability that distinguishes between a tumor invasion lymphocyte region and a tumor adjacent lymphocyte region. However, different features can also be inherently co variant in a non-diagnostic manner. For example, tumors exhibiting high density stromal invasion and a tumor invasion lymphocyte region will also necessarily typically exhibit a high density of tumor invasion lymphocyte regions. Thus, adding a class-based probability that a tumor belongs to a class based on stromal invasion and a tumor invasion lymphocyte region to a class-based probability that a tumor belongs to a class based on a tumor invasion lymphocyte region can exceed the likelihood of the tumor in the class, as the inherent covariance of the features is not accounted for. Due to these complex features of the data, it can be difficult to determine distinguishing features of each class of tumors, particularly in high-dimensional feature sets where many features can be available.

[0065] To classify data sets exhibiting such overlapping class data, various machine learning models can be used. The respective machine learning models can provide different capabilities for classifying overlapping data sets, for example, based on uniqueness, tolerance to false positives, tolerance to false negatives, scalability to large numbers of features, and avoidance of characteristics such as overfitting and underfitting.

[0066] Figure 6A to Figure 6C An example of a Gaussian mixture model is shown together that can be developed to classify overlapping class data, such as classifying tumors into low-risk tumors and high-risk tumors based on two features, which can be used in some example embodiments.

[0067] Figure 6A is a plot of a set of samples arranged in a two-dimensional feature space 606. In Figure 6AIn the two-dimensional feature space 606, samples 600-1 of a first class 602-1 (denoted as a circle) and samples 600-2 of a second class 602-2 (denoted as a cross) are involved. Each sample 600 can be evaluated and quantified for the first feature 604-1 and the second feature 604-2, which allows each sample to be located within the two-dimensional feature space 606, where the vertical axis represents the first feature 604-1 and the horizontal axis represents the second feature 604-2. Within the two-dimensional feature space 606, samples 600 of each class 602 can be clearly clustered, but clusters can also overlap, such that samples within overlapping regions can belong to either class 602. For a particular sample 600, it may be desirable to determine the probability that sample 600 belongs to each class 602 based on feature 604 of sample 600, especially in overlapping regions associated with samples 600 of multiple classes 602. Although clustering in Figure 6A While this may be obvious in a simple illustration, such clustering may be more difficult to determine, for example, in a feature space 606 with higher dimensions, in a dataset characterized by a category 602 with greater overlap, and / or in a dataset where two or more features 604 covariate, for determining the diagnosis or inherent covariance of feature 604 for those two or more features.

[0068] Various machine learning models can be used to classify overlapping datasets, such as Figure 6A As shown. Some such models include, for example, Bayesian (including Naive Bayes) classifiers; Gaussian classifiers; probabilistic classifiers; principal component analysis (PCA) classifiers; linear discriminant analysis (LDA) classifiers; quadratic discriminant analysis (QDA) classifiers; single-layer or multi-layer perceptron networks; convolutional neural networks; recurrent neural networks; nearest neighbor classifiers; linear SVM classifiers; radial basis function kernel (RBF) SVM classifiers; Gaussian process classifiers; decision tree classifiers, including random forest classifiers; and / or restricted or unrestricted Boltzmann machines, etc.

[0069] Figure 6B It is configured to Figure 6A The illustration shows a Gaussian mixture model classifying the group of samples 600 as a cluster of probability distributions within a two-dimensional feature space 606. In some example embodiments, this Gaussian mixture model can be used to distinguish different categories of tumors. Figure 6BIn this model, a first Gaussian probability distribution 608-1 can be identified for samples 600-1 of the first category 602-1, and a second Gaussian probability distribution 608-2 can be identified for samples 600-2 of the second category 602-2. For example, based on the mean and variance of samples 600 for each feature 604, a Gaussian probability distribution 608 for each category 602 can be fitted to samples 600 of each category 602. The selection of the Gaussian probability distribution 608 can also take into account other factors, such as avoiding false negatives (e.g., samples 600 of category 602 are incorrectly excluded from category 602) and / or avoiding false positives (e.g., samples 600 of different categories 602 are incorrectly included in category 602). Furthermore, the Gaussian probability distribution 608 can be selected to model covariance, for example, by associating the distribution of the Gaussian probability distribution 608 of the first feature 604-1 with the distribution of the Gaussian probability distribution 608 of the second feature 604-2. For example, a first feature 604-1 and a second feature 604-2 can be selected, or a similar bias and / or Gaussian probability distribution 608 can be selected independently for each feature 604. For a specific sample 600 (such as an image of a tumor of unknown category), features 604 of sample 600 can be evaluated to locate sample 600 within feature space 606, and the relative probabilities within the Gaussian probability distribution 608 of the corresponding category 602 can be compared to determine the possible category 602 of the tumor.

[0070] like Figure 6B Furthermore, the goodness of fit of the selected Gaussian mixture model can also be evaluated, for example, as an estimate of the diagnostic properties of the Gaussian mixture model. For example, a contour score 610 can be determined for each Gaussian probability distribution 608, where the contour score indicates the contour coefficient 612 (e.g., the number of samples 600 of category 602 within a selected distance from the mean or centroid of the Gaussian probability distribution 608). The discriminative properties of the Gaussian mixture model can be improved by selecting a Gaussian probability distribution 608 with similar contour scores 610. Figure 6B As shown, the contour scores of Gaussian probability distribution 608 are dissimilar. For example, because the first Gaussian probability distribution 608-1 of the first class 602-1 represents a greater number of samples 600-1 than the second Gaussian probability distribution 608-2 of samples 600-2 of the second class 602-2 (i.e., the first Gaussian probability distribution 608-1 has a higher contour than the second Gaussian probability distribution 608-2), and also because the distances of samples 600-1 of the first class 602-1 are more widely distributed in feature space 606 than those of samples 600-2 of the second class 602-2, a larger range of contour coefficients 612 results (i.e., the contour of the first Gaussian probability distribution 608-1 is longer than the contour of the second Gaussian probability distribution 608-2). Therefore, improvements can be made by selecting different mixtures of Gaussian probability distribution 608. Figure 6Ba Gaussian mixture model.

[0071] Figure 6C is another illustration of a Gaussian mixture model configured to classify the set of samples as a set of clusters of probability distributions within a two-dimensional feature space, and in some example embodiments, the Gaussian mixture model can be used to distinguish between different classes of tumors. In Figure 6C , the Gaussian probability distribution of the first class 602 is instead identified as a first Gaussian probability distribution 608-4 of a first cluster of samples 600-1 of the first class 602-1 and a second Gaussian probability distribution 608-5 of a second cluster of samples 600-1 of the first class 602-1. Further, for each Gaussian probability distribution 608, a mixing parameter indicative of a proportion of samples 600 represented by the Gaussian probability distribution 608 that are of the class 602 can be identified. For example, the first Gaussian probability distribution 608-4 of the first class 602-1 can fit a smaller number of samples 600-1 of the first class 602-1 than the second Gaussian probability distribution 608-5 of the first class 602-1, and thus can have a first mixing parameter 614-1 that is smaller than a second mixing parameter 614-2 of the second Gaussian probability distribution 608-5. When a sample 600 of an unknown class 602 is positioned within the feature space 606, the probability that the sample 600 is classified as each class 602 can be determined as a sum of a product of a probability distribution of a location of the sample 600 with a mixing parameter 614 of each Gaussian probability distribution 608. Further, the classification ability of the Gaussian mixture model can be assessed based on the silhouette scores 610 of the respective Gaussian probability distributions 608; for example, a similarity of a sample size and a silhouette coefficient 612 of each Gaussian probability distribution 608 can indicate a more reliable and more predictive classifier than a Gaussian mixture model. Figure 6B

[0072] Alternatively or in addition to Figure 6B and Figure 6C ​Beyond the contour scores shown, other measures can be used to determine the classification ability of a Gaussian mixture model. As an example, for tumors in different tumor categories (e.g., low-risk and high-risk categories) respectively associated with survival ability, a consistency index (“C-index”) can be developed to indicate the degree of consistency between the predicted survival time of an individual with a tumor in a tumor category and the actual survival time of an individual with a tumor in a tumor category. The consistency index can be determined based on each feature of the tumor category to determine the degree to which the feature corresponds to the predicted survival rate of an individual with a tumor in that tumor category, where a high consistency index indicates a highly predictive feature of the Gaussian mixture model and a low consistency index indicates a poorly predictive feature of the Gaussian mixture model. Because the consistency index for each feature depends on the selected Gaussian mixture model, it may be desirable to limit the number of features to those exhibiting a high consistency index, alternatively or additionally limiting it to the contour scores of the corresponding Gaussian probability distributions for each category. Choosing such features can reduce the dimension of the feature space 606 of the dataset to a smaller set of features with higher discriminative power for the corresponding categories 602, which can result in a more accurate, precise, and / or efficient classification process.

[0073] Figure 7 This is an illustration of a process 700 in which a subset of clinical features is selected for a classifier from a set of clinical features within a feature space based on the correlation between corresponding clinical features and corresponding categories, according to some example embodiments. A Gaussian mixture model is developed for a set of nine clinical features, such as... Figure 5 The nine clinical features are shown. A set of contour scores and concordance indices can be determined for each clinical feature. In a set of available clinical features 704, in a first selection step 702-1, a first clinical feature 706-1, such as the concentration (specifically, percentage) of lymphocytes in the tumor-adjacent lymphocyte region, can be selected from the available clinical features 704 to provide the highest contour score and / or concordance index. In the remaining clinical features (i.e., all clinical features other than the first selected clinical feature), in a second selection step 702-2, a second Gaussian mixture model can be developed, and a second clinical feature 706-2, such as the concentration (specifically, percentage) of the stroma and tumor region, can be selected from the remaining clinical features to provide the highest contour score and / or concordance index. Similar selection steps 702-3, 702-4 can be performed to select a third clinical feature 706-3 (such as the concentration of tumor invading the stroma region) and a fourth clinical feature 706-4 (such as the concentration of lymphocytes), each of which provides an improved concordance score compared to the previously selected clinical features, which indicates the complementary classification ability of the selected clinical features compared to the other remaining clinical features. The selection process can continue up to the fifth selection step 702-5, in which the selected clinical features are determined. NoThe consistency index of previously selected clinical features is improved, and further clinical features can not be selected for the subset of clinical features.

[0074] In some example embodiments, a classifier for tumors can include a Gaussian mixture model configured to determine, for a respective class, a probability distribution of features of tumors of the class within a feature space, the features can be selected from a feature set including measurements of tumor regions of images, measurements of interstitial regions of images, measurements of lymphocyte regions of images, measurements of tumor infiltrating lymphocyte regions of images, measurements of tumor adjacent lymphocyte regions of images, measurements of interstitial infiltrating lymphocyte regions of images, measurements of interstitial adjacent lymphocyte regions of images, measurements of tumor infiltrating interstitial regions of images, and measurements of tumor adjacent interstitial regions of images. In some example embodiments, a subset of features of the Gaussian mixture model can be selected based on a relevance of the respective class to the respective features of the subset, where the relevance can be based on one or both of a silhouette score or a consistency index of the feature space. In some example embodiments, the subset of features can consist essentially of measurements of lymphocyte regions of images, measurements of tumor infiltrating lymphocyte regions of images, measurements of tumor adjacent lymphocyte regions of images, and measurements of tumor infiltrating interstitial regions of images. D. Image-based tumor assessment

[0075] In some example embodiments, a clinical value (e.g., prognosis) of an individual can be determined based on a tumor shown in an image by determining a lymphocyte distribution of lymphocytes in the tumor based on the image, applying a classifier to the lymphocyte distribution to classify the tumor, the classifier having been trained to classify tumors into classes selected from at least two classes respectively associated with lymphocyte distributions, and determining the clinical value (e.g., prognosis) of the individual based on a prognosis of the individual having the tumor classified by the classifier into the class. The classifier can be invoked to determine a lymphocyte distribution of lymphocytes in a respective region of a tumor image.

[0076] Figure 8 is an illustration of classification of different classes of tumors based on a subset of features in accordance with some example embodiments. Figure 8A comparison 800 of a subset of selected features 806 to images 804 of tumors in a low risk tumor class 802-1 and a high risk tumor class 802-2 is presented, i.e., a percentage of area in each image 804 of each class 802 corresponding to each feature 806 in the subset of features. The high risk tumor class can be associated with a first survival probability and the low risk tumor class can be associated with a second survival probability that is longer than the first survival probability. For example, the percentages of the respective features 806 of the images 804 of the tumor classes 802 can be compared to determine a degree of diagnosis of the features 806 to the respective tumor classes 802. For example, the images 804 of the high risk tumor class 802-2 can exhibit smaller and more consistent ranges of values for the first feature 806-1 and the third feature 806-2 compared to the images 804 of the low risk tumor class 802-1. Also, the values of the third feature 806-3 can generally be higher in the images of the tumors of the low risk tumor class 802-1 than in the images of the tumors of the high risk tumor class 802-2. The measurements determined based on the selected features 706 of the subset of features of the tumor classes 802 can present clinically meaningful findings in the pathology of the tumors in the different tumor classes 802 and can be used by clinicians and automated processes, such as diagnostic and / or prognostic machine learning processes, to classify tumors into the different tumor classes 802. Figure 7 The selection process 700 can be repeated for different subsets of features of the tumor classes 802 to determine different measurements based on different selected features 706 of the subsets of features of the tumor classes 802. The measurements determined based on the selected features 706 of the subset of features of the tumor classes 802 can present clinically meaningful findings in the pathology of the tumors in the different tumor classes 802 and can be used by clinicians and automated processes, such as diagnostic and / or prognostic machine learning processes, to classify tumors into the different tumor classes 802.

[0077] Figure 9 is a plot of Kaplan Meier survival curves based on image analysis according to some example embodiments. In Figure 9 , first and second Kaplan Meier survival curves 900-1 and 900-2, respectively, are generated (e.g., a percentage of surviving individual populations measured in days after diagnosis) for training and test sets of individual populations of tumors based on low risk tumor classes 802-1 and high risk tumor classes 802-2 having Figure 8 Further, a set of tumors for which images and data are available are divided into training and test sets. Machine learning models, including convolutional neural networks and / or Gaussian mixture models, are trained on the images of the training set to a point of convergence in which the machine learning models produce outputs within an accuracy range of expected outputs. The test set is then used to test the machine learning models to determine whether the machine learning models produce outputs of new data that are consistent with the expected outputs. Such validation can include a cross-validation process in which the set of tumors is first divided into multiple subsets and repeated training and testing is performed using a selection from a subset of the training set and a remaining subset of the test set.

[0078] As Figure 9As shown, the image-based tumor assessment techniques presented herein performed classification on a training dataset with a hazard ratio (HR) of 0.5117, a statistical P-value of 0.0570, and a concordance index of 0.6667, and demonstrated performance on a test set with a hazard ratio of 0.5154, a statistical P-value of 0.3405, and a concordance index of 0.5964. According to some example embodiments, many such machine learning models can be trained to classify tumors and determine clinical values (such as prognosis and / or survivability) of individuals. E. Cox Proportional Hazards Model

[0079] In some example embodiments, the image-based prognosis determination techniques can be combined with a Cox proportional hazards model, which can improve the prognostic capabilities of tumor analysis. The Cox proportional hazards model is a regression model that correlates clinical features, such as demographic features of individuals, clinical observations of individuals and tumors, and pathology measurements, with different tumor categories to determine the contribution of each clinical feature to tumor classification. For example, the regression model can determine that an individual within a particular age range, with particular personal habits (such as smoking or drinking), and with a cancer staging score based on the American Joint Committee on Cancer (AJCC) cancer staging system is more likely to be classified as having a tumor within a low-risk tumor category, while an individual within another age range, with other personal habits, and with other cancer staging scores is more likely to be classified as having a tumor within a high-risk tumor category.

[0080] The Cox proportional hazards model can be developed using a training set featuring tumors with known clinical features. Stepwise selection can be performed to select a subset of clinical features that significantly contribute to classification, for example, by removing clinical features that do not significantly improve the predictability of other clinical features. The Cox proportional hazards model can also be trained on two or more categories of tumors to determine different proportional survivability of tumors in different tumor categories, such as a low-risk tumor category with one set of shared characteristics and / or similar survivability metrics, and a high-risk tumor category with another set of shared characteristics and / or other similar survivability metrics.

[0081] Figure 10 is an illustration of selecting a subset of features for a Cox proportional hazards model from a feature set 1000 of features within a feature space according to some example embodiments. In Figure 10In this study, for a set of tumors from an individual and pathologically evaluated, values ​​for a clinical feature set 1000 are identified. This clinical feature set includes: the initial diagnosis of the tumor (e.g., AJCC stage score for T-class); the tumor's measurement outcome (e.g., AJCC stage score for N-class); the tumor's treatment; the location of the tumor; the individual's smoking frequency; the tumor's metastatic status; the individual's prior cancer history; the individual's duration of smoking; the individual's initial diagnosis; the individual's alcohol consumption history; and the individual's sex. The first step 1002-1 of the regression analysis can determine the degree to which each feature of the feature set 1000 distinguishes between tumor categories (e.g., low-risk and high-risk), and the features can be ranked, for example, by statistical p-values. Features with p-values ​​within a certain range (e.g., a statistical significance threshold below 0.05) can be selected as a subset of features, and other features can be excluded. Additional steps 1002-2, 1002-3, and 1004 of the regression analysis can be performed to exclude other features and retain other features in feature set 1000 until no features can be excluded without significantly reducing the classification accuracy of the Cox proportional hazards model. Feature set 1004, obtained based on the correlation between the corresponding category and the corresponding feature of the subset, can be identified as the retained features of the Cox proportional hazards model.

[0082] like Figure 10 As shown, the Cox proportional hazards model developed in this manner identifies a subset of features consisting of tumor measurements and tumor metastasis status. For the training set, the Cox proportional hazards model exhibits a hazard ratio of 0.2182 and a statistical p-value of 0.0200, and for the test set, it exhibits a hazard ratio of 0.4065 and a statistical p-value of 0.2855. Based on some example embodiments, many such Cox proportional hazards models can be identified for tumor classification. F. Combinatorial Model

[0083] In some example embodiments, image-based classification (e.g., based on convolutional neural networks and Gaussian mixture models) can be combined with a Cox proportional hazards model to classify tumors based on image features and clinical features. That is, the at least two categories are low-risk tumor categories and high-risk tumor categories; determining lymphocyte distribution can further include applying a convolutional neural network to the image, the convolutional neural network being configured to measure the lymphocyte distribution of different regions of the image; the classifier can be a bidirectional Gaussian mixture model, the bidirectional Gaussian mixture model being configured to determine the probability distribution of features of tumors of that category in a feature space for the corresponding category; the Cox proportional hazards model can be applied to the clinical features of the tumor to determine the tumor category; and determining an individual's clinical value (e.g., prognosis) can be further based on the category determined by the Cox proportional hazards model.

[0084] Figure 11 is a plot of Kaplan Meier survival capability based on image analysis and Cox proportional hazards model according to some example embodiments. In Figure 11 , first and second Kaplan Meier survival capability plots 900-3 and 900-4, respectively, are generated for a training set and a test set of individual populations of tumors based on low risk tumor category 802-1 and high risk tumor category 802-2 having Figure 8 Further, a set of tumors for which images and data including clinical features are available are divided into a training set and a test set. Machine learning models including a convolutional neural network, a Gaussian mixture model, and a Cox proportional hazards model are trained on images of the training set to a point of convergence at which the machine learning models produce outputs within an accuracy range of expected outputs. The test set is then used to test the machine learning models to determine whether the machine learning models produce outputs of new data that are consistent with expected outputs. Such validation can include a cross-validation process in which the set of tumors is first divided into a plurality of subsets and repeated training and testing is performed using selections from a subset of the training set and a remaining subset of the test set.

[0085] As shown in Figure 11 , the image-based tumor assessment techniques presented herein performed classification on a training data set having a hazard ratio (HR) of 0.2545, a statistical P-value of 0.0065, and a concordance index of 0.7141, and demonstrated performance on a test set having a hazard ratio of 0.3742, a statistical P-value of 0.0696, and a concordance index of 0.6120.

[0086] Figure 12 is a plot 1200 of results sets of classification of a tumor training data set and a tumor test data set based on image analysis and Cox proportional hazards model according to some example embodiments. As shown in Figure 12 , the classification results of a combined model featuring image-based analysis and statistical analysis of clinical features demonstrated higher classification accuracy than either model used alone. According to some example embodiments, many such machine learning models can be trained to classify tumors and determine clinical values (such as prognosis and / or survival capability) of individuals. G. Tumor Assessment and Output

[0087] In some example embodiments, the tumor analysis models disclosed herein can be used to determine and output, for a user, an individual’s clinical value based on a tumor shown in an image. The user can be, for example, the individual with the tumor; a family member or guardian of the individual; or a medical service provider, including a physician, a nurse, or a clinical pathologist. The clinical value and / or output can be, for example, one or more of: a diagnosis of the individual, a prognosis of the individual, a survivability of the individual, a classification of the tumor, a diagnosis and / or treatment recommendation for the individual, and the like.

[0088] Some example embodiments can use the determination of the tumor analysis models to display a visualization of the individual’s clinical value, such as a prognosis. For example, a terminal can receive an image of a tumor of an individual, and optionally a set of clinical features, such as demographic features of the individual, clinical observations of the individual and the tumor, and pathological measurements. The terminal can apply the tumor analysis models (e.g., process the image by a convolutional neural network and a Gaussian mixture model, and optionally, process the clinical features by a Cox proportional hazards model) to determine a class of the tumor (such as a low-risk tumor class and a high-risk tumor class) and a prognosis associated with individuals having a tumor of the tumor class. For example, the clinical value can be determined as a survivability (such as a predicted survival duration and probability), optionally including a confidence or accuracy for each probability. In some example embodiments, the clinical value can be presented as a visualization, such as a Kaplan Meier survivability projection of the tumor. In some example embodiments, the visualization can include additional information about the tumor, such as one or more of: a mask 302 indicating a region type of a region of the image 102; a measurement 304 of the image 102, such as a concentration of each region type (e.g., a concentration of lymphocytes in one or more regions determined by binning), and / or a percentage of the region of the region type compared to the entire image 102. In some example embodiments, the visualization can include additional information about the individual, such as clinical features of the individual, and can indicate how the respective clinical features contributed to determining the clinical value (such as the prognosis) of the individual.

[0089] Some example embodiments can use the determination of the tumor analysis models to determine and display, for a user, a diagnostic test of the tumor based on the individual’s clinical value, such as a prognosis. For example, based on the tumor analysis model classifying the tumor as a low-risk class, the apparatus can recommend a less invasive test (such as a blood test or imaging) to further characterize the tumor. Based on the tumor analysis model classifying the tumor as a high-risk class, the apparatus can recommend a more invasive test (such as a biopsy) to further characterize the tumor. Some example embodiments can also display, for the user, an explanation of the basis of the determination; a set of options for further testing; and / or a recommendation of one or more options to consider by the individual and / or a medical service provider.

[0090] Some example embodiments can use the determination of the tumor analysis model to determine and display to the user a treatment for the individual based on the individual’s clinical value (e.g., prognosis). For example, based on the tumor analysis model classifying the tumor as a low-risk category, the device can recommend a less aggressive treatment for the tumor, such as less aggressive chemotherapy. Based on the tumor analysis model classifying the tumor as a high-risk category, the device can recommend a more aggressive treatment for the tumor, such as more aggressive chemotherapy and / or surgical removal. Some example embodiments can also display to the user an explanation of the basis for the determination; a set of options for further testing; and / or a recommendation of one or more options for consideration by the individual and / or a medical service provider.

[0091] Some example embodiments can use the determination of the tumor analysis model to determine and display to the user a schedule of therapeutic agents for treating the tumor based on the individual’s clinical value (e.g., prognosis). For example, based on the tumor analysis model classifying the tumor as a low-risk category, the device can recommend chemotherapy at a lower frequency, later date, and / or lower dosage. Based on the tumor analysis model classifying the tumor as a high-risk category, the device can recommend a more aggressive treatment for the tumor, such as chemotherapy at a higher frequency, earlier date, and / or higher dosage. Some example embodiments can also display to the user an explanation of the basis for the determination; a set of options for further testing; and / or a recommendation of one or more options for consideration by the individual and / or a medical service provider. In some example embodiments, many such types of classifications and outputs of the individual’s clinical value and information about the tumor can be provided. H. Technical Effects

[0092] Some example embodiments of feature analysis using distribution-based machine learning classifiers can exhibit a variety of technical effects.

[0093] A first example of a technical effect that some example embodiments can exhibit is novel input classification based on distributions, which can be difficult to achieve with other machine learning models. For example, as shown in Figure 9 the image-based tumor classification models disclosed herein are able to classify tumors with reasonable accuracy. As shown in Figure 11 and Figure 12Further shown, a combined model including image-based analysis (e.g., classification based on convolutional neural networks and Gaussian mixture models) and regression-based clinical feature analysis (e.g., based on Cox proportional hazards models) can have higher classification accuracy than either model used alone. In some scenarios, the use of machine learning models, including visualization and / or interpretation of the basis of such determinations of clinical values for individuals (e.g., indications of image features and clinical features that contribute to prognosis determinations) can provide an automated process for providing clinical values that provide diagnostic, prognostic, and / or therapeutic information, and which can be leveraged by care personnel to select medical service regimens for individuals.

[0094] A second example of a technical effect that can be exhibited by some example embodiments is more efficient resource allocation based on such analysis. For example, tumor classification based on automated techniques can reduce the amount and / or dependency on clinical and pathological resources applied to diagnose and classify tumors and determine clinical values (e.g., prognoses) for individuals. Such resource economy can also involve faster classification processes that are systematically achieved than classification processes performed by individuals. I. EXAMPLE EMBODIMENTS

[0095] Figure 13 is a flowchart of a first example method 1300 according to some example embodiments.

[0096] The first example method 1300 can be implemented, for example, as a set of instructions that, when executed by processing circuitry of an apparatus, cause the apparatus to perform each of the elements of the first example method 1300. The first example method 1300 can also be implemented, for example, as a set of instructions that, when executed by processing circuitry of an apparatus, cause the apparatus to provide a system for components including an image evaluator, a classifier, and a tumor evaluator that interoperate to provide a system for classifying tumors.

[0097] The first example method 1300 includes instructions executed by processing circuitry of an apparatus 1304 causing the apparatus to perform a set of elements.

[0098] For example, execution of the instructions can cause the apparatus to determine 1306 a lymphocyte distribution of lymphocytes in the tumor based on the image.

[0099] For example, execution of the instructions can cause the apparatus to apply 1308 the classifier to the lymphocyte distribution to classify the tumor, the classifier having been trained to classify tumors into a class selected from at least two classes that are respectively associated with lymphocyte distributions.

[0100] For example, execution of the instructions can cause the apparatus to determine 1310 a clinical value (e.g., a prognosis) for the individual based on a prognosis of individuals having tumors classified by the classifier into the class to which the tumor is classified.

[0101] In this way, execution of the instructions by the processing circuitry can cause the apparatus to perform elements of the first example method 1300, and thus the first example method 1300 ends.

[0102] Figure 14 is a flowchart of a second example method according to some example embodiments.

[0103] The second example method 1400 can be implemented, for example, as a set of instructions 1402 that, when executed by processing circuitry of an apparatus, cause the apparatus to perform each of the elements of the second example method 1400. The second example method 1400 can also be implemented as, for example, a set of instructions that, when executed by processing circuitry of an apparatus, cause the apparatus to provide a system for the components including an image evaluator, a classifier, and a tumor evaluator that interoperate to provide a system for classifying tumors.

[0104] The second example method 1400 includes execution 1404 of instructions by processing circuitry of an apparatus causing the apparatus to perform a set of elements.

[0105] For example, execution of the instructions can cause the apparatus to apply 1406 a convolutional neural network to the image to determine a lymphocyte distribution of lymphocytes in the tumor, wherein the convolutional neural network is configured to measure the lymphocyte distribution of lymphocytes for different region types of the image.

[0106] For example, execution of the instructions can cause the apparatus to apply 1408 the classifier to the lymphocyte distribution to classify the tumor, wherein the classifier has been trained to classify the tumor into a class selected from a low-risk class and a high-risk class, the classes being respectively associated with the lymphocyte distribution, and the classifier comprises a bi-directional Gaussian mixture model configured to determine, for the respective class, a probability distribution of features of tumors of that class within a feature space.

[0107] For example, execution of the instructions can cause the apparatus to apply 1410 a Cox proportional hazards model to clinical features of the tumor to determine the class of the tumor.

[0108] For example, execution of the instructions can cause the apparatus to determine 1412 a clinical value (such as a prognosis) for the individual based on a prognosis of the individual of the tumor classified by the classifier into the class and the class determined by the Cox proportional hazards model.

[0109] In this way, execution of the instructions by the processing circuitry can cause the apparatus to perform elements of the second example method 1400, and thus the second example method 1400 ends.

[0110] Figure 15 is a component block diagram of an example apparatus according to some example embodiments.

[0111] AsFigure 15 As shown, the example apparatus 1500 can include processing circuitry 1502 and memory 1504. According to some example embodiments, the memory 1504 can store instructions 1506 that, when executed by the processing circuitry 1502, cause the example apparatus 1500 to determine a clinical value (such as a prognosis) for an individual based on a tumor shown in the image 102. In some example embodiments, execution of the instructions 1506 can cause the example apparatus 1500 to instantiate and / or use a set of components of a system 1508. While Figure 15 One such system 1508 is shown, but some example embodiments can embody any of the methods disclosed herein.

[0112] Figure 15 The example system 1508 includes an image evaluator 1510 configured to determine a lymphocyte distribution of lymphocytes in the image 102. For example, a set of categories 1516 can associate respective lymphocyte distributions 1520-1, 1520-2 with different categories 1518-1, 1518-2 of tumors, each category 1518 associated with a prognosis 1522-1, 1522-2.

[0113] Figure 15 The example system 1508 includes a tumor classifier 1512 configured to classify the tumor into a category selected from at least two categories 1518 associated with the lymphocyte distribution 1520.

[0114] Figure 15 The example system 1508 includes a tumor evaluator 1514 configured to determine a clinical value (such as a prognosis) for an individual based on a tumor in the image 102 by: invoking the image evaluator 1510 with the image 102 to determine a lymphocyte distribution 1520-3 of lymphocytes in the tumor; invoking the tumor classifier 1512 to classify the tumor into a category 1518 based on the lymphocyte distribution 1520-3, and outputting the clinical value (such as a prognosis) for the individual to a user 1524 based on the prognosis 1522 of the tumor in the category 1518-3 into which the tumor classifier 1512 classified the tumor.

[0115] In this way, the example apparatus 1500 and the example system 1508 provided thereon can classify tumors according to some example embodiments.

[0116] Figure 16 is a component block diagram of another example apparatus according to some example embodiments.

[0117] As Figure 16As shown, the example apparatus 1600 can include processing circuitry 1502 and memory 1504. According to some example embodiments, the memory 1504 can store instructions 1506 that, when executed by the processing circuitry 1502, cause the example apparatus 1600 to determine a clinical value (e.g., a prognosis) for an individual based on a tumor shown in an image 102. In some example embodiments, execution of the instructions 1506 can cause the example apparatus 1600 to instantiate and / or use a set of components of a system 1602.

[0118] Figure 16 The example system 1602 includes a convolutional neural network 110 as an image evaluator configured to determine a lymphocyte distribution of lymphocytes in the image 102 by measuring the lymphocyte distribution of lymphocytes of different region types of the image 102. For example, a set of classes 1516 can associate respective lymphocyte distributions 1520-1, 1520-2 with different classes 1518-1, 1518-2 of tumors, including a low-risk tumor class and a high-risk tumor class, each class 1518 associated with a prognosis 1522-1, 1522-2.

[0119] Figure 16 The example system 1602 includes a bidirectional Gaussian mixture model 1604 as a tumor classifier configured to determine a probability distribution of features of tumors in a class 1518 within a feature space 606 for the respective class 1518.

[0120] Figure 16 The example system 1602 includes a Cox proportional hazards model 1608 configured for use in determining a set of clinical features 1606 of a class of tumors.

[0121] Figure 16 The example system 1602 includes a tumor evaluator 1514 configured to determine a clinical value (e.g., a prognosis) for an individual based on a tumor in the image of the tumor 102 by: invoking the convolutional neural network 110 with the image 102 to determine a lymphocyte distribution 1520-3 of lymphocytes in the tumor; invoking the Gaussian mixture model 1604 to classify the tumor into a class 1518-3 based on the lymphocyte distribution 1520-3; invoking the Cox proportional hazards model 1608 with the set of clinical features 1606 of the tumor to determine a class 1518-4 of the tumor, and outputting the clinical value (e.g., the prognosis) for the individual based on a prognosis of the individual for the tumor classified into the class 1518-3 by the Gaussian mixture model 1604 and the class 1518-4 of the tumor determined by the Cox proportional hazards model 1608 based on the tumor classified into the class 1518-5 by the tumor classifier 1512.

[0122] In this way, the example apparatus 1600 and example system 1602 provided thereon can classify tumors in accordance with some example embodiments.

[0123] As Figure 15 and Figure 16 further shown, the example apparatus 1500, 1600 can include processing circuitry 1502 capable of executing instructions. The processing circuitry 1502 can include: hardware, such as logic circuits, including a processor; a hardware / software combination, such as a processor executing software; or a combination thereof. For example, the processor can include, but is not limited to, a central processing unit (CPU), a graphics processing unit (GPU), an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a system on a chip (SoC), a programmable logic unit, a microprocessor, an application-specific integrated circuit (ASIC), and the like.

[0124] As Figure 15 and Figure 16 further shown, the example apparatus 1500, 1600 can include a memory 1504 storing instructions 1506. The memory 1504 can include, for example, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), and the like. The memory 1504 can be volatile, such as a system memory that is directly coupled to the processor, and / or non-volatile, such as a hard disk drive, a solid-state storage device, a flash memory, or a magnetic tape. The instructions 1506 stored in the memory 1504 can be specified in a native instruction set architecture of a processor, such as a variation of the IA-32 instruction set architecture or a variation of the ARM instruction set architecture, as: assembly and / or machine language (e.g., binary) instructions; instructions of a high-level imperative and / or declarative language that are compilable and / or interpretable for execution on the processor; and / or instructions that are compilable and / or interpretable for execution by a virtual processor of a virtual machine, such as a web browser. A non-limiting example set of such high-level languages can include, for example: C, C++, C#, Objective-C, Swift, Haskell, Go, SQL, R, Lisp, Java®, Fortran, Perl, Pascal, Curl, OCaml, HTML5 (Hypertext Markup Language, 5th revision), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Swift, Eiffel, Smalltalk, Erlang, Ruby, Visual Basic®, Visual Basic.NET, MATLAB, SIMULINK, and the like. Lua, MATLAB, SIMULINK, and ​​​Such instructions 1506 can also include instructions for libraries, resources, platforms, application programming interfaces (APIs), etc., that are used to determine a clinical value (e.g., prognosis) of an individual based on a tumor shown in an image.

[0125] As Figure 15 and Figure 16 shown, example systems 1508, 1602 can be organized in a particular way, e.g., to allocate some functionality to each component of the system. Some example embodiments can implement each such component in various ways, e.g., as software, hardware (e.g., processing circuitry), or a combination thereof. In some example embodiments, the organization of the system can be different than that of example systems 1508, 1602 as shown in Figure 15 and Figure 16 shown. For example, some example embodiments can include systems featuring different organizations of components, e.g., renaming, rearranging, adding, dividing, duplicating, merging, and / or removing components, groups of components, and relationships therebetween, without departing from the scope of the present disclosure. All such variations that are technically and logically possible and do not contradict other statements are intended to be included in the present disclosure, the scope of which should be understood to be limited only by the claims.

[0126] Figure 17 is an illustration of an example computer-readable medium 1700 according to some example embodiments.

[0127] As Figure 17 shown, according to some example embodiments, non-transitory computer-readable medium 1700 can store binary data 1702 encoding a set of instructions 1704 that, when executed by processing circuitry 1502 of example apparatus 1500, 1600, cause example apparatus 1500, 1600 to determine a clinical value (e.g., prognosis) of an individual based on a tumor shown in an image. As a first such example, instructions 1704 can encode elements of example method 1706 (e.g., first example method 1300 of Figure 13 As a second such example, instructions 1704 can encode elements of second example method 1400 of Figure 14 As a third such example, instructions 1704 can encode components of first example system 1508 of Figure 15 As a fourth such example, instructions 1704 can encode components of second example system 1602 of Figure 16 As a fourth such example, instructions 1704 can encode components of second example system 1602 of

[0128] In some example embodiments, the system can include an image evaluation device for determining a lymphocyte distribution of lymphocytes in the image. For example, the image evaluation device can be or can include one or more convolutional neural networks and / or any other image evaluation model discussed herein. The system can include a classification device for classifying the tumor into a class selected from at least two classes that are each associated with a lymphocyte distribution. For example, the classification device can be or can include one or more Gaussian mixture models and / or any other classifier discussed herein.

[0129] The system can include a tumor evaluator device for determining a clinical value (such as a prognosis) for the individual based on the tumor in the image by: invoking the image evaluation device with the image to determine a lymphocyte distribution of lymphocytes in the tumor; invoking the classifier to classify the tumor into a class based on the lymphocyte distribution, and outputting the clinical value (such as a prognosis) for the individual based on a prognosis for the individual having a tumor classified into the class by the classifier. For example, the tumor evaluator device can be or can include the classifier, such as a neural network, a soft-margin or hard-margin support vector machine, and / or any other classifier discussed herein. For example, the tumor evaluator device can be or can include a display device, such as a liquid crystal display (LCD), light emitting diode (LED), or organic light emitting diode (OLED) display; a communication interface, such as a web server, email server, or text message server; and / or any other output device disclosed herein. J. Variations

[0130] Some example embodiments of the present disclosure can include variations in many aspects, and some variations can present additional advantages and / or reduce drawbacks relative to these and other variations of the technology. Moreover, some variations can be implemented in combination, and some combinations can have additional advantages and / or reduced drawbacks through synergistic cooperation. These variations can be incorporated into some example embodiments (e.g., Figure 13 a first example method of the Figure 14 a second example method of the Figure 15 and Figure 16 example devices 1500, 1600 and example systems 1508, 1602, and / or Figure 17 example non-transitory computer-readable medium 1700 of the F1. Scenarios

[0131] Some example embodiments can be used in various scenarios involving input analysis using distribution-based machine learning models. For example, some example embodiments can use the disclosed techniques to classify tumors for use in various areas of life sciences, including medical services and biomedical research. The tumor classification techniques disclosed herein can be applicable to a variety of cancer types, including (but not limited to) lung cancer tumors, pancreatic cancer tumors, and / or breast cancer tumors. Clinical pathology laboratories can use such techniques to determine tumor classes for tumor samples, and / or to compare or validate determinations of tumor classes by individuals and / or other automated processes. Researchers can use such techniques to determine tumor classes for tumors in images of research data sets, which can be from human patients or from human or non-human experimental subjects, where such research can involve additional techniques for classifying tumors, identifying prevalence and classes of tumors in different demographics, identifying risk factors associated with tumors of different tumor classes, predicting survivability, for determining or comparing effectiveness of treatment options. Clinicians can use results of classifications to evaluate diagnoses, prognoses, and / or treatment options for individuals with tumors, and / or to explore and understand correlations of various risk factors with different tumor classes and prognoses for individuals with such tumors. Many such scenarios can be designed that can utilize the disclosed techniques.

[0132] F2. Determine feature presence and distribution

[0133] In some example embodiments, machine learning models, including deep learning models, can be used to detect various features of various inputs. In various example embodiments, such machine learning models can be used to: determine feature maps 118 of images 102, e.g., by generating and applying masks 302 of mask sets 300; determine distributions of features, such as clusters 204; perform classifications 402 of image regions, such as region types based on anatomical features and / or tissue types; perform measurements 304 of features, such as concentrations of features or region types (e.g., percentage area of entire image), e.g., lymphocyte density range estimates 404, using techniques such as binning; generate density maps, such as lymphocyte density maps 406; select sets of features for use in classifying images 102 or particular images 102, such as performing feature selection 408; and / or perform other tasks. Figure 7The selection process 700 can include selecting a subset of features; classifying the image 102 of the tumor based on the set or subset of image features; selecting a subset of clinical features from the values of the clinical features 706 in the set of clinical features of the individual or tumor; determining the values of the clinical features 706 of the set of clinical features of the individual and / or tumor; determining the class of the tumor, such as by preparing and applying a Cox proportional hazards model to the clinical features of the tumor; determining the class of the tumor based on the image features of the tumor (such as the output of a Gaussian mixture model) and / or the clinical features of the tumor or individual (such as the output of a Cox proportional hazards model); predicting the survivability of the individual based on the classification of the tumor of the individual; and / or generating one or more outputs of such determinations, including visualizations. Each of these features and other features of some example embodiments can be performed, for example, by a machine learning model; multiple machine learning models of the same or similar type, such as a random forest, or a convolutional neural network that evaluates different portions of an image or performs different tasks on an image; and / or a combination of machine learning models of different types. As one such example, in an ensemble, a first machine learning model performs a classification based on the output of other machine learning models.

[0134] As a first such example, the presence of a feature (e.g., activation within a feature map, and / or biological activation of a lymphocyte) can be determined in various ways. For example, where the input further comprises an image 102, the presence of a feature can be determined by applying the at least one convolutional neural network 110 to the image 102 and receiving a feature map 118 from the at least one convolutional neural network 110 that indicates the presence of a feature of the input. That is, the convolutional neural network 110 can be applied to the image 102 to identify clusters of pixels 108 for which a feature is apparent. For example, a cytometry convolutional neural network 110 can be applied to count cells in a tissue sample, where such cells can be lymphocytes. In such a scenario, the tissue sample can be assayed, such as with a dye or a luminescent (e.g., fluorescent) agent, and a set of images 102 of the tissue sample can be selective for the cells and thus can not include other visible components of the tissue sample. The images 102 of the tissue sample can then be subjected to a machine learning model (such as a convolutional neural network 110) that can be configured (e.g., trained) to detect shapes, such as circles that are indicative of the selected cells, and can output a count of the cells in and / or across different regions of the image 102. Notably, in such a case, the convolutional neural network 110 can not be configured to and / or used to further detect an arrangement of such features, e.g., a number, orientation, and / or positioning of lymphocytes relative to other lymphocytes; rather, the counts of the respective portions of the image 102 can be compiled into a profile that can be processed with a tumor mask to determine whether the distribution of lymphocytes in the tissue sample is tumor-invasive, tumor-adjacent, or elsewhere. Some example embodiments can use machine learning models other than convolutional neural networks to detect the presence of a feature, such as (e.g.) non-convolutional neural networks, such as fully-connected networks or perceptron networks or Bayesian classifiers.

[0135] As a second such example, the distribution of the feature can be determined in various ways. As a first example, where the input further includes an image 102 showing a tissue region of the individual, determining the distribution of the feature can further include determining a region type (e.g., tumor or non-tumor) of each region of the image, and determining the distribution based on the region type of each region of the image. The distribution of the detected lymphocytes, including the lymphocyte count, can then be determined based on the tissue type for which such counting occurs. That is, the distribution can be determined by tabulating the counts of lymphocytes for tumor regions, tumor-adjacent regions (such as stroma), and non-tumor regions of the image 102. As another such example, the determining can include determining boundaries of the tissue region within the image 102, and determining the distribution based on the boundaries of the tissue region within the image 102. That is, the boundaries of the regions of the image 102 that are classified as tumor (e.g., by the convolutional neural network 110 and / or a human) can be determined, and the example embodiments can tabulate the counts of lymphocytes for all regions of the image that are within the determined tumor boundaries. As yet another example, the tissue region of the image includes at least two regions, and determining the distribution can include determining a count of lymphocytes within each region of the tissue region and determining the distribution based on the count within each region of the tissue region. For example, determining the count within each tissue region can include determining a density of the count of lymphocytes within each region, and then determining the distribution based on the count within each region. Figure 3 An example is shown in which a first convolutional neural network 110-1 is provided to classify regions as tumor regions versus non-tumor regions and a second convolutional neural network 110-2 is provided to estimate a density range of lymphocytes.

[0136] As a third such example, processing of the image 102 to determine the presence of features and / or the distribution of features can occur in several ways. For example, some example embodiments can be configured to divide the image 102 into a set of regions of the same or different size and / or shape, such as based on a number of pixels or a corresponding physical size (e.g., regions of 100 square microns), and / or based on similarity grouping (e.g., identifying regions of similar appearance within the image 102). Some example embodiments can be configured to classify each region (e.g., as tumor, tumor proximity, or non-tumor), and / or to determine the distribution by tabulating the presence of features (e.g., counts) within each region of a certain region type to determine the distribution of lymphocytes. Alternatively, a counting process can be applied to each region, and each region can be classified based on the counts (e.g., high lymphocyte regions versus low lymphocyte regions). As yet another example, the distribution can be determined parametrically, such as according to a selected distribution type or kernel that can fit to the distribution of features in the input according to a machine learning model (e.g., a Gaussian mixture model can be applied to determine a Gaussian distribution of a subset of features). Other distribution models can be applied, including parametric distribution models such as chi-squared fits, Poisson distributions, and beta distributions, and non-parametric distribution models such as histogram, binning, and kernel methods.

[0137] As a fourth such example, many forms of classifiers can be used, such as Bayesian (including Naive Bayes) classifiers; Gaussian classifiers; probabilistic classifiers; principal component analysis (PCA) classifiers; linear discriminant analysis (LDA) classifiers; quadratic discriminant analysis (QDA) classifiers; single- or multi-layer perceptron networks; convolutional neural networks; recurrent neural networks; nearest neighbor classifiers; linear SVM classifiers; radial basis function kernel (RBF) SVM classifiers; Gaussian process classifiers; decision tree classifiers, including random forest classifiers; and / or Boltzmann machines, with or without restrictions, and the like. Examples of convolutional neural network classifiers include, but are not limited to, LeNet, ZfNet, AlexNet, BN-Inception, CaffeResNet-101, DenseNet-121, DenseNet-169, DenseNet-201, DenseNet-161, DPN-68, DPN-98, DPN-131, FBResNet-152, GoogLeNet, Inception-ResNet-v2, Inception-v3, Inception-v4, MobileNet-vl, MobileNet-v2, NASNet-A-Large, NASNet-A-Mobile, ResNet-101, ResNet-152, ResNet-18, ResNet-34, ResNet-50, ResNext-101, SE-ResNet-101, SE-ResNet-152, SE-ResNet-50, SE-ResNeXt-101, SE-ResNeXt-50, SENet-154, ShuffleNet, SqueezeNet-vl.0, SqueezeNet-vl. l, VGG-11, VGG-11_BN, VGG-13, VGG-13_BN, VGG-16, VGG-16_BN, VGG-19, VGG-19_BN, Xception, DelugeNet, FractalNet, WideResNet, PolyNet, PyramidalNet, and U-net.

[0138] In some example embodiments, the classification can include regression, and as used herein the term“classification” is intended to include some example embodiments that perform regression as an alternative or in addition to class selection. For example, some example embodiments can feature regression as an alternative or in addition to classification. As a first such example, the determination of the presence of a feature can include regression of the presence of the feature, e.g., a numerical value indicative of the density of the feature in the input. As a second such example, the determination of the distribution of a feature can include regression of the distribution of the feature, e.g., a variance of the regression-based density determined for the input. As a third such example, the selection can include performing regression on the distribution of the feature and selecting a regression value for the distribution of the feature. Such regression aspects can be performed in lieu of classification or in addition to classification (e.g., determining the presence of a feature in an image region and the density of the feature). Some example embodiments can involve regression-based machine learning models, such as Bayesian linear or nonlinear regression, regression-based artificial neural networks, such as convolutional neural network regression, support vector regression, and / or decision tree regression.

[0139] Each classifier can be linear or nonlinear; for example, a nonlinear classifier can be provided (e.g., trained) to perform linear classification based on a kernel transformation in a nonlinear space (i.e., transforming linear values of a feature vector into nonlinear features). The classifiers can include various techniques to facilitate accurate generalization and classification (such as input normalization, weight regularization) and / or output processing (such as softmax activation output). The classifiers can use various techniques to facilitate efficient training and / or classification. For example, a bi-directional Gaussian mixture model can be used, in which the same size Gaussian distribution is selected for each dimension of the feature space, which can reduce the search space compared to other Gaussian mixture models in which the distribution size can vary for different dimensions of the feature space.

[0140] Each classifier can be trained to perform classification in a particular manner, such as supervised learning, unsupervised learning, and / or reinforcement learning. Some example embodiments can include additional training techniques to facilitate generalization, accuracy, and / or convergence, such as validation, training data augmentation, and / or dropout regularization. Ensembles of such classifiers can also be utilized, where such ensembles can be homogeneous or heterogeneous, and where classification 122 based on the outputs of the classifiers can be produced in various manners, such as by consensus, based on the confidence of each output (e.g., as a weighted combination), and / or via a stacked architecture, such as based on one or more mixers. Ensembles can be trained independently (e.g., bootstrap aggregating training models, or random forest training models) and / or sequentially (e.g., boosting training models, such as Adaboost). As an example of boosting training models, in some support vector machine ensembles, at least some of the support vector machines can be trained based on the errors of previously trained support vector machines; for example, each successive support vector machine can be trained particularly on inputs of the training dataset 100 that were misclassified by a previously trained support vector machine.

[0141] As a fifth such example, the one or more classifiers can produce various forms of classification 122. For example, a perceptron or binary classifier can output a value indicating whether an input is classified as a first class 106 or a second class 106, such as whether a region of an image 102 is a tumor region or a non-tumor region. As another example, a probabilistic classifier can be configured to output a probability that an input is classified as each class 106 of the set of classes 104. Alternatively or additionally, some example embodiments can be configured to determine a probability that an input is classified as each class 106 of the set of classes 104, and to select a class 106 of the input from the set of classes 104 that includes at least two classes 106, the selection based on the probabilities that the input is classified as each class 106 of the set of classes 104. For example, a classifier can output a confidence of a classification 122, e.g., a probability of classification error, and / or can withhold outputting a classification 122 based on a difference confidence, e.g., a minimum risk classifier. For example, in a region of an image 102 that cannot be clearly identified as a tumor or a non-tumor, a classifier can be configured to withhold classifying the region in order to improve the accuracy of a computed distribution of lymphocytes in regions that can be identified as tumors and non-tumors with an acceptable confidence.

[0142] As a sixth such example, some example embodiments can be configured to process a distribution of features with a linear or non-linear classifier, and can receive a classification 122 of a class 106 of an input from the linear or non-linear classifier. For example, the linear or non-linear classifier can include a support vector machine ensemble of at least two support vector machines, and some example embodiments can be configured to receive the classification 122 by receiving a candidate classification 122 from each of the at least two support vector machines, and determine the classification 122 based on a consistency of the candidate classifications 122 between the at least two support vector machines.

[0143] As a seventh such example, some example embodiments can be configured to use a linear or non-linear classifier (including a classifier set or ensemble) to perform a classification 122 of an input in various ways. For example, an input can be divided into regions, and for each input portion, example embodiments can use a classifier to classify the input portion according to an input portion type selected from a set of input portion types (e.g., classifying portions 122 of an image 102 of a tumor as tumor regions and non-tumor regions, as shown by the example of FIG. 3). Example embodiments can then be configured to select a class 106 of the input from a set of classes including at least two classes 106, the selection based on a distribution of features of each input portion of the input portion types and a distribution of features of each input portion type in the set of input portion types. Figure 3

[0144] As an eighth such example, some example embodiments can be configured to perform a distribution classification by determining a variance of a distribution of features of an input (e.g., a variance of a distribution over a region of an input such as an image 102). Some example embodiments can then be configured to perform a classification 122 by selecting a class 106 of the input from a set 104 of classes 106 including at least two classes 106, the selection based on the variance of the distribution of features of the input and a variance of a distribution of features of each class 106 in the set 104 of classes 106. For example, some example embodiments can be configured to determine a class of a tumor based at least in part on a variance of a distribution of lymphocytes over different regions of an image.

[0145] ​As a ninth such example, some example embodiments may use different training and / or testing methods to generate and validate machine learning models. For example, training may be performed using heuristics such as stochastic gradient descent, nonlinear conjugate gradient, or simulated annealing. Training may be performed offline (e.g., based on a fixed training dataset 100) or online (e.g., continuous training using new training data). Training may be evaluated based on various metrics such as perceptron error, Kullback-Leibler (KL) divergence, precision, and / or recall. Training may be performed for a fixed period (e.g., a selected number of epochs or generations) until training fails to produce additional improvements, and / or until a convergence point is reached (e.g., when classification accuracy reaches a target threshold). The machine learning model may be tested in various ways (e.g., k-fold cross-validation) to determine its proficiency with previously unseen data. In some example embodiments, many such forms of classification 122, classifiers, training, testing, and validation may be included and used. K. Example Computing Environment

[0146] Figure 18 This is an illustration of an example device in which some example embodiments may be implemented.

[0147] Figure 18 The following discussion provides a brief overview of suitable computing environments for implementing one or more embodiments of the provisions set forth herein. Figure 18 The operating environment described is merely an example of a suitable operating environment and is not intended to impose any limitations on the scope or functionality of the operating environment. Example computing devices include, but are not limited to: personal computers; server computers; handheld or laptop devices; mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.); multiprocessor systems; media devices such as televisions; consumer electronics; embedded devices; microcomputers; mainframe computers; distributed computing environments including any of the above systems or devices; wearable computing devices (such as glasses, headphones, watches, rings, pendants, handheld and / or body-mounted cameras, clothing-integrated devices, and implantable devices); autonomous vehicles; extended reality (XR) devices, such as augmented reality (AR) and / or virtual reality (VR) devices; Internet of Things (IoT) devices, etc.

[0148] Some example embodiments can include combinations of the same and / or different types of components, such as multiple processors and / or processing cores in a single processor or multi-processor computer; two or more processors operating in tandem, such as a CPU and a GPU; a CPU utilizing an ASIC; and / or software executed by processing circuitry. Some example embodiments can include components of a single device, such as a computer including one or more CPUs that store, access, and manage a cache. Some example embodiments can include components of multiple devices, such as two or more devices with CPUs in communication to access and / or manage a cache. Some example embodiments can include one or more components contained in a server computing device, a server computer, a series of server computers, a server farm, a cloud computer, a content platform, a mobile computing device, a smartphone, a tablet computer, or a set-top box. Some example embodiments can include components in direct communication (e.g., two or more cores of a multi-core processor) and / or indirect communication (e.g., via a bus, via over a wired or wireless channel or network, and / or via an intermediary component such as a microcontroller or arbitrator). Some example embodiments can include multiple instances of a system or instance executed by a device or component, respectively, where such system instances can be executed simultaneously, consecutively, and / or in an interleaved manner. Some example embodiments can feature distribution of an instance or system over two or more devices or components.

[0149] While not necessarily required, some example embodiments are described in the general context of "computer readable instructions" being executed by one or more computing devices. Computer readable instructions can be distributed via computer readable media as discussed herein. Computer readable instructions can be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), data structures, and the like, that perform particular tasks or implement particular abstract data types. Typically, the functionality of the computer readable instructions can be combined or distributed as desired in various environments.

[0150] Figure 18 An example of an example apparatus 1800 configured to or including one or more example embodiments, such as example embodiments provided herein, is shown. In one apparatus configuration 1802, the example apparatus 1800 can include processing circuitry 1502 and a memory 1804. Depending on the exact configuration and type of computing device, the memory 1804 can be volatile (such as RAM), non-volatile (such as ROM, flash memory, etc.) or some combination of the two.

[0151] In some example embodiments, the example apparatus 1800 can include additional features and / or functionality. For example, the example apparatus 1800 can also include additional storage (e.g., removable and / or non-removable) including, but not limited to, magnetic storage, optical storage, and the like. Such additional storage is illustrated in FIG. 18 by the additional storage 1806.Figure 18 The storage device 1806 is also shown in FIG. 18. In some example embodiments, computer-readable instructions for implementing one or more embodiments provided herein can be stored in the memory 1804 and / or the storage device 1806.

[0152] In some example embodiments, the storage device 1806 can be configured to store further computer-readable instructions for implementing operating systems, applications, etc. For example, computer-readable instructions can be loaded into the memory 1804 for execution by the processing circuitry 1502. The storage device can include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions or other data. The storage device can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by the example device 1800. Any such computer storage media can be part of the example device 1800.

[0153] In some example embodiments, the example device 1800 can include input device(s) 1810, such as keyboard, mouse, pen, voice input device, touch input device, infrared camera, video input device, and / or any other input device. Output device(s) 1808, such as one or more displays, speakers, printers, and / or any other output device, can also be included in the example device 1800. The input device(s) 1810 and the output device(s) 1808 can be connected to the example device 1800 via wired connections, wireless connections, or any combination

[0154] In some example embodiments, the example device 1800 can be connected through various interconnections, such as a bus. Such interconnection can include a peripheral component interconnect (PCI) (e.g., PCI Express), a universal serial bus (USB), Firewire (IEEE 1394), an optical bus structure, etc. In other example embodiments, components of the example device 1800 can be interconnected through a network. For example, the memory 1804 can include multiple physical memory units located at different physical locations that are interconnected through a network.

[0155] In some example embodiments, the example apparatus 1800 can include one or more communication devices 1812 through which the example apparatus 1800 can communicate with other devices. The communication device(s) 1812 can include, for example, a modem, a network interface card (NIC), an integrated network interface, a radio-frequency transmitter / receiver, an infrared port, a USB connection, or other interface to connect the example apparatus 1800 to other computing devices, including a remote device 1816. The communication device(s) 1812 can include a wired connection or a wireless connection. The communication device(s) 1812 can be configured to transmit and / or receive communication media.

[0156] Those skilled in the art will realize that the storage devices used to store the computer-readable instructions can be distributed across several computers over a network. For example, the example apparatus 1800 can communicate with a remote device 1816 via a network 1814 to store and / or retrieve computer-readable instructions to implement one or more example embodiments provided herein. For example, the example apparatus 1800 can be configured to access the remote device 1816 to download some or all of the computer-readable instructions for execution. Alternatively, the example apparatus 1800 can be configured to download portions of the computer-readable instructions as needed, wherein some instructions can be executed at or by the example apparatus 1800, and some other instructions can be executed at or by the remote device 1816.

[0157] In this application, including the definitions below, the term “module” or the term “controller” can be replaced with the term “circuit.” The term “module” can refer to the processing circuitry 1502 (shared, dedicated, or group) that executes code to perform the functions of a module, the memory hardware (shared, dedicated, or group) that stores the code executed by the processing circuitry 1502, a portion thereof, or includes the processing circuitry and the memory hardware.

[0158] A module can include one or more interface circuits. In some examples, the interface circuit(s) can implement a wired or wireless interface to connect to a local area network (LAN) or a wireless personal area network (WPAN). Examples of a LAN are the Institute of Electrical and Electronics Engineers (IEEE) Standard 802.11-2016 (also known as the WIFI wireless networking standard) and the IEEE Standard 802.3-2015 (also known as the Ethernet wired networking standard). Examples of a WPAN are the IEEE Standard 802.15.4 (including the ZIGBEE standard from the ZigBee Alliance) and the BLUETOOTH wireless networking standard from the Bluetooth Special Interest Group (SIG) (including the Core Specification Versions 3.0, 4.0, 4.1, 4.2, 5.0, and 5.1 from the Bluetooth SIG).

[0159] Modules can communicate with other modules using interface circuitry. Although modules can be depicted in the disclosure as being in electrical communication with other modules, in various implementations, modules can actually communicate over a communication system. The communication system includes physical and / or virtual networking equipment such as hubs, switches, routers, and gateways. In some implementations, the communication system connects to or traverses a wide area network (WAN) such as the Internet. For example, the communication system can include multiple LANs connected to each other over the Internet or point-to-point leased lines using technologies including Multiprotocol Label Switching (MPLS), as well as virtual private networks (VPNs).

[0160] In various implementations, the functionality of a module can be distributed among multiple modules connected via the communication system. For example, multiple modules can implement the same functionality distributed by a load balancing system. In further examples, the functionality of a module can be divided between a server (also referred to as a remote or cloud) module and a client (or user) module.

[0161] The term code, as used in the specification above, can include software, firmware, and / or microcode, and can refer to programs, routines, functions, classes, data structures, and / or objects. Shared processing circuitry 1502 can encompass a single microprocessor executing some or all of the code from multiple modules. Group processing circuitry 1502 can encompass a microprocessor combined with additional microprocessors to execute some or all of the code from one or more modules. References to a plurality of microprocessors encompass a plurality of microprocessors on discrete dies, a plurality of microprocessors on a single die, a plurality of cores of a single microprocessor, multiple threads of a single microprocessor, or a combination thereof.

[0162] Shared memory hardware encompasses a single memory device that stores some or all of the code from multiple modules. Group memory hardware encompasses a memory device in combination with other memory devices to store some or all of the code from one or more modules.

[0163] The term memory hardware is a subset of the term computer-readable medium. The term computer-readable medium, as used herein, does not encompass transitory propagating signals per se (e.g., waves such as those used to communicate data asynchronously over the Internet or point-to-point lines); thus, the term computer-readable medium is therefore considered tangible and non-transitory. Non-limiting examples of non-transitory computer-readable media are nonvolatile memory devices (e.g., flash memory devices, erasable programmable read-only memory devices, or mask read-only memory devices); volatile memory devices (e.g., static random access memory devices or dynamic random access memory devices); magnetic storage media (e.g., analog or digital tapes or hard disk drives); and optical storage media (e.g., CDs, DVDs, or Blu-ray discs).

[0164] Example embodiments of the apparatus and methods described herein can be implemented in part or in whole using specialized computer equipment, which is created by configuring general-purpose computer equipment to perform one or more particular functions embodied in computer programs. The functional blocks and flowchart elements described herein can be used as software specifications to create computer programs, which can be translated into computer programs by the routine work of skilled technicians or programmers.

[0165] A computer program includes processor-executable instructions stored on at least one non-transitory computer-readable medium. A computer program can also include or rely on stored data. A computer program can encompass a Basic Input / Output system (BIOS) that interacts with the hardware of the special purpose computer, device drivers that interact with particular devices of the special purpose computer, one or more operating systems, user applications, background services, background applications, etc.

[0166] A computer program can include: (i) descriptive text to be parsed, such as HTML (HyperText Markup Language), XML (Extensible Markup Language), or JSON (JavaScript Object Notation), (ii) assembly code, (iii) object code that results from Fortran, Perl, Pascal, Curl, OCaml, HTML5 (HyperText Markup Language, 5th revision), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Eiffel, Smalltalk, Erlang, Ruby, Visual Lua, MATLAB, SIMULINK, and syntax of languages from the C family of programming languages. L. Use of Terms

[0167] The foregoing description is merely illustrative in nature and is in no way intended to limit the disclosure, its application, or uses. The broad teachings of the disclosure can be implemented in a variety of forms. Therefore, while this disclosure includes particular examples, the true scope of the disclosure should not be so limited since other modifications will become apparent upon a study of the drawings, specification, and appended claims. It should be understood that one or more steps within a method can be executed in different order (or concurrently) without altering the principles of the disclosure. Further, although each of the embodiments describes at least one feature, any single implementation can include one or more of the features described. Still further, any one or more feature described in relation to any one embodiment can be implemented with any other embodiment described or can be implemented apart from the other embodiments. In other words, the described embodiments are not mutually exclusive, and permutations of one or more embodiments' features with one another are within the scope of this disclosure.

[0168] Various terminology can be used in describing spatial and functional relationships between elements (for example, modules) in the above description. Such terminology includes "connected," "engaged," "interfaced," and "coupled." As used in the above description, the relationship between the first element and the second element can be a direct relationship, in which the first element and the second element are directly connected, engaged, interfaced, or coupled to each other, or an indirect relationship in which one or more other elements are present between the first element and the second element. As used in the above description, the phrase at least one of A, B, and C should be construed to mean a logical (A OR B OR C), using a non- exclusive logical OR, and should not be construed to mean "at least one of A and / or B and / or C." As used in the above description, the phrase "at least one of A, B, and C" should be construed to mean A or B or C or any combination thereof (A and B and C).

[0169] In the drawings, the arrow direction generally indicates the direction of information flow (e.g., data or instructions) that is of interest to the illustration. For example, when elements A and B exchange a variety of information but the information of interest to the illustration is that which is transmitted from A to B, the arrow can point from A to B. This unidirectional arrow does not imply that no other information is being transmitted from B to A. Further, for information sent from A to B, B can send requests for, or receive acknowledgements of, the information from A. The term subset is not necessarily meant to be a proper subset. In other words, a first subset of a first set can be coextensive (equal to) the first set.

[0170] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

[0171] As used herein, the terms “component,” “module,” “system,” “interface,” and the like are generally intended to refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to being, a process running on the processing circuitry 1502, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a controller and the controller can be a component. One or more components can reside within a process and / or thread of execution and a component can be localized on one computer and / or distributed between two or more computers.

[0172] Further, some example embodiments can include a method of using standard programming and / or engineering techniques to produce software, firmware, hardware, or any combination thereof to control a computer to implement the disclosed subject matter. The term “article of manufacture” as used herein is intended to encompass a computer program accessible from any computer-readable device, carrier, or media. Of course, those skilled in the art will recognize that many modifications can be made to this configuration without departing from the scope or spirit of the claimed subject matter.

[0173] Various operations of embodiments are provided herein. In some example embodiments, one or more of the described operations can constitute computer- readable instructions stored on one or more computer-readable media that, when executed by a computing device, will cause the computing device to perform the described operations. The order in which some or all of the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or omitted. Alternative ordering will be appreciated by skilled artisans per the description herein. Further, it should be understood that not all operations are necessarily present in every example embodiment provided herein.

[0174] As used herein, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise, or clear from context, “X employs A or B” is intended to mean any of the natural inclusive permutations. That is, if X employs A; X employs B; or X employs both A and B, then “X employs A or B” is satisfied under any of the foregoing instances. As used herein and in the appended claims, the article “a” and “an” as used with “comprising” and / or “including” can generally be interpreted to mean “one or more” unless otherwise indicated or clear from context to be directed to a singular form.

[0175] While the present disclosure has been shown and described with respect to certain example embodiments thereof, it will be apparent to other skilled in the art that various alterations and modifications can be made to the present disclosure based on reading and understanding the specification and the appended drawings. The present disclosure includes all such modifications and alterations and is only limited by the scope of the appended claims. In particular, with respect to the various functions described above with regard to the components (e.g., elements, resources, etc.) described above, unless otherwise specified, the terminology used to describe such components is intended to correspond to any component (e.g., functionally equivalent) that performs the specified function of the described component, even if not structurally equivalent to the disclosed structure that performs the function in some of the example embodiments of the present disclosure illustrated herein. Moreover, while a particular feature of the present disclosure can have been disclosed with respect to only one of several implementations, such feature can be combined with one or more other features of the other implementations as can be desired and advantageous for any given or particular application. Furthermore, to the extent that the terms "includes", "including", "has", "having", "with" or variants thereof are used in either the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "comprising".

Claims

1. A method of operating an apparatus comprising a processing circuit system, the method comprising: The processing circuit system executes instructions that cause the device to: Receive an image depicting at least a portion of the tumor; Based on the image, the lymphocyte distribution of lymphocytes in the tumor was determined; A Gaussian mixture model is applied to the lymphocyte distribution to classify the tumor, wherein... The Gaussian mixture model was trained to classify tumors into categories selected from at least two categories associated with lymphocyte distribution. The Gaussian mixture model is configured to determine the probability distribution of features of tumors of that category in the feature space for that category, and The features of the feature space of the Gaussian mixture model are selected from a feature set, which includes at least one of the following: measurement results of the tumor region of the image, measurement results of the stroma region of the image, measurement results of the lymphocyte region of the image, measurement results of the tumor-infiltrating lymphocyte region of the image, measurement results of the tumor-adjacent lymphocyte region of the image, measurement results of the stroma-infiltrating lymphocyte region of the image, measurement results of the stroma-adjacent lymphocyte region of the image, measurement results of the tumor-infiltrating stroma region of the image, or measurement results of the tumor-adjacent stroma region of the image, and The feature subset is selected from the feature set based on the relevance between the corresponding category and the corresponding feature of the subset. The correlation between the corresponding category and the corresponding feature is based on at least one of the contour score and consistency index of the feature space; as well as Based on the prognostic dataset corresponding to individuals with tumors classified into the category by the classifier, the individual's clinical value is determined.

2. The method of claim 1, wherein the feature subset consists essentially of the following: The measurement results of the lymphocyte region in the image, The measurement results of the tumor-infiltrating lymphocyte region in the image. The measurement results of the tumor-adjacent lymphocyte region in the image, and The measurement results of the tumor-infiltrating stroma region in the image.

3. The method of claim 1, wherein the tumor is one of the following: Pancreatic adenocarcinoma tumor; or Breast cancer tumor.

4. The method according to claim 1, wherein: The device also includes a convolutional neural network trained to determine the lymphocyte distribution of lymphocytes in regions of an image; and The instructions cause the device to invoke the convolutional neural network to determine the lymphocyte distribution of lymphocytes in the corresponding region of the image of the tumor.

5. The method of claim 4, wherein the convolutional neural network is further trained to classify regions of the image into one or more region types selected from a set of region types, the set of region types including: Tumor area; Lymphocyte region; as well as Interstitial region.

6. The method according to claim 5, wherein: Determining the lymphocyte distribution in the tumor includes, for the corresponding lymphocyte regions in the image: Determine the distance from the lymphocyte region to one or both of the tumor region and the stromal region; and Based on the distance, the lymphocyte region can be characterized as one of the following: Tumor-infiltrating lymphocyte region The tumor is located near the lymphocyte region. Interstitial infiltrating lymphocyte areas, and Interstitial adjacent lymphocyte regions; and The classifier also classifies the tumor based on the characteristics of the lymphocyte regions.

7. The method according to claim 5, wherein: Determining the lymphocyte distribution in the tumor includes, for the corresponding stromal region of the image, Determine the distance from the stromal region to the tumor region, and Based on the distance, the interstitial region can be characterized as one of the following: Tumor infiltration of the stroma, and The stromal region adjacent to the tumor; and The classifier also classifies the tumor based on the characterization of the interstitial region.

8. The method of claim 1, wherein the at least two categories comprise: High-risk cancer categories associated with first survival probability; as well as Low-risk tumor categories associated with a second survival probability that is longer than the first survival probability.

9. The method of claim 1, wherein the instructions further cause the device to display the Kaplan Meier survival projection of the individual's clinical value.

10. The method of claim 1, wherein the instructions further cause the apparatus to determine at least one of the following: Based on the individual's clinical values, a diagnostic test for the tumor is determined; Based on the individual's clinical values, determine the individual's treatment; or Based on the individual's clinical values, a schedule for therapeutic agents to treat the tumor is determined.

11. A method of operating an apparatus comprising a processing circuit system, the method comprising: The processing circuit system executes instructions that cause the device to: Receive an image depicting at least a portion of the tumor; Based on the image, the lymphocyte distribution of lymphocytes in the tumor was determined; A classifier is applied to the lymphocyte distribution to classify the tumor, the classifier being trained to classify the tumor into a category selected from at least two categories associated with the lymphocyte distribution; The Cox proportional hazards model is applied to the clinical characteristics of the tumor to determine the tumor category, wherein, The clinical characteristics of the tumor in the Cox proportional hazards model are selected from a set of clinical characteristics, which includes the initial diagnosis of the tumor, the location of the tumor, the treatment of the tumor, the measurement results of the tumor, the metastatic status of the tumor, the individual's initial diagnosis, the individual's previous cancer history, the individual's race, the individual's ethnicity, the individual's sex, the individual's smoking frequency, the individual's smoking duration, and the individual's alcohol consumption history. The clinical feature subset is selected from the clinical feature set based on the correlation between the corresponding category and the corresponding clinical feature in the clinical feature subset, and for the Cox proportional hazards model. The subset of clinical features consists of the measurements of the tumor and the metastatic status of the tumor; and Based on the prognostic dataset corresponding to individuals suffering from tumors of the category to which the classifier classifies the tumor, and based on the prognosis of the individuals suffering from tumors of the category to which the classifier classifies the tumor and the category of tumors determined by the Cox proportional hazards model, the individual's clinical value is determined.

12. The method according to claim 11, wherein, The tumor is one of the following: Pancreatic adenocarcinoma tumor, and Breast cancer tumor.

13. The method according to claim 11, wherein: The device also includes a convolutional neural network trained to determine the lymphocyte distribution of lymphocytes in regions of an image; and The instructions cause the device to invoke the convolutional neural network to determine the lymphocyte distribution of lymphocytes in the corresponding region of the image of the tumor.

14. The method of claim 13, wherein the convolutional neural network is further trained to classify regions of the image into one or more region types selected from a set of region types, the set of region types including: Tumor area; Lymphocyte region; as well as Interstitial region.

15. The method of claim 14, wherein: Determining the lymphocyte distribution in the tumor includes, for the corresponding lymphocyte regions in the image: Determine the distance from the lymphocyte region to one or both of the tumor region and the stromal region; and Based on the distance, the lymphocyte region can be characterized as one of the following: Tumor-infiltrating lymphocyte region The tumor is located near the lymphocyte region. Interstitial infiltrating lymphocyte areas, and Interstitial adjacent lymphocyte regions; and The classifier also classifies tumors based on the characteristics of the lymphocyte regions.

16. The method of claim 14, wherein: Determining the lymphocyte distribution in the tumor includes, for the corresponding stromal region of the image, Determine the distance from the stromal region to the tumor region, and Based on the distance, the interstitial region can be characterized as one of the following: Tumor infiltration of the stroma, and The stromal region adjacent to the tumor; and The classifier also classifies the tumor based on the characterization of the interstitial region.

17. The method of claim 11, wherein the at least two categories comprise: High-risk cancer categories associated with first survival probability; as well as Low-risk tumor categories associated with a second survival probability that is longer than the first survival probability.

18. The method of claim 11, wherein the classifier further comprises a Gaussian mixture model configured to determine a probability distribution of features of the tumor in a feature space for a given category, and the features of the feature space of the Gaussian mixture model are selected from a feature set including at least one of the following: Measurement results of the tumor region in the image; The measurement results of the interstitial region of the image; Measurement results of the lymphocyte region in the image; Measurement results of the tumor-infiltrating lymphocyte region in the image; Measurement results of the tumor-adjacent lymphocyte region in the image; Measurement results of the interstitial infiltrated lymphocyte region in the image; The measurement results of the interstitial adjacent lymphocyte regions in the image; The measurement results of the tumor-infiltrating stroma region in the image; or The image shows measurements of the stroma region adjacent to the tumor.

19. The method according to claim 11, wherein, The feature subset is selected from the feature set based on the correlation between the corresponding category and the corresponding feature of the subset, and the correlation between the corresponding category and the corresponding feature is based on at least one of the contour score and consistency index of the feature space.

20. A method of operating an apparatus comprising a processing circuit system, the method comprising: The processing circuit system executes instructions that cause the device to: Receive an image depicting at least a portion of the tumor; By applying a convolutional neural network to the image, the lymphocyte distribution of lymphocytes in the tumor is determined based on the image; the convolutional neural network is configured to measure the lymphocyte distribution of different types of lymphocytes in different regions of the image. A bidirectional Gaussian mixture model is applied to the lymphocyte distribution to classify the tumor, wherein the bidirectional Gaussian mixture model is trained to classify the tumor into categories selected from at least two categories associated with the lymphocyte distribution. The at least two categories are low-risk tumor categories and high-risk tumor categories, and The bidirectional Gaussian mixture model is configured to determine the probability distribution of features of tumors of the corresponding category in the feature space. The Cox proportional hazards model is applied to the clinical characteristics of the tumor to determine the tumor category; and Based on the prognostic dataset of individuals with tumors corresponding to the categories into which the tumors are classified by the classifier, and based on the categories determined by the Cox proportional hazards model, the individual's clinical value is determined.