Deep learning model for tumor assessment
Through deep learning models, the lymphocyte distribution in tumor images was analyzed, combined with Gaussian mixed model and Cox proportional hazards model, the complex factors and difficulties in tumor prognosis assessment were solved, and more accurate individual prognosis prediction and treatment plan guidance were achieved.
Patent Information
- Application Number
- CN202180008820.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-11
- Filing Date
- 2021-01-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-01-11
AI Technical Summary
When evaluating the prognosis of cancer tumors, the prior art faces the problem that due to the large number and variety of related factors and the complex correlation, it is difficult to accurately determine the clinical value of an individual.
The deep learning model was used, combined with convolutional neural network and Gaussian hybrid model, and the tumor category was determined by analyzing the lymphocyte distribution in the tumor image, and the Cox proportional hazards model was used to predict individual prognosis.
It improves the accuracy of tumor prognosis and predicts individual clinical values, providing more accurate survival prediction and treatment plan guidance.
Smart Images

Figure CN115039125B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 959,931, filed on January 11, 2020, the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] This disclosure relates to the field of using machine - learning models for image analysis, and more particularly to using deep - learning models to determine tumor - related clinical variables. Background Art
[0004] The background art description provided herein is for the purpose of generally presenting the context of the present disclosure. The work of the current inventors within the scope described in this background art section and aspects that may not qualify as prior art at the time of filing are neither expressly nor impliedly admitted to be prior art with respect to the present disclosure.
[0005] In the medical field, many scenarios involve the analysis of cancer tumors to evaluate tumor classes based on characteristics such as location, size, shape, and composition. The evaluation can enable predictions such as tumor behavior and likely aggressiveness, such as the probability and rate of growth and / or metastasis. These characteristics of the tumor can in turn determine the clinical values (such as prognosis) of an individual, such as the likely survival rate, and can guide medical decisions such as the selection, type, and / or timing of chemotherapy, surgery, and palliative care. However, the determination of prognosis (including survival ability) is made difficult by the number and variety of relevant factors that can affect the determination and the factor correlations.
[0006] Multiple diagnostic and prognostic techniques can be used to perform tumor class evaluation. For example, a data set can include characteristics of individuals with tumors (such as the age, physiology, medical history, and / or behavioral habits such as smoking of each individual), which can be correlated with the prognostic data of the individual (such as typical survival rates). The Cox proportional hazards model can be applied to determine the relevance of features in a clinical feature set based on the collected data set and clinical data, which can support some conclusions about the relevance of individual risk factors for clinical values (such as prognosis) based on the tumor. Thereafter, clinicians can use this information to guide the determination or prediction of the diagnosis, prognosis, and / or effective care options for individuals with similar tumors. Also, similar risk factors for these individuals can be collected and processed through the Cox proportional hazards model to predict the clinical values (such as prognosis) of individuals based on similar tumors, where the Cox proportional hazards model is developed based on these similar tumors.
[0007] Other techniques for assessing an individual's tumor-based clinical values, such as prognosis, can utilize one or more machine learning models. For example, a training data set of tumor samples with known characteristics can be generated, such as data about the tumor, the individual from whom the tumor was removed, and / or the individual's clinical value (such as prognosis). A machine learning classifier can be trained using the labeled training data set, where the labels indicate the class represented by each input, such as whether each tumor represents a high-risk tumor with a poor prognosis or a low-risk tumor with a good prognosis. The training process can result in a trained machine learning classifier that classifies new inputs in a manner consistent with the examples in the training data set.
[0008] A variety of different machine learning models can be selected as classifiers for tumors, such as Bayesian classifiers, artificial neural networks, and support vector machines (SVMs). As a first such example, a convolutional neural network (CNN) can process an n-dimensional input to detect features of a tumor image that may be present therein. A feature vector, such as an array of pixels of a tumor image, can be provided to the convolutional neural network that includes a sequence of neuron convolutional layers. Each convolutional layer can produce a feature map indicating some of the tumor image features detected at a level of detail, and the feature map can be processed by the next convolutional layer in the sequence. The feature map produced by the final convolutional layer of the convolutional neural network can be processed by a classifier for the tumor, and the classifier can be trained to indicate whether the feature map is similar to the feature map of an object in an image of the training data set. For example, a CNN can be trained to recognize tumor visual features associated with high-risk and low-risk prognoses.
[0009] As a second such example, a Gaussian mixture model (GMM) can be generated to classify data about tumors into different tumor clusters with representative characteristics. For each sample representing a tumor in the training data set, a set of features can be identified, such as location, size, shape, and composition. The samples of the training data set can be located within a multi-dimensional feature space, where each feature is represented along a dimensional axis. Machine learning techniques can be applied to identify clusters of tumors within the feature space with similar prognoses, such as a first cluster representing high-risk tumors and a second cluster representing low-risk tumors, where each cluster is represented as a set of Gaussian probability distributions of the corresponding features within the feature space. Even if some portions of the clusters overlap (e.g., even if a tumor with a particular set of features can be included in the high-risk or low-risk cluster), the Gaussian probability distributions of the clusters can enable a probability prediction regarding the likelihood of a tumor belonging to each cluster. In this manner, the Gaussian mixture model can achieve individual prognosis prediction based on the clustering of similar tumor samples in the training data set. Summary of the Invention
[0010] Some example embodiments may include a method of operating an apparatus including processing circuitry, where the method includes executing instructions by the processing circuitry that cause the apparatus to: receive an image depicting at least a portion of a tumor; determine a lymphocyte distribution of lymphocytes in the tumor based on the image; apply a classifier to the lymphocyte distribution to classify the tumor, the classifier having been trained to classify tumors into categories selected from at least two categories respectively associated with lymphocyte distributions; and determine a clinical value of an individual based on a prognosis data set corresponding to individuals with tumors of the category into which the classifier classifies the tumor.
[0011] In some example embodiments, the tumor is a pancreatic cancer tumor or a breast cancer tumor.
[0012] In some example embodiments, the apparatus may further include a convolutional neural network that is trained to determine a lymphocyte distribution of lymphocytes in an image region, and the instructions may cause the apparatus to invoke the convolutional neural network to determine the lymphocyte distribution of lymphocytes in a corresponding region of the image of the tumor. In some example embodiments, the convolutional neural network may further be trained to classify the region of the image into one or more region types selected from a set of region types including a tumor region, a lymphocyte region, or a stromal region. In some example embodiments, determining the lymphocyte distribution of lymphocytes in the tumor may include, for a corresponding lymphocyte region of the image: determining a distance of the lymphocyte region to one or both of the tumor region or the stromal region; characterizing the lymphocyte region as one of a tumor-infiltrating lymphocyte region, a tumor-adjacent lymphocyte region, a stromal-infiltrating lymphocyte region, or a stromal-adjacent lymphocyte region based on the distance, and the classifier may further classify the tumor based on the characterization of the lymphocyte region. In some example methods, determining the lymphocyte distribution of lymphocytes in the tumor may include, for a corresponding stromal region of the image: determining a distance of the stromal region to the tumor region; and characterizing the stromal region as one of a tumor-infiltrating stromal region or a tumor-adjacent stromal region based on the distance, and the classifier may further classify the tumor based on the characterization of the stromal region.
[0013] In some example embodiments, the at least two categories may include a high-risk tumor category associated with a first survival probability, and a low-risk tumor category associated with a second survival probability that is longer than the first survival probability.
[0014] In some example embodiments, the classifier may further include a Gaussian mixture model configured to determine a probability distribution of features of tumors of a corresponding class within a feature space for the corresponding class. In some example embodiments, features of the feature space of the Gaussian mixture model may be selected from a feature set including: measurements of a tumor region of the image, measurements of an interstitial region of the image, measurements of a lymphocyte region of the image, measurements of a tumor-infiltrating lymphocyte region of the image, measurements of a tumor-adjacent lymphocyte region of the image, measurements of an interstitial-infiltrating lymphocyte region of the image, measurements of an interstitial-adjacent lymphocyte region of the image, measurements of a tumor-infiltrating interstitial region of the image, and measurements of a tumor-adjacent interstitial region of the image. In some example embodiments, a feature subset may be selected from the feature set based on a correlation of the corresponding classes with corresponding features of the subset. In some example embodiments, the correlation of the corresponding classes with the corresponding features may be based on one or both of a contour score or a consistency index of the feature space. In some example embodiments, the feature subset may consist essentially of measurements of a lymphocyte region of the image, measurements of a tumor-infiltrating lymphocyte region of the image, measurements of a tumor-adjacent lymphocyte region of the image, and measurements of a tumor-infiltrating interstitial region of the image.
[0015] In some example embodiments, the instructions may further cause the device to apply a Cox proportional hazards model to clinical features of the tumor to determine a class of the tumor, and to determine a clinical value (such as a prognosis) of an individual based on a prognosis of an individual having a tumor classified into a class by the classifier and a class determined by the Cox proportional hazards model. In some example embodiments, clinical features of the Cox proportional hazards model of the tumor may be selected from a clinical feature set including: a primary diagnosis of the tumor, a location of the tumor, a treatment of the tumor, a measurement of the tumor, a metastasis status of the tumor, a primary diagnosis of the individual, a previous cancer history of the individual, a gender of the individual, a frequency of a smoking habit of the individual, a duration of a smoking habit of the individual, and a history of alcohol consumption of the individual. In some example embodiments, a clinical feature subset of the clinical features may be selected from the clinical feature set for the Cox proportional hazards model based on a correlation of the corresponding classes with corresponding clinical features of the subset. In some example embodiments, the clinical feature subset may consist of a measurement of the tumor and a metastasis status of the tumor.
[0016] In some example embodiments, the instructions may further cause the device to display a visualization of the individual's clinical values (such as prognosis). In some example embodiments, the visualization is a Kaplan Meier survival projection of the tumor. In some example embodiments, the instructions may further cause the device to determine a diagnostic test for the tumor based on the individual's clinical values (such as prognosis). In some example embodiments, the instructions may further cause the device to determine a treatment for the individual based on the individual's clinical values (such as prognosis). In some example embodiments, the instructions may further cause the device to determine a schedule of therapeutic agents for treating the tumor based on the individual's clinical values (such as prognosis).
[0017] In some example embodiments, the at least two categories are a low-risk tumor category and a high-risk tumor category, determining the lymphocyte distribution further includes applying a convolutional neural network to the image, the convolutional neural network being configured to measure the lymphocyte distribution of lymphocytes of different region types of the image, the classifier being a bidirectional Gaussian mixture model, the bidirectional Gaussian mixture model being configured to determine a probability distribution of features of tumors of the corresponding category within a feature space for the corresponding category, the method further includes applying a Cox proportional hazards model to the clinical features of the tumor to determine the category of the tumor, and determining the individual's clinical values (such as prognosis) is further based on the category predicted by the Cox proportional hazards model.
[0018] Some example embodiments may include a system, the system includes: memory hardware, the memory hardware being configured to store instructions embodying any of the above methods; and processing hardware, the processing hardware being configured to execute the instructions stored by the memory hardware.
[0019] Some example embodiments may include a system that includes: an image evaluator configured to determine a lymphocyte distribution of lymphocytes in an image; a classifier configured to classify a tumor into a category selected from at least two categories respectively associated with the lymphocyte distribution; and a tumor evaluator configured to determine a clinical value (such as prognosis and / or viability) of an individual based on the tumor in the image by: invoking the image evaluator with the image to determine the lymphocyte distribution of lymphocytes in the tumor, invoking the classifier to classify the tumor into a certain category based on the lymphocyte distribution, and outputting the clinical value (such as prognosis) of the individual based on the prognosis of individuals with tumors of the category into which the classifier classifies the tumor. In some example embodiments, the at least two categories are a low-risk tumor category and a high-risk tumor category, the image evaluator is a convolutional neural network configured to measure the lymphocyte distribution of lymphocytes of different region types of the image, the classifier is a bidirectional Gaussian mixture model configured to determine a probability distribution of features of tumors of the corresponding category within a feature space, the system further includes a Cox proportional hazards model for the clinical features of the tumor to determine the category of the tumor, and the tumor evaluator is further configured to determine the clinical value (such as prognosis) of the individual based on the prognosis of individuals with tumors of the category into which the classifier classifies the tumor and the category determined by the Cox proportional hazards model.
[0020] Some example embodiments may include a system that includes: an image evaluation device configured to determine a lymphocyte distribution of lymphocytes in an image; a classification device configured to classify a tumor into a category selected from at least two categories respectively associated with the lymphocyte distribution; and a tumor evaluator device configured to determine a clinical value (such as prognosis) of an individual based on the tumor in the image by: invoking the image evaluation device with the image to determine the lymphocyte distribution of lymphocytes in the tumor, invoking the classifier to classify the tumor into a certain category based on the lymphocyte distribution, and outputting the clinical value (such as prognosis) of the individual based on the prognosis of individuals with tumors of the category into which the classifier classifies the tumor.
[0021] Some example embodiments may include an apparatus that includes a memory storing instructions and processing circuitry configured to determine a clinical value (such as prognosis) of an individual based on a tumor in an image by performing the following operations: determining a lymphocyte distribution of lymphocytes in the tumor based on an image of the tumor; applying a classifier to the lymphocyte distribution to classify the tumor, the classifier being configured to classify the tumor into a category selected from at least two categories respectively associated with the lymphocyte distribution; and outputting the clinical value (such as prognosis) of the individual based on the prognosis of an individual having a tumor of the category into which the classifier classifies the tumor. In some example embodiments, the at least two categories are a low-risk tumor category and a high-risk tumor category, determining the lymphocyte distribution further includes applying a convolutional neural network to the image, the convolutional neural network being configured to measure the lymphocyte distribution of lymphocytes of different region types of the image, the classifier includes a two-way Gaussian mixture model configured to determine a probability distribution of features of tumors of the category in a feature space for the corresponding category, and the instructions further cause the processing circuitry to apply a Cox proportional hazards model to the clinical features of the tumor to determine the category of the tumor, and determine the clinical value (such as prognosis) of the individual based on the prognosis of an individual having a tumor of the category into which the classifier classifies the tumor and the category determined by the Cox proportional hazards model.
[0022] Some example embodiments may include a non-transitory computer-readable medium storing instructions that, when executed by processing circuitry, cause the processing circuitry to determine a clinical value (such as prognosis) of an individual based on a tumor in an image by performing the following operations: determining a lymphocyte distribution of lymphocytes in the tumor based on an image of the tumor; applying a classifier to the lymphocyte distribution to classify the tumor, the classifier being configured to classify the tumor into a category selected from at least two categories respectively associated with the lymphocyte distribution; and outputting the clinical value (such as prognosis) of the individual based on the prognosis of an individual having a tumor of the category into which the classifier classifies the tumor. In some example embodiments, the at least two categories are a low-risk tumor category and a high-risk tumor category, determining the lymphocyte distribution further includes applying a convolutional neural network to the image, the convolutional neural network being configured to measure the lymphocyte distribution of lymphocytes of different region types of the image, the classifier includes a two-way Gaussian mixture model configured to determine a probability distribution of features of tumors of the category in a feature space for the corresponding category, and the instructions further cause the processing circuitry to apply a Cox proportional hazards model to the clinical features of the tumor to determine the category of the tumor, and determine the clinical value (such as prognosis) of the individual based on the prognosis of an individual having a tumor of the category into which the classifier classifies the tumor and the category determined by the Cox proportional hazards model. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The present disclosure will be more fully understood from the detailed description and the accompanying drawings. In the drawings, reference numerals may be repeated to identify similar and / or identical elements.
[0024] Figure 1 is an illustration of an example convolutional neural network.
[0025] Figure 2A is an illustration of an example image analysis for identifying region types and lymphocyte distributions in tumor images according to some example embodiments.
[0026] Figure 2B is an illustration of an example image analysis for classifying lymphocyte distributions in tumor images according to some example embodiments.
[0027] Figure 3 is an illustration of a set of masks for a lung tissue sample including lymphocytes of lymphocyte distributions by an example machine learning model according to some example embodiments.
[0028] Figure 4 is an illustration of an example machine learning model for classifying tumors according to some example embodiments.
[0029] Figure 5 is an illustration of a characterization of a set of images of a pancreatic cancer tissue sample according to some example embodiments.
[0030] Figure 6A is an illustration of a set of samples arranged in a two - dimensional feature space.
[0031] Figure 6B is an illustration of a Gaussian mixture model configured to classify the set of samples into a set of probability distribution clusters within the two - dimensional feature space.
[0032] Figure 6C is another illustration of a Gaussian mixture model configured to classify the set of samples into a set of probability distribution clusters within the two - dimensional feature space.
[0033] Figure 7 is an illustration of selecting a feature subset for a classifier from a set of features within a feature space based on the correlation between corresponding features and corresponding classes according to some example embodiments.
[0034] Figure 8 is an illustration of the classification of different types of tumors based on a feature subset according to some example embodiments.
[0035] Figure 9 is an illustration of a Kaplan Meier survival plot based on image analysis according to some example embodiments.
[0036] Figure 10 It is a diagram showing the selection of a subset of clinical features for a Cox proportional hazards model from a set of clinical features in a feature space based on the correlation between corresponding features and corresponding classes according to some example embodiments.
[0037] Figure 11 It is a diagram showing a Kaplan Meier survival plot based on image analysis and a Cox proportional hazards model according to some example embodiments.
[0038] Figure 12 It is a diagram showing a result set of the classification of a tumor training data set and a tumor test data set based on image analysis and a Cox proportional hazards model according to some example embodiments.
[0039] Figure 13 It is a flowchart of a first example method according to some example embodiments.
[0040] Figure 14 It is a flowchart of a second example method according to some example embodiments.
[0041] Figure 15 It is a block diagram of components of an example apparatus according to some example embodiments.
[0042] Figure 16 It is a block diagram of components of another example apparatus according to some example embodiments.
[0043] Figure 17 It is a diagram of an example computer-readable medium according to some example embodiments.
[0044] Figure 18 It is a diagram of an example apparatus in which some example embodiments can be implemented. Detailed implementation
[0045] A. Introduction
[0046] The following introduction aims to provide an overview of some machine learning features related to some example embodiments.
[0047] Figure 1 It is an example of a convolutional neural network (CNN) 110 trained to process an n-dimensional input to detect multiple features.
[0048] In Figure 1 example, the convolutional neural network 110 processes the image 102 into a two-dimensional array of pixels 108 of one or more colors, but some such convolutional neural networks 110 can process other forms of data, such as sound, text, or signals from sensors.
[0049] In Figure 1In the example, the training dataset 100 is provided as a set of images 102 each associated with a class 106 from the class set 104. For example, the training dataset 100 may include a first vehicle image 102-1 associated with a first class 106-1 for vehicle images; a second house image 102-2 associated with a second class 106-2 for house images; and a third cat image 102-3 associated with a third class 106-3 for cat images. The association of the images 102 with the corresponding classes 106 is sometimes referred to as the label of the training dataset 100.
[0050] As Figure 1 As further shown, each image 102 can be processed by a convolutional neural network 110 organized as a series of convolutional layers 112, each convolutional layer having a set of neurons 114 and one or more convolutional filters 116. In the first convolutional layer 112-1, each neuron 114 can apply the first convolutional filter 116-1 to a certain region of the image and can output an activation indicating whether the pixels in that region correspond to the first convolutional filter 116-1. The set of activations (referred to as the feature map 118-1) generated by the neurons 114 of the first convolutional layer 112-1 can be received as input by the second convolutional layer 112-2 in the sequence of convolutional layers 112 of the convolutional neural network 110, and the neurons 114 of the second convolutional layer 112-2 can apply the second convolutional filter 116-2 to the feature map 118-1 generated by the first convolutional layer 112-1 to produce a second feature map 118-2. Similarly, the second feature map 118-2 can be received as input by the third convolutional layer 112-3, and the neurons 114 of the third convolutional layer 112-3 can apply the third convolutional filter 116-3 to the feature map 118-2 generated by the second convolutional layer 112-2 to produce a third feature map 118-3. Such a machine learning model including a large number of multi-layer or more complex layer architectures is sometimes referred to as a deep learning model.
[0051] As Figure 1Further shown, the third feature map 118-3 produced by the third and final convolutional layer 112-3 can be received by a classification layer 120 (such as a "dense" or fully connected layer), which can perform classification of the third feature map 118-3 to determine the classification of the content of the image 102-1. For example, each neuron 114 of the classification layer 120 can apply weights to each activation of the third feature map 118-3. Each neuron 114 outputs an activation as the sum of the products of each activation of the third feature map and the weights connecting the neuron 114 to the activation. Thus, each neuron 114 outputs an activation indicating the degree to which the activations included in the third feature map 118-3 match the corresponding weights of the neuron 114. Further, the weights of each neuron 114 are selected based on the activations of the third feature map 118-3 produced by an image 102 of one of the categories 106 in the category set 104. That is, each neuron 114 outputs an activation based on the similarity between the third feature map 118-3 of the currently processed image 102-1 and the third feature map 118-3 produced by an image 102 of one of the categories 106 in the category set 104 by the convolutional neural network 110. Comparison of the outputs of the neurons 114 of the classification layer 120 can allow the convolutional neural network 110 to perform classification 122 by selecting the category 106-4 with the highest probability corresponding to the third feature map 118-3. In this way, the convolutional neural network 110 can perform classification 122 of the image 102-1 as the category 106 of the most similar image 102 among the multiple images 102.
[0052] As Figure 1 Further shown, a training process can be applied to train the convolutional neural network 110 to recognize the category set 104 represented by a specific set of images 102 of the training dataset 100. During the training process, each image 102 of the training dataset 100 can be processed by the convolutional neural network 110, producing a classification 122 of the image 102. If the classification 122 is incorrect, the convolutional neural network 110 can be updated by adjusting the weights of the neurons 114 of the classification layer 120 and the filters 116 of the convolutional layer 112 such that the classification 122 of the convolutional neural network 110 is closer to the correct classification 122 of the image 102 being processed. Repeating the training of the convolutional neural network 110 on the training dataset 100 while incrementally adjusting the convolutional neural network 110 to produce the correct classification 122 for each image 102 can result in the convergence of the convolutional neural network 110, where the convolutional neural network 110 correctly classifies the images 102 of the training dataset 100 within an acceptable error range. Examples of convolutional neural network architectures include ResNet and Inception.
[0053] As regarding Figure 1As discussed, machine learning models such as convolutional neural network 110 can classify inputs such as image 102 based on the arrangement of features relative to each other (such as the number, orientation, and location of characteristics of recognizable features). For example, convolutional neural network 110 can classify image 102 by generating: a first feature map 118-1 indicating the detection of certain geometric shapes (such as curves and lines) occurring at various positions within image 102; a second feature map 118-2 indicating that the geometric shapes are arranged to produce certain higher-level features (such as a set of curves arranged as a circle or a set of lines arranged as a rectangle); and a third feature map 118-3 indicating that the higher-level features are arranged to produce even higher-level features (such as a set of circles arranged as wheels and a set of rectangles arranged as doorframes). Neurons 114 of classification layer 120 of convolutional neural network 110 can determine that the features of the third feature map 118-3 (such as two wheels positioned between two doorframes) are arranged in a manner that depicts the side of a vehicle (such as a car). Similar classifications 122 can be performed by other neurons 114 of classification layer 120 to classify image 102 as belonging to other categories 106 within category set 104, such as the arrangement of two eyes, two triangular ears, and a nose depicting a cat, or the arrangement of windows, doorframes, and a roof depicting a house. In this way, machine learning models such as convolutional neural networks can classify inputs (such as image 102) based on feature arrangements corresponding to similar feature arrangements depicted in inputs (such as image 102) of a training dataset 100. Additional details regarding convolutional neural networks and other machine learning models (including support vector machines) can be found in U.S. Patent Application 62 / 959,931, which is hereby incorporated by reference as if fully rewritten.
[0054] B. Distribution-Based Classification
[0055] In some machine learning scenarios, the classification of an input (such as an image) can occur based on the arrangement of recognizable features relative to each other (such as the number, orientation, and / or location of recognizable pixel patterns detected in the feature maps 118 of convolutional neural network 110). However, in some other scenarios, the classification may not be based on the features relative to each other Arrangement , but rather on the features in the input Distribution, such as whether the density change of the features on the input region corresponds to the recognizable density change that is a characteristic of class 106. That is, when a set of lower-level features corresponds to the recognized arrangement of higher-level features relative to each other (e.g., quantity, orientation, and / or positioning) (the recognized arrangement corresponding to class 106), class 106 in class set 104 may not be recognizable. Instead, each class 106 can be recognized as the correspondence between the activation distribution of the features and some characteristics of the input. This distribution may not reflect any specific quantity, orientation, and / or positioning of the activation of the features of the input, but can indicate whether the distribution of the activation of the features corresponds to the distribution of the activation of the features of the corresponding class 106. In this scenario, the input for each class 106 (such as the training data set) can be associated with the characteristic distribution of the activation of the features, and the classification of the input can be based on whether the distribution of the activation of the features of the input corresponds to the distribution of the activation of the features in the input for each class 106. This distribution-based classification may also occur in various scenarios.
[0056] Figure 2A and Figure 2B Together show examples of several types of image analysis that can be used to identify region types and lymphocyte distribution in tumor images in some example embodiments.
[0057] Figure 2A Is an illustration of an example image analysis for identifying region types and lymphocyte distribution in tumor images according to some example embodiments. As Figure 2A shown, the data set can include an image 102 of tissue of an individual suffering from one type of tumor, and stroma including connective tissue and tumor support. The classification of the features of image 102 can enable the determination of regions, for example, parts of regions of image 102 having similar features. Further identification of regions of the feature map 118 of a certain feature including filter 116 can enable the determination of the region types of the corresponding regions 200, such as a first region depicting a tumor and a second region depicting stroma in image 102. Thus, each filter 116 of the convolutional neural network 110 can be regarded as a mask of the region of image 102 indicating a specific region type, such as the presence, size, shape, and extent of a tumor or stroma adjacent to the tumor.
[0058] Further, the image 102 may show the presence of lymphocytes, which may be distributed with respect to tumors, stroma, and other tissues. Further analysis of the image may enable the determination 202 of lymphocyte clusters 204 as contiguous regions and / or regions of high lymphocyte concentration (e.g., compared to the lymphocyte concentration in other parts of the image, or compared to a concentration threshold) by, for example, counting the number of lymphocytes present in a particular region of the image 102. Thus, the lowest convolutional layer 112 and the filter 116 of the convolutional neural network 110 are capable of identifying features indicative of tumors, stroma, and lymphocytes.
[0059] Figure 2B is an illustration of an example image analysis for classifying lymphocyte distribution in a tumor image according to some example embodiments. In Figure 2B this, a first image analysis 206 may be performed by first dividing the image 102 into a set of regions and classifying each region of the image 102 as tumor, tumor adjacent, stroma, stroma adjacent, or elsewhere. Thus, each region including a lymphocyte cluster 204 may further characterize the lymphocyte cluster 204 based on the region type, e.g., a first lymphocyte cluster 204-3 that appears within a tumor and a second lymphocyte cluster 204-4 that appears within the stroma.
[0060] A second image analysis 208 can be performed to further compare the location of the lymphocyte clusters 204 with the locations of different region types, thereby further characterizing the lymphocyte clusters 204. For example, the first lymphocyte cluster 204-3 can be identified as appearing within the central portion of the tumor region and / or within a first threshold distance of the location identified as the tumor centroid, and can thus be characterized as a tumor-infiltrating lymphocyte (TIL) cluster. Similarly, the second lymphocyte cluster 204-4 can be identified as appearing within the central portion of the stromal region and thus represents a stromal-infiltrating lymphocyte cluster. However, the third cluster 204-5 can be identified as appearing within the peripheral portion of the tumor region and / or within a second threshold distance of the tumor (the second threshold distance being greater than the first threshold distance), and can thus be characterized as a tumor-adjacent lymphocyte cluster. Alternatively or additionally, the third cluster 204-5 can be identified as appearing within the peripheral portion of the stromal region and / or within a second threshold distance of the stroma (the second threshold distance being greater than the first threshold distance), and can thus be characterized as a stroma-adjacent lymphocyte cluster. Some example embodiments can classify regions as tumor, stroma, or lymphocyte; have two labels, such as tumor and lymphocyte, stroma and lymphocyte, or tumor and stroma; and / or have three labels, such as tumor, stroma, and lymphocyte. Then, some example embodiments can be configured to identify the lymphocyte clusters 204 that appear in each region of the image 102 and tabulate these regions to determine the distribution. In this way, in some example embodiments, image analysis of the image 102, including the feature maps 118 provided by different filters 116 of the convolutional neural network 110, can be used to identify and characterize the distribution and / or concentration of lymphocytes in a tumor image.
[0061] Figure 3 is an illustration of a mask set 300 of masks 302 of a lung tissue sample of a lymphocyte distribution including lymphocytes by an example machine learning model according to some example embodiments. As Figure 3As shown, a mask 302 of the image 102 can be prepared, and each mask 302 indicates regions in the image 102 corresponding to one or more region types. For example, the first mask 302-1 can indicate regions in the image 102 identified as tumor regions. The second mask 302-2 can indicate regions in the image 102 identified as stromal regions. The third mask 302-3 can indicate regions in the image 102 identified as lymphocyte regions. Other masks can also be characterized based on the feature distribution in the feature map 118. For example, the fourth mask 302-4 can indicate regions in the image 102 identified as tumor-infiltrating lymphocyte regions. The fifth mask 302-5 can indicate regions in the image 102 identified as tumor-adjacent lymphocyte regions. The sixth mask 302-6 can indicate regions in the image 102 identified as stromal-infiltrating lymphocyte regions. The seventh mask 302-7 can indicate regions in the image 102 identified as stromal-adjacent lymphocyte regions. The eighth mask 302-8 can indicate regions in the image 102 identified as overlapping stromal and tumor regions. The ninth mask 302-9 can indicate regions in the image 102 identified as tumor-adjacent stromal regions.
[0062] In some example embodiments, the image 102 can also be processed to determine various measurements 304 of the corresponding region types of the image 102. For example, the concentration of each region type (as a percentage of the image 102) can be calculated (e.g., the number of pixels 108 corresponding to each region compared to the total number of pixels in the image 102, optionally considering the apparent concentration of features, such as the density or count of lymphocytes in the corresponding regions of the image 102). In this way, in some example embodiments, the image analysis of the image 102 based on Figure 2A and Figure 2B the distribution analysis shown can be aggregated into the mask 302 and / or a mask group 300 of quantities.
[0063] Figure 4 is an illustration of an example machine learning model for classifying tumors according to some example embodiments. Figure 4 The example machine learning model of includes a first convolutional neural network 112-1, which is configured to perform region classification 402 of the corresponding regions 400 of the tumor image 102 according to different categories such as tumor regions, stromal regions, lymphocyte regions, tumor-infiltrating lymphocyte regions, etc. Based on the region classification 402, a mask group 300 of the mask 302 can be generated. For example, the first mask 302-1 indicating the region 400 as a tumor region in the image 102, the second mask 302-2 indicating the region 400 as a stromal region in the image 102, and the third mask 302-3 indicating the region 400 as a lymphocyte region in the image 102. Figure 4An example system includes a second convolutional neural network 112-2 configured to determine the density or concentration of various features, such as a lymphocyte density range estimate 404 indicating the count or percentage of lymphocytes in a corresponding region 400 of an image 102. Based on the lymphocyte density range estimate 404, a lymphocyte density map 406 indicating regions 400 in the image 102 with a high density of lymphocytes (such as lymphocyte clusters) can be generated. Based on the mask 302 in the mask set 300, the region classification 402, and / or the lymphocyte density map 406 of the lymphocyte density range estimate 404, the image evaluator 408 can identify: aggregated regions of the image 102, such as a tumor region 410-1, a stromal region 410-2, and a lymphocyte region 410-3; one or more region measurements, such as a tumor measurement 304-1, a stromal measurement 304-2, and a lymphocyte measurement 304-3; and / or one or more regions indicating the feature distribution, such as a tumor-infiltrating lymphocyte region, a tumor-adjacent lymphocyte region 412-1, a tumor and stromal region, and a tumor-adjacent stromal region 412-2.
[0064] Generally speaking, in some such example embodiments, one or more convolutional neural networks can be trained to determine the lymphocyte distribution of lymphocytes in an image region, for example, to classify the image region into one or more region types selected from a set of region types including a tumor region, a lymphocyte region, or a stromal region. In some example embodiments, the convolutional neural network can determine the lymphocyte distribution of lymphocytes in a tumor, including for a corresponding lymphocyte region of an image: determining the distance from the lymphocyte region to one or both of the tumor region or the stromal region; and characterizing the lymphocyte region as one of a tumor-infiltrating lymphocyte region, a tumor-adjacent lymphocyte region, a stroma-infiltrating lymphocyte region, or a stroma-adjacent lymphocyte region based on the distance. In some example embodiments, the convolutional neural network can determine the lymphocyte distribution of lymphocytes in a tumor by performing the following operations on a corresponding stromal region of the image: determining the distance from the stromal region to the tumor region, and characterizing the stromal region as one of a tumor-infiltrating stromal region or a tumor-adjacent stromal region based on the distance, and the classifier further classifying the tumor based on the characterization of the stromal region. Thus, the classifier can further classify the tumor based on the characterization of the lymphocyte region. In some example embodiments, many such convolutional neural networks can perform various analyses of an image, which can inform the determination of a clinical value (such as a prognosis) of an individual.
[0065] C. Learning Parameter Determination
[0066] Figure 5 is an illustration of a characterization 500 of a set of images of a pancreatic cancer tissue sample according to some example embodiments. In Figure 5 the graph, a set of tumors is composed ofFigure 4 The system shown is characterized to determine the distribution of detected tumor features, such as tumor region 502-1, stromal region 502-2, lymphocyte region 502-3, tumor-infiltrating lymphocyte region 502-4, tumor-adjacent lymphocyte region 502-5, stromal and tumor-infiltrating lymphocyte region 502-6, stromal-adjacent lymphocyte region 502-7, tumor and stromal region 502-8, and tumor-adjacent stromal region 502-9. Each feature can be evaluated in terms of the density or concentration (vertical axis) and percentage (horizontal axis) of each feature in the tumor image 102. The set of tumor images can be further divided into a training image subset and a test image subset. The training image subset can be used to train machine learning classifiers such as neural networks 112-1, 112-2, etc. to determine features, and the test image subset can be used to evaluate the effectiveness of the machine learning classifier in determining features in previously unseen tumor images. In this way, the machine learning classifier can be verified to determine the consistency of the underlying logic when applied to new data. For example, Figure 5 the graphs in were developed based on diagnostic hematoxylin and eosin staining (H&E staining) of pathological images of pancreatic cancer patients who received chemotherapy.
[0067] Based on factors that are characteristics of tumors of each category, it may be further necessary to characterize a tumor as one of several categories, such as low-risk tumors and high-risk tumors. For example, tumors of corresponding categories can also be associated with different features, such as the concentration and / or percentage of a specific type of tumor region (e.g., tumor-infiltrating lymphocytes), and such differences in characteristic features can enable the distinction of tumors of one category from those of another category. Further, different tumor categories can be associated with different clinical characteristics, such as responsiveness to various treatment options and prognosis, such as viability. To determine such clinical characteristics of a specific tumor in an individual, it may be necessary to determine the tumor category of the tumor in order to guide the selection of the individual's diagnostic and / or treatment regimen.
[0068] However, in many diagnostic scenarios, it may be difficult to associate features that are characteristics of different categories, such as different tumor categories. As a first such example, the characteristics of tumors in one category can differ within a probability range from the characteristics of tumors in another category, and the probability ranges can overlap substantially. For example, the density and percentage of tumor-infiltrating lymphocytes in a high-risk tumor category and a low-risk tumor can each fit a bell curve of probabilities within the tumor category, and the means of the bell curves are only slightly offset, such that the probability distributions may overlap. Thus, it may be difficult to determine whether a tumor exhibiting a feature within the overlapping region belongs to the high-risk category or the low-risk category. As a second such example, different features of a tumor may be covariant; for example, high-risk tumors and low-risk tumors can be distinguished based on a combined probability that differentiates the distribution of tumor-infiltrating lymphocyte regions and tumor-adjacent lymphocyte regions. However, different features may also inherently covary in a non-diagnostic manner. For example, a tumor that exhibits a high density of stromal invasion and a tumor-infiltrating lymphocyte region will typically also necessarily exhibit a high density of the tumor-infiltrating lymphocyte region. Thus, adding the category-based probabilities that a tumor belongs to a category based on stromal invasion and tumor-infiltrating lymphocyte regions and the category-based probabilities that a tumor belongs to a category based on the tumor-infiltrating lymphocyte region may exceed the likelihood of the tumor within the category, because the inherent covariance of the features is not taken into account. Due to these complex characteristics of the data, it may be difficult to determine the distinguishing features of tumors in each category, particularly in high-dimensional feature sets where there may be many features available.
[0069] To classify a dataset exhibiting such overlapping category data, various machine learning models can be used. The corresponding machine learning models can provide different capabilities for classifying the overlapping dataset, e.g., based on uniqueness, tolerance for false positives, tolerance for false negatives, scalability to a large number of features, and characteristics such as avoiding overfitting and underfitting.
[0070] Figures 6A to 6C An example of a Gaussian mixture model is shown together, which can be developed to classify overlapping category data, such as classifying tumors as low-risk tumors and high-risk tumors based on two features, which can be used in some example embodiments.
[0071] Figure 6A is an illustration of a set of samples arranged in a two-dimensional feature space 606. In Figure 6AIn it, the feature space 606 involves samples 600-1 of the first class 602-1 (represented as circles) and samples 600-2 of the second class 602-2 (represented as crosses). Each sample 600 can be evaluated and quantified for the first feature 604-1 and the second feature 604-2, which can enable each sample to be positioned within the two-dimensional feature space 606, where the vertical axis represents the first feature 604-1 and the horizontal axis represents the second feature 604-2. Within the two-dimensional feature space 606, the samples 600 of each class 602 can be clearly clustered, but the clusters can also overlap such that the samples within the overlapping portion can belong to either class 602. For a particular sample 600, it may be desirable to determine the probability that the sample 600 belongs to each class 602 based on the features 604 of the sample 600, especially in the overlapping regions associated with samples 600 of multiple classes 602. Although clustering may be obvious in Figure 6A a simple illustration, such clustering may be more difficult to determine, for example, in a feature space 606 with a higher dimension, in a dataset characterized by classes 602 with a greater degree of overlap, and / or in a dataset in which two or more features 604 co-vary, for which the diagnostic or intrinsic covariance of the features 604 is determined for the two or more features.
[0072] Various machine learning models can be used to classify overlapping datasets, as Figure 6A shown. Some such models include, for example, Bayesian (including naive Bayesian) classifiers; Gaussian classifiers; probabilistic classifiers; principal component analysis (PCA) classifiers; linear discriminant analysis (LDA) classifiers; quadratic discriminant analysis (QDA) classifiers; single-layer or multi-layer perceptron networks; convolutional neural networks; recurrent neural networks; nearest neighbor classifiers; linear SVM classifiers; radial basis function kernel (RBF) SVM classifiers; Gaussian process classifiers; decision tree classifiers, including random forest classifiers; and / or restricted or unrestricted Boltzmann machines, etc.
[0073] Figure 6B is an illustration of a Gaussian mixture model configured to classify the set of samples 600 as shown in Figure 6A into a set of probability distribution clusters within the two-dimensional feature space 606, and in some example embodiments, this Gaussian mixture model can be used to distinguish different types of tumors. In Figure 6BIn [the system], the first Gaussian probability distribution 608-1 of the sample 600-1 of the first category 602-1 can be recognized, and the second Gaussian probability distribution 608-2 of the sample 600-2 of the second category 602-2 can be recognized. For example, based on the mean and variance of the sample 600 of each feature 604, the Gaussian probability distribution 608 of each category 602 can be fitted to the sample 600 of each category 602. The selection of the Gaussian probability distribution 608 can also consider other factors, such as avoiding false negatives (e.g., the sample 600 of the category 602 is wrongly excluded from the category 602) and / or avoiding false positives (e.g., the sample 600 of different categories 602 is wrongly included in the category 602). Further, the Gaussian probability distribution 608 can be selected to model the covariance. For example, by correlating the distribution of the Gaussian probability distribution 608 of the first feature 604-1 with the distribution of the Gaussian probability distribution 608 of the second feature 604-2. For example, similar variances and / or Gaussian probability distributions 608 can be selected for the first feature 604-1 and the second feature 604-2, or can be independently selected for each feature 604. For a specific sample 600 (such as an image of a tumor of an unknown category), the features 604 of the sample 600 can be evaluated to locate the sample 600 within the feature space 606, and the relative probabilities within the Gaussian probability distribution 608 of the corresponding category 602 can be compared to determine the possible category 602 of the tumor.
[0074] As Figure 6B As further shown, the goodness of fit of the selected Gaussian mixture model can also be evaluated, for example, as an estimate of the diagnostic characteristics of the Gaussian mixture model. For example, a silhouette score 610 can be determined for each Gaussian probability distribution 608, where the silhouette score indicates the silhouette coefficient 612 (e.g., the number of samples 600 within a selected distance from the mean or centroid of the Gaussian probability distribution 608 of the category 602). The discriminative characteristics of the Gaussian mixture model can be improved by selecting Gaussian probability distributions 608 with similar silhouette scores 610. As Figure 6B shown, the silhouette scores of the Gaussian probability distributions 608 are dissimilar. For example, because the first Gaussian probability distribution 608-1 of the first category 602-1 represents a larger number of samples 600-1 than the second Gaussian probability distribution 608-2 of the samples 600-2 of the second category 602-2 (i.e., the first Gaussian probability distribution 608-1 has a higher silhouette than the second Gaussian probability distribution 608-2), and also because the distances of the samples 600-1 of the first category 602-1 are more widely distributed in the feature space 606 than the samples 600-2 of the second category 602-2, resulting in a larger range of silhouette coefficients 612 (i.e., the silhouette of the first Gaussian probability distribution 608-1 is longer than the silhouette of the second Gaussian probability distribution 608-2). Therefore, the improvement can be achieved by selecting different mixtures of the Gaussian probability distribution 608 Figure 6BGaussian mixture model
[0075] Figure 6C is another illustration of a Gaussian mixture model configured to classify the set of samples into a set of probability distribution clusters within a two-dimensional feature space, and in some example embodiments, the Gaussian mixture model can be used to distinguish different classes of tumors. In Figure 6C , the Gaussian probability distribution of the first class 602 is alternatively identified as the first Gaussian probability distribution 608-4 of the first cluster of samples 600-1 of the first class 602-1 and the second Gaussian probability distribution 608-5 of the second cluster of samples 600-1 of the first class 602-1. Further, for each Gaussian probability distribution 608, a mixing parameter indicating the proportion of samples 600 represented by the Gaussian probability distribution 608 for class 602 can be identified. For example, the first Gaussian probability distribution 608-4 of the first class 602-1 can fit a smaller number of samples 600-1 of the first class 602-1 than the second Gaussian probability distribution 608-5 of the first class 602-1, and thus can have a first mixing parameter 614-1 smaller than the second mixing parameter 614-2 of the second Gaussian probability distribution 608-5. When a sample 600 of an unknown class 602 is located within the feature space 606, the probability that the sample 600 is classified into each class 602 can be determined as the sum of the products of the probability distribution of the position of the sample 600 for each Gaussian probability distribution 608 and the mixing parameter 614 of the Gaussian probability distribution 608. Further, the classification ability of the Gaussian mixture model can be evaluated based on the profile score 610 of the corresponding Gaussian probability distribution 608; for example, the similarity of the sample size and the silhouette coefficient 612 of each Gaussian probability distribution 608 can indicate a classifier that is more reliable and predictive than Figure 6B the Gaussian mixture model
[0076] Alternatively or in addition to Figure 6B and Figure 6CIn addition to the silhouette scores shown, other measurements can be used to determine the classification ability of the Gaussian mixture model. As an example, for tumors of different tumor classes (such as low-risk classes and high-risk classes) respectively associated with viability, a concordance index ("C-index") can be developed that indicates the degree of concordance between the predicted survival time of an individual with a tumor of a tumor class and the actual survival time of an individual with a tumor of a tumor class. The concordance index can be determined based on each feature of the tumor class to determine the extent to which the feature corresponds to the predicted survival rate of an individual with a tumor of a tumor class, where a high concordance index indicates a highly predictive feature of the Gaussian mixture model and a low concordance index indicates a poor predictive feature of the Gaussian mixture model. Since the concordance index for each feature depends on the selected Gaussian mixture model, it may be desirable to limit the number of features to those that exhibit a high concordance index, alternatively or additionally limited to the silhouette scores of the corresponding Gaussian probability distributions for each class. Selecting such features can reduce the dimensionality of the feature space 606 of the data set to a smaller feature set that is more discriminative for the corresponding class 602, which can result in a more precise, accurate, and / or efficient classification process.
[0077] Figure 7 is an illustration of a process 700 for selecting a subset of clinical features for a classifier from a set of clinical features of clinical features within a feature space based on the correlation of the corresponding clinical features with the corresponding class. A Gaussian mixture model is developed for a set of nine clinical features, such as Figure 5 the nine clinical features shown. A set of silhouette scores and concordance indices can be determined for each clinical feature. Among a set of available clinical features 704, in a first selection step 702-1, a first clinical feature 706-1 that provides the highest silhouette score and / or concordance index can be selected among the available clinical features 704, such as the concentration (specifically, percentage) of tumor adjacent to the lymphocyte region. Among the remaining clinical features (i.e., all clinical features other than the first selected clinical feature), in a second selection step 702-2, a second Gaussian mixture model can be developed, and a second clinical feature 706-2 that provides the highest silhouette score and / or concordance index can be selected among the remaining clinical features, such as the concentration (specifically, percentage) of stroma and tumor regions. Similar selection steps 702-3, 702-4 can be performed to select a third clinical feature 706-3 (such as the concentration of tumor invading the stroma region) and a fourth clinical feature 706-4 (such as the concentration of lymphocytes), each of these clinical features providing an improved concordance score compared to the previously selected clinical features, and the concordance score indicating a complementary classification ability of the selected clinical feature compared to the other remaining clinical features. The selection process can continue until a fifth selection step 702-5, in which the selected clinical features are determined NoneImprove the concordance index of previously selected clinical features and may not select further clinical features for a subset of clinical features.
[0078] Generally speaking, in some example embodiments, a classifier for tumors may include a Gaussian mixture model configured to determine, for a corresponding class, the probability distribution of the features of tumors of that class within a feature space, where the features may be selected from a feature set including: measurements of the tumor region of an image, measurements of the stromal region of the image, measurements of the lymphocyte region of the image, measurements of the tumor-infiltrating lymphocyte region of the image, measurements of the tumor-adjacent lymphocyte region of the image, measurements of the stromal-infiltrating lymphocyte region of the image, measurements of the stromal-adjacent lymphocyte region of the image, measurements of the tumor-infiltrating stromal region of the image, and measurements of the tumor-adjacent stromal region of the image. In some example embodiments, a feature subset of the Gaussian mixture model may be selected based on the correlation between the corresponding class and the corresponding features of the subset, where the correlation may be based on one or both of the contour score or the concordance index of the feature space. In some example embodiments, the feature subset may consist essentially of measurements of the lymphocyte region of the image, measurements of the tumor-infiltrating lymphocyte region of the image, measurements of the tumor-adjacent lymphocyte region of the image, and measurements of the tumor-infiltrating stromal region of the image.
[0079] D. Image-based Tumor Assessment
[0080] In some example embodiments, the clinical value (such as prognosis) of an individual based on a tumor shown in an image may be determined by: determining the lymphocyte distribution of lymphocytes in the tumor based on the image; applying a classifier to the lymphocyte distribution to classify the tumor, where the classifier has been trained to classify the tumor into a class selected from at least two classes respectively associated with the lymphocyte distribution; and determining the clinical value (such as prognosis) of the individual based on the prognosis of individuals having tumors of the class into which the classifier classifies the tumor. The classifier may be invoked to determine the lymphocyte distribution of lymphocytes in a corresponding region of the tumor image.
[0081] Figure 8 Is an illustration of the classification of different classes of tumors based on a feature subset according to some example embodiments. Figure 8Presents a comparison 800 of a subset of features 806 of selected features with images 804 of tumors in a low-risk tumor class 802-1 and a high-risk tumor class 802-2, i.e., the percentage of the area in each image 804 of each class 802 corresponding to each feature 806 in the subset of features. The high-risk tumor class may be associated with a first survival probability, and the low-risk tumor class may be associated with a second survival probability longer than the first survival probability. For example, the percentages of the corresponding features 806 of the images 804 of the tumor class 802 may be compared to determine the degree of diagnosis of the feature 806 for the corresponding tumor class 802. For example, compared with the image 804 of the low-risk tumor class 802-1, the image 804 of the high-risk tumor class 802-2 may show a smaller and more consistent value range of the first feature 806-1 and the third feature 806-2. Moreover, the value of the third feature 806-3 may generally be higher in the images of tumors in the low-risk tumor class 802-1 than in the images of tumors in the high-risk tumor class 802-2. These measurements determined based on Figure 7 the selection process 700 based on the selected features 706 of the subset of features of the tumor class 802 may present clinically significant findings in the pathology of tumors in different tumor classes 802 and may be used by clinicians and automated processes (such as diagnostic and / or prognostic machine learning processes) to classify tumors into different tumor classes 802.
[0082] Figure 9 is an illustration of a Kaplan Meier survival plot based on image analysis according to some example embodiments. In Figure 9 , for a training set and a test set of an individual population of tumors with Figure 8 a low-risk tumor class 802-1 and a high-risk tumor class 802-2, a first Kaplan Meier survival plot 900-1 and a second Kaplan Meier survival 900-2 (e.g., the percentage of the surviving individual population measured in days after diagnosis) are generated respectively. Further, the tumor set with available images and data is divided into a training set and a test set. A machine learning model including a convolutional neural network and / or a Gaussian mixture model is trained on the images of the training set to a convergence point where the machine learning model produces an output within the accuracy range of the expected output. Then the test set is used to test the machine learning model to determine whether the machine learning model produces an output of new data consistent with the expected output. Such verification may include a cross-validation process in which the tumor set is first divided into multiple subsets and repeated training and testing are performed using the selection of subsets from the training set and the remaining subsets of the test set.
[0083] As Figure 9As shown, the image-based tumor assessment techniques presented herein perform classification on a training dataset with a hazard ratio (HR) of 0.5117, a statistical P-value of 0.0570, and a concordance index of 0.6667, and demonstrate performance on a test set with a hazard ratio of 0.5154, a statistical P-value of 0.3405, and a concordance index of 0.5964. According to some example embodiments, many such machine learning models can be trained to classify tumors and determine clinical values (such as prognosis and / or viability) of an individual.
[0084] E. COX proportional hazards model
[0085] In some example embodiments, the image-based prognosis determination techniques can be combined with a Cox proportional hazards model, which can improve the prognostic ability of tumor analysis. The Cox proportional hazards model is a regression model that associates clinical features such as an individual's demographic characteristics, clinical observations of the individual and the tumor, and pathological measurement results with different tumor classes to determine the contribution of each clinical feature to tumor classification. For example, the regression model can determine that an individual within a specific age range, with specific personal habits (such as smoking or drinking), and with a cancer stage score based on the American Joint Committee on Cancer (AJCC) cancer staging system is more likely to be classified as having a tumor within a low-risk tumor class, while an individual within another age range, with other personal habits, and with other cancer stage scores is more likely to be classified as having a tumor within a high-risk tumor class.
[0086] A training set characterized by tumors with known clinical features can be used to develop the Cox proportional hazards model. Stepwise selection can be performed to select a subset of clinical features that significantly contribute to classification, for example, by removing clinical features that do not significantly improve the predictability of other clinical features. The Cox proportional hazards model can also be trained on tumors of two or more classes to determine the different proportional viabilities of tumors in different tumor classes (such as a low-risk tumor class with a set of shared characteristics and / or similar viability metrics and a high-risk tumor class with another set of shared characteristics and / or other similar viability metrics).
[0087] Figure 10 is an illustration of selecting a feature subset of the Cox proportional hazards model from a feature set 1000 of features within a feature space based on the correlation of corresponding features with corresponding classes. In Figure 10In , for the corresponding tumors in a set of tumors taken from an individual and pathologically evaluated, the values of the clinical feature set 1000 are identified, which includes: the initial diagnosis of the tumor (e.g., the T - category AJCC stage score); the measurement results of the tumor (e.g., the N - category AJCC stage score); the treatment of the tumor; the location of the tumor; the frequency of the individual's smoking habit; the metastasis status of the tumor; the individual's prior cancer history; the duration of the individual's smoking habit; the initial diagnosis of the individual; the individual's drinking history; and the gender of the individual. The first step 1002 - 1 of the regression analysis can determine the degree to which each feature of the feature set 1000 differentiates between tumor categories (e.g., low - risk and high - risk), and can rank the features, for example, by a statistical P - value. Features with P - values within a certain range (e.g., a statistical significance threshold below 0.05) can be selected as a feature subset, and other features can be excluded. Additional steps 1002 - 2, 1002 - 3, 1004 of the regression analysis can be performed to exclude other features and retain other features in the feature set 1000 until no more features can be excluded without significantly reducing the classification accuracy of the Cox proportional hazards model. The feature set 1004 obtained based on the correlation between the corresponding categories and the corresponding features of the subset can be identified as the retained features of the Cox proportional hazards model.
[0088] As Figure 10 shown, the Cox proportional hazards model developed in this way identifies a feature subset consisting of the measurement results of the tumor and the metastasis status of the tumor. For the training set, the Cox proportional hazards model shows a hazard ratio of 0.2182 and a statistical P - value of 0.0200, and for the test set, the Cox proportional hazards model shows a hazard ratio of 0.4065 and a statistical P - value of 0.2855. According to some example embodiments, many such Cox proportional hazards models can be determined to classify tumors.
[0089] F. Combined Model
[0090] In some example embodiments, image - based classification (e.g., based on convolutional neural networks and Gaussian mixture models) can be combined with the Cox proportional hazards model to classify tumors based on image features and clinical features. That is, the at least two categories are the low - risk tumor category and the high - risk tumor; determining the lymphocyte distribution can further include applying a convolutional neural network to the image, the convolutional neural network being configured to measure the lymphocyte distribution of lymphocytes in different region types of the image; the classifier can be a bidirectional Gaussian mixture model, the bidirectional Gaussian mixture model being configured to determine the probability distribution of the features of the tumors of the corresponding category within the feature space for the corresponding category; the Cox proportional hazards model can be applied to the clinical features of the tumor to determine the category of the tumor; and determining the clinical value of the individual (such as prognosis) can be further based on the category determined by the Cox proportional hazards model.
[0091] Figure 11 is an illustration of a Kaplan Meier survival plot based on image analysis and Cox proportional hazards model according to some example embodiments. In Figure 11 , for a training set and a test set of an individual population of tumors for low-risk tumor class 802-1 and high-risk tumor class 802-2 having Figure 8 , a first Kaplan Meier survival plot 900-3 and a second Kaplan Meier survival 900-4 (e.g., percentage of the surviving individual population measured according to the number of days after diagnosis) are respectively generated. Further, a tumor set with available images and data including clinical features is divided into a training set and a test set. A machine learning model including a convolutional neural network, a Gaussian mixture model, and a Cox proportional hazards model is trained on the images of the training set to a convergence point, in which the machine learning model produces an output within the accuracy range of the expected output. Then the test set is used to test the machine learning model to determine whether the machine learning model produces an output of new data consistent with the expected output. Such verification may include a cross-validation process, in which the tumor set is first divided into multiple subsets, and repeated training and testing are performed using a selection of subsets from the training set and the remaining subsets of the test set.
[0092] As Figure 11 shown, the image-based tumor assessment technique presented herein performs classification on a training data set with a hazard ratio (HR) of 0.2545, a statistical P-value of 0.0065, and a concordance index of 0.7141, and demonstrates performance on a test set with a hazard ratio of 0.3742, a statistical P-value of 0.0696, and a concordance index of 0.6120.
[0093] Figure 12 is an illustration 1200 of a result set of classification of a tumor training data set and a tumor test data set based on image analysis and Cox proportional hazards model according to some example embodiments. As Figure 12 shown, the classification results of a combined model characterized by image-based analysis and statistical analysis of clinical features show a higher classification accuracy than using either model alone. According to some example embodiments, many such machine learning models can be trained to classify tumors and determine the clinical values (such as prognosis and / or survival ability) of individuals.
[0094] G. Tumor Assessment and Output
[0095] In some example embodiments, the tumor analysis model disclosed herein can be used to determine and output to a user clinical values of an individual based on a tumor shown in an image. The user can be, for example, an individual with a tumor; a family member or guardian of the individual; or a healthcare provider, including a doctor, a nurse, or a clinical pathologist. The clinical values and / or outputs can be, for example, one or more of the following: a diagnosis of the individual, a prognosis of the individual, the viability of the individual, a classification of the tumor, a diagnosis and / or treatment recommendation for the individual, and the like.
[0096] Some example embodiments can use the determination of the tumor analysis model to display a visualization of a clinical value (such as a prognosis) of an individual. For example, a terminal can receive a tumor image of an individual and, optionally, a set of clinical features, such as the demographic features of the individual, the clinical observations of the individual and the tumor, and the pathological measurement results. The terminal can apply the tumor analysis model (e.g., process the image through a convolutional neural network and a Gaussian mixture model, and optionally, process the clinical features through a Cox proportional hazards model) to determine the category of the tumor (such as a low-risk tumor category and a high-risk tumor category) and the prognosis associated with an individual having a tumor of that tumor category. For example, the clinical value can be determined as viability (such as an expected survival duration and probability), optionally including the confidence or accuracy of each probability. In some example embodiments, the clinical value can be presented as a visualization, such as a Kaplan Meier viability projection of the tumor. In some example embodiments, the visualization can include additional information about the tumor, such as one or more in a mask 302 indicating the region type of a region of the image 102; the measurement results 304 of the image 102, such as the concentration of each region type (e.g., the concentration of lymphocytes in one or more regions determined by binning), and / or the percentage of the region of that region type compared to the entire image 102. In some example embodiments, the visualization can include additional information about the individual, such as the clinical features of the individual, and can indicate how the corresponding clinical features contribute to determining the clinical value (such as a prognosis) of the individual.
[0097] Some example embodiments can use the determination of the tumor analysis model to determine and display to a user a diagnostic test for a tumor based on the clinical value (such as a prognosis) of the individual. For example, based on the tumor analysis model classifying the tumor as a low-risk category, the device may recommend a less invasive test (such as a blood test or imaging) to further characterize the tumor. Based on the tumor analysis model classifying the tumor as a high-risk category, the device may recommend a more invasive test (such as a biopsy) to further characterize the tumor. Some example embodiments can also display to the user an explanation of the basis for the determination; a set of options for further testing; and / or a recommendation of one or more options to be considered by the individual and / or the healthcare provider.
[0098] Some example embodiments may use the determination of a tumor analysis model to determine and display to a user a treatment for an individual based on the individual's clinical values (such as prognosis). For example, based on classifying a tumor into a low-risk category using a tumor analysis model, the device may recommend a less invasive treatment for the tumor, such as less invasive chemotherapy. Based on classifying a tumor into a high-risk category using a tumor analysis model, the device may recommend a more invasive treatment for the tumor, such as more invasive chemotherapy and / or surgical removal. Some example embodiments may also display to the user an explanation of the basis for the determination; a set of options for further testing; and / or recommendations for one or more options to be considered by the individual and / or healthcare provider.
[0099] Some example embodiments may use the determination of a tumor analysis model to determine and display to a user a schedule of therapeutic agents for treating a tumor based on the individual's clinical values (such as prognosis). For example, based on classifying a tumor into a low-risk category using a tumor analysis model, the device may recommend a lower frequency, later date, and / or lower dose of chemotherapy. Based on classifying a tumor into a high-risk category using a tumor analysis model, the device may recommend a more invasive treatment for the tumor, such as a higher frequency, earlier date, and / or higher dose of chemotherapy. Some example embodiments may also display to the user an explanation of the basis for the determination; a set of options for further testing; and / or recommendations for one or more options to be considered by the individual and / or healthcare provider. In some example embodiments, many such classifications and outputs of the individual's clinical values, as well as information about the tumor, may be provided.
[0100] H. Technical Effects
[0101] Some example embodiments using feature analysis with a distribution-based machine learning classifier may exhibit a variety of technical effects.
[0102] A first example of a technical effect that some example embodiments may exhibit is a novel input classification based on distribution, which may be difficult to achieve with other machine learning models. For example, as Figure 9 shown, the image-based tumor classification model disclosed herein is capable of classifying tumors with reasonable accuracy. As Figure 11 and Figure 12Further shown, a combined model including image-based analysis (e.g., classification based on convolutional neural networks and Gaussian mixture models) and regression-based clinical feature analysis (e.g., based on Cox proportional hazards models) may have higher classification accuracy than using either model alone. In some scenarios, using a machine learning model, including visualization and / or interpretation of the basis for such determination of an individual's clinical values (such as indications of image features and clinical features that contribute to prognosis determination) can provide an automated process for providing clinical values that provide diagnostic, prognostic, and / or treatment information, and caregivers can utilize these clinical values to select a healthcare service plan for an individual.
[0103] A second example of a technical effect that can be demonstrated by some example embodiments is more efficient resource allocation based on such analysis. For example, tumor classification based on automated techniques can reduce the amount and / or dependence of clinical and pathological resources applied to diagnose and classify tumors and determine an individual's clinical values (such as prognosis). This resource economy may also involve a faster classification process than that systematically achieved by classification processes performed by individuals.
[0104] I. Example Embodiments
[0105] Figure 13 is a flowchart of a first example method 1300 according to some example embodiments.
[0106] The first example method 1300 can be implemented as, for example, an instruction set that, when executed by a processing circuitry of a device, causes the device to perform each element of the first example method 1300. The first example method 1300 can also be implemented as, for example, an instruction set that, when executed by a processing circuitry of a device, causes the device to provide a system for components including an image evaluator, a classifier, and a tumor evaluator that interoperate to provide a system for classifying tumors.
[0107] The first example method 1300 includes instructions that, when executed by a processing circuitry of a device at 1304, cause the device to perform a set of elements.
[0108] For example, execution of the instructions can cause the device to determine 1306 the lymphocyte distribution of lymphocytes in a tumor based on an image.
[0109] For example, execution of the instructions can cause the device to apply 1308 a classifier to the lymphocyte distribution to classify the tumor, where the classifier has been trained to classify tumors into categories selected from at least two categories respectively associated with the lymphocyte distribution.
[0110] For example, execution of the instructions can cause the device to determine 1310 an individual's clinical value (such as prognosis) based on the prognosis of an individual with a tumor classified into the category by the classifier.
[0111] In this manner, the processing circuitry's execution of the instructions can cause the apparatus to perform the elements of the first example method 1300, and thus the first example method 1300 ends.
[0112] Figure 14 is a flowchart of a second example method according to some example embodiments.
[0113] The second example method 1400 can be implemented as, for example, an instruction set that, when executed by the processing circuitry of the apparatus, causes the apparatus to perform each element of the second example method 1400. The second example method 1400 can also be implemented as, for example, an instruction set that, when executed by the processing circuitry of the apparatus, causes the apparatus to provide a system for a component including an image evaluator, a classifier, and a tumor evaluator, which interoperate to provide a system for classifying a tumor.
[0114] The second example method 1400 includes instructions that, when executed 1404 by the processing circuitry of the apparatus, cause the apparatus to perform a set of elements.
[0115] For example, the execution of the instructions can cause the apparatus to apply a convolutional neural network to 1406 an image to determine the lymphocyte distribution of lymphocytes in a tumor, where the convolutional neural network is configured to measure the lymphocyte distribution of lymphocytes of different region types of the image.
[0116] For example, the execution of the instructions can cause the apparatus to apply a classifier to 1408 the lymphocyte distribution to classify the tumor, where the classifier has been trained to classify the tumor into a category selected from a low-risk category and a high-risk category, which are respectively associated with the lymphocyte distribution, and the classifier includes a two-way Gaussian mixture model that is configured to determine the probability distribution of the characteristics of the tumors of the corresponding category within a feature space.
[0117] For example, the execution of the instructions can cause the apparatus to apply a Cox proportional hazards model to 1410 the clinical characteristics of the tumor to determine the category of the tumor.
[0118] For example, the execution of the instructions can cause the apparatus to determine 1412 the clinical value (such as prognosis) of an individual based on the prognosis of an individual with a tumor classified by the classifier and the category determined by the Cox proportional hazards model.
[0119] In this manner, the processing circuitry's execution of the instructions can cause the apparatus to perform the elements of the second example method 1400, and thus the second example method 1400 ends.
[0120] Figure 15 is a component block diagram of an example apparatus according to some example embodiments.
[0121] As Figure 15 shown, example device 1500 may include processing circuitry 1502 and memory 1504. According to some example embodiments, memory 1504 may store instructions 1506 that, when executed by processing circuitry 1502, cause example device 1500 to determine a clinical value (such as a prognosis) of an individual based on a tumor shown in image 102. In some example embodiments, execution of instructions 1506 may cause example device 1500 to instantiate and / or use a set of components of system 1508. Although Figure 15 one such system 1508 is shown, some example embodiments may System 1508 may embody any method disclosed herein.
[0122] Figure 15 Example system 1508 of includes an image evaluator 1510 configured to determine the lymphocyte distribution of lymphocytes in image 102. For example, class set 1516 may associate respective lymphocyte distributions 1520-1, 1520-2 with different classes 1518-1, 1518-2 of tumors, each class 1518 being associated with a prognosis 1522-1, 1522-2.
[0123] Figure 15 Example system 1508 of includes a tumor classifier 1512 configured to classify a tumor into a class selected from at least two classes 1518 respectively associated with lymphocyte distribution 1520.
[0124] Figure 15 Example system 1508 of includes a tumor evaluator 1514 configured to determine a clinical value (such as a prognosis) of an individual based on a tumor in image 102 by: invoking image evaluator 1510 with image 102 to determine the lymphocyte distribution 1520-3 of lymphocytes in the tumor; invoking tumor classifier 1512 to classify the tumor into a class 1518 based on lymphocyte distribution 1520-3, and outputting, for user 1524, the clinical value (such as a prognosis) of the individual based on the prognosis 1522 of the tumor in the class 1518-3 into which the tumor classifier 1512 classifies the tumor.
[0125] In this way, example device 1500 and example system 1508 provided thereon may classify tumors according to some example embodiments.
[0126] Figure 16 is a component block diagram of another example device according to some example embodiments.
[0127] As Figure 16As shown, example apparatus 1600 may include processing circuitry 1502 and memory 1504. According to some example embodiments, memory 1504 may store instructions 1506 that, when executed by processing circuitry 1502, cause example apparatus 1600 to determine a clinical value (such as a prognosis) of an individual based on a tumor shown in image 102. In some example embodiments, execution of instructions 1506 may cause example apparatus 1600 to instantiate and / or use a set of components of system 1602.
[0128] Figure 16 Example system 1602 includes convolutional neural network 110 as an image evaluator, which is configured to determine a lymphocyte distribution of lymphocytes in image 102 by measuring lymphocyte distributions of lymphocytes of different region types of image 102. For example, class set 1516 may associate corresponding lymphocyte distributions 1520-1, 1520-2 with different classes 1518-1, 1518-2 of tumors (including a low-risk tumor class and a high-risk tumor class), and each class 1518 is associated with a prognosis 1522-1, 1522-2.
[0129] Figure 16 Example system 1602 includes bidirectional Gaussian mixture model 1604 as a tumor classifier, which is configured to determine a probability distribution of features of tumors in class 1518 within feature space 606 for corresponding class 1518.
[0130] Figure 16 Example system 1602 includes Cox proportional hazards model 1608, which is configured to determine a set of clinical features 1606 of clinical features of a class of tumors.
[0131] Figure 16 Example system 1602 includes tumor evaluator 1514, which is configured to determine a clinical value (such as a prognosis) of an individual based on a tumor in tumor image 102 by: invoking convolutional neural network 110 with image 102 to determine a lymphocyte distribution 1520-3 of lymphocytes in the tumor; invoking Gaussian mixture model 1604 to classify the tumor into class 1518-3 based on lymphocyte distribution 1520-3; invoking Cox proportional hazards model 1608 with set of clinical features 1606 of the tumor to determine tumor class 1518-4, and outputting a clinical value (such as a prognosis) of the individual for user 1524 based on a prognosis 1522 of a tumor of class 1518-5 into which the tumor is classified by tumor classifier 1512 based on the class 1518-3 into which the tumor is classified by Gaussian mixture model 1604 and the tumor class 1518-4 determined by Cox proportional hazards model 1608.
[0132] In this manner, the example apparatus 1600 and the example system 1602 provided thereon can classify tumors according to some example embodiments.
[0133] As Figure 15 and Figure 16 shown, the example apparatuses 1500, 1600 can include processing circuitry 1502 capable of executing instructions. The processing circuitry 1502 can include: hardware such as logic circuits; a hardware / software combination such as a processor executing software; or a combination thereof. For example, the processor can include, but is not limited to, a central processing unit (CPU), a graphics processing unit (GPU), an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a system on a chip (SoC), a programmable logic unit, a microprocessor, an application specific integrated circuit (ASIC), etc.
[0134] As Figure 15 and Figure 16 Further shown, the example apparatuses 1500, 1600 can include a memory 1504 storing instructions 1506. The memory 1504 can include, for example, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), etc. The memory 1504 can be volatile, such as system memory, and / or non-volatile, such as a hard disk drive, a solid state storage device, flash memory, or magnetic tape. The instructions 1506 stored in the memory 1504 can be specified according to the native instruction set architecture of a processor such as a variant of the IA-32 instruction set architecture or a variant of the ARM instruction set architecture as: assembly and / or machine language (e.g., binary) instructions; high-level imperative and / or declarative language instructions that can be compiled and / or interpreted to execute on the processor; and / or instructions that can be compiled and / or interpreted to be executed by a virtual processor of a virtual machine such as a web browser. A non-limiting set of examples of such high-level languages can include, for example: C, C++, C#, Objective-C, Swift, Haskell, Go, SQL, R, Lisp, Fortran, Perl, Pascal, Curl, OCaml, HTML5 (HyperText Markup Language 5th Revision), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Swift, Eiffel, Smalltalk, Erlang, Ruby, Visual Lua, MATLAB, SIMULINK and Such instruction 1506 may also include instructions for libraries, resources, platforms, application programming interfaces (APIs), etc., which are used to determine the clinical value (such as prognosis) of the tumor shown in the image for an individual.
[0135] Such as Figure 15 and Figure 16 As shown, example systems 1508, 1602 may be organized in a particular manner, for example, to allocate some functions to each component of the system. Some example embodiments may implement each such component in various ways, such as software, hardware (e.g., processing circuitry), or a combination thereof. In some example embodiments, the organization of the system may be different compared to some other example embodiments of the example systems 1508, 1602 including Figure 15 and Figure 16 shown. For example, without departing from the scope of the present disclosure, some example embodiments may include systems characterized by different organizations of components, such as renaming, rearranging, adding, dividing, duplicating, combining, and / or removing components, multiple sets of components, and the relationships therebetween. All such variations that are technically and logically feasible and do not contradict other statements are intended to be included in the present disclosure, and the scope of the present disclosure should be understood to be limited only by the claims.
[0136] Figure 17 is an illustration of an example computer-readable medium 1700 according to some example embodiments.
[0137] Such as Figure 17 shown, according to some example embodiments, the non-transitory computer-readable medium 1700 may store binary data 1702 encoding an instruction set 1704, which, when executed by the processing circuitry 1502 of the example devices 1500, 1600, causes the example devices 1500, 1600 to determine the clinical value (such as prognosis) of the tumor shown in the image for an individual. As a first such example, the instructions 1704 may encode the elements of an example method 1706 (such as Figure 13 the first example method 1300). As a second such example, the instructions 1704 may encode the elements of Figure 14 the second example method 1400. As a third such example, the instructions 1704 may encode the components of Figure 15 the first example system 1508. As a fourth such example, the instructions 1704 may encode the components of Figure 16 the second example system 1602.
[0138] In some example embodiments, the system may include an image evaluation device configured to determine the lymphocyte distribution of lymphocytes in an image. For example, the image evaluation device may be or may include one or more convolutional neural networks and / or any other image evaluation model discussed herein. The system may include a classification device configured to classify a tumor into a category selected from at least two categories respectively associated with lymphocyte distributions. For example, the classification device may be or may include one or more Gaussian mixture models and / or any other classifier discussed herein.
[0139] The system may include a tumor evaluator device configured to determine a clinical value (such as a prognosis) of an individual based on a tumor in an image by: invoking the image evaluation device with the image to determine the lymphocyte distribution of lymphocytes in the tumor; invoking a classifier to classify the tumor into a certain category based on the lymphocyte distribution, and outputting a clinical value (such as a prognosis) of the individual based on the prognosis of individuals having tumors classified into the category by the classifier. For example, the tumor evaluator device may be or may include a classifier such as a neural network, a soft margin or hard margin support vector machine, and / or any other classifier discussed herein. For example, the tumor evaluator device may be or may include a display device such as a liquid crystal display (LCD), a light emitting diode (LED), or an organic light emitting diode (OLED) display; a communication interface such as a web server, an email server, or a text message server; and / or any other output device disclosed herein.
[0140] J. Variations
[0141] Some example embodiments of the present disclosure may include variations in many aspects, and some variations may present additional advantages and / or reduce disadvantages relative to other variations of these and other technologies. Additionally, some variations may be implemented in combination, and some combinations may have additional advantages and / or reduced disadvantages through synergy. These variations may be incorporated into some example embodiments (e.g., Figure 13 the first example method of Figure 14 the second example method of Figure 15 and Figure 16 example devices 1500, 1600 and example systems 1508, 1602, and / or Figure 17 example non-transitory computer-readable medium 1700 of
[0142] F1. Scenarios
[0143] Some example embodiments can be used in various scenarios involving input analysis using distribution-based machine learning models. For example, some example embodiments can use the disclosed techniques to classify tumors for various fields in life sciences, including healthcare services and biomedical research. The tumor classification techniques disclosed herein can be applicable to a variety of cancer types, including (but not limited to) lung cancer tumors, pancreatic cancer tumors, and / or breast cancer tumors. Clinical pathology laboratories can use such techniques to determine the tumor class of a tumor sample, and / or compare or validate the determination of the tumor class by an individual and / or other automated processes. Researchers can use such techniques to determine the tumor class of tumors in images of a research dataset, which can be from human patients or from human or non-human subjects, where such research can involve additional techniques for classifying tumors, identifying the prevalence and class of tumors in different demographics, identifying risk factors associated with tumors of different classes, predicting survival ability, and determining or comparing the effectiveness of treatment options. Clinicians can use the results of the classification to evaluate the diagnosis, prognosis, and / or treatment options of an individual with a tumor, and / or explore and understand the correlation between various risk factors and different tumor classes and the prognosis of an individual with such a tumor. Many such scenarios can be designed that can utilize the disclosed techniques.
[0144] F2. Determine the presence and distribution of features
[0145] In some example embodiments, a machine learning model, including a deep learning model, can be used to detect various features of various inputs. In various example embodiments, such a machine learning model can be used to: determine a feature map 118 of an image 102, for example, by generating and applying a set of masks 300 of masks 302; determine the distribution of features, such as clusters 204; perform classification 402 of image regions, such as region types based on anatomical features and / or tissue types; perform measurements 304 of features, such as the concentration of a feature or region type (e.g., percentage area of the entire image), using techniques such as binning, for example, lymphocyte density range estimation 404; generate a density map, such as a lymphocyte density map 406; select a feature set for classifying the image 102 or a particular image 102, such as performing Figure 7The selection process 700 selects a feature subset; classifies the image 102 of the tumor based on the image feature set or the image feature subset; selects a clinical feature subset from the values of the clinical features 706 in the clinical feature set of the individual or the tumor; determines the values of the clinical features 706 of the clinical feature set of the individual and / or the tumor; determines the class of the tumor, such as by preparing a Cox proportional hazards model and applying it to the clinical features of the tumor; determines the class of the tumor based on the image features of the tumor (such as the output of a Gaussian mixture model) and / or the clinical features of the tumor or the individual (such as the output of a Cox proportional hazards model); predicts the survival ability of the individual based on the tumor classification of the individual; and / or generates one or more outputs of such determination, including visualization. Each of these features and other features of some example embodiments can be performed, for example, by: a machine learning model; multiple machine learning models of the same or similar type, such as a random forest, or a convolutional neural network that evaluates different parts of the image or performs different tasks on the image; and / or a combination of different types of machine learning models. As one such example, in an enhanced ensemble, a first machine learning model performs classification based on the outputs of other machine learning models.
[0146] As a first such example, the presence of a feature (e.g., an activation within a feature map, and / or a biological activation of lymphocytes) can be determined in various ways. For example, in the case where the input further includes an image 102, the presence of the feature can be determined by applying at least one convolutional neural network 110 to the image 102 and receiving from the at least one convolutional neural network 110 a feature map 118 indicative of the presence of the feature in the input. That is, the convolutional neural network 110 can be applied to the image 102 to identify clusters of pixels 108 that are feature prominent. For example, a cell counting convolutional neural network 110 can be applied to count cells in a tissue sample, where such cells can be lymphocytes. In such a scenario, the tissue sample can be assayed, such as with a dye or a luminescent (e.g., fluorescent) agent, and a set of images 102 of the tissue sample can be selective for the cells and thus may not include other visible components of the tissue sample. The images 102 of the tissue sample can then be subjected to a machine learning model, such as the convolutional neural network 110, which can be configured (e.g., trained) to detect shapes, such as the circular shape indicative of the selected cells, and can output a count of the image 102 and / or different regions of the image 102. It is noted that in such a case, the convolutional neural network 110 may not be configured to and / or for further detecting the arrangement of such features, e.g., the number, orientation, and / or positioning of lymphocytes relative to other lymphocytes; rather, the counts of the corresponding portions of the image 102 can be compiled into a distribution map, which can be processed together with a tumor mask to determine whether the distribution of lymphocytes in the tissue sample is tumor-invasive, tumor-adjacent, or elsewhere. Some example embodiments can use machine learning models other than convolutional neural networks to detect the presence of features, such as (e.g.) non-convolutional neural networks, such as fully connected networks or perceptron networks or Bayesian classifiers.
[0147] As a second such example, the distribution of the features can be determined in various ways. As a first example, in the case where the input further includes an image 102 showing an organizational region of an individual, determining the distribution of the features can further include determining the region type (e.g., tumor or non-tumor) of each region of the image, and determining the distribution based on the region type of each region of the image. Then the distribution of the detected lymphocytes, including lymphocyte counts, can be determined based on the tissue type in which such counts occur. That is, the distribution can be determined by tabulating the counts of lymphocytes in the tumor regions, tumor adjacent regions (such as stroma), and non-tumor regions of the image 102. As another such example, the determination can include determining the boundaries of the tissue regions within the image 102, and determining the distribution based on the boundaries of the tissue regions within the image 102. That is, the boundaries of the regions of the image 102 classified as tumors can be determined (e.g., by the convolutional neural network 110 and / or a person), and an example embodiment can tabulate the counts of lymphocytes in all regions within the determined tumor boundaries of the image. As yet another example, the tissue regions of the image include at least two regions, and determining the distribution can include determining the counts of lymphocytes within each region of the tissue regions and determining the distribution based on the counts within each region of the tissue regions. For example, determining the counts within each tissue region can include determining the density of the counts of lymphocytes within each region, and then determining the distribution based on the counts within each region. Figure 3 An example is shown where a first convolutional neural network 110-1 is provided to classify regions as tumor regions and non-tumor regions and a second convolutional neural network 110-2 is provided to estimate the density range of lymphocytes.
[0148] As a third such example, processing of the image 102 to determine the presence and / or distribution of features can occur in several ways. For example, some example embodiments can be configured to divide the image 102 into a set of regions of the same or different sizes and / or shapes, such as based on the number of pixels or the corresponding physical size (e.g., a region of 100 square microns), and / or grouped based on similarity (e.g., identifying regions of similar appearance within the image 102). Some example embodiments can be configured to classify each region (e.g., as tumor, tumor adjacent, or non-tumor), and / or determine the distribution by tabulating the presence of features (e.g., counting) within each region of a certain region type to determine the distribution of lymphocytes. Alternatively, a counting process can be applied to each region, and each region can be classified based on the count (e.g., high lymphocyte region versus low lymphocyte region). As yet another example, the distribution can be determined parametrically, such as according to a selected distribution type or kernel to which a machine learning model can be fit to the distribution of features in the input (e.g., a Gaussian mixture model can be applied to determine the Gaussian distribution of a subset of features). Other distribution models can be applied, including parametric distribution models such as chi-square fitting, Poisson distribution, and beta distribution, and non-parametric distribution models such as histograms, binning, and kernel methods.
[0149] As a fourth such example, many forms of classifiers can be used, such as Bayesian (including Naive Bayes) classifiers; Gaussian classifiers; probabilistic classifiers; principal component analysis (PCA) classifiers; linear discriminant analysis (LDA) classifiers; quadratic discriminant analysis (QDA) classifiers; single-layer or multi-layer perceptron networks; convolutional neural networks; recurrent neural networks; nearest neighbor classifiers; linear SVM classifiers; radial basis function kernel (RBF) SVM classifiers; Gaussian process classifiers; decision tree classifiers, including random forest classifiers; and / or restricted or unrestricted Boltzmann machines, etc. Examples of convolutional neural network classifiers include, but are not limited to, LeNet, ZfNet, AlexNet, BN-Inception, CaffeResNet-101, DenseNet-121, DenseNet-169, DenseNet-201, DenseNet-161, DPN-68, DPN-98, DPN-131, FBResNet-152, GoogLeNet, Inception-ResNet-v2, Inception-v3, Inception-v4, MobileNet-v1, MobileNet-v2, NASNet-A-Large, NASNet-A-Mobile, ResNet-101, ResNet-152, ResNet-18, ResNet-34, ResNet-50, ResNext-101, SE-ResNet-101, SE-ResNet-152, SE-ResNet-50, SE-ResNeXt-101, SE-ResNeXt-50, SENet-154, ShuffleNet, SqueezeNet-v1.0, SqueezeNet-v1.1, VGG-11, VGG-11_BN, VGG-13, VGG-13_BN, VGG-16, VGG-16_BN, VGG-19, VGG-19_BN, Xception, DelugeNet, FractalNet, WideResNet, PolyNet, PyramidalNet, and U-net.
[0150] In some example embodiments, classification can include regression, and the term "classification" as used herein is intended to include some example embodiments that perform regression as an alternative or supplement to class selection. For example, some example embodiments can be characterized by regression as an alternative or supplement to classification. As a first such example, determination of the presence of a feature can include regression of the presence of the feature, e.g., a numerical value indicating the density of the feature in the input. As a second such example, determination of the distribution of a feature can include regression of the distribution of the feature, such as the variance of a regression-based density determined for the input. As a third such example, selection can include performing regression on the distribution of a feature and selecting a regression value for the distribution of the feature. Such regression aspects can be performed instead of or as a supplement to classification (e.g., determining the presence and density of a feature in an image region). Some example embodiments can relate to regression-based machine learning models, such as Bayesian linear or nonlinear regression, regression-based artificial neural networks, such as convolutional neural network regression, support vector regression, and / or decision tree regression.
[0151] Each classifier can be linear or nonlinear; for example, a nonlinear classifier can be provided (e.g., trained) to perform linear classification based on a kernel transformation in a nonlinear space (i.e., transforming linear values of a feature vector into nonlinear features). Classifiers can include various techniques to facilitate accurate generalization and classification (such as input normalization, weight regularization) and / or output processing (such as softmax activation output). Classifiers can use various techniques to facilitate efficient training and / or classification. For example, a two-way Gaussian mixture model can be used, where Gaussian distributions of the same size are selected for each dimension of the feature space, and this Gaussian distribution of the same size can reduce the search space compared to other Gaussian mixture models where the distribution sizes of different dimensions of the feature space can vary.
[0152] Each classifier can be trained to perform classification in a specific manner, such as supervised learning, unsupervised learning, and / or reinforcement learning. Some example embodiments can include additional training techniques to facilitate generalization, accuracy, and / or convergence, such as validation, training data augmentation, and / or dropout regularization. An ensemble of such classifiers can also be utilized, where such an ensemble can be homogeneous or heterogeneous, and where classification 122 based on the outputs of the classifiers can be produced in various ways, such as by consensus, based on the confidence of each output (e.g., as a weighted combination), and / or via a stacked architecture, such as based on one or more mixers. The ensemble can be trained independently (e.g., bootstrap aggregating training models, or random forest training models) and / or sequentially (e.g., boosting training models, such as Adaboost). As an example of a boosting training model, in some support vector machine ensembles, at least some of the support vector machines can be trained based on the errors of previously trained support vector machines; for example, each successive support vector machine can be specifically trained on the inputs of the training data set 100 that were misclassified by the previously trained support vector machine.
[0153] As a fifth such example, the one or more classifiers can produce various forms of classification 122. For example, a perceptron or binary classifier can output a value indicating whether the input is classified as a first class 106 or a second class 106, such as whether a region of the image 102 is a tumor region or a non-tumor region. As another example, a probability classifier can be configured to output the probability that the input is classified as each class 106 in the class set 104. Alternatively or additionally, some example embodiments can be configured to determine the probability that the input is classified as each class 106 in the class set 104, and select the class 106 of the input from the class set 104 that includes at least two classes 106, the selection being based on the probability that the input is classified as each class 106 in the class set 104. For example, the classifier can output the confidence of the classification 122, e.g., the probability of the classification error, and / or can suppress the output classification 122 based on a poor confidence, e.g., a minimum risk classifier. For example, in a region of the image 102 that cannot be clearly identified as a tumor or non-tumor, the classifier can be configured to suppress classifying the region, in order to improve the accuracy of the computed distribution of lymphocytes in regions that can be identified as tumors and non-tumors with an acceptable confidence.
[0154] As a sixth such example, some example embodiments may be configured to process a distribution of features using a linear or non-linear classifier, and may receive a classification 122 of an input class 106 from the linear or non-linear classifier. For example, the linear or non-linear classifier may include a support vector machine ensemble of at least two support vector machines, and some example embodiments may be configured to receive the classification 122 by receiving a candidate classification 122 from each of the at least two support vector machines, and determining the classification 122 based on a consistency of the candidate classifications 122 between the at least two support vector machines.
[0155] As a seventh such example, some example embodiments may be configured to perform a classification 122 of an input in various ways using a linear or non-linear classifier (including a classifier set or ensemble). For example, the input may be partitioned into regions, and for each input portion, an example embodiment may use a classifier to classify the input portion according to an input portion type selected from a set of input portion types (e.g., classifying 122 a portion of an image 102 of a tumor as a tumor region versus a non-tumor region, as Figure 3 shown in the example of). An example embodiment may then be configured to select a class of the input from a class set that includes at least two classes, the selection being based on a feature distribution of each input portion of the input portion type and a feature distribution of each input portion type in the input portion type set (e.g., classifying regions of tumor infiltrating lymphocytes (TIL), such as tumor adjacent lymphocytes in the stroma, highly activated lymphocyte regions, and / or low activated lymphocyte regions).
[0156] As an eighth such example, some example embodiments may be configured to perform a distribution classification by determining a variance of a feature distribution of the input (e.g., a variance of a distribution over a region of an input such as an image 102). Some example embodiments may then be configured to perform a classification 122 by selecting a class 106 of the input from a class set 104 that includes at least two classes 106, the selection being based on the variance of the feature distribution of the input and a variance of the feature distribution of each class 106 in the class set 104. For example, some example embodiments may be configured to determine a class of a tumor based at least in part on a variance of lymphocyte distribution over different regions of an image.
[0157] As a ninth such example, some example embodiments may use different training and / or testing to generate and validate a machine learning model. For example, heuristic methods such as stochastic gradient descent, nonlinear conjugate gradient, or simulated annealing may be used to perform training. Training may be performed offline (e.g., based on a fixed training dataset 100) or online (e.g., using new training data for continuous training). Training may be evaluated based on various metrics such as perceptron error, Kullback-Leibler (KL) divergence, precision, and / or recall. Training may be performed for a fixed time (e.g., a selected number of epochs or generations), until training fails to produce additional improvement, and / or until a convergence point is reached (e.g., when classification accuracy reaches a target threshold). The machine learning model may be tested in various ways (such as k-fold cross-validation) to determine the proficiency of the machine learning model with respect to previously unseen data. In some example embodiments, many such forms of classification 122, classifiers, training, testing, and validation may be included and used.
[0158] K. Example Computing Environments
[0159] Figure 18 is a diagram of an example apparatus in which some example embodiments may be implemented.
[0160] Figure 18 And the following discussion provides a brief, general description of a suitable computing environment for embodiments that implement one or more of the provisions set forth herein. Figure 18 The operating environment of is only one example of a suitable operating environment and is not intended to impose any limitation as to the scope of use or functionality of the operating environment. Example computing devices include, but are not limited to: personal computers; server computers; hand-held or laptop devices; mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.); multiprocessor systems; media devices such as televisions; consumer electronics; embedded devices; minicomputers; mainframe computers; distributed computing environments including any of the above systems or devices; wearable computing devices (such as glasses, headphones, watches, rings, pendants, hand-held and / or body-mounted cameras, clothing-integrated devices, and implantable devices); self-driving vehicles; extended reality (XR) devices such as augmented reality (AR) and / or virtual reality (VR) devices; Internet of Things (IoT) devices, etc.
[0161] Some example embodiments may include combinations of the same and / or different types of components, such as multiple processors and / or processing cores in a single-processor or multi-processor computer; two or more processors operating in series, such as a CPU and a GPU; a CPU utilizing an ASIC; and / or software executed by processing circuitry. Some example embodiments may include components of a single device, such as a computer including one or more CPUs that store, access, and manage a cache. Some example embodiments may include components of multiple devices, such as two or more devices having CPUs that communicate to access and / or manage a cache. Some example embodiments may include one or more components included in a server computing device, a server computer, a series of server computers, a server farm, a cloud computer, a content platform, a mobile computing device, a smart phone, a tablet computer, or a set-top box. Some example embodiments may include components that communicate directly (e.g., two or more cores of a multi-core processor) and / or indirectly (e.g., via a bus, via a wired or wireless channel or network, and / or via an intermediate component such as a microcontroller or arbiter). Some example embodiments may include multiple instances of a system or instance executed by a device or component, where such system instances may execute simultaneously, sequentially, and / or in an interleaved manner. Some example embodiments may be characterized by the distribution of an instance or system across two or more devices or components.
[0162] Although not required, some example embodiments are described in the general context of "computer-readable instructions" executed by one or more computing devices. The computer-readable instructions may be distributed via a computer-readable medium (discussed below). The computer-readable instructions may be implemented as program modules, such as functions, objects, application programming interfaces (APIs), data structures, etc. that perform particular tasks or implement particular abstract data types. Generally, the functionality of the computer-readable instructions may be combined or distributed as desired in various environments.
[0163] Figure 18 An example of an example apparatus 1800 is shown, which is configured to or includes one or more example embodiments, such as the example embodiments provided herein. In one apparatus configuration 1802, the example apparatus 1800 may include processing circuitry 1502 and a memory 1804. Depending on the exact configuration and type of the computing device, the memory 1804 may be volatile (e.g., RAM), non-volatile (e.g., ROM, flash memory, etc.), or some combination of the two.
[0164] In some example embodiments, the example apparatus 1800 may include additional features and / or functionality. For example, the example apparatus 1800 may also include additional storage devices (e.g., removable and / or non-removable), including but not limited to magnetic storage devices, optical storage devices, etc. Such additional storage devices areFigure 18 is presented by the storage device 1806. In some example embodiments, the computer-readable instructions for implementing one or more embodiments provided herein may be stored in the memory 1804 and / or the storage device 1806.
[0165] In some example embodiments, the storage device 1806 may be configured to store other computer-readable instructions for implementing an operating system, applications, etc. For example, the computer-readable instructions may be loaded into the memory 1804 for execution by the processing circuitry 1502. The storage device may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions or other data. The storage device may include, but is not limited to: RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage devices, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by the example device 1800. Any such computer storage medium may be part of the example device 1800.
[0166] In some example embodiments, the example device 1800 may include one or more input devices 1810 such as a keyboard, mouse, pen, voice input device, touch input device, infrared camera, video input device, and / or any other input device. One or more output devices 1808 such as one or more displays, speakers, printers, and / or any other output device may also be included in the example device 1800. The one or more input devices 1810 and the one or more output devices 1808 may be connected to the example device 1800 via a wired connection, a wireless connection, or any combination thereof. In some example embodiments, an input device or an output device from another computing device may be used as the one or more input devices 1810 or the one or more output devices 1808 of the example device 1800.
[0167] In some example embodiments, the example device 1800 may be connected through various interconnections such as a bus. Such interconnections may include Peripheral Component Interconnect (PCI) (such as PCI Express), Universal Serial Bus (USB), FireWire (IEEE 1394), optical bus architectures, etc. In other example embodiments, the components of the example device 1800 may be interconnected through a network. For example, the memory 1804 may include multiple physical memory units located at different physical locations interconnected through a network.
[0168] In some example embodiments, example apparatus 1800 may include one or more communication devices 1812 through which example apparatus 1800 may communicate with other devices. The (one or more) communication devices 1812 may include, for example, a modem, a network interface card (NIC), an integrated network interface, a radio frequency transmitter / receiver, an infrared port, a USB connection, or other interfaces for connecting example apparatus 1800 to other computing devices including remote device 1816. The (one or more) communication devices 1812 may include a wired connection or a wireless connection. The (one or more) communication devices 1812 may be configured to transmit and / or receive communication media.
[0169] Those skilled in the art will recognize that storage devices for storing computer-readable instructions may be distributed across a network. For example, example apparatus 1800 may communicate with remote device 1816 via network 1814 to store and / or retrieve computer-readable instructions to implement one or more example embodiments provided herein. For example, example apparatus 1800 may be configured to access remote device 1816 to download some or all of the computer-readable instructions for execution. Alternatively, example apparatus 1800 may be configured to download portions of the computer-readable instructions as needed, where some instructions may be executed at or by example apparatus 1800 and some other instructions may be executed at or by remote device 1816.
[0170] In this application, including the following definitions, the term "module" or the term "controller" may be replaced with the term "circuit". The term "module" may refer to processing circuitry 1502 (shared, dedicated, or grouped) that executes code and memory hardware (shared, dedicated, or grouped) that stores the code executed by the processing circuitry 1502, is part of, or includes the processing circuitry and the memory hardware.
[0171] A module may include one or more interface circuits. In some examples, the (one or more) interface circuits may implement a wired or wireless interface connected to a local area network (LAN) or a wireless personal area network (WPAN). Examples of a LAN are Institute of Electrical and Electronics Engineers (IEEE) standard 802.11-2016 (also known as the WIFI wireless network standard) and IEEE standard 802.3-2015 (also known as the Ethernet wired network standard). Examples of a WPAN are IEEE standard 802.15.4 (including the ZIGBEE standard from the ZigBee Alliance) and the Bluetooth wireless network standard from the Bluetooth Special Interest Group (SIG) (including core specification versions 3.0, 4.0, 4.1, 4.2, 5.0, and 5.1 from the Bluetooth SIG).
[0172] A module can communicate with other modules using an (optional) interface circuit. Although a module may be depicted in this disclosure as directly communicating logically with other modules, in various embodiments, a module may actually communicate via a communication system. The communication system includes physical and / or virtual network devices such as hubs, switches, routers, and gateways. In some embodiments, the communication system is connected to or traverses a wide area network (WAN) such as the Internet. For example, the communication system may include multiple LANs and virtual private networks (VPNs) that are connected to each other via the Internet or point-to-point leased lines using technologies including Multiprotocol Label Switching (MPLS).
[0173] In various embodiments, the functionality of a module may be distributed among multiple modules connected via a communication system. For example, multiple modules may implement the same functionality distributed by a load balancing system. In another example, the functionality of a module may be divided between a server (also referred to as a remote or cloud) module and a client (or user) module.
[0174] The term code as used above may include software, firmware, and / or microcode and may refer to a program, routine, function, class, data structure, and / or object. The shared processing circuitry 1502 may encompass a single microprocessor that executes some or all of the code from multiple modules. The grouped processing circuitry 1502 may encompass a microprocessor that, in combination with additional microprocessors, executes some or all of the code from one or more modules. References to multiple microprocessors encompass multiple microprocessors on discrete die, multiple microprocessors on a single die, multiple cores of a single microprocessor, multiple threads of a single microprocessor, or a combination of the above.
[0175] The shared memory hardware encompasses a single memory device that stores some or all of the code from multiple modules. The grouped memory hardware encompasses a memory device that, in combination with other memory devices, stores some or all of the code from one or more modules.
[0176] The term memory hardware is a subset of the term computer-readable medium. The term computer-readable medium as used herein does not encompass transient electrical or electromagnetic signals propagated through a medium (such as on a carrier wave); thus, the term computer-readable medium is considered tangible and non-transitory. Non-limiting examples of non-transitory computer-readable media are non-volatile storage devices (such as flash memory devices, erasable programmable read-only storage devices, or mask read-only memory devices), volatile memory devices (such as static random access memory devices or dynamic random access memory devices), magnetic storage media (such as analog or digital tape or hard disk drives), and optical storage media (such as CDs, DVDs, or Blu-ray discs).
[0177] Example embodiments of the apparatus and methods described herein may be implemented in whole or in part by a special purpose computer created by configuring a general purpose computer to execute one or more specific functions embodied in a computer program. The functional blocks and flowchart elements described herein may serve as a software specification, which can be translated into a computer program by the routine work of a skilled technician or programmer.
[0178] The computer program includes processor-executable instructions stored on at least one non-transitory computer-readable medium. The computer program may also include or rely on stored data. The computer program may encompass a basic input / output system (BIOS) that interacts with the hardware of the special purpose computer, device drivers that interact with specific devices of the special purpose computer, one or more operating systems, user applications, background services, background applications, and the like.
[0179] The computer program may include: (i) descriptive text to be parsed, such as HTML (HyperText Markup Language), XML (eXtensible Markup Language), or JSON (JavaScript Object Notation), (ii) assembly code, (iii) object code generated from source code by a compiler, (iv) source code executed by an interpreter, (v) source code compiled and executed by a just-in-time compiler, etc. By way of example only, the source code may be written using the syntax of languages including C, C++, C#, Objective-C, Swift, Haskell, Go, SQL, R, Lisp, Fortran, Perl, Pascal, Curl, OCaml, HTML5 (HyperText Markup Language 5th Revision), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Eiffel, Smalltalk, Erlang, Ruby, Visual Lua, MATLAB, SIMULINK, and the like.
[0180] L. Use of Terms
[0181] The foregoing description is merely illustrative in nature and is in no way intended to limit the disclosure, its application, or uses. The broad teachings of the disclosure can be implemented in a variety of forms. Thus, while the disclosure includes specific examples, the true scope of the disclosure should not be so limited since other modifications will become apparent upon study of the drawings, the specification, and the appended claims. It should be understood that one or more steps within a method can be executed in a different order (or concurrently) without altering the principles of the disclosure. Further, while each embodiment has been described above as having certain features, any one or more of those features described with respect to any embodiment of the disclosure can be implemented in and / or combined with the features of any other exemplary embodiment, even if the combination is not explicitly described. In other words, the described embodiments are not mutually exclusive, and arrangements of one or more of the embodiments with each other remain within the scope of the disclosure.
[0182] Various terms are used to describe spatial and functional relationships between elements (e.g., between modules), including “connected,” “engaged,” “interface,” and “coupled.” Unless explicitly described as “direct,” when describing a relationship between a first element and a second element in the foregoing disclosure, the relationship encompasses both a direct relationship where no other intervening elements exist between the first element and the second element and an indirect relationship where one or more intervening elements (spatially or functionally) exist between the first element and the second element. As used herein, the phrase “at least one of A, B, and C” should be construed to mean a logical (A OR B OR C), using inclusive logical OR, and should not be construed to mean “at least one of A, at least one of B, and at least one of C.”
[0183] In the figures, the arrow direction indicated by an arrow generally shows the information flow (such as data or instructions) of interest in the illustration. For example, when element A and element B exchange various information but the information transmitted from element A to element B is relevant to the illustration, the arrow may point from element A to element B. This one-way arrow does not mean that no other information is transmitted from element B to element A. Further, for the information sent from element A to element B, element B can send a request for the information or receive an acknowledgement to element A. The term subset does not necessarily require a proper subset. In other words, a first subset of a first set can be coextensive with (equal to) the first set.
[0184] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
[0185] As used herein, the terms "component", "module", "system", "interface", etc. generally intend to refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a processing circuitry 1502, a process running on the processing circuitry 1502, an object, an executable, an execution thread, a program, and / or a computer. By way of illustration, an application running on a controller and the controller can both be components. One or more components can reside within a process and / or an execution thread and a component can be located on one computer and / or distributed between two or more computers.
[0186] In addition, some example embodiments may include using standard programming and / or engineering techniques to generate software, firmware, hardware, or any combination thereof to control a computer to implement the methods, apparatuses, or articles of manufacture disclosed herein. The term "article of manufacture" as used herein is intended to cover a computer program accessible from any computer-readable device, carrier, or medium. Of course, those skilled in the art will recognize that many modifications can be made to this configuration without departing from the scope or spirit of the claimed subject matter.
[0187] Various operations of embodiments are provided herein. In some example embodiments, one or more of the described operations may constitute computer-readable instructions stored on one or more computer-readable media that, when executed by a computing device, will cause the computing device to perform the described operations. The order in which some or all of the operations are described should not be construed as implying that these operations must be order-dependent. Those skilled in the art benefiting from this specification will understand alternative orderings. Further, it should be understood that not all operations must be present in each example embodiment provided herein.
[0188] As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X employs A or B" is intended to mean any natural inclusive permutation. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" holds true in any of the foregoing instances. As used herein and in the appended claims, the articles "a" and "an" generally may be construed to mean "one or more" unless otherwise specified or clearly indicated to be in the singular form from the context.
[0189] Although the present disclosure has been shown and described with respect to certain example embodiments, equivalent changes and modifications will occur to others upon reading and understanding this specification and the drawings. The present disclosure includes all such modifications and changes and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise specified, the terms used to describe such components are intended to correspond to any component that performs the specified function of the described component (e.g., functionally equivalent), even if not structurally equivalent to the disclosed structure that performs the functions in some of the example embodiments shown herein. Further, although a particular feature of the present disclosure may have been disclosed only with respect to one of several embodiments, such feature may be combined with one or more other features of the other embodiments, which may be desirable and advantageous for any given or particular application. Additionally, with respect to the terms "includes", "having", "has", "with" or variants thereof as used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "comprising".
Claims
1. A method of operating a device including a processing circuitry, the method comprising: executing, by the processing circuitry, instructions that cause the device to: receive an image depicting at least a portion of a tumor, invoke a convolutional neural network to determine a lymphocyte distribution of lymphocytes in respective regions of the tumor based on the image, the convolutional neural network being trained to determine a lymphocyte distribution of lymphocytes in an image region, apply a classifier to the lymphocyte distribution to classify the tumor, the classifier being trained to classify a tumor into a category selected from at least two categories associated with the lymphocyte distribution, and determine a clinical value of an individual based on a prognosis data set corresponding to an individual having a tumor of the category into which the classifier classifies the tumor; wherein the convolutional neural network is further trained to classify regions of the image into one or more region types selected from a set of region types including: a tumor region, a lymphocyte region, and a stromal region; wherein determining the lymphocyte distribution of lymphocytes in the tumor includes, for respective lymphocyte regions of the image: determining a distance of the lymphocyte region to one or both of a tumor region or a stromal region, and characterizing the lymphocyte region, based on the distance, as one of a tumor-infiltrating lymphocyte region, a tumor-adjacent lymphocyte region, a stromal-infiltrating lymphocyte region, and a stromal-adjacent lymphocyte region; and wherein the classifier is configured to further classify the tumor based on the characterization of the lymphocyte region.
2. The method according to claim 1, wherein The tumor is one of the following: a pancreatic cancer tumor, and a breast cancer tumor.
3. The method of claim 1, wherein: determining the lymphocyte distribution of lymphocytes in the tumor includes, for respective stromal regions of the image: determining a distance of the stromal region to the tumor region, and characterizing the stromal region, based on the distance, as one of the following: a tumor-infiltrating stromal region, and a tumor-adjacent stromal region; and the classifier further classifies the tumor based on the characterization of the stromal region.
4. The method according to claim 1, wherein, The at least two categories include: a high-risk tumor category associated with a first survival probability, and a low-risk tumor category associated with a second survival probability, the second survival probability being longer than the first survival probability.
5. The method according to claim 1, wherein, The classifier further includes a Gaussian mixture model configured to determine a probability distribution of features of tumors of a respective category within a feature space, and features of the feature space of the Gaussian mixture model are selected from a set of features including the following: measurements of the tumor region of the image, measurements of the stromal region of the image, including measurements of the tumor-infiltrating stromal region of the image and measurements of the tumor-adjacent stromal region of the image; and measurements of the lymphocyte region of the image, including measurements of the tumor-infiltrating lymphocyte region of the image, measurements of the tumor-adjacent lymphocyte region of the image, measurements of the stromal-infiltrating lymphocyte region of the image, and measurements of the stromal-adjacent lymphocyte region of the image.
6. The method according to claim 5, wherein The feature subset is selected from the feature set based on the correlation of these corresponding categories with the corresponding features of the subset, and the correlation of the corresponding categories with the corresponding features is based on at least one of the profile score and the consistency index of the feature space.
7. The method according to claim 6, wherein The feature subset consists of: Measurements of the tumor-infiltrating lymphocyte region of the image, Measurements of the tumor-adjacent lymphocyte region of the image, and Measurements of the tumor-infiltrating stromal region of the image.
8. The method according to claim 1, wherein These instructions further cause the device to perform the following operations: Apply a Cox proportional hazards model to the clinical features of the tumor to determine the category of the tumor, and Determine the clinical value of the individual based on the prognosis of individuals with the category into which the classifier classifies the tumor and the category determined by the Cox proportional hazards model.
9. The method according to claim 8, wherein, The clinical features of the tumor for the Cox proportional hazards model are selected from a set of clinical features that includes: The initial diagnosis of the tumor, The location of the tumor, The treatment of the tumor, The measurements of the tumor, The metastatic status of the tumor, The initial diagnosis of the individual, The individual's prior cancer history, The gender of the individual, The frequency of the individual's smoking habit, The duration of the individual's smoking habit, and The individual's alcohol history.
10. The method according to claim 9, wherein, The clinical feature subset of the features is selected from the set of clinical features for the Cox proportional hazards model based on the correlation of these corresponding categories with the corresponding clinical features in the clinical feature subset, and the clinical feature subset consists of the measurements of the tumor and the metastatic status of the tumor.
11. The method according to claim 1, wherein: The at least two categories are a low-risk tumor category and a high-risk tumor category, Determining the lymphocyte distribution further includes applying a convolutional neural network to the image, the convolutional neural network being configured to measure the lymphocyte distribution of lymphocytes in different region types of the image, The classifier is a two-way Gaussian mixture model, the two-way Gaussian mixture model being configured to determine the probability distribution of the features of the tumors of the corresponding category within the feature space for the corresponding category, The method further includes applying a Cox proportional hazards model to the clinical features of the tumor to determine the category of the tumor, and Determining the clinical value of the individual is further based on the category determined by the Cox proportional hazards model.
12. The method according to claim 1, wherein These instructions further cause the device to display a Kaplan Meier survival projection of the clinical value of the individual.
13. The method according to claim 1, wherein These instructions further cause the device to determine at least one of the following: A diagnostic test for the tumor based on the clinical value of the individual; and The treatment of the individual based on the clinical value of the individual, including a schedule of therapeutic agents for treating the tumor based on the clinical value of the individual.
14. A device for tumor assessment, comprising: A memory that stores instructions: and A processing circuitry that is further configured to determine the clinical value of an individual based on a tumor in an image by performing the following operations by executing the instructions stored in the memory: Receiving an image depicting at least a portion of a tumor, Invoke a convolutional neural network to determine the lymphocyte distribution of lymphocytes in various regions of the tumor based on the image, the convolutional neural network being trained to determine the lymphocyte distribution of lymphocytes in an image region, apply a classifier to the lymphocyte distribution to classify the tumor, the classifier being trained to classify the tumor into a category selected from at least two categories associated with the lymphocyte distribution, and determine a clinical value of an individual based on a prognostic data set corresponding to an individual with a tumor of the category into which the classifier classifies the tumor; wherein the convolutional neural network is further trained to classify regions of the image into one or more region types selected from a set of region types, the set of region types including: a tumor region, a lymphocyte region, and a stromal region; wherein determining the lymphocyte distribution of lymphocytes in the tumor includes, for a corresponding lymphocyte region of the image: determine the distance of the lymphocyte region to one or both of a tumor region or a stromal region, and based on the distance, characterize the lymphocyte region as one of a tumor-infiltrating lymphocyte region, a tumor-adjacent lymphocyte region, a stromal-infiltrating lymphocyte region, and a stromal-adjacent lymphocyte region; and wherein the classifier is configured to further classify the tumor based on the characterization of the lymphocyte region; wherein the at least two categories are a low-risk tumor category and a high-risk tumor category, the classifier includes a two-way Gaussian mixture model configured to determine a probability distribution of the characteristics of tumors of the corresponding category within a feature space for the corresponding category, and the instructions further cause the processing circuitry to perform the following operations: apply a Cox proportional hazards model to the clinical characteristics of the tumor to determine the category of the tumor, and determine a clinical value of the individual based on the prognosis of an individual with a tumor of the category into which the classifier classifies the tumor and the category determined by the Cox proportional hazards model.
15. A non-transitory computer-readable medium storing instructions that, when executed by a processing circuitry, cause the processing circuitry to determine a clinical value of an individual based on a tumor in an image by: receiving an image depicting at least a portion of a tumor, invoke a convolutional neural network to determine the lymphocyte distribution of lymphocytes in various regions of the tumor based on the image, the convolutional neural network being trained to determine the lymphocyte distribution of lymphocytes in an image region, apply a classifier to the lymphocyte distribution to classify the tumor, the classifier being trained to classify the tumor into a category selected from at least two categories associated with the lymphocyte distribution, and determine a clinical value of an individual based on a prognostic data set corresponding to an individual with a tumor of the category into which the classifier classifies the tumor; Among them, the convolutional neural network is further trained to classify regions of the image into one or more region types selected from a set of region types, the set of region types including: a tumor region, a lymphocyte region, and a stromal region; wherein determining the lymphocyte distribution of lymphocytes in the tumor includes, for a corresponding lymphocyte region of the image: Determine the distance from the lymphocyte region to one or both of the tumor region or the stromal region, and Based on this distance, characterize the lymphocyte region as one of a tumor-infiltrating lymphocyte region, a tumor-adjacent lymphocyte region, a stromal-infiltrating lymphocyte region, and a stromal-adjacent lymphocyte region; and wherein the classifier is configured to further classify the tumor based on the characterization of the lymphocyte region; wherein at least two classes are a low-risk tumor class and a high-risk tumor class, The classifier includes a two-way Gaussian mixture model, which is configured to determine the probability distribution of the characteristics of the tumors of the corresponding class within the feature space for each class, and These instructions further cause the processing circuitry to perform the following operations: Apply the Cox proportional hazards model to the clinical characteristics of the tumor to determine the class of the tumor, and Based on the prognosis of the individual with the tumor classified into the class by the classifier and the class determined by the Cox proportional hazards model, determine the clinical value of the individual.
16. A method of operating a device including a processing circuit, the method comprising: Execute instructions by the processing circuit, the instructions causing the device to perform the following operations: Receive an image depicting at least a portion of a tumor, Invoke a convolutional neural network to determine the lymphocyte distribution of lymphocytes in the respective regions of the tumor based on the image, the convolutional neural network being trained to determine the lymphocyte distribution of lymphocytes in an image region, Apply a classifier to the lymphocyte distribution to classify the tumor, the classifier being trained to classify the tumor into a class selected from at least two classes related to the lymphocyte distribution, and Determine the clinical value of an individual based on prognosis data corresponding to an individual with a tumor classified into the class by the classifier, wherein the convolutional neural network is further trained to classify the regions of the image into one or more region types selected from a set of region types, the set of region types including: a tumor region, a lymphocyte region, and a stromal region, wherein determining the lymphocyte distribution of lymphocytes in the tumor includes, for the respective stromal regions of the image, Determine the distance from the stromal region to the tumor region, and Based on this distance, characterize the stromal region as one of a tumor-infiltrating stromal region and a tumor-adjacent stromal region, and wherein the classifier is configured to further classify the tumor based on the characterization of the stromal region.
Citation Information
Patent Citations
Computerized analysis of computed tomography (CT) imagery to quantify tumor infiltrating lymphocytes (TILS) in non-small cell lung cancer (NSCLC)
US20170352157A1
Scoring of tumor infiltration by lymphocytes
US20170365053A1