Histological image analysis
A computer-implemented system processes histological images with varying resolutions to combine classifiers, addressing inconsistencies in histopathological analysis and improving prognostic accuracy for colorectal cancer treatment decisions.
Patent Information
- Application Number
- JP2025067662
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-20
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-20
Smart Images

Figure 2025121912000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the analysis of histological images, and in particular to the use of machine learning algorithms to perform such analysis and to training machine learning algorithms to perform this analysis. [Background technology]
[0002] Biomarkers are increasingly being used to match anti-cancer therapies to specific tumor genotype, protein, and RNA expression profiles, typically found in patients with advanced disease (La Thangue & Kerr, Nat Rev Clin Oncol, 2011;8:587-96; Van Allen et al., Nat Med, 2014;20:682-8; Moscow et al., Nat Rev Clin Oncol 2018;15:183-92).
[0003] One example of this is the selection of KRAS wild-type colorectal cancer (CRC) for treatment with epidermal growth factor receptor inhibitors (Karapetis et al., N Engl J Med, 2008;359:1757-65). However, in the adjuvant setting of CRC, the primary question is binary: whether to offer treatment, and the subsequent agent choice, dose, and schedule are determined primarily by stage, not by the presence of a companion diagnostic. If prognostic models could be further refined, this would allow for a more targeted approach by defining subgroups for whom the absolute benefit of adjuvant chemotherapy compared with surgery alone is minimal and, at the other end of the spectrum, who may benefit from long-term combination chemotherapy (Kerr & Shi, Nat Rev Clin Oncol,2013;10:429-30; Hutchins et al., J Clin Oncol,2011;29:1261-70; Salazar et al., J Clin Oncol,2011;29:17-24; Gray et al., J Clin Oncol,2011;29:4611-9).
[0004] Over 20 years of adjuvant trials in patients with early-stage CRC using fluoropyrimidines in combination with cytotoxic agents such as oxaliplatin have improved overall survival (OS) for patients with stage II or IIIA CRC by approximately 3–5%, with the majority (approximately 80%) being cured with surgery alone. Despite adjuvant chemotherapy, approximately 20% will relapse, chemotherapy-related mortality is likely to be 0.5–1%, and 20% of patients will suffer from significant side effects. Although the risk / benefit ratio is quite small, it could be much lower if subgroups at increased risk of recurrence and cancer-specific mortality could be defined (Group QC, Lancet, 2000;355:1588-96, Quasar Collaborative G, Gray R, Barnwell J, et al., Lancet, 2007;370:2020-9, Andre et al., J Clin Oncol, 2009;27:3109-16, Andre et al., J Clin Oncol, 2015;33:4176-87).
[0005] Clinically validated prognostic biomarkers facilitate adjuvant treatment decisions, but few have been validated reliably enough for routine clinical application. Patients with mismatch repair (MMR)-deficient tumors tend to have a favorable prognosis, making a case for routine assessment of MMR status (Sinicrope, Nat Rev Clin Oncol, 2010;7:174-7; Mouradov et al., Am J Gastroenterol, 2013;108:1785-93). We recently reported that measurement of tumor cell DNA content (ploidy) combined with stromal fractionation can stratify stage II patients into very good, intermediate, and poor prognosis groups (Danielsen et al., Ann Oncol, 2018;29:616-23). Interestingly, analysis of driver mutations and RNA signatures has shown that they are individually weak prognostic markers and cannot guide clinical decision-making (Grey et al., 2011, supra; Mouradov et al., 2013, supra).
[0006] Thus, there is a need to provide improved means for assessing biomarkers in biological materials and further develop capabilities to provide useful and efficient means for classifying biological materials, such as for prognostic and diagnostic approaches.
[0007] Deep learning has already been shown to be suitable for the detection and delineation of several tumor types (Ehteshami Bejnordi et al. JAMA, 2017;318:2199-210) and various cancer classifications have been reported (Coudray et al. Nat Med, 2018;24:1559-67). However, we have not yet seen a validated system for directly predicting patient outcomes based on histological images.
[0008] The aim of this study is to facilitate the use of deep learning and digital analytics to develop a fully automated system for histological image analysis, which is tested and validated in predicting the prognosis of primary CRC patients using conventional whole slide images (WSI).
[0009] As used herein, a "histological image" refers to an image showing the microscopic structure of a biological material. A "histological feature of interest" refers to a feature of this microscopic structure. The feature may be of interest, for example, for prognostic, diagnostic, or therapeutic purposes, or for scientific research purposes.
[0010] Histological specimens are typically used to confirm structure and determine a diagnosis or to attempt to determine a prognosis.
[0011] When the histological image relates to pathology, the term "histopathological image" may be used.
[0012] At the microscopic scale, many of the interesting features of cells are invisible because they are transparent and colorless. To reveal these features, specimens are typically stained with one or more markers before being imaged under a microscope. Markers include one or more stains (dyes or pigments) designed to specifically bind to particular components of cellular structure, thereby revealing the histological features of interest.
[0013] One commonly used staining system is called H&E (hematoxylin and eosin). H&E contains two dyes: hematoxylin and eosin. Eosin is an acid dye and is negatively charged. Eosin stains basic (or acidophilic) structures red or pink. Hematoxylin can be considered a basic dye. Hematoxylin is used to stain acidophilic (or basophilic) structures a purplish-blue color.
[0014] DNA in the nucleus (heterochromatin and nucleoli) and RNA in ribosomes and the rough endoplasmic reticulum are both acidic, so hematoxylin binds to them and stains them purple. Some extracellular substances (i.e., carbohydrates in cartilage) are also basophilic. Most proteins in the cytoplasm are basic, so eosin binds to these proteins and stains them pink. This includes cytoplasmic filaments in muscle cells, intracellular membranes, and extracellular fibers.
[0015] Examples of some alternative staining methods that may be used, as will be recognized by those skilled in the art, are discussed further in this application.
[0016] Such histological image can be used to evaluate the tissue that may be diseased, for example, the tissue that may be cancerous.Therefore, the image can be histopathological image.It is useful to classify histological (for example, histopathological) image, for example, for the purpose of diagnosis, prognosis and / or stratification of the subject that histological image is obtained, to determine expected results, to make treatment decisions for the subject, or to evaluate the effect of the treatment that the subject is undergoing and / or has undergone.
[0017] Traditionally, histologic features of interest are identified in histologic images by histopathologists (specialized medical professionals skilled in interpreting these images).
[0018] However, experiments have been performed and classification by histopathologists is inconsistent, often limiting prognostic value both when comparing the identifications of different histopathologists and even when presenting the same images to the same histopathologist on different occasions. Such inconsistencies, as well as inter- and intra-observer variability, can have serious implications.
[0019] Therefore, there is a need for improved automated histopathological image analysis methods and devices. Summary of the Invention [Means for solving the problem]
[0020] The invention is defined by the claims.
[0021] A computer-implemented system for determining a global classifier for one or more source histological images is disclosed, the system comprising: a first tile generator configured to generate a plurality of first tiles from one or more source histological images, each of the plurality of first tiles including a plurality of pixels representing a region of the one or more source histological images having a first area and a first resolution; a second tile generator configured to generate a plurality of second tiles from the one or more source histological images, each of the plurality of second tiles including a plurality of pixels representing a region of the one or more source histological images having a second area and a second resolution; a first area of the first tile is greater than a second area of the second tile; a second tile generator, wherein the second resolution of the second tiles is higher than the first resolution of the first tiles; a machine learning network configured to process the plurality of first tiles to determine a first classifier for the one or more source histological images; a machine learning network configured to process the plurality of second tiles to determine a second classifier for the one or more source histological images; a classifier combiner configured to combine the first classifier and the second classifier to determine an overall classifier for the one or more source histological images.
[0022] The classifier combiner applying a thresholding function to the first classifier to determine a thresholded first classifier; applying a thresholding function to the second classifier to determine a thresholded second classifier; and combining the thresholded first classifier and the thresholded second classifier to determine an overall classifier.
[0023] The machine learning network may be configured to process the plurality of first tiles to determine a plurality of first classifiers for the one or more source histological images. The machine learning network may be configured to process the plurality of second tiles to determine a plurality of second classifiers for the one or more source histological images.
[0024] The classifier combiner applying a statistical function to the plurality of first classifiers to determine a combined first classifier; applying a statistical function to the plurality of second classifiers to determine a combined first classifier; and combining the combined first classifier and the combined second classifier to determine an overall classifier.
[0025] The classifier combiner may be configured to perform a logical combination of the first classifier and the second classifier to determine an overall classifier for the one or more source histological images.
[0026] Also disclosed is a computer-implemented method for processing histological images, the method comprising: receiving one or more source histological images; generating a plurality of first tiles from the one or more source histological images, each of the plurality of first tiles including a plurality of pixels representing a region of the one or more source histological images having a first area and a first resolution; generating a plurality of second tiles from the source histological images, each of the plurality of second tiles including a plurality of pixels representing a region of the one or more source histological images having a second area and a second resolution; a first area of the first tile is greater than a second area of the second tile; generating a second resolution of the second tile that is higher than the first resolution of the first tile; applying a machine learning network to the plurality of first tiles to determine a first classifier for the one or more source histological images; applying the machine learning network to the plurality of second tiles to determine a second classifier for the one or more source histological images; combining the first classifier and the second classifier to determine an overall classifier for the one or more source histological images.
[0027] Also disclosed is a computer-implemented system for determining a global classifier for one or more source histological images, the system comprising: a tile generator configured to generate a plurality of tiles from one or more source histological images, each of the plurality of tiles including a plurality of pixels representing a region of the one or more source histological images; a first neural network configured to process the plurality of tiles to determine tile features for each of the plurality of tiles; a pooling function configured to combine subsets of the tile features to generate a bag feature for each of the subsets; and a second neural network configured to process the bag features to determine a classifier for the one or more source histological images. The second neural network may be a classification network.
[0028] The system is comparing the classifier determined by the second neural network with ground truth represented by the truth data; and setting trainable parameters for the first neural network, the pooling function, and the second neural network based on a result of the comparison.
[0029] The system is The method further comprises a segmentation block configured to apply an image segmentation method to the whole slide image histological image to provide a source histological image.
[0030] Also disclosed is a computer-implemented method for processing histological images, said method comprising: receiving one or more source histological images; generating a plurality of tiles from one or more source histological images, each of the plurality of tiles including a plurality of pixels representing a region of the source histological image; applying the plurality of tiles to a first neural network to determine tile features for each of the plurality of tiles; combining subsets of the tile features to generate a bag feature for each of the subsets; and applying the bag features to a second neural network to determine a classifier for the one or more source histological images. The second neural network may be a classification network.
[0031] Also disclosed is a computer-implemented system for determining a global classifier for one or more source histological images, the system comprising: a first tile generator configured to generate a plurality of first tiles from one or more source histological images, each of the plurality of first tiles including a plurality of pixels representing a region of the one or more source histological images having a first area and a first resolution; a second tile generator configured to generate a plurality of second tiles from the one or more source histological images, each of the plurality of second tiles including a plurality of pixels representing a region of the one or more source histological images having a second area and a second resolution; a first area of the first tile is greater than a second area of the second tile; a second tile generator, wherein the second resolution of the second tiles is higher than the first resolution of the first tiles; a machine learning network configured to process a plurality of first tiles to determine a first classifier for one or more source histological images; a first neural network configured to process the plurality of first tiles to determine tile features for each of the plurality of first tiles; a pooling function configured to combine subsets of the tile features to generate a bag feature for each of the subsets; a second neural network, the second neural network being a classification network, configured to process the bag features to determine a first classifier for the one or more source histological images; a machine learning network configured to process a plurality of second tiles to determine a second classifier for the one or more source histological images; a first neural network configured to process the plurality of second tiles to determine tile features for each of the plurality of second tiles; a pooling function configured to combine subsets of the tile features to generate a bag feature for each of the subsets; a second neural network, the second neural network being a classification network, configured to process the bag features to determine a second classifier for the one or more source histological images; a classifier combiner configured to combine the first classifier and the second classifier to determine an overall classifier for the one or more source histological images.
[0032] Also disclosed is a computer-implemented method for determining a global classifier for one or more source histological images, the method comprising: generating a plurality of first tiles from the one or more source histological images, each of the plurality of first tiles including a plurality of pixels representing a region of the one or more source histological images having a first area and a first resolution; generating a plurality of second tiles from the one or more source histological images, each of the plurality of second tiles including a plurality of pixels representing a region of the one or more source histological images having a second area and a second resolution; a first area of the first tile is greater than a second area of the second tile; generating a second resolution of the second tile that is higher than the first resolution of the first tile; applying a machine learning network to the plurality of first tiles to determine a first classifier for the one or more source histological images, wherein applying the machine learning network includes: applying a first neural network to the plurality of first tiles to determine tile features for each of the plurality of first tiles; combining subsets of the tile features to generate a bag feature for each of the subsets; applying a second neural network (which may be a classification network) to the bag features to determine a first classifier for the one or more source histological images; applying a machine learning network to the plurality of second tiles to determine a second classifier for the one or more source histological images, wherein applying the machine learning network includes: applying a first neural network to the plurality of second tiles to determine tile features for each of the plurality of second tiles; combining subsets of the tile features to generate a bag feature for each of the subsets; applying a second neural network (which may be a classification network) to the bag features to determine a second classifier for the one or more source histological images; combining the first classifier and the second classifier to determine an overall classifier for the one or more source histological images.
[0033] Any first tile generator and second tile generator disclosed herein may be configured to generate their respective tiles independently of each other.
[0034] Any system disclosed herein may include: The system may further comprise a segmentation block configured to apply an image segmentation method to the whole slide image histological images to provide a plurality of source histological images.
[0035] Any of the first neural networks disclosed herein may be trained using training histological images and associated ground truth. Any of the second neural networks disclosed herein may be trained using training histological images and associated ground truth.
[0036] Any of the machine learning networks disclosed herein may be trained using training histological images and associated ground truth.
[0037] Any of the methods disclosed herein may be a method of generating a diagnosis and / or prognosis for a subject, The method includes receiving one or more source histological images obtained from one or more histological samples obtained from a subject, and the method includes: determining a classifier for one or more source histological images (102) according to any suitable method disclosed herein, and / or determining an overall classifier for one or more source histological images according to any suitable method disclosed herein; and and attributing a diagnosis and / or prognosis assessment to the classifier and / or the overall classifier.
[0038] The subject may be a human.
[0039] The subject may have a pathological condition, may have been diagnosed with a pathological condition, may be suspected of having a pathological condition, may be being treated for a pathological condition, may have previously been treated for a pathological condition, and / or may have previously had a pathological condition.
[0040] The or each histological sample obtained from a subject may be obtained from a part of the subject's body that has a pathological condition, is suspected of having a pathological condition, is being treated for a pathological condition, has previously been treated for a pathological condition, and / or has previously had a pathological condition.
[0041] The pathological condition can be a cancer, for example, a cancer selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed type cancers.
[0042] The cancer may be colon cancer.
[0043] The method includes evaluating a plurality of source histological images obtained from a plurality of histological samples obtained from the subject to determine a plurality of classifiers and / or an overall classifier; Optionally, attributing a diagnosis and / or prognosis assessment to a plurality of classifiers and / or an overall classifier; Optionally, the subject has, has been diagnosed with, is suspected of having, is being treated for, has previously been treated for, and / or has previously had a pathological condition, such as a cancer, e.g., a cancer selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed cancers, and optionally, the cancer can be colon cancer.
[0044] The method may include evaluating one or more additional diagnostic and / or prognostic markers for the pathological condition, The step of attributing a diagnostic and / or prognostic assessment to the classifier and / or overall classifier may comprise evaluation of the results of the or each evaluation of the or each further diagnostic and / or prognostic marker.
[0045] The method may further include making a subject treatment decision based on the diagnostic and / or prognostic assessment; Optionally, the treatment decision relates to a diagnosed or prognosticated pathological condition, such as a cancer, e.g., a cancer selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed cancers, and optionally, the cancer is colorectal cancer.
[0046] Also disclosed are methods of treating in a subject in need thereof, wherein a diagnostic and / or prognostic assessment has been imparted to the subject by any suitable method disclosed herein, said method comprising treating the subject by surgery and / or non-surgical therapy; Optionally, the treatment of the diagnosed or prognosed pathological condition is for a cancer selected from the group consisting of, for example, carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed cancers, and optionally, the cancer is colon cancer.
[0047] The subject may be a human. In some examples, the subject is (a) has a pathological condition, has been diagnosed with a pathological condition, is suspected of having a pathological condition, is being treated for a pathological condition, has previously been treated for a pathological condition, and / or has previously had a pathological condition; and / or (b) A diagnosis and / or prognosis assessment of the pathological condition is imputed to the subject by any suitable method disclosed herein.
[0048] The pathological condition can be a cancer, for example, a cancer selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed type cancers, and optionally, the cancer is colon cancer.
[0049] The method may include adapting one or more parameters of surgical and / or non-surgical therapy in light of the diagnostic and / or prognostic assessment ascribed to the subject by any suitable method disclosed herein, and optionally, The one or more parameters of the surgical and / or non-surgical therapy are selected from the group consisting of: the nature of the surgical and / or non-surgical therapy, the timing of the surgical and / or non-surgical therapy, the duration of the surgical and / or non-surgical therapy, the dosage of the therapy, the route of administration of the non-surgical therapy, and the site in the body targeted by the surgical and / or non-surgical therapy.
[0050] The diagnostic and / or prognostic assessment of a subject may include evaluating the effect of previous or ongoing treatment with surgical and / or non-surgical therapy on the subject; For example, to monitor the progress and / or effectiveness of such treatment, and further optionally: The method includes making a further treatment decision, such as discontinuing, continuing, repeating or modifying a previous or ongoing treatment and / or the implementation of a different treatment modality; and, optionally, and implementing further treatment decisions regarding the subject, optionally including: The diagnosis and / or prognosis assessment, treatment and / or therapy decision relates to a pathological condition such as a cancer, e.g., a cancer selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed cancers, and optionally the cancer is colorectal cancer.
[0051] A computer program may be provided which, when executed on a computer, causes the computer to configure any device, including the system, block, or module disclosed herein, or to perform any method disclosed herein. The computer program may be implemented in software, and the computer may be considered to be any suitable hardware, including, but not limited to, a digital signal processor, a microcontroller, and implementations with read-only memory (ROM), erasable programmable read-only memory (EPROM), and erasable programmable read-only memory (EEPROM). The software may also be an assembly program.
[0052] The computer program may be provided on a computer-readable medium, which may be a physical computer-readable medium such as a disk or memory device, or may be embodied as a transitory signal, which may be a network download, including an internet download.
[0053] The invention will now be described, by way of example only, with reference to the accompanying drawings in which: [Brief explanation of the drawings]
[0054] [Figure 1] 1 illustrates a computer-implemented system for determining a classifier for a source histological image, such as a source histopathological image. [Figure 2] 1 illustrates a computer-implemented system for determining a global classifier for a source histological image, such as a source histopathological image. [Figure 3]1 illustrates a particular implementation of a system for determining a global classifier for a source histological image, such as a source histopathological image. [Figure 4] 1 shows a diagram of an example segmentation network architecture. [Figure 5] 1 shows the implementation of the segmentation method on a training set. [Figure 6] A diagram of this implementation of the machine learning network architecture is shown in Figure 3. [Figure 7A] Figure 1 shows the c-index of the 21 candidate 10x models of machine learning network 3 for patients with uncertain prognosis in the training cohort. [Figure 7B] Figure 1 shows the c-index of the 21 candidate 10x models of machine learning network 3 for patients with uncertain prognosis in the training cohort. [Figure 7C] Figure 1 shows the c-index of the 21 candidate 10x models of machine learning network 3 for patients with uncertain prognosis in the training cohort. [Figure 8A] Figure 1 shows the c-index of the 21 candidate 40x models of machine learning network 3 for patients with uncertain prognosis in the training cohort. [Figure 8B] Figure 1 shows the c-index of the 21 candidate 40x models of machine learning network 3 for patients with uncertain prognosis in the training cohort. [Figure 8C] Figure 1 shows the c-index of the 21 candidate 40x models of machine learning network 3 for patients with uncertain prognosis in the training cohort. [Figure 9A] Figure 1 shows the c-index of 21 candidate 10x models of the Inception v3 network for patients with uncertain prognosis in the training cohort. [Figure 9B] Figure 1 shows the c-index of 21 candidate 10x models of the Inception v3 network for patients with uncertain prognosis in the training cohort. [Figure 9C]Figure 1 shows the c-index of 21 candidate 10x models of the Inception v3 network for patients with uncertain prognosis in the training cohort. [Figure 10A] Figure 1 shows the c-index of 21 candidate 40x models of the Inception v3 network for patients with uncertain prognosis in the training cohort. [Figure 10B] Figure 1 shows the c-index of 21 candidate 40x models of the Inception v3 network for patients with uncertain prognosis in the training cohort. [Figure 10C] Figure 1 shows the c-index of 21 candidate 40x models of the Inception v3 network for patients with uncertain prognosis in the training cohort. [Figure 11A] For patients with uncertain prognosis in the training cohort, the c-index of the predicted probability of poor prognosis for the ensemble model is shown, thresholded at thresholds including 0.01, 0.02, etc. up to a maximum of 0.99. [Figure 11B] For patients with uncertain prognosis in the training cohort, the c-index of the predicted probability of poor prognosis for the ensemble model is shown, thresholded at thresholds including 0.01, 0.02, etc. up to a maximum of 0.99. [Figure 12] FIG. 1 shows a diagram identifying inclusions and exclusions of patients, slides, and slide images from the Ahus cohort. [Figure 13] FIG. 1 shows a diagram identifying inclusions and exclusions of patients, slides, and slide images from the Aker cohort. [Figure 14] FIG. 1 shows a diagram identifying the inclusion and exclusion of patients, slides, and slide images from the Gloucester cohort. [Figure 15] A diagram identifying the inclusion and exclusion of patients, slides, and slide images from the VICTOR cohort is shown. [Figure 16A]1 shows the primary and stage-specific analyses of DoMore-v1-CRC markers in the validation cohort, where (A) shows the results for all patients evaluated using Aperio AT2 images, (B) shows the results for all patients evaluated using NanoZoomer XR images, (C) shows the results for stage II patients evaluated using Aperio AT2 images, (D) shows the results for stage III patients evaluated using Aperio AT2 images, (E) shows the results for pN2 patients evaluated using Aperio AT2 images, and (F) shows the results for PT4 patients evaluated using Aperio AT2 images. [Figure 16B] 1 shows the primary and stage-specific analyses of DoMore-v1-CRC markers in the validation cohort, where (A) shows the results for all patients evaluated using Aperio AT2 images, (B) shows the results for all patients evaluated using NanoZoomer XR images, (C) shows the results for stage II patients evaluated using Aperio AT2 images, (D) shows the results for stage III patients evaluated using Aperio AT2 images, (E) shows the results for pN2 patients evaluated using Aperio AT2 images, and (F) shows the results for PT4 patients evaluated using Aperio AT2 images. [Figure 16C] 1 shows the primary and stage-specific analyses of DoMore-v1-CRC markers in the validation cohort, where (A) shows the results for all patients evaluated using Aperio AT2 images, (B) shows the results for all patients evaluated using NanoZoomer XR images, (C) shows the results for stage II patients evaluated using Aperio AT2 images, (D) shows the results for stage III patients evaluated using Aperio AT2 images, (E) shows the results for pN2 patients evaluated using Aperio AT2 images, and (F) shows the results for PT4 patients evaluated using Aperio AT2 images. [Figure 17A]Results of the DoMore-v1-CRC marker evaluated on Aperio AT2 slide images in the study cohort are shown. The conventional DoMore-v1-CRC marker is evaluated in A, C, and D, where A) is associated with all patients evaluated by DoMore-v1-CRC, C) is associated with stage II patients evaluated by DoMore-v1-CRC, and D) is associated with stage III patients evaluated by DoMore-v1-CRC. The binary DoMore-v1-CRC marker evaluated in B was created by averaging the predicted probabilities of two ensemble models of the DoMore v1 network (one at 10x and one at 40x) and thresholding the average at 0.58; the threshold was calculated using the same method applied to create the two ensemble markers from the two ensemble models (see the classification section in Example 1). [Figure 17B] Results of the DoMore-v1-CRC marker evaluated on Aperio AT2 slide images in the study cohort are shown. The conventional DoMore-v1-CRC marker is evaluated in A, C, and D, where A) is associated with all patients evaluated by DoMore-v1-CRC, C) is associated with stage II patients evaluated by DoMore-v1-CRC, and D) is associated with stage III patients evaluated by DoMore-v1-CRC. The binary DoMore-v1-CRC marker evaluated in B was created by averaging the predicted probabilities of two ensemble models of the DoMore v1 network (one at 10x and one at 40x) and thresholding the average at 0.58; the threshold was calculated using the same method applied to create the two ensemble markers from the two ensemble models (see the classification section in Example 1). [Figure 18] A diagram identifying the inclusion and exclusion of patients, slides, and slide images from the QUASAR 2 cohort is shown. [Figure 19]Kaplan-Meier curves using hazard ratios for patient groups predicted as having a good prognosis, an uncertain prognosis, and a poor prognosis are shown. A) Results of scans using the Aperio AT2, B) Results of scans using the NanoZoomer XR. Note: Cancer-specific survival is not a percentage, even though the label includes (%). DETAILED DESCRIPTION OF THE INVENTION
[0055] 1. Computer System of the Present Disclosure 1.1 First Example Computer System FIG. 1 illustrates a computer-implemented system 100 for determining a classifier 118 for one or more source histological images 102, such as source histopathological images.
[0056] As discussed below, the system 100 includes a network architecture 111 that applies a machine learning algorithm, including two neural networks 108, 116. When the system 100 is being trained, the received source histological images 102 include training histological images, such as training histopathological images, and the system 100 also processes received truth data 120 that represents known results ("ground truth") associated with the training histological images. For training purposes, the goal of applying the system 100 is to properly configure various trainable parameters (including those associated with the two neural networks 108, 116) that are used to accurately classify subsequently received source histological images 102, such as source histopathological images.
[0057] As also discussed below, the system 100 can be used to process one or more source histological images 102, such as source histopathological images for which the outcome ("ground truth") is not known. In this case, the system can apply neural networks 108, 116, constructed using training histological images (such as training histopathological images), such that the output of the system 100 is a classifier 118 for the received source histological images 102.
[0058] The specific descriptions in Sections 1.1 and 1.2 (and other sections) herein describe applications in which a single source histological image is divided into multiple tiles, with each tile being a subset of pixels from the single source histological image. In other examples, multiple source histological images can be processed so that a tile can be the entire source histological image or a subset of pixels from the source histological image. Optionally, multiple source histological images are processed so that each source histological image is divided into multiple tiles, with each tile being a subset of pixels from each single source histological image.
[0059] For example, without limitation, the multiple source histological images can be from the same histological specimen and / or from different histological specimens obtained from the same biological source. For example, the multiple source histological images can be from different histological specimens obtained from the same organism, where the different histological specimens can be, for example, from the same tissue or from different tissues within the same organism, from the same organ or from different organs within the same organism, or from the same structure or from different structures within the same organism.
[0060] Processing multiple source histological images allows the system to consider features present at different locations in a single biological source. Such locations can optionally represent different planes (e.g., parallel or substantially parallel planes, or intersecting planes) present in a single biological source. Thus, such an approach can be used to generate information about different locations within a single biological source. One such example is using images from multiple parallel or substantially parallel planes to generate three-dimensional information related to that biological source. Such information can be considered a "three-dimensional source histological image" and can also be considered an optional form of source histological image in the practice of the present invention.
[0061] Optionally, the three-dimensional source histological image may be constructed from multiple physically separated (typically serial parallel or essentially parallel) sections of biological material obtained from a biological source, e.g., as discussed in Section 2.4 of this application. Another means of obtaining a three-dimensional source histological image is to obtain multiple histological images from a "thick" histological section using selective focusing techniques, e.g., as discussed in Section 2.4 of this application.
[0062] Another such example is generating information about multiple distinct locations within a single biological source, for example to determine tumor heterogeneity. Further details about such processing are provided below.
[0063] 1.1.1 Training machine learning algorithms This section relates to training the machine learning algorithms of the system 100 of FIG.
[0064] The system 100 includes a tile generator 104 that receives one or more source histological images 102, such as one or more source histopathological images. When the system is being trained, the source histological images 102 may be referred to as training histological images. When the system is being trained using one or more source histopathological images 102, each image may be referred to as a training histopathological image. Various examples of source histological images are described herein, and it will be understood that the system 100 can process any type of histological image, including, but not limited to, any type of histopathological image. In some examples, the or each one or more source histological images 102 may be segmented before being provided to the tile generator 104. For example, as described in detail below with reference to FIG. 3, an optional segmentation block 122 may process a WSI histological image 124, such as a WSI histopathological image, to provide the source histological image 102. (WSI stands for whole slide image.) In some examples, the segmentation block 122 may itself be a neural network.
[0065] The tile generator 104 generates a plurality of tiles 106 from one or more source histological images 102, such as one or more source histopathological images. Each of the plurality of tiles 106 includes a plurality of pixels representing a region of the source histological image 102. In some examples, the plurality of tiles may be rectangular and may correspond to adjacent regions of one or more source histological images 102. In some examples, pixels of the or each source histological image 102 are included in only a single tile. As desired, tiles may be spaced apart, adjacent to one another, or overlapping one another. In some applications, portions of the or each source histological image 102 may not be included in any one pixel, such as when insufficient pixels are arranged around the source histological image 102 to form a complete tile.
[0066] The tile generator 104 (or another component of the system 100) then assigns a subset of the tiles 106 to a known "bag" in multiple-instance learning. The tiles may be assigned randomly to the bags. In one example, the tiles may be drawn randomly and uniformly without replacement. If a bag can fit all tiles in the image, all tiles are sampled in order. Otherwise, a sampling scheme may be applied that gives some tiles more weight than others based on some criteria (which may be changed during training).
[0067] Each bag represents a subset of tiles 106. The collection of all tiles in an image is denoted by I. A collection of tiles from the same image is called a bag, and a collection of bags is called a batch (or mini-batch). Individual tiles are not assigned ground truth (labels); instead, a bag of tiles inherits the ground truth of the original image. We denote a bag as a collection of tiles by B ⊆ I.
[0068] The first neural network 108, the pooling function 112, and the second neural network 116 can be collectively considered as a network architecture 111 for applying machine learning algorithms using multiple instance learning. One update step of training is summarized below and discussed in detail in the following description. 1. A batch of bags is input into the network architecture 111. 2. A first neural network (representation network) maps each tile 106 to a representation of the tile (tile features 110). 3. Aggregate the tile features 100 by a pooling function 112. 4. A second neural network (classification network) takes the pooled tile features (bag features 114) as input and generates predictions (classifier 18). 5. This prediction (classifier 118) is compared to the reference classification (truth data 120) using a loss function. 6. The derivative of the loss function with respect to the network's parameters is used to update each parameter of the network architecture 111. The first neural network 108 (representation network), the pooling function 112, and the second neural network 116 (classification network) can all have trainable parameters that are updated based on the derivative of the loss function. The entire network 111 can be trained end-to-end.
[0069] For the remainder of the explanation, we will ignore the batch dimension (implicitly assume a batch size of 1). Extending to batch sizes larger than 1 works as expected in typical deep learning settings with neural networks.
[0070] The first neural network 108 processes at least some of the tiles 106 to determine tile features 110 for these tiles 106. The tile features 110 are representations of the associated input tiles 106. The first neural network 108 may also be referred to as a representation network and may include a relatively lightweight neural network applied to each of the tiles 106. In one application, the first neural network 108 is implemented using the known MobileNetV2 network, as described further herein (e.g., in Section 1.3.4 of this application). It will be appreciated that any neural network capable of extracting features from the tiles 106 can be used as the first neural network 108. For example, any convolutional network can be used, although the choice of network may contribute to overall classification performance. Well-known examples of networks that can be used as the first neural network 108 include the VGG family, the Inception family, and ResNet.
[0071] A mechanism for training the first neural network 108, for example, by adjusting the weight values (or other trainable parameters) of the first neural network 108, is described below. In this example, the output of the first neural network 108 is not directly compared to truth data 120 representing the ground truth (i.e., known true results) associated with the source histological image 102.
[0072] The first neural network (which may also be called a representation network) is a network that computes the function
number
[0073] In this example, the first neural network 108 is applied to all tiles in bag B, and the expression R={f r (x;θ r ):x∈B}. Within the same update, all tiles in the bag, and all bags in the batch, are updated with θ r Note that we use the exact same first neural network 108 with the same values of . All representations in the batch can be computed and stored before the next step.
[0074] The tile features 110 for each of the plurality of tiles 106 are output by the first neural network 108 and provided to a pooling function 112. The pooling function 112 may combine the tile features 110 associated with the tiles for a bag of tiles to provide a bag feature 114 for each of the bags.
[0075] The pooling function 112 can reduce the set of tile features R 110 into a single representation in one bag B, typically using the function
number
[0076] These bag features 114 are suitable for processing by downstream multiple-instance learning (MIL) algorithms. In one application, the pooling function 112 can apply the well-known Noisy-AND pooling function to generate the bag features 114. In another example, the pooling function 112 can apply any reduction function, such as sum, mean, median, etc., to the tile features 110 (which can be considered as the input tiled representation). Other, more sophisticated examples that can be used include Noisy-OR, ISR, generalized mean, and LSE.
[0077] The second neural network 116 processes each of the bag features 114 to determine a classifier 118 for one or more source histological images 102, such as one or more source histopathological images. The second neural network 116 may be referred to as a classification network. In some examples, the second neural network 116 may be provided as a fully connected neural network.
[0078] The second neural network 116 (classification network) is characterized by:
number
number
number
[0079] During the training phase, the system 100 also receives truth data 120. The truth data 120 represents the ground truth (i.e., the known true result) associated with the or each source histological image 102, such as a source histopathological image. For example, without limitation, if the training phase involves the use of multiple source histological images from the same histological specimen and / or from different histological specimens obtained from the same biological source, the truth data 120 may represent the same ground truth (i.e., the same known true result) associated with the same biological source from which the or each histological specimen was obtained.
[0080] The full network 111 functions are:
number
[0081] Ideally, we want the tiles in the bag to span the entire image, but this can be limited by hardware constraints in some applications, often necessitating subsampling. In principle, tiles can be sampled from an image in many different ways, but assuming no prior knowledge, uniform random sampling without replacement is sufficient. With random subsampling of tiles, it is unlikely that an image will be represented by the same configuration of tiles each time. This has a regularizing effect on training, which may aid in generalization.
[0082] As shown above, bags are annotated with labels from the original image, and this assignment may not be fully justified if the tiles in the bag do not span the entire image. However, it has been found that the error introduced by assigning image labels to bags of tiles is smaller than assigning image labels to a single tile. This assumption implies that the approximation error decreases as the area represented by the tiles in the bag increases, from a bag containing one tile to a bag containing all tiles in the image.
[0083] It may be desirable to use as many tiles as possible to represent an image during training of the network 111; for large images, the number of tiles per bag may be limited by the memory of the hardware on which the method is executed. The first neural network 108 (representation network) may be the largest consumer of memory in this framework. In forward propagation, all tiles in a bag may be processed by the first neural network, but only the tile's representation (i.e., tile features 110) is further used in forward propagation. Storing and processing the tile features 110 may require significantly less memory than storing and processing the associated tile image 106.
[0084] Gradient-based optimization methods can use intermediate representations of each tile "within" the network 111 to update the network's parameters. This means that these intermediate representations (including the tile features 110) are stored until the associated gradients are calculated based on the output of the loss function 126. Reducing the number of tiles used in backpropagation can significantly reduce the memory footprint. Thus, in one application, the network 111 can use the entire bag B in the forward propagation of the second neural network 116, but only use a subset of the bag G ⊆ B in the backpropagation. Note that in such an application, it may be only the second neural network that employs this truncation of the gradient contribution. All tile features 110 from the bag can be used by the pooling function 112 and, therefore, the second neural network 116 to update the parameters (θ ) associated with the pooling function 112 and the first neural network 108. p,θc ) updates are not affected by the truncation in the first neural network.
[0085] All trainable parameters in the network 111 can be iteratively updated using this feedback mechanism.
[0086] As described in more detail below, the system may apply multiple iterations around a training loop to select final weighting values or other trainable parameters to be used as part of the machine learning algorithm 108. Thus, the purpose of applying training to the network 111 is to properly configure the network 111, as described above, so that it can accurately classify source histological images 102 for which the results ("ground truth") are not known, such as subsequently received source histopathological images.
[0087] In some examples, the training phase may involve processing data from one or more training cohorts. It will be appreciated that using a larger training cohort may result in a better system 100 for later determining a classifier 118 for source histological images 102 for which the outcome is not known. For example, the training cohort may include one or more source histological images from each of several different biological sources, which may be selected from at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 200, 2500, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or more different biological sources, where the ground truth is known for each biological source in the training cohort.
[0088] Following the processing of the training cohort, the system can apply an adjustment process using data from the training cohort and / or apply a validation process using data from the validation cohort. The adjustment cohort and / or validation cohort can, for example, include one or more source histological images from each of several different biological sources, and the number of different biological sources can be, for example, approximately the same as (e.g., ±50%, 40%, 30%, 20%, 10%, 5% or less), at least the same as, or greater than (e.g., at least 50%, 60%, 70%, 80%, 90%, 100% or more) the number of different biological sources in the training cohort. Exemplary details are described below, but these are not limiting, and those skilled in the art can use their well-known general knowledge to select suitable training cohorts, adjustment cohorts, and / or validation cohorts for use in this process.
[0089] 1.1.2 Using trained machine learning algorithms This section relates to applying network 111 of FIG. 1 to one or more unclassified source histological images 102, such as one or more unclassified source histopathological images for which the outcome ("ground truth") is not known. This is also referred to as inference. For example, network 111 may be applied to multiple source histological images present in a three-dimensional source histological image, as discussed further above.
[0090] That is, the system 100 can apply a network 111 configured using training histological images (e.g., training histopathological images) to determine a classifier 118 for the or each received source histological image 102. Processing blocks that function similarly to the training phase will not be described in similar detail here.
[0091] In this example, and other examples described herein, the system 100 used to process unclassified images is trained using training histological images and associated ground truth. More specifically, one or more of the first neural network, the pooling function, and the second neural network are trained using the training histological images and associated ground truth. As known in the art, training can be performed, by way of non-limiting example, until a maximum number of iterations around the loop have been performed or until the loss function reaches an acceptably low value. At that point, training can be considered complete. (This does not mean that training cannot be resumed at a later date, if desired.) For example, optimization can be performed for a fixed set of epochs (complete traversal of the dataset). Thus, if there are 100 images in the training set and the algorithm trains for 10 epochs with a batch size of 5, 200 optimization steps will be performed before termination. More generally, in numerical optimization, optimization can be terminated based on the loss function value. For example, (1) the value is below a certain threshold, (2) the absolute change in the value over the last k steps is below a certain threshold, or (3) the change in the value over the last k steps relative to the current value is below a certain threshold.
[0092] In the same manner as above, the tile generator 104 receives one or more source histological images 102 and generates a plurality of tiles 106 from the or each source histological image 102 .
[0093] The first neural network 108 processes the multiple tiles 106 to determine tile features 110 for the multiple tiles 106. The first neural network 108 may utilize a neural network that applies weighting values determined during a training phase. The first neural network 108 may be applied to all tiles in the image. This may be done one tile at a time, and each tile feature 110 may be stored until all tiles have been processed by the first neural network 108. Because each tile representation 110 may require only a small amount of memory, the number of tiles per image in inference is, for all practical purposes, nearly unlimited in terms of memory requirements.
[0094] The pooling function 112 aggregates the tile features 110 to provide bag features 114. Optionally, particularly when processing unclassified source histological images 102, the pooling function 112 may generate bags of bag features 114, with all tile features 110 included between them. That is, none of the tile features 110 may be excluded from the bag features 114.
[0095] A second neural network 116 processes the bag features 114 to determine a classifier 118 for the or each source histological image 102 .
[0096] Although bag sizes can vary between training and inference (when applied to unclassified image data), they have been found to produce reasonable results even in successfully trained networks. For applications in which the network 111 uses all tiles in an image for inference, the image is typically better represented in inference than in training. Having a relatively large bag size for training can be advantageous because it can enable the network 111 to learn features that can generalize across the entire image. In some examples, a bag containing more than 5%, 8%, 10%, 15%, or 20% of the tiles from the entire image can be considered a large bag. Some applications may also use relatively large bags of tiles during inference. For example, an algorithm may subsample tiles during inference to generate large bags, thereby speeding up classification.
[0097] It will be appreciated that when processing source histological images 102 that are unclassified due to the lack of associated truth data 120, the loss function 126 is not applied.
[0098] Optionally, as described in detail with reference to FIG. 3, one or more portions of the system 100 of FIG. 1 may be applied multiple times to determine multiple classifiers 118 for an image or collection of images, which may be combined to determine an overall classifier.
[0099] Advantageously, the system of FIG. 1 can accurately and efficiently classify histological images and can be used to classify the prognosis of, among other things, colon cancer.
[0100] 1.2 Second Computer System Example 2 illustrates a computer-implemented system 200 for determining an overall classifier 232 for one or more source histological images 202, such as one or more source histopathological images, which may optionally include determining an overall classifier 232 for a collection of source histological images 202, such as multiple source histological images present in a three-dimensional source histological image, as further discussed above.
[0101] As discussed below, the system 200 can be used to process one or more source histological images 202, such as one or more source histopathological images with unknown outcomes. During a training phase, the system can apply machine learning networks 211, 215, appropriately configured using training histological images (e.g., training histopathological images). The machine learning networks 211, 215 can be trained in any manner known in the art. In some examples, each of the machine learning networks 211, 215 can include the network 111 described with reference to FIG. 1 and can be trained similarly to that described with reference to FIG. 1. That is, each of the machine learning networks 211, 215 can include a first neural network, a pooling function, and a second neural network. Figure 3 below illustrates such a combined example of the systems of Figures 1 and 2.
[0102] In one embodiment, system 200 includes a first tile generator 204 and a second tile generator 205, both of which receive the same one or more source histological images 202. Similar to Figure 1, system 200 of Figure 2 can process any type of histological image, such as any type of histopathological image. Additionally, one or more source histological images 202 may be provided by an optional segmentation block 222 that processes one or more WSI histological images 224.
[0103] In this embodiment, a first tile generator 204 generates a plurality of first tiles 206 from the or each source histological image 202, and a second tile generator 205 generates a plurality of second tiles 207 from the or each source histological image 202.
[0104] In an alternative embodiment, system 200 includes a first tile generator 204 and a second tile generator 205 that receive different source histological images 202 (first source histological image and second source histological image, respectively) that were generated by obtaining histological images at different magnifications from the same histological source sample. First source histological image 202 and second source histological image 202 may differ at least with respect to (and typically only with respect to) the magnification applied to imaging the same histological sample during image generation.
[0105] In this alternative embodiment, a first tile generator 204 generates a plurality of first tiles 206 from a first source histological image 202, and a second tile generator 205 generates a plurality of second tiles 207 from a second source histological image 202.
[0106] In either embodiment, each of the plurality of first tiles 206 includes a plurality of pixels representing a region of the source histological image 202 having a first area and a first resolution, and each of the plurality of second tiles 207 includes a plurality of pixels representing a region of the source histological image 202 having a second area and a second resolution. The first area of the first tile 206 is greater than the second area of the second tile 207. In this manner, the first tile 206 represents a larger region of the source histological image 202, and therefore a larger region of the histological specimen shown in the source histological image 202, than the second tile 207.
[0107] The first tile 206 has a first resolution that is lower than the second resolution of the second tile 207. In this manner, the second tile 207 shows a finer, more detailed view of the source histological image 202 than the first tile 206. Thus, the second tile 207 shows a greater detail of the histological specimen shown in the source histological image 202.
[0108] In this manner, the first tile 206 can represent a sufficiently large area of a histological specimen to include structural information of the specimen, such as the general structure and / or orientation of tissues within the specimen, the size and / or shape of individual cells within the specimen, and / or the relative position of individual cells and / or groups of cells relative to other individual cells and / or groups of cells. More generally, the first tile 206 can be sized to include the organization of solid multicellular structures, such as organs, tissues, or other solid structures. This embodiment is therefore particularly well-suited for evaluating solid biological samples (e.g., without limitation, solid tumor samples) in which such supracellular structures are present. Such structural information can therefore include information such as tissue structure and / or differentiation patterns relevant to the classification (e.g., diagnosis or prognosis) of the specimen.
[0109] Because the area represented by the second tile 207 is smaller than the area represented by the first tile 206, some or all of this structural information may not necessarily be visible in the second tile 207. Furthermore, the second tile 207 may contain sufficient detail of the histological specimen to include information at the individual cell level, and in particular to indicate subcellular structures such as the presence, size, location, shape, density, and / or one or more other characteristics of the subcellular structures. Such subcellular structures include, but are not limited to, nuclei within cells; for example, the size and / or density of nuclei within cells observable in the second tile 207 may be of particular interest in the diagnosis and / or prognosis of a condition such as cancer, particularly solid cancers. Additional and / or alternative subcellular structures of interest in the second tile 207 may include one or more organelles and / or one or more other cellular components, such as individual proteins, DNA molecules, RNA molecules, lipids, and / or membranes.
[0110] Exemplary organelles and other macromolecules include, but are not limited to, endoplasmic reticulum (rough and / or smooth), Golgi apparatus, mitochondria, vacuoles, chloroplasts, acrosomes, autophagosomes, centrioles, cilia, cnidocysts, oculus apparatus, glycosomes, glyoxysomes, hydrogenosomes, lysosomes, melanosomes, mitosomes, myofibrils, nucleoli, ocelli, parentesomes, peroxisomes, proteasomes, ribosomes (80S), stress granules, TIGER domains, and / or vesicles. Some or all of these subcellular structures may not be visible in the first tile 206, typically because they do not have the required resolution.
[0111] In one embodiment, the first tile 206 contains information equivalent to that obtained by typical low-magnification light microscopy of histological images. In this example, the level of magnification used in the low-magnification light microscopy is 10x magnification, although other low-magnification levels can be used. For example, the level of magnification used in the low-magnification light microscopy can be at least 4x and less than 40x magnification, or can be approximately 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, 20x, 25x, 30x, or 35x magnification. In this context, the term "about" refers to a value of ±1x, 2x, 3x, 4x, or 5x of the stated magnification level, provided that the magnification level is suitable to provide some or all of the structural information for specimens of the types discussed above. In principle, the low magnification factor should simply be "low" in a relative sense compared to the magnification factor used for the second tile 207.
[0112] Additionally or alternatively, in this embodiment, the second tile 207 contains information equivalent to that obtained by typical high-magnification optical microscopy of histological images. In this example, the level of magnification used in the high-magnification optical microscopy analysis is 40x magnification, although other high-magnification levels can also be used. For example, the level of magnification used in the high-magnification optical microscopy analysis can be between about 20x and about 100x magnification, such as about 20x, 30x, 40x, 50x, 60x, 70x, 80x, 90x, or 100x magnification. In this context, the term "about" refers to a value that is ±1x, 2x, 3x, 4x, or 5x the stated magnification level, provided that the magnification is sufficient to provide information at the individual cell level and is capable of showing some or all of the subcellular structures, particularly those of the type discussed above. As a general rule, the high magnification need only be "high" in a relative sense, compared to the magnification used in the first tile 206.
[0113] It will be appreciated that the information observable in a histological image will depend not only on the level of magnification used to generate the image and / or the resolution of the image generated, but also on the staining and / or imaging technique used. For example, the staining used will affect the physical structures that are labeled and visible in the image, as discussed further in Section 2.4.1 of this application.
[0114] Importantly, both the first tile 206 and the second tile 207 can be represented by a sufficiently small amount of computer data that they can be properly processed without requiring too much processing power or too slowly by downstream processing blocks of the system 200. If the tiles were generated differently, with both the second tile 207 having a higher resolution and the first tile 206 having a larger area, the downstream processing blocks may require unacceptably high processing resources to properly perform.
[0115] In some applications, the first tile generator 204 and the second tile generator 205 can generate their respective tiles independently of each other. For example, the first tile 206 does not necessarily have to be centered on the same point on the source histological image 202. In fact, the first tile generator 204 and the second tile generator 205 do not need to have information about which other tile generators generate their tiles. In some examples, both the first tile 206 and the second tile 207 are each randomly oriented in a different area of the source histological image 202. Thus, the region of the source histological image 202 on which each of the first slides 206 is centered can be randomly selected, and the region of the source histological image 202 on which each of the second tiles 207 is centered can also be randomly selected, thereby making the random selection of the first tile 206 and the second tile 207 independent of each other.
[0116] Subsequent application of the two machine learning networks 211, 215 does not require coordination of how the two sets of tiles are generated or how they relate to each other to generate accurate classifiers 218, 219. Thus, advantageously, the system 200 of FIG. 2 can avoid the need for coordination between the two tile generators 204, 205 and the two machine learning networks 211, 215. The system 200 includes a machine learning network 211 that processes a plurality of first tiles 206 to determine a first classifier 218 for the or each source histological image 202. The system 200 also includes a machine learning network 215 that processes a plurality of second tiles 207 to determine a second classifier 219 for the or each source histological image 202. In this manner, the second classifier 219 can be based on information represented in the second tiles 207, and the first classifier 218 can be based on information represented in the first tiles 206. These machine learning networks 211, 215 can each include a single neural network or multiple neural networks as shown in FIG.
[0117] The system also includes a classifier combiner 230 that combines the first classifier 218 and the second classifier 219 to determine an overall classifier 232 for the or each source histological image 202. Optionally, this may include determining an overall classifier 232 for a collection of source histological images 202, such as multiple source histological images, present in a three-dimensional source histological image, as further discussed above.
[0118] The combination can be implemented in several ways, including mathematical combination, logical combination, or a combination of mathematical and logical combination. For example, the first classifier 218 and the second classifier 219 can have numerical values, and the overall classifier 232 can be the average of the first classifier 218 and the second classifier 219. Alternatively, the first classifier 218 and the second classifier 219 can have logical values, and the overall classifier 232 can be a logical combination of the first classifier 218 and the second classifier 219. Such a logical function can be an AND function, which sets the overall classifier 232 to the same logical value as the first classifier 218 and the second classifier 219 only if they have the same logical value. Further details regarding a specific implementation of the classifier combiner 230 are described with reference to FIG. 3.
[0119] 2 can classify one or more source histological images 202 in an improved manner based on different characteristics that are well represented in only one of the set of tiles (either the large area first tile 206 or the high resolution second tile 207). Also, because relevant information and features can be extracted from the or each source histological image 202 without requiring tiles with both large area and high resolution, the classification can be considered efficient in terms of the amount of processing resources required.
[0120] 1.3 Third Computer System Example 3 illustrates a particular implementation of a system 300 for determining an overall classifier 332 for one or more source histological images 302, such as one or more source histopathological images, e.g., multiple source histological images present in a three-dimensional source histological image, as discussed further above. As discussed below, similar to the system of FIG. 2, the system 300 of FIG. 3 generates a first tile 306 having a relatively large area (compared to a second tile 307), generates a second tile 307 having a relatively high resolution (compared to the first tile 306), and also has a classifier composer 330. The system 300 also includes a machine learning network 311, which includes a representation network 308 (which may also be referred to as a first neural network) and a classification network 316 (which may also be referred to as a second neural network), which are similar to the corresponding components of FIG. 1.
[0121] The following discussion relates to the use of the system 300 of FIG. 3 for externally evaluating a deep learning model for predicting cancer-specific survival from colon cancer tissue sections represented by WSI histopathological images 324.
[0122] 1.3.1 Training Cohort This study utilized four training cohorts: the Ahus, Aker, Gloucester, and VICTOR cohorts, which are further described in the following subsections. Patients in the training cohorts were labeled as having a definite or unclear prognosis depending on their age at surgery and follow-up data. Patients with a definite prognosis consisted of patients defined as having a favorable prognosis and those defined as having a poor prognosis. Patients were defined as having a favorable prognosis if they were younger than 85 years at the time of surgery, had more than 6 years of follow-up after surgery, and had no documented cancer-specific death or recurrence. The availability of recurrence data varied between cohorts and was particularly limited in the Gloucester cohort. For the Ahus cohort, favorable prognosis patients were required to have no documented metastases (local recurrence records were unavailable), whereas for Aker, Gloucester, and VICTOR patients, documented local or metastatic recurrence was not required. Patients were defined as having a poor prognosis if they were under 85 years of age at the time of surgery and died of cancer-specific disease between 100 days (inclusive) and 2.5 years (exclusive) after surgery. Patients who did not meet the criteria for a good or poor prognosis were defined as having an unclear prognosis.
[0123] 1.3.1.1 Ahus cohort Figure 12 shows a diagram specifying the inclusion and exclusion of patients, slides, and slide images from the Ahus cohort, as well as the prognosis of included patients. CSS, cancer-specific survival; CSD, cancer-specific death.
[0124] From a consecutive series of 219 patients with colorectal adenocarcinoma treated at Akershus University Hospital, Norway, between 1988 and 2000, 172 patients had stage I, II, or III disease and had formalin-fixed, paraffin-embedded (FFPE) tissue blocks available. Three-micrometer sections of each FFPE tumor tissue block were stained with hematoxylin and eosin (H&E) and prepared as histological slides by laboratory technicians at Cancer Genetics and Informatics (ICGI), Oslo University Hospital, Norway. A pathologist confirmed the presence of tumor in each tissue section, and 12 patients without tumor slides were excluded (shown in Figure 12). Tumor tissue slides were scanned using two scanners: the Aperio AT2 (Leica Biosystems, Germany) and the NanoZoomer XR (Hamamatsu Photonics, Japan). Scans were read using the Python interface (version 1.1.1) of OpenSlide 3.4.1, available at https: / / openslide.org / . An automated segmentation method (discussed below) was applied to identify tumors in the 320 slide images, and each slide image was divided into multiple non-overlapping regions, called tiles, using two resolutions, designated 40x and 10x (see the Tiling section below). The 160 patients whose tiles were included within the tumor segmentation were defined as the Ahus cohort, and Figure 12 shows the prognosis of these patients (see definitions of clear and non-clear prognosis above).
[0125] 1.3.1.2 Aker cohort Figure 13 shows a diagram identifying the inclusion and exclusion of patients, slides, and slide images from the Aker cohort, as well as the prognosis of included patients. CSS, cancer-specific survival; CSD, cancer-specific death.
[0126] One slide from each of 578 stage I, II, or III patients treated for primary colorectal cancer at Aker University Hospital, Norway, and analyzed by Danielsen and colleagues between 1993 and 2003 was processed in the same manner as the Aker cohort. Three slides had damaged coverslips that prevented them from being scanned with the NanoZoomer XR scanner, and the automated segmentation method did not identify tumors in three Aperio AT2 slide images and two NanoZoomer XR slide images; the other patients comprised the Aker cohort (Figure 13).
[0127] 1.3.1.3 Gloucester cohort Figure 14 shows a diagram identifying patient inclusions and exclusions, slides and slide images from the Gloucester cohort, and the outcomes of included patients. CSS, cancer-specific survival; CSD, cancer-specific mortality.
[0128] The Gloucester Colorectal Cancer Study recruited 1,036 patients between 1988 and 1996, of which 19 were excluded due to synchronous multiple cancers (Figure 14). The remaining 1,017 patients were processed in the same manner as the Ahus cohort, resulting in 969 patients using Aperio AT2 segmentation and 967 patients using NanoZoomer XR segmentation (Figure 14). These patients comprised the Gloucester cohort, except for one patient who was excluded from the Aperio AT2 10x tile set due to a missing tile in the tumor segmentation (Figure 14).
[0129] 1.3.1.3 VICTOR cohort Figure 15 shows a diagram identifying patient inclusions and exclusions, slides and slide images from the VICTOR cohort, and the prognosis of included patients. CSS, cancer-specific survival; CSD, cancer-specific mortality.
[0130] The VICTOR trial randomized patients with stage II and III colorectal cancer who received rofecoxib or placebo after first-line treatment to investigate cardiovascular adverse events. For 795 patients recruited between 2002 and 2004, H&E-stained 3-μm sections were collected from FFPE tissue blocks, some of which were sectioned at ICGI and some elsewhere. Sections were processed in the same manner as in the Ahus cohort. The VICTOR cohort consisted of 767 patients using the Aperio AT2 40x tile, 764 patients using the Aperio AT2 10x tile, 761 patients using the NanoZoomer XR 40x tile, and 756 patients using the NanoZoomer XR 10x tile (Figure 15).
[0131] 1.3.2 Segmentation As shown schematically in FIG. 3, a source histological image 302 is generated from a WSI histological image 324, such as a WSI histological image, by applying an image segmentation method.
[0132] The segmentation method in this example includes a process for creating a probability map from an input image, and a different process for creating an image segmented into foreground and background regions based on the input image and the corresponding probability map.
[0133] Figure 4 shows a diagram of an example segmentation network architecture. Each layer is represented by its name, output height, output width, and number of output channels. Progression is downward from the input image at the top to the predicted output at the bottom. The probability map can be generated by the segmentation network in Figure 4, which is based on the DeepLab network (Chen et al., IEEE Trans Pattern Anal Mach Intell, 2018, 40:834-848). Final segmentation can be achieved using dense conditional random fields (Krahenbuhl, & Koltun, Adv Neural Inf Process Syst, 2011;24:109-117).
[0134] The method was first trained on 1,077 images with corresponding annotations from the Aker cohort (670 images) and the VICTOR cohort (407 images). Images were obtained from slides scanned with a NanoZoomer (Hamamatsu Photonics, Japan) scanner, and annotations were hand-drawn by a pathologist. This trained method was then applied to images from the Aker, Ahus, and Gloucester cohorts. The resulting segmentations were validated by a pathologist, who corrected any insufficient segmentations. This set of images with corresponding (potentially corrected) masks constitutes the development dataset for the image segmentation method.
[0135] From the 1717-patient development dataset, 25% (429 patients) were randomly and uniformly sampled to form the calibration set, and the remaining 1288 patients comprised the training set. The training set included images from 358 patients with cancer-specific events and 930 patients without. The calibration set included images from 128 patients with cancer-specific events and 301 patients without. Slides from patients in the segmentation development set were scanned with both the Aperio AT2 and NanoZoomer XR scanners. Thus, the development set for the segmentation task consisted of 3430 scans (4 scans missing from the NanoZoomer XR scanner), the training set consisted of 2573 scans, and the calibration set consisted of 857 scans.
[0136] Each scan was digitally resized (see the Tiling section below) to a size corresponding to 2.5x the resolution and stored as a PNG image. Each image was then resized using a Catmull-Rum cubic filter to fit within a 1600 x 1600 pixel frame. This was done by resizing the image while maintaining its aspect ratio until the largest dimension (height or width) was 1600 pixels. A new image was then formed by padding the resized image along its shortest dimension to 1600 pixels. The center of the resized image was aligned with the center of the padded image, and the padded image was used further.
[0137] The segmentation network is trained over 100,000 update steps (training iterations), each using 16 images (this collection is called a mini-batch) distributed across four GPUs. Every image in the development dataset is used once before being used twice, meaning that each image is processed approximately 622 times during training (one progression through the dataset is called an epoch). In each epoch, the same image is used once, with slight variations each time. First, a 641x641 pixel slice is cropped at a random position within the image. Then, a series of directional distortions are applied in the following order: 1.50% chance to flip the image horizontally (mirror it along the horizontal axis). 2.50% chance to flip the image vertically (mirror it along the vertical axis). 3.50% chance to rotate the image once by one of the following angles: 0, 90, 180, or 270.
[0138] Finally, we center the image to its mean and standard deviation (see https: / / www.tensorflow.org / versions / r1.10 / api_docs / python / tf / image / per_image_standardization). We send the resulting image as an RGB image to the segmentation network.
[0139] Trainable parameters are initialized using the Xavier weight initialization scheme and updated using standard stochastic gradient descent optimization (Glorot & Bengio, Understanding the difficulty of training deep feedforward neural networks. In Proc 13th Int Conf Artif Intell Stat, Vol. 9 249-256 (2010)). The optimization step length is initialized to 0.05 and reduced by a factor of 0.1 over 96488 iterations (approximately 600 training epochs).
[0140] When the trained network is applied to an image, it generates a probability map with the same spatial shape as the image: a single-channel grayscale image with intensity values of 0, 1, ..., 255. The method assigns high values to regions it deems likely to depict cancerous tissue.
[0141] For each image, we create additional versions of the image by rotating and flipping the original image and then applying the trained network to all the different versions. There are eight versions, which we obtain from the original image by the following operations: 1. Do nothing (this is the original image) 2. Flip the image around the horizontal axis 3. Flip the image around the vertical axis 4. Rotate the image 90 degrees clockwise 5. Rotate the image 180 degrees clockwise 6. Rotate the image 270 degrees clockwise 7. Rotate the image 90 degrees clockwise and flip the result around its horizontal axis 8. Rotate the image 270 degrees clockwise and flip the result around its horizontal axis
[0142] The resulting probability map is restored to its original orientation and the average image of all the different versions is calculated and used further in the processing.
[0143] For inference, we apply the trained network to one image at a time (i.e., using a batch size of 1) and, contrary to the training phase, apply no cropping or orientation distortion. However, it is important to center all images around their mean and standard deviation, as we did in training. The network was implemented and run in Python 3.5 (https: / / www.python.org) using TensorFlow 1.10 (https: / / www.tensorflow.org).
[0144] Probability map segmentation was performed using the Python library pydensecrf v1.0rc3 (https: / / github.com / lucasb-eyer / pydensecrf). The model used a unary potential (probability map), a Gaussian pairwise potential (addPairwiseGaussian(sxy=1,compat=1)), and a bilateral pairwise potential (addPairwiseBilateral(sxy=30,srgb=3,compat=100)). The resulting float-valued (0, 1) image was thresholded at 0.5 to create a binary mask, where pixels with values below 0.5 were labeled as background and the rest as foreground.
[0145] The resulting segmentation is smoothed with a 5x5 averaging filter, and then foreground regions with eight connected neighbors less than 20,000 pixels are removed. Background regions that are completely contained within a foreground region are marked as foreground.
[0146] The method was applied to the training set for every 4,000 iterations, and the predicted segmentations were assessed against the reference segmentations. The model with the highest average bookmakers informedness score was then selected as the model to use in the remaining experiments.
[0147] Figure 5 shows the performance of the segmentation method on the training set. As shown above, the method is evaluated over multiple training iterations, evenly spaced throughout the training progression, with each iteration being 4,000 iterations. Figure 5 shows that the model at iteration 88,000 (reference number 534 in Figure 5) achieved the best score of 0.902.
[0148] Returning to FIG. 3, the output of the segmentation method is a source histopathological image 302 .
[0149] 1.3.3 Tiling The regions identified as tumors by the segmentation method are source histopathological images 302, which in this example are not directly suitable for use as input to a convolutional neural network (CNN) due to limited GPU memory on commonly available hardware. Therefore, in this process, multiple non-overlapping regions of fixed size (i.e., regions segmented as tumors in each slide image), called tiles, were generated from within the source histological image 302. It will be appreciated that a comparable approach can be employed to generate tiles from any region of interest within the source histological image 302 (e.g., in source histological images taken or not taken from histological samples of any other type of cancer and / or any other pathological condition).
[0150] Figure 3 shows a plurality of first tiles 306 and a plurality of second tiles 307. Similar to Figure 2, the first tiles 306 have a relatively large area and a relatively low resolution, and the second tiles 307 have a relatively small area and a relatively high resolution.
[0151] Because the physical area represented by a pixel can vary depending on factors such as the scanner used to acquire the image, tiles representing the same physical area were created by including slightly different numbers of pixels in tiles from the Aperio AT2 and NanoZoomer XR slide images. At the highest resolution, designated 40x, the physical size of a pixel in the Aperio AT2 slide image was 0.253 μm / pixel in both the vertical and horizontal directions, while the physical size of a pixel in the NanoZoomer XR slide image was 0.227 μm / pixel in both the vertical and horizontal directions. To create the 40x tile (the second tile at higher resolution 307), a 486 × 486 pixel tile was extracted from within the tumor segmentation in the Aperio AT2 slide image, and a 542 × 542 pixel tile was used for the NanoZoomer XR slide image. Similarly, a tile size of 1942 x 1942 pixels was used for the Aperio AT2 slide images, and 2166 x 2166 pixels was used for the NanoZoomer XR slide images to create the 10x tiles (lower resolution first tiles 306). Each of these raw tiles was then resampled to 512 x 512 pixels, making the physical area of each pixel similar for both scanners: 0.240 x 0.240 μm for the 40x tiles and 0.960 x 0.960 μm for the 10x tiles.
[0152] In this example, tiling was performed by defining a grid of candidate tiles from the upper left corner of the WSI histological image 324, including the area outside the tumor segmentation (represented in the source histological image 302). Candidate tiles whose four corners along the edge and their midpoints were within the boundary of the segmentation were included as either the first tile 306 or the second tile 307. Tiles were extracted from level 0 in OpenSlide, converted to numpy arrays, resized in OpenCV using the resize() function with interpolation set to cv2.INTER_CUBIC for upsampling or cv2.INTER_AREA for downsampling (https: / / docs.opencv.org / 3.4.0 / da / d54 / group__imgproc__transform.html), and saved in a lossless format (as PNG files).
[0153] 1.3.4 Patient Survival Prediction Method - Using Machine Learning Networks 311 The machine learning network 311 was trained using all patients in the training cohort with a clear prognosis. As described in detail below, the machine learning network 311 was trained five times with a 40x tile (models 1-5 in the second tile 307 of Figure 3) and five more times with a 10x tile (models 6-10 in the first tile 306), each time using resampled tiles of 512 x 512 pixels. The ground truth (i.e., the true outcome represented by the truth data 320) applied by these supervised classification methods was the clear prognosis for the patient, either a good prognosis or a poor prognosis (as defined in the Training Cohort section above).
[0154] The machine learning network 311 is a multi-instance classification method that includes a representation network 308 (corresponding to the first neural network in FIG. 1), a pooling function 312 (corresponding to the pooling function in FIG. 1), and a classification network 316 (corresponding to the second neural network in FIG. 1).
[0155] Figure 6 illustrates this implementation of the machine learning network architecture of Figure 3. The left side shows an overview of the progression from the input bag of tiles 614 to the bag prediction (classifier 618 output by the machine learning network). The right side shows the architecture of the representation network 608 (corresponding to the first neural network in Figure 1), with each layer represented by its name, output height, output width, and number of output channels.
[0156] Returning to Figure 3, the machine learning network 311 does not classify a single tile, but rather a collection of tiles, called a bag, where all tiles in the bag originate from the same scanned image (WSI histological image 324 / source histological image 302). Each tile in the bag is applied to a representation network 308 to generate tile features 310 for the tile (note that within a single update step, all tiles can use the same representation network with the same parameter values). All tile features can be aggregated, and a pooling function 312 can generate a single value for each class. A final classification network 316 is then applied to generate predictions. These predictions are compared using a loss function to ground truth (represented by truth data 320) corresponding to the source histological image 302 from which the bag originated.
[0157] In this example, a gradient-based optimization routine is used to optimize the loss function, and at each training iteration, the trainable parameters of the machine learning network 311 are updated according to this optimization method. In this example, only a randomly selected subset of the tiles in the bag are used to update the machine learning network 311. This asymmetric forward and back propagation reduces the memory footprint of the machine learning network 311 and allows the bag of tiles to grow during training.
[0158] As noted above, the representation network 308 in this implementation is based on the MobileNetV2 network, the details of which are shown in Figure 6 (Sandler, M., Howard, A., Zhu, M., Zhmoginov, A. & Chen, L. MobileNetV2: Inverted Residuals and Linear Bottlenecks. IEEE Conf Comput Vis Pattern Recognit, pp. 4510-4520 (2018)). The first convolutional layer of the representation network uses a 3x3 convolution kernel with a stride of 2. The activation function is the ReLU activation function (Glorot, X., Bordes, A. & Bengio, Y. Deep Sparse Rectifier Neural Networks. Proc 14th Int Conf Artif Intell Stat, Vol. 15, pp. 315-323 (2011)). Within each inverse bottleneck module, the first convolutional layer uses a 1x1 convolutional kernel with stride 1 and a ReLU6 activation function (Krizhevsky, A. Convolutional Deep Belief Networks on CIFAR-10. Available at https: / / www.cs.toronto.edu / ~kriz / conv-cifar10-aug2010.pdf (2010)). The depthwise separable convolutional layer uses a 3x3 convolutional kernel. If the spatial size is halved, the stride is 2, and otherwise it is 1. The activation function is a ReLU6 function. The last convolutional layer uses a 1x1 convolutional kernel with stride 1 and an identity activation function. If the number of input channels to an inverse bottleneck module is equal to the number of output channels within the same module, the input to the first convolutional layer within the module is added to the result of the last convolutional layer within the module. The convolutional layer after the inverse bottleneck module uses 1x1 convolution with stride 1 and ReLU activation function.All convolutional and separable convolutional layers described above employ batch normalization on the convolution results before applying the activation function (Ioffe, S. & Szegedy, C. Batch normalization: accelerating deep network training by reducing internal covariate shift. Proc. 32nd Int. Conf. Mach. Learn, Vol. 37, pp. 448-456 (2015)). All kernel weights are initialized with Xavier initialization, and no bias parameters are used. The final convolutional layer uses a 1x1 convolution kernel with stride 1. This layer does not use batch normalization, and the activation is the identity function. The remainder of the network consists of a Noisy-AND pooling function followed by a softmax classification, following the design of Kraus and colleagues (Kraus, O. Z., Ba, J. L. & Frey, B. J. Classifying and segmenting microscopy images with deep multiple instance learning. Bioinformatics 32, i52-i59 (2016)). There is one cross-entropy loss function associated with the output of the pooling function, and one cross-entropy loss function associated with the classification output.
[0159] The machine learning network 311 network was trained with a batch size of 32 bags and distributed across eight GPUs, each with four bags. Each bag consisted of 64 tiles, each 512 × 512 × 3 pixels in size, with values 0, 1, . . . , 255. Eight tiles contributed to the gradient calculation. To update the network parameters, the Adam optimization method was used with an initial step size of 0.001 (Kingma, D.P. & Ba, J. Adam: A Method for Stochastic Optimization. Available at https: / / arxiv.org / abs / 1412.6980 (2015)). When training with 10x tiles (the first tile 306), the training rate was initially set to 0.001, then reduced to 0.1 at iteration 6,000, then reduced again at iteration 12,000, and finally finished at iteration 15,000. Training was performed on a 40x tile (second tile 307) using twice the number of iterations, i.e., the training rate started at 0.001, decreased to 0.1x at 12,000 iterations, decreased again at 24,000 iterations, and training was terminated after 30,000 iterations.
[0160] At each step, each tile is warped and normalized before entering the machine learning network 311. First, the tile is randomly cropped to a size of 448x448, then its orientation is warped. The tile is randomly flipped from left to right (around the central vertical axis), then randomly flipped from top to bottom (around the central horizontal axis), and finally randomly rotated by either 0°, 90°, 180°, or 270°. Next, the value is scaled to (0,1) by casting it to a 32-bit floating-point number, and then the entire tile is divided by 255.0. Next, the tile is converted from RGB color space to HSV color space, and then each channel is scaled by a uniformly distributed value from 1 / 1.1 to 1.1. The tile is then converted back to RGB. Finally, we normalize the tiles to have zero mean and unit norm (see rgb_to_hsv, hsv_to_rgb, per_image_standardization at https: / / www.tensorflow.org / versions / r1.10 / api_docs / python / tf / image for details).
[0161] For inference, no cropping is applied, so the entire tile, 512 × 512 × 3 pixels in size, is evaluated by the machine learning network 311. Also, orientation and color distortions are not applied. As in training, each tile is normalized to have a zero mean and unit norm before entering the machine learning network 311. The network was implemented and run in Python 3.5 (https: / / www.python.org) using TensorFlow 1.10 (https: / / www.tensorflow.org). To account for class imbalance in the training set, minority classes within each cohort and scanner combination were oversampled, ensuring an equal number of images labeled as good prognosis and poor prognosis across all cohort and scanner combinations. Images were randomly and uniformly sampled without replacement within each cohort and scanner combination.
[0162] 1.3.5 Another network used for comparison - Inception v3 network Another network, the Inception v3 network (Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J. & Wojna, Z. Rethinking the Inception Architecture for Computer Vision. Proc 2016 IEEE Conf Comput Vis Pattern Recognit, pp. 2818-2826 (2016)), was used to obtain classification results for comparison with the results of the machine learning network 311 in Figure 3.
[0163] An Inception v3 network was trained in Keras (2.1.6) using the Tensorflow Docker image (tensorflow / tensorflow:1.9.0-gpu-py3). The input image size was 512 × 512, and the output was two classes: the first class was the probability of a good prognosis, and the second class was the probability of a poor prognosis. The binary cross-entropy loss function was used, and optimization was performed with keras.optimizers.Adam using default arguments, except for the initial training rate, which was set to 0.0001. To account for class imbalance between tiles with good and poor prognosis, minority-class tiles were oversampled for each cohort before training, and the file paths were saved as a list. As a result, each cohort contained an equal number of tiles with good and poor prognosis, although some tiles could be included twice. Before training, we loaded a list of tiles, randomly shuffled them, and then loaded a batch of images using 16 worker threads using a modified version of keras.preprocessing.image.ImageDataGenerator. We modified ImageDataGenerator to perform color distortion in the following way: 1. Convert the tiles to HSV color space, 2. Enhance the hue by adding randomly uniformly sampled values between ±0.05, 3. Scale the saturation by a randomly and uniformly sampled value between 1 / 1.1 and 1.1. 4. Shift the saturation between ±0.1 with randomly and uniformly sampled values. 5. Scale the value by a randomly and uniformly sampled value between 1 / 1.1 and 1.1. 6. Shift the value by a random, uniformly sampled value between ±0.1. 7. Convert the tiles back to RGB color space.
[0164] Next, tiles were normalized by subtracting the mean color value and dividing by the standard deviation across all tiles used for training, i.e., all tiles for patients with a clear prognosis in the training cohort. Due to GPU memory constraints, a batch size of 16 tiles was used for each training iteration. When training with 10x tiles, the learning rate was initially set to 0.0001 and then halved every 25,000 iterations, starting at 25,000 iterations, until training ended after 150,000 iterations. Training with 40x tiles (the second tile) was performed using twice the number of iterations, i.e., the learning rate started at 0.0001 and halved every 50,000 iterations, starting at 50,000 iterations, until training ended after 300,000 iterations. The network output was the predicted probability of a poor prognosis for the tile. The predicted probability of a poor prognosis for a patient was calculated by averaging the predicted probabilities for all tiles for that patient.
[0165] 1.3.6 Individual Models Twenty training runs were performed by training each of the two networks ((i) the machine learning network 311 in Figure 3 and (ii) the Inception v3 network) five times at each of the two resolutions. For each of these 20 training runs, 21 models were evaluated for all patients with uncertain prognosis in the training cohort. The 21 models evaluated in each training run were uniformly distributed from one-third of the iterations until the end of training (inclusive). Each 10x model of the machine learning network 311 was evaluated for 5,000, 5,500, etc. iterations up to a maximum of 15,000 iterations. Each 40x model of the machine learning network 311 was evaluated for 10,000, 11,000, etc. iterations up to a maximum of 30,000 iterations. Each 10x model of the Inception v3 network was evaluated for 50,000, 55,000, etc. iterations up to a maximum of 150,000 iterations. Each 40x model of the Inception v3 network was evaluated for 100,000 iterations, 110,000 iterations, etc. up to a maximum of 300,000 iterations.
[0166] To reduce evaluation time for the 40x models, a random sample of 2,000 40x tiles was selected for each slide containing more than 2,000 40x tiles. The same tiles were evaluated for all models. To further reduce evaluation time for the 40x models of machine learning network 311, patients with more than 50 tiles were evaluated using 50 tiles at a time. These evaluations ignored tiles ordered after the last multiple of 50, i.e., a maximum of 49 tiles for each patient. Note that these speedups were applied only during model selection; all tiles were evaluated for all applications of the selected models, including the external evaluation described in this paper.
[0167] For each training run, we selected the model that maximized Harrell's concordance index (c-index) (Harrell, FE, Jr, Califf, RM, Pryor, DB, Lee, KL, & Rosati, RA. Evaluating the yield of medical tests. J Am Med Assoc. 247, 2543-2546 (1982)). The c-index compared the observation time until cancer-specific death or censoring with the model's predicted probability of poor prognosis for patients with an uncertain prognosis in the training cohort. Thus, the model with the largest c-index appeared to provide the most prognostic information in terms of predicted probability when the prognosis was assessed as uncertain in the training cohort.
[0168] Figures 7-10 show the c-index of all candidate models and indicate the model selected for each of the 21 training runs.
[0169] Figure 7 shows the c-index of the 21 candidate 10x models of machine learning network 311 for patients with uncertain prognosis in the training cohort. Subplots a–e show training runs 1–5. The points referenced 736a–e indicate the selected models. Points including and following 738a–e (excluding 736a–e, of course) indicate models that were not selected. For comparison, the c-index of nine models from the first third of each training run (models prior to those referenced 738a–e) is shown.
[0170] Figure 8 shows the c-index of the 21 candidate 40x models of machine learning network 311 for patients with uncertain prognosis in the training cohort. Subplots a–e show training runs 1–5. The points referenced 836a–e indicate the selected models. Points including and following 838a–e (excluding 836a–e, of course) indicate models that were not selected. For comparison, the c-index of nine models from the first third of each training run (models prior to those referenced 838a–e) is shown.
[0171] Figure 9 shows the c-index of 21 candidate 10x models for the Inception v3 network for patients with uncertain prognosis in the training cohort. Subplots a–e show training runs 1–5. Points referenced as 936a–e indicate the selected models. Points including and following 938a–e (excluding 936a–e, of course) indicate models that were not selected. For comparison, we show the c-index of nine models from the first third of each training run (models prior to those referenced as 938a–e).
[0172] Figure 10 shows the c-index of 21 candidate 40x models for the Inception v3 network for patients with uncertain prognosis in the training cohort. Subplots a–e show training runs 1–5. The points referenced 1036a–e indicate the selected models. Points including and following 1038a–e (excluding 1036a–e, of course) indicate models that were not selected. For comparison, we show the c-index of nine models from the first third of each training run (models prior to those referenced 1038a–e).
[0173] 1.3.6 Ensemble Models Ensemble models were created for each network and resolution by averaging the predicted probabilities of poor patient outcomes for the five selected models. Four ensemble models were created: 10x and 40x ensemble models for machine learning network 311, and similar 10x and 40x ensemble models for the Inceptionv3 network.
[0174] Referring to FIG. 3, five instances of a machine learning network 311 are shown as models 6-10. Each of these instances includes a first neural network 308 and a second neural network 316. Each of these instances is then applied to a plurality of first tiles 306 to determine five instances of a first classifier 318 for the source histological image 302. In FIG. 3, the values of the five instances of the first classifier 318 are 0.3142, 0.1930, 0.2533, and 0.2451, respectively. The classifier combiner 330 calculates a statistical representation of these five instances of the first classifier 318—in this example, the average value—to determine an averaged first classifier 340. In the example of FIG. 3, the value of the averaged first classifier 340 is 0.2468, referred to as the 10x ensemble prediction.
[0175] Similarly, five instances of machine learning network 311 are shown as Models 1-5. Each of these instances includes a first neural network 308 and a second neural network 316. Each of these instances is then applied to multiple second tiles 307 to determine five instances of second classifier 319 for the source histological image 302. In FIG. 3, the values of the five instances of second classifier 319 are 0.2972, 0.3325, 0.3025, 0.5958, and 0.3112, respectively. System 300 then calculates a statistical representation of these five instances of second classifier 319—in this example, the average value—to determine averaged second classifier 341. In the example of FIG. 3, the value of averaged second classifier 341 is 0.3678, referred to as the 40x ensemble prediction.
[0176] In this manner, the classifier combiner 330 can apply a statistical function (such as an averaging function) to multiple classifiers to determine a combined classifier (such as an averaged classifier 340, 341).
[0177] In this example, the classifier combiner 330 applies a first thresholding function 342 to the averaged first classifiers 340 to determine a thresholded first classifier. In this example, the first thresholding function 342 applies a single threshold to the averaged first classifiers 340 so that the thresholded first classifiers have a binary value. Examples of how the single threshold can be set are described below. In FIG. 3, the thresholded first classifiers can have a value of "predicted good" or "predicted poor." This process can therefore be considered as dichotomizing the predicted probability of poor prognosis for the ensemble model. In other examples, the first thresholding function 342 can apply multiple thresholds to the averaged first classifiers 340 so that the thresholded first classifiers can take on one of three or more discrete values.
[0178] Similarly, the classifier combiner 330 applies a second thresholding function 343 to the averaged second classifier 341 to determine a thresholded second classifier. In this example, as described above, the second thresholding function 343 applies a single threshold to the averaged second classifier 341 such that the thresholded second classifier has a binary value. Examples of how the single threshold may be set are described below. Again, the thresholded second classifier may have a value of "predicted good" or "predicted bad." In other examples, the second thresholding function 343 may apply multiple thresholds to the averaged second classifier 341 such that the thresholded second classifier can take on one of three or more discrete values.
[0179] In this manner, the classifier combiner 330 may apply a thresholding function (including one or more thresholds) to the classifiers (optionally averaged classifiers) to determine a thresholded classifier.
[0180] In this example, the thresholded classifier determined was a prognostic indicator for the patient from which the evaluated histological image was obtained. A comparable approach can be employed to create an ensemble model that determines a classifier (optionally averaged) for any other outcome, and it will be understood that the result will depend on the ground truth associated with the source histological images used during the training phase.
[0181] In one implementation, to determine a suitable threshold for dichotomizing each ensemble model's predicted probability of poor outcome, the c-index of the dichotomized ensemble model predictions is predicted for patients with uncertain outcomes in the training cohort against thresholds including 0.01, 0.02, etc. up to 0.99. The threshold for obtaining the maximum c-index can be selected for each ensemble model.
[0182] Figure 11 shows the c-index of the predicted probability of poor prognosis for the ensemble models, thresholded at thresholds including 0.01, 0.02, etc. up to a maximum of 0.99, for patients with an uncertain prognosis in the training cohort. Plot a shows the 10x ensemble model of the machine learning network 311 in Figure 3. If the ensemble model's predicted probability of poor prognosis is greater than 0.51, the predicted outcome is poor prognosis. Otherwise, the predicted probability is less than or equal to 0.51, and the predicted outcome is good prognosis. This threshold (which may be referred to as a dichotomous marker) may be referred to as the 10x ensemble marker of the machine learning network 311. Plot b shows the 40x ensemble model of the machine learning network 311 of Figure 3. The threshold identified in plot b may be referred to as the 40x ensemble marker of the machine learning network 311, which in this example is defined as a threshold of 0.56. Plot c shows the 10x ensemble model of the Inception v3 network. The 10x ensemble marker for the Inception v3 network is defined as a threshold of 0.54. Plot d shows the 40x ensemble model of the Inception v3 network. The 40x ensemble marker of the Inception v3 network is also defined as a threshold of 0.54.
[0183] Returning to FIG. 3 , the classifier combiner 330 then combines the thresholded first classifier and the thresholded second classifier to determine an overall classifier 332 for the histological source image 202. In this example, the classifier combiner 330 performs a logical combination of the thresholded first classifier and the thresholded second classifier. If both the thresholded first classifier and the thresholded second classifier represent the same result, the classifier combiner sets the overall classifier 332 to a value that is the same as the thresholded first classifier and the thresholded second classifier. If the thresholded first classifier and the thresholded second classifier represent different results, the classifier combiner sets the overall classifier 332 to a value that is different from both the thresholded first classifier and the thresholded second classifier. As shown in FIG. 3 , If the thresholded first classifier and the thresholded second classifier represent a good prediction, the classifier combiner 330 sets the overall classifier 332 to "agreement with good prognosis." If the thresholded first classifier and the thresholded second classifier represent a poor prediction, the classifier combiner 330 sets the overall classifier 332 to "agree on poor prognosis." If one of the thresholded first classifier and the thresholded second classifier represents a predicted good, and the other represents a predicted bad, the classifier combiner 330 sets the overall classifier 332 to "mismatch."
[0184] In this way, if the 10x ensemble marker and the 40x ensemble marker predict the same outcome, the patient is defined as predicted to have a good prognosis (if both ensemble markers predict a good prognosis) or a poor prognosis (if both ensemble markers predict a poor prognosis), and if the 10x ensemble marker and the 40x ensemble marker predict different outcomes, the patient is defined as predicted to have an uncertain prognosis, thereby creating a 10x ensemble model and a 40x ensemble model for the machine learning network 311 (and also the Inception v3 network). Optionally, if one of the ensemble markers for a patient cannot be analyzed, for example, because there is no 10x tile, the combination of the 10x ensemble marker and the 40x ensemble marker is also not defined. Therefore, such patients are excluded from the analysis of the combined model.
[0185] This results in two combinations of the 10x ensemble markers and the 40x ensemble markers, one for the machine learning network 311 and one for the Inception v3 network. These three variables grouped together can be referred to as DoMore v1 markers and Inception v3 markers.
[0186] The computer systems disclosed herein may include a computer-readable storage medium, memory, a processor, and one or more interfaces, all linked together via one or more communication buses. Exemplary computer systems may take the form of conventional computer systems, such as desktop computers, personal computers, laptops, tablets, smartphones, smartwatches, virtual reality headsets, servers, mainframe computers, and the like. In some embodiments, exemplary computer systems may be embedded in microscopy devices, such as virtual slide microscopes capable of whole-slide imaging.
[0187] The computer-readable storage medium and / or memory may store one or more computer programs (or software or code) and / or data. The computer programs stored on the computer-readable storage medium may include an operating system for execution by a processor to enable the computer system to function. The computer programs stored on the computer-readable storage medium and / or memory may include computer programs according to embodiments of the present invention or computer programs that, when executed by a processor, cause the processor to perform methods according to embodiments of the present invention.
[0188] A processor may be any data processing unit suitable for executing one or more computer-readable program instructions, such as those belonging to a computer program stored on a computer-readable storage medium and / or memory. As part of the execution of one or more computer-readable program instructions, the processor may store data in and / or read data from a computer-readable storage medium and / or memory. A processor may include a single data processing unit or multiple data processing units operating in parallel or in cooperation with each other. In a particularly preferred embodiment, the processor may include one or more graphics processing units (GPUs). GPUs are suitable for the types of calculations involved in training and using machine learning algorithms such as those disclosed herein. As part of the execution of one or more computer-readable program instructions, the processor may store data in and / or read data from a computer-readable storage medium and / or memory.
[0189] The one or more interfaces may include a network interface that enables the computer system to communicate with other computer systems over a network. The network may be any type of network suitable for transmitting or communicating data from one computer system to another. For example, the network may include one or more of a local area network, a wide area network, a metropolitan area network, the Internet, a wireless communication network, etc. The computer system may communicate with other computer systems through the network via any suitable communication mechanism / protocol. The processor may communicate with the network interface via one or more communication buses and cause the network interface to transmit data and / or commands to another computer system over the network. Similarly, the one or more communication buses may enable the processor to operate on data and / or commands received by the computer system via the network interface from other computer systems over the network.
[0190] The interface may alternatively or additionally include a user input interface and / or a user output interface. The user input interface may be arranged to receive input from a user or operator of the system. The user may provide this input via one or more user input devices (not shown), such as a mouse (or other pointing device, trackball, or keyboard). The user output interface may be arranged to provide graphical / visual output to the user or operator of the system on a display (or monitor or screen) (not shown). The processor may direct the user output interface to form image / video signals that cause the display to display the desired graphical output. The display may be touch-sensitive, allowing the user to provide input by touching or pressing the display.
[0191] According to embodiments of the present invention, the interface may alternatively or additionally include an interface to a digital microscope or other microscope system. For example, the interface may include an interface to a virtual microscope device capable of Whole Slide Imaging (WSI). In WSI, a virtual slide is generated by high-resolution scanning of a glass slide with a slide scanner. The scan is typically performed piecewise, and the resulting images are stitched together to form a single, very large image at the maximum magnification possible by the scanner. These images may be on the order of 100,000 x 200,000 pixels in size, or may contain billions of pixels. According to some embodiments, the computer system can control the microscope device via the interface to scan a slide containing a specimen. Thus, the computer system can acquire microscopic images of a histological specimen from the microscope device received via the interface.
[0192] It will be understood that the computer system architectures described above are merely exemplary, and that systems using alternative components or having different architectures using more (or fewer) components may be used instead.
[0193] 2. Histological images and their sources As used herein, the term "histological image" refers to an image of a histological specimen showing the microscopic structure of biological material. A "source histological image" is a histological image of a histological specimen obtained from a defined source of biological material. The defined source of biological material may be, for example, an ex vivo sample of biological material. Histological images may be used in accordance with the present invention, for example, for training purposes or for inference purposes, as further described herein.
[0194] Histological images can be obtained, for example, by light microscopy of the histological specimen, at magnification levels equivalent to low or high magnification, e.g., as discussed in more detail in Section 1.2 of this application. However, those skilled in the art will understand that other means of imaging a histological specimen, including other forms of microscopy, may be used to generate a histological image.
[0195] Histological images can also be conveniently generated by generating whole slide images (WSIs), such as by techniques conventional in the art, as discussed in, for example, Farahani et al., Pathology and Laboratory Medicine International, 2015, 7:23-33, the contents of which are incorporated herein by reference. WSIs, also commonly referred to as "virtual microscopes," typically aim to emulate traditional optical microscopes in a computer-generated manner. In practice, WSIs typically involve two processes. The first process typically uses dedicated hardware (scanners) to digitize images of histological specimens (usually provided on glass slides) and generate large, graphical digital images (so-called "digital slides"). The second process typically employs dedicated software (e.g., so-called virtual slide viewers) to view and / or analyze the resulting large digital files. As discussed in Farahani et al. 2015 (supra), various commercially available WSI instruments have been developed over the past decade.A list of common WSI systems and their respective vendors includes 3DHistech (Pannoramic SCAN II, 250 Flash), DigiPath (PathScope), Hamamatsu (NanoZoomer RS, HT, and XR), Huron (TISSUEscope 4000, 4000XT, HS), Leica (ScanScope AT, AT2, CS, FL, SCN400), formerly known and operated as Aperio, Mikroscan (D2), Olympus (V S120-SL), Omnyx (VL4, VL120), PerkinElmer (Lamina), Philips (Ultra-Fast Scanner), Sakura Finetek (VisionTek), Unic (Precice 500, Precice 600x), Ventana (iScan Coreo, iScan HT), formerly known and operated as Bioimagene, and Zeiss (Axio). These devices are intended to meet the needs of a diverse user base. A list of differences between selected WSI systems is provided in Table 2 of Farahani et al. 2015 (supra). Preferred devices include the device used in this example, including the NanoZoomer XR scanner and / or the Apiero AT2 scanner.
[0196] Thus, in a preferred embodiment, the or each source histological image may be a WSI.
[0197] In a preferred embodiment, the methods, particularly the training methods, of the present invention include acquiring multiple source histological images, such as WSI, of a histological specimen stained with a marker using at least two different pieces of image scanning equipment. By avoiding the use of a single image scanning equipment, the trained machine learning algorithm is not trained to use only images from that single image scanning equipment, and therefore should be better able to process images from a variety of image scanning equipment. Introducing more scanners in training can advantageously improve the generalization of subsequent inference.
[0198] If different pieces of image scanning devices are used, the method of the present invention may further include the step of aligning the images, which may be performed, for example, by a scale-invariant feature transform, commonly referred to as a SIFT transform.
[0199] Each microscope image is preferably a grayscale or color image consisting of one, two, or three color channels. Most preferably, each microscope image is a color image consisting of three color channels. Thus, three samples are provided per pixel. The samples are coordinates in a three-dimensional color space. Suitable 3D color spaces include, but are not limited to, RGB, HSV, YCbCr, and YUV.
[0200] A "histological feature of interest" in a histological image refers to a microstructural feature present in the histological image, such as a WSI. Without limitation, the feature may be of interest, for example, for diagnostic or therapeutic purposes, or for scientific research.
[0201] Histological specimens are generally used to review the structure and determine a diagnosis or prognosis for the subject from whom the histological specimen was taken.
[0202] Histological specimens can be obtained from any biological source. They can be, for example, from any organism, from any tissue, organ, or other structure within an organism, and from healthy and / or pathological samples. Sources of particular interest are discussed further below.
[0203] As discussed in more detail in Section 1 of this application, and as further defined by the claims, this application describes a computer-implemented system (100, 200, 300) for determining a classifier (118, 318), or an overall classifier (232, 332), for a source histological image (102, 202, 302), which may be a source histopathological image.
[0204] The source histological images (e.g., WSIs) used in the training phase are images of histological specimens obtained from sources for which the ground truth is known. Each of the source histological images is then paired with associated truth data that represents the ground truth for each source from which each histological specimen was obtained. In the examples described herein, the histological specimens were obtained from cancer patients, and the ground truth was related to the categorization of each patient into a prognostic group, as further described herein.
[0205] When using a computer-implemented system (100, 200, 300) that has already been trained, during the training phase, the source histological images (e.g., WSI) may be images of histological specimens obtained from sources for which ground truth is unknown. The unknown ground truth may be, for example, but is not limited to, diagnostic or prognostic information associated with the source (e.g., subject) from which the histological specimen was obtained. The computer-implemented system (100, 200, 300) can then be used to determine a classifier (118, 318) or overall classifier (232, 332) for the source histological images (102, 202, 302) associated with the unknown ground truth, for example, by performing a diagnostic or prognostic assessment of the source from which the source histological images were obtained.
[0206] 2.1 Biology The biological source of the histological specimen from which the source histological image (102, 202, 302) is obtained can be, for example, from any organism, preferably a cellular organism, more preferably a multicellular organism. The biological source can be, for example, an animal, such as a human or a non-human animal, such as a primate, non-human primate, laboratory animal, farm animal, livestock, or household pet.
[0207] It may be most preferred that the biological source is a human subject.
[0208] Exemplary non-human animals include, optionally, poultry (such as chickens, turkeys, geese, quails, or ducks), livestock (such as cattle, sheep, goats, or pigs, alpacas, bantengs, bison, camels, cats, deer, dogs, donkeys, gayal, guinea pigs, horses, llamas, mules, rabbits, reindeer, buffalo, and yaks), and other animals, including zoo animals, captive animals, and game animals; fish (including freshwater and saltwater fish, farmed fish, and ornamental animals); and other animals, including zoo animals, captive animals, and game animals. Also included are marine and aquatic animals (including, but not limited to, crustaceans such as oysters, mussels, clams, shrimp, prawns, lobsters, spiny lobsters, crabs, cuttlefish, octopus, squid, etc.), domestic animals (such as cats and dogs), rodents (such as mice, rats, guinea pigs, hamsters, etc.), and horses, as well as other domestic, wild, and farm animals, such as mammals, marine animals, amphibians, birds, reptiles, insects, and other invertebrates.
[0209] In alternative embodiments, the organisms used in the biological source may include biological material obtained from non-animal sources, such as plants, fungi, or Monera (e.g., bacteria or archaea).
[0210] 2.2 Tissues, organs and other structures The biological source may be derived from any tissue type, organ or other structure of interest derived from a selected organism, such as from a human subject or any other multicellular organism, as discussed above in Section 2.1 of this application.
[0211] For example, the biological source may be derived from tissue present within a human or non-human animal, including, for example, one or more of the epithelium, connective tissue, muscle, and / or nervous system of a selected organism. It may be particularly preferred that the biological source be derived from tissue present within a human subject.
[0212] Thus, the biological source may be derived from or include tissue residing within the epithelium of a selected organism, such as from a human subject.
[0213] Epithelium is a continuous sheet of cells (one or more layers thick) that covers the exterior of the body and lines bodily tracts that connect internal closed cavities with the external environment (digestive, respiratory, and urogenital tracts), constitutes the secretory portion of glands and their ducts, and is found in the sensory receptor areas of certain sensory organs (e.g., ear and nose). Epithelium covering and lining surfaces (e.g., skin) are involved in absorption (e.g., intestine), secretion (e.g., glands), and can be sensory (e.g., neuroepithelium) or contractile (e.g., myoepithelial cells). Epithelium is typically a continuous sheet of cells that covers the surfaces of the body. There are two main types of epithelium: covering epithelium and glandular epithelium.
[0214] The covering epithelium can include squamous epithelium (e.g., the endothelial lining of blood vessels and the mesothelial lining of body cavities), cuboidal epithelium (e.g., tissue lining small ducts and / or tubules such as the salivary glands or kidneys), columnar epithelium (such as cells lining the stomach, cervix, and / or intestine), pseudostratified epithelium, stratified epithelium (such as keratinized (i.e., skin) or non-keratinized (i.e., esophagus) forms of stratified epithelium).
[0215] Glandular epithelium is present in glands, which are organized collections of secretory epithelial cells. Most glands form during development by proliferation of epithelial cells so that they project into the underlying connective tissue. Some glands maintain continuity with the surface through ducts and are known as exocrine glands. Other glands lose this direct continuity with the surface as their ducts degenerate during development. These glands are known as endocrine glands.
[0216] Thus, the biological source of a source histological image for use in the present invention may consist of, consist essentially of, or include any one or more of the aforementioned types of epithelium, or any other type of epithelium of interest.
[0217] Additionally and / or alternatively, the biological source of the histological image may be derived from or include tissues present within the connective tissue of a selected organism, such as a human subject. Connective tissue is composed of cells and extracellular matrix. The extracellular matrix is composed of fibers of protein and polysaccharide matrix, secreted and organized by cells in the extracellular matrix. Changes in the composition of the extracellular matrix determine the properties of connective tissue. For example, when the matrix mineralizes, it can form bone or teeth. Specialized forms of extracellular matrix also constitute tendons, cartilage, and the cornea of the eye. Typical connective tissues are either loose or dense, depending on the arrangement of the fibers. Cells reside within a matrix composed of glycoproteins, fibrous proteins, and glycosaminoglycans secreted by fibroblasts, and the primary component of the matrix is, in fact, water.
[0218] Connective tissue can be, for example, in the form of proper connective tissue (e.g., loose irregular connective tissue and / or dense irregular connective tissue) or in the form of specialized connective tissue, examples of which include dense regular connective tissue found in tendons and ligaments, cartilage, adipose tissue, hematopoietic tissue (such as bone marrow, lymphatic tissue), blood, and bone.
[0219] Thus, the biological source of a source histological image for use in the present invention may consist of, consist essentially of, or include any one or more of the aforementioned types of connective tissue, or any other type of connective tissue of interest.
[0220] Additionally and / or alternatively, the biological source of the source histological image may be derived from or include tissue present within the muscle of a selected organism, such as from a human subject. The muscle tissue may be either striated or smooth muscle. The muscle tissue may be in the form of skeletal or cardiac muscle (both of which are striated), or smooth muscle (such as the muscle tissue found in the walls of most blood vessels and in tubular organs such as the intestine).
[0221] Thus, the biological source of a source histological image for use in the present invention may consist of, consist essentially of, or include any one or more of the aforementioned types of muscle tissue, or any other type of muscle tissue of interest.
[0222] Additionally and / or alternatively, the biological source of the source histological image may be derived from or include tissue present within the nervous system of a selected organism, such as from a human subject. The nervous system includes the central nervous system (CNS), which is made up of the brain and spinal cord, as well as the peripheral nervous system (PNS), which is made up of all nervous tissue outside the CNS, including cranial nerves from the brain, spinal nerves from the spinal cord, and nodules known as ganglia, which contain neuronal cell bodies.
[0223] Thus, the biological source of a source histological image for use in the present invention may consist of, consist essentially of, or include any one or more of the aforementioned types of nervous system, or any other type of tissue of interest within the nervous system.
[0224] Optionally, the biological source of a source histological image for use in the present invention can consist of, consist essentially of, or include biological material obtained from any one or more of the following organs of a selected organism, such as from a human subject: -Skeletal system, joints, ligaments, musculature, and / or tendons. - The digestive system, including the mouth (e.g., teeth and / or tongue), salivary glands (e.g., parotid, submandibular, and / or sublingual glands), pharynx, esophagus, stomach, small intestine (e.g., duodenum, jejunum, and / or ileum), large intestine, liver, esophagus, mesentery, pancreas, anal canal, and / or anus. -Respiratory system, including the nasal cavity, pharynx, larynx, trachea, bronchi, lungs and / or diaphragm. - Urinary system, including kidneys, ureters, bladder, and / or urethra. Reproductive organs, such as the female or male reproductive organs. The female reproductive system includes the internal reproductive organs (such as the ovaries, fallopian tubes, uterus, and vagina), the external reproductive organs (such as the vulva and clitoris), and the placenta. The male reproductive system includes the internal reproductive organs (such as the testes, epididymis, vas deferens, seminal vesicles, prostate, and bulbourethral glands) and the external reproductive organs (such as the penis and scrotum). -The endocrine system, including the pituitary gland, pineal gland, thyroid gland, parathyroid glands, adrenal glands, and pancreas. -Circulatory system including the heart, patent foramen ovale, arteries, veins, and capillaries. -Lymphatic system, including lymphatic vessels, lymph nodes, bone marrow, gut-associated lymphoid tissue including the thymus, spleen, and tonsils. - the brain (including the cerebrum (e.g., cerebral hemispheres) and diencephalon), brainstem (including the midbrain, pons, and medulla oblongata), cerebellum, spinal cord, and ventricular system, including the choroid plexus; the peripheral nervous system, such as nerves (e.g., cranial nerves, spinal nerves, ganglia, and enteric nervous system). - the eye and its components (e.g., the cochlea, iris, cilia, lens and / or retina), the ear or its components (e.g., the outer ear such as the earlobe, the middle ear such as the tympanic membrane and ossicles, the inner ear such as the cochlea, vestibule, and / or semicircular canals), the olfactory epithelium, and the tongue (including taste buds). - Integumentary system, such as mammary glands, skin, and / or subcutaneous tissue.
[0225] 2.3 Healthy or diseased samples Histological specimens can be obtained from healthy or diseased samples.
[0226] The biological source of a source histological image for use in the present invention may consist of, consist essentially of, or include biological material obtained from a healthy or pathological sample.
[0227] When biological material is obtained from a pathological sample, the histological specimen obtained therefrom may be referred to as a "histopathological" sample, and the image obtained from the defined source histopathological sample may be referred to as a "source histopathological image." As used herein, the term "pathological" refers to an unhealthy state, including any medical condition, disorder, or disease.
[0228] Such histopathological samples may contain or be suspected of containing biological material that constitutes a pathological condition. For example, the biological source of a histopathological sample may be a subject, such as a human subject, that has a pathological condition, has been diagnosed with a pathological condition, is suspected of having a pathological condition, is being treated for a pathological condition, has previously been treated for a pathological condition, and / or has previously had a pathological condition.
[0229] For example, if the pathological condition is cancer and the biological source is a subject, such as a human subject, that has cancer, has been diagnosed with cancer, is suspected of having cancer, is being treated for cancer, has previously been treated for cancer, and / or has previously had cancer, the histopathological sample may contain or be suspected of containing biological material including cancerous cells.
[0230] A source histopathological image obtained from a histopathological sample may be identified as including an image containing biological material, where the biological material (a) comprises, consists essentially of, or consists of biological material having a pathological condition, and / or (b) comprises, consists essentially of, or consists of biological material altered by a pathological condition.
[0231] This confirmation step may be performed in a variety of ways. Those skilled in the art are aware of many different approaches for confirming the presence of a pathological condition within a biological material in a histological sample obtained from the biological material and / or in a histological image obtained therefrom. For example, but not by way of limitation, the confirmation step may be performed by human evaluation (e.g., by a trained pathologist). As exemplified herein, but not by way of limitation, in the context of source histopathological images obtained from a human subject with cancer, a pathologist was used to confirm whether each tissue section contained tumors, although it will be understood that an equivalent computer-implemented evaluation of the presence of tumor material in a tissue sample can also be used.
[0232] Thus, a histological specimen may preferably comprise biological material obtained from a subject having or suspected of having a pathology.
[0233] The medical condition may be, for example, a disease. More specifically, the disease may be a disease selected from the group consisting of an infectious disease, a deficiency disease, a genetic disease (including both genetic and non-genetic genetic diseases), and a physiological disease.
[0234] The disease may be a communicable disease. Alternatively, the disease may be a non-communicable disease.
[0235] As described above in Section 2.2 of this application, the disease may optionally be present in one or more tissues and / or organs, or other body parts.
[0236] Exemplary diseases include, but are not limited to, diseases of genetic origin, diseases resulting from chemical and / or physical injury, diseases of immune origin (including immunodeficiencies and immune responses in the absence of infection), diseases of biological origin (e.g., viral, rickettsial, bacterial, and diseases caused by fungi and other parasites), diseases associated with abnormal growth of cells, particularly cancer (including, but not limited to, hyperplasia, benign tumors, and malignant tumors), diseases of metabolic-endocrine origin, diseases of nutrition (e.g., including diseases of nutritional over- and / or under-nutrition), diseases of neuropsychiatric origin (including, for example, neurological disorders such as Alzheimer's disease, Huntington's chorea, and Parkinson's disease), and diseases of aging.
[0237] Diseases may optionally be acute, chronic, malignant, or benign. Acute disease processes usually begin suddenly and end quickly. Chronic diseases often begin very slowly and then persist for a long period of time. The terms benign and malignant are most often used to describe tumors and may be used in a more general sense. Benign diseases are generally uncomplicated and usually have a good prognosis (outcome). Malignant tumors refer to processes that, if left untreated, cause fatal illness. Cancer is a general term for all malignant tumors.
[0238] 2.3.1 Cancer In one embodiment of particular interest to the present invention, the biological source of the source histological image for use in the present invention may consist of, consist essentially of, or include biological material obtained from a subject, such as a human subject, who has cancer, has been diagnosed with cancer, is suspected of having cancer, is being treated for cancer, has previously been treated for cancer, and / or has previously had cancer.
[0239] The biological material may be obtained from the site of a primary, secondary, or any other tumor within the subject's body, or may be obtained from a site that is local, regional, or distal to the site of such a known tumor.
[0240] A histopathological sample may contain or be suspected of containing biological material that includes one or more cancerous cells.
[0241] A source histopathological image obtained from a histopathological sample can be confirmed to contain an image of biological material containing one or more cancerous cells prior to use in the present invention.
[0242] The present invention can be used to evaluate any type of cancer.
[0243] Tumors are often assigned a grade and a stage. The stage of a solid tumor refers to its size or extent and whether it has spread to other organs or tissues. The tumor grade (cancer grade) indicates how quickly the tumor is likely to grow and spread.
[0244] For example, a cancer may be stage 0, stage I, stage II, stage III, or stage IV, or any one or more subdivisions thereof. The characteristics of these different stages and their subdivisions are well known in the art. Generally, however, stage 0 indicates that the cancer is local (in situ) and has not spread. Stage I indicates that the cancer is small and has not spread anywhere else. Stage II indicates that the cancer has grown but has not spread. Stage III indicates that the cancer is larger and may have spread to surrounding tissues and / or lymph nodes (part of the lymphatic system). Stage IV indicates that the cancer has spread from where it began to at least one other organ in the body and is also known as "secondary" or "metastatic" cancer.
[0245] In one embodiment, the cancer may be staged according to the TNM staging system, which is a system developed and maintained by the AJCC and the Union for International Cancer Control (UICC). This is the most commonly used staging system by medical professionals worldwide. The TNM staging system was developed as a tool for physicians to stage various types of cancer based on specific standardized criteria. The TNM staging system is based on the extent of the tumor (T), the extent of lymph node metastasis (N), and the presence of metastases (M). -T category describes the original (primary) tumor. TX refers to the inability to assess the primary tumor. TO refers to no evidence of a primary tumor. Tis refers to carcinoma in situ (early cancer that has not spread to adjacent tissues). T1, T2, T3, and T4 relate to the size and / or extent of the primary tumor. -N category describes whether the cancer has reached nearby lymph nodes. NX indicates that the regional lymph nodes cannot be assessed. N0 indicates no regional lymph node spread (no cancer found in the lymph nodes). N1, N2, and N3 indicate regional lymph node involvement (number and / or extent of spread). - The M category indicates whether there is distant metastasis (spread of cancer to other parts of the body). M0 indicates no distant metastasis (cancer has not spread to other parts of the body). M1 indicates distant metastasis (cancer has spread to distant parts of the body).
[0246] Each type of cancer has its own classification system, so letters and numbers do not always mean the same thing for all types of cancer. Once the T, N, and M are determined, they can be combined to assign an overall stage: 0, I, II, III, or IV. These stages may also be subdivided using letters, such as IIIA and IIIB. Further guidance can be found at www.https: / / cancerstaging.org.
[0247] For some cancer types, nonanatomic factors can be taken into account when assigning an anatomic stage / prognostic group. These are clearly defined in the chapters of the AJCC Cancer Staging Manual (e.g., Gleason Score in Prostate). These factors remain purely anatomical and are collected separately from the T, N, and M used to assign stage groups. When nonanatomic factors are used for grouping, a grouping definition is provided for cases where nonanatomic factors are unavailable (X) or where it is desirable to assign a group ignoring nonanatomic factors.
[0248] Stage I cancer is the slowest growing and often has a better prognosis. Higher stage cancers are often more aggressive but can still often be successfully treated.
[0249] Additionally and / or alternatively, the cancer may be of a particular grade. The determination of malignancy is typically based on cell differentiation (the degree of similarity to normal cells). The particular grade of cancer may be, for example, grade I, grade II, grade III, or grade IV cancer, or a combination of two or three categories. The characteristics of these different grades are well known in the art. Generally, however, grade I cancer is a type of cancer in which cells resemble normal cells and do not grow rapidly; grade II cancer is a type of cancer in which cancer cells do not look like normal cells and grow faster than normal cells; and grades III and IV are types of cancer in which cancer cells look abnormal and may grow or spread more aggressively. Growth characteristics can be assessed in some cancers, for example, based on the frequency of cell division.
[0250] The cancer can be, for example, a type of cancer selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed type cancer.
[0251] Carcinoma refers to a malignant neoplasm of epithelial origin, or a cancer of the internal or external layers of the body. Carcinomas, which are malignant tumors of epithelial tissue, account for 80 to 90 percent of all cancer cases. Epithelial tissue is found throughout the body. It is present in the skin, as well as in the covering and lining of organs and internal passageways such as the gastrointestinal tract, as discussed further above.
[0252] Carcinomas can be divided into two main subtypes: adenocarcinoma, which arises in glandular organs, and squamous cell carcinoma, which arises in squamous epithelium.
[0253] Adenocarcinomas generally arise in the mucous membranes and initially appear as thickened, plaque-like white mucosa. Adenocarcinomas often spread easily through the soft tissues in which they develop. Squamous cell carcinomas can occur in many parts of the body.
[0254] Most carcinomas affect organs or glands that can secrete fluid, such as the breast, which produces milk, or the lungs, which secrete mucus, or the colon, or the prostate, or the bladder.
[0255] In one embodiment, the cancer may be a cancer arising from epithelial cells selected from breast cancer, basal cell carcinoma, adenocarcinoma, gastrointestinal cancer, lip cancer, mouth cancer, esophageal cancer, small intestine and stomach cancer, colon cancer, liver cancer, bladder cancer, pancreatic cancer, ovarian cancer, cervical cancer, lung cancer, and skin cancer, including squamous cell and basal cell carcinoma, prostate cancer, renal cell carcinoma, and other known cancers that affect epithelial cells throughout the body.
[0256] Sarcomas are cancers that begin in the supportive and connective tissues, such as bone, tendons, cartilage, muscle, and fat. Sarcomas, which typically occur in young adults, often present as painful masses on the bones. Sarcomas usually resemble the tissues in which they grow.
[0257] Examples of sarcomas include osteosarcoma or osteogenic sarcoma (bone), chondrosarcoma (cartilage), leiomyosarcoma (smooth muscle), rhabdomyosarcoma (skeletal muscle), mesothelioma or mesothelioma (membranous lining of body cavities), fibrosarcoma (fibrous sarcoma) hemangioendothelioma (blood vessels), liposarcoma (fatty tissue), glioma or astrocytoma (neurogenic connective tissue found in the brain), myxosarcoma (primitive embryonic connective tissue), mesenchymal or mixed mesodermal tumor (mixed connective tissue type).
[0258] Myeloma is a cancer that begins in the plasma cells of the bone marrow, which produce some of the proteins found in the blood.
[0259] The cancer may be solid or liquid.
[0260] For example, leukemia ("liquid cancer" or "blood cancer") is cancer of the bone marrow (the site of blood cell production). The disease is often associated with the overproduction of immature white blood cells. These immature white blood cells do not function as well as they should, and patients are often susceptible to infections. Leukemia can also affect red blood cells, causing poor blood clotting and fatigue due to anemia. Examples of leukemia include: -Myeloid or granulocytic leukemia (malignant tumors of the myeloid and granulocytic leukemia lineage) - Lymphocytic, lymphocytic, or lymphoblastic leukemia (malignant tumors of the lymphocyte and lymphoid blood lineage) -Polycythemia vera or erythrocytemia (malignancy of various blood cell products, but with a predominance of red blood cells)
[0261] Lymphomas originate in the lymphatic system, a network of glands or nodes, blood vessels, nodes, and organs (especially the spleen, tonsils, and thymus) that produce white blood cells, or lymphocytes, that purify body fluids and fight infection. Unlike leukemia, often referred to as a "liquid cancer," lymphomas are "solid cancers." Lymphomas can also develop in certain organs, such as the stomach, breast, and brain. These lymphomas are called extranodal lymphomas. Lymphomas are divided into two categories: Hodgkin lymphoma and non-Hodgkin lymphoma. The presence of Reed-Sternberg cells in Hodgkin lymphoma distinguishes it diagnostically from non-Hodgkin lymphoma.
[0262] Mixed-type cancers can include cancers in which the type components are within one cancer category or belong to different cancer categories. Some examples include adenosquamous carcinoma, mixed mesodermal tumor, carcinosarcoma, and teratoma.
[0263] In some cases, cancer can be primary cancer or metastatic cancer. Primary cancer refers to cancer cells of a primary tumor, which is a tumor that first appears in a subject, and can be distinguished from metastatic tumors, which appear in a subject's body at a site distant from the primary tumor. Metastatic cancer arises from metastasis, which refers to the spread of cancer from the organ of origin to additional distant sites in the patient.
[0264] In one preferred embodiment, the histological specimen used in the present invention comprises biological material obtained from a subject having or suspected of having any one or more types of cancer selected from the following list: Acute lymphoblastic leukemia (ALL), Acute myeloid leukemia (AML), Cancer in adolescents (e.g., adolescents aged 12-18 years), Adrenal cortical carcinoma, including, for example: Pediatric adrenocortical cancer AIDS-related cancers, including, for example: Kaposi's sarcoma (soft tissue sarcoma) AIDS-related lymphoma (lymphoma) Primary CNS lymphoma (lymphoma) Anal cancer Appendix cancer Pediatric astrocytoma (brain tumor) Atypical malformation / rhabdoid tumor, childhood, central nervous system (brain tumor) Basal cell carcinoma of the skin Bile duct cancer Bladder cancer, including, for example: Pediatric bladder cancer Osteosarcoma (Ewing's sarcoma, osteosarcoma, malignant fibrous histiocytoma, etc.) Brain tumors Breast cancer, including, for example: Pediatric breast cancer Pediatric bronchial tumors Burkitt's lymphoma Carcinoid tumors (gastrointestinal), including, for example: Pediatric carcinoid tumor Carcinoma of unknown primary origin, including, for example: Childhood cancer of unknown primary origin Pediatric heart (heart) tumors Central nervous system, including, for example: Pediatric atypical malformation / rhabdoid tumor (brain tumor) Pediatric embryonal tumors (brain tumors) Pediatric germ cell tumors (brain tumors) Primary CNS lymphoma Cervical cancer, including, for example: Childhood cervical cancer Childhood cancer (e.g., under 18 years of age, preferably under 12 years of age, such as 16, 14, or in the range of 1-12 years of age), · Abnormal childhood cancers, Bile duct cancer Pediatric chordoma, Chronic lymphocytic leukemia (CLL) Chronic myeloid leukemia (CML) Chronic myeloproliferative neoplasms Colon cancer, including, for example: Pediatric colorectal cancer Pediatric craniopharyngioma (brain tumor) ·Cutaneous T-cell lymphoma Ductal carcinoma in situ (DCIS) Embryonal tumors, central nervous system, childhood (brain tumors) Endometrial cancer (uterine cancer) Pediatric ependymoma (brain tumor) Esophageal cancer, including, for example: Pediatric esophageal cancer Olfactory neuroblastoma (head and neck cancer) Ewing's sarcoma (osteosarcoma) Pediatric extracranial germ cell tumors, Extragonadal germ cell tumors Eye cancer, including: Pediatric intraocular melanoma ○Intraocular melanoma Retinoblastoma Fallopian tube cancer Fibrous histiocytoma of bone, malignant, osteosarcoma Gallbladder cancer Gastric (stomach) cancer, including: Pediatric gastric (stomach) cancer Gastrointestinal carcinoid tumor Gastrointestinal stromal tumors (GIST) (soft tissue sarcomas), including, for example: Pediatric gastrointestinal stromal tumors Germ cell tumors, including, for example: Pediatric central nervous system germ cell tumors (brain tumors) Pediatric extracranial germ cell tumors Extragonadal germ cell tumor Ovarian germ cell tumor Testicular tumor ·Gestational trophoblastic disease ·hairy cell leukemia Head and neck cancer Pediatric heart tumors Hepatocellular (liver) cancer Histiocytosis, Langerhans cells Hodgkin's lymphoma Hypopharyngeal cancer (head and neck cancer) Intraocular melanoma, including, for example: Pediatric intraocular melanoma Pancreatic islet cell tumors, pancreatic neuroendocrine tumors Kaposi's sarcoma (soft tissue sarcoma) Kidney (renal cell) cancer Langerhans cell histiocytosis Laryngeal cancer (head and neck cancer) ·leukemia Lip and oral cancer (head and neck cancer) Liver cancer Lung cancer (non-small cell and small cell), including, for example: Pediatric lung cancer Lymphoma Male breast cancer Osteosarcoma and osteosarcoma malignant fibrous histiocytoma Melanoma, including, for example: Pediatric melanoma Intraocular (eye) melanoma, including: Pediatric intraocular melanoma Merkel cell carcinoma (skin cancer) Malignant mesothelioma, including, for example: Childhood mesothelioma Metastatic cancer Metastatic squamous cell carcinoma (head and neck cancer) with a latent primary tumor Midline tract cancer with NUT gene alterations Oral cancer (head and neck cancer) Multiple endocrine neoplasia syndrome Multiple myeloma / plasma cell neoplasm ·Mycosis fungoides (lymphoma) Myelodysplastic syndromes, myelodysplastic / myeloproliferative neoplasms Chronic myeloid leukemia (CML) Acute myeloid leukemia (AML) Chronic myeloproliferative neoplasms Nasal cavity and paranasal sinus cancer (head and neck cancer) Nasopharyngeal cancer (head and neck cancer) Neuroblastoma Non-Hodgkin's lymphoma Non-small cell lung cancer Oral cavity cancer, lip and oral cavity cancer, and oropharyngeal cancer (head and neck cancer) Osteosarcoma and malignant fibrous histiocytoma of bone Ovarian cancer, including, for example: Pediatric ovarian cancer Pancreatic cancer, including, for example: Pediatric pancreatic cancer Pancreatic neuroendocrine tumors (islet cell tumors) Papillomatosis (pediatric larynx) Paragangliomas, including, for example: Pediatric paraganglioma Paranasal sinus and nasal cavity cancer (head and neck cancer) Parathyroid cancer Penile cancer Pharyngeal cancer (head and neck cancer) Pheochromocytoma, including, for example: Pediatric pheochromocytoma Pituitary tumor Plasma cell neoplasms / multiple myeloma ·Pleural pulmonary blastoma Pregnancy and breast cancer (i.e., breast cancer in pregnant women) Primary central nervous system (CNS) lymphoma Primary peritoneal cancer Prostate cancer Rectal cancer Recurrent cancer Renal cell (kidney) cancer Retinoblastoma Pediatric rhabdomyosarcoma (soft tissue sarcoma) Salivary gland cancer (head and neck cancer) Sarcomas, including, for example: Pediatric rhabdomyosarcoma (soft tissue sarcoma) Pediatric vascular tumors (soft tissue sarcomas) Ewing's sarcoma (osteosarcoma) Kaposi's sarcoma (soft tissue sarcoma) Osteosarcoma (bone cancer) Soft tissue sarcoma Uterine sarcoma Sézary syndrome (lymphoma) Skin cancer, including, for example: Pediatric skin cancer Small cell lung cancer Small intestine cancer Soft tissue sarcoma Squamous cell carcinoma of the skin Metastatic squamous cell carcinoma (head and neck cancer) with a latent primary tumor Gastric (stomach) cancer, including: Pediatric gastric (stomach) cancer ·Cutaneous T-cell lymphoma Testicular cancer, including, for example: Pediatric testicular cancer Throat cancer (head and neck cancer), including: Nasopharyngeal cancer Oropharyngeal cancer Hypopharyngeal cancer Thymoma and thymic carcinoma Thyroid cancer Transitional cell carcinoma of the renal pelvis and ureter (kidney (renal cell) cancer) Carcinoma of unknown primary origin, including, for example: Childhood cancer of unknown primary origin · Abnormal childhood cancers ·Ureter ·renal pelvis, transitional cell carcinoma (kidney (renal cell) cancer) Urethral cancer Endometrial cancer Uterine sarcoma Vaginal cancer, including, for example: Childhood vaginal cancer Vascular tumors (soft tissue sarcomas) Vulvar cancer Wilms tumor and other childhood kidney tumors Cancer in young adults (e.g., young adults aged 16-30, such as 18-30, optionally 28, 26, 24, 22, under 20)
[0265] More specifically, in preferred embodiments, the biological source of a source histological image for use in the present invention may consist of, consist essentially of, or comprise biological material obtained from a subject (such as a human subject) that has, has been diagnosed with, is suspected of having, is being treated for, has previously been treated for, and / or has previously had any one or more of the types of cancer selected from the list below: Skin Cancer: There are three main types of skin cancer: basal cell, squamous cell, and melanoma. These cancers originate in the epidermal layer of the same name. Melanoma originates in melanocytes, or pigment cells, which are found in the deepest level of the epidermis. Basal cell and squamous cell carcinomas usually occur on parts of the body that are exposed to the sun, such as the face, ears, and extremities. Lung cancer: Lung cancer is very difficult to detect in the early stages because symptoms often do not appear until the disease is advanced. Symptoms include persistent cough, bloody sputum, chest pain, and recurring bouts of pneumonia and bronchitis. Breast cancer in women or men: In the United States, it is estimated that about 1 in 8 women will eventually develop breast cancer in her lifetime. Most breast cancers are ductal carcinomas. Women most likely to develop the disease are those over 50, those who have already had cancer in one breast, those whose mother or sister had breast cancer, those who have never had children, and those who have their first child after age 30. Other risk factors include obesity, a high-fat diet, early menarche (the age at which menstruation begins), and late menopause (the age at which menstruation ends). Prostate Cancer: Cancer of the prostate gland is primarily seen in older men. As men age, the prostate gland can enlarge and block the urethra or bladder. This can make urination difficult or interfere with sexual function. This condition is called benign prostatic hyperplasia (BPH). BPH is not cancerous, but surgery may be needed to correct it. Symptoms of BPH, or other problems with the prostate gland, can be similar to those of prostate cancer. Colon and / or rectal cancer: Colorectal cancer (CRC) is a disease that usually originates in the epithelial cells lining the colon or rectum of the digestive tract. Of cancers that affect the large intestine, approximately 70% begin in the colon and approximately 30% begin in the rectum. These cancers are the third most common cancer overall. Symptoms include blood in the stool, which can be tested for with a fecal occult blood test, or changes in bowel habits, such as severe constipation or diarrhea. Uterine (corpus uteri) cancer: The uterus is a sac in a woman's pelvis that can develop a baby from fertilized egg to protect it until birth. Cancer of the uterus is the most common gynecological malignancy. This cancer rarely occurs in women under 40 years of age. It occurs most frequently after age 60. The main symptom is usually abnormal uterine bleeding. An endometrial biopsy or D&C is often performed to confirm the diagnosis.
[0266] In addition to the cancer types named after their primary site discussed above, there are many other examples, such as brain tumors, testicular cancer, and bladder cancer.
[0267] 2.3.2 Treatment As discussed above, the biological source of the histopathological sample from which the source histopathological image is generated may be a subject (e.g., a human subject) being treated for a pathological condition (e.g., cancer) and / or having previously had a pathological condition (e.g., cancer).
[0268] Additionally and / or alternatively, the subject may be, for example, a subject having a pathological condition of interest (including, but not limited to, cancer), for which it is desirable to obtain further information about the pathological condition to aid in making decisions about the need for, nature of, and / or potential benefits of future treatment for the pathological condition.
[0269] The subject may be, for example, a subject who has been diagnosed with a pathological condition of interest (including, but not limited to, cancer) and has already received one or more treatments for that pathological condition, and for whom it is desirable to know whether the subject still has the pathological condition and / or to obtain further information regarding the biological material in the part of the subject's body where the pathological condition was previously treated or in other parts of the subject's body to help make a decision regarding the need for, nature and / or potential benefit of future treatment for the pathological condition.
[0270] Such forms of treatment for pathological conditions (including cancer) are included, and those skilled in the art are familiar with surgical and / or non-surgical treatment approaches (e.g., administration of one or more compositions, where the or each composition comprises one or more active agents that provide a therapeutic and / or prophylactic effect with respect to the pathological condition), as well as how to match the form of treatment to the type of pathological condition.
[0271] More specifically, cancer treatments include, but are not limited to, surgery, radiation therapy, chemotherapy, bisphosphonates, gene therapy, immunotherapy, targeted therapy, hormone therapy, stem cell and / or bone marrow transplantation, and / or precision medicine. Additional cancer treatments include, but are not limited to, radiofrequency ablation, laser therapy, high-intensity focused ultrasound (HIFU), photodynamic therapy, cryotherapy, ultraviolet light therapy, and / or electrochemotherapy. In one embodiment, cancer treatment (e.g., in the context of making a decision about the necessity, nature, and / or potential benefit of future treatment for the condition) can be a form of adjuvant therapy. Adjuvant therapy is a form of treatment administered in addition to primary (initial) treatment. Adjuvant cancer therapy includes forms of treatment following primary (e.g., surgical) treatment and can be any form of cancer treatment (e.g., chemotherapy or radiation therapy, or other forms of cancer treatment) aimed at treating cancer cells remaining in a subject after primary treatment and / or reducing the risk of cancer recurrence.
[0272] 2.3.3 Target Characteristics As discussed above, the subject (such as a human subject) that serves as the source of the biological material present in the histological specimen used in accordance with the present invention may be a healthy subject, or more typically, a diseased subject.
[0273] The subject may optionally be male, such as a male human.
[0274] The subject may optionally be female, such as a female human.
[0275] The subject may be an adult, adolescent, juvenile, child, infant, newborn, fetus, or embryo, such as, for example, a human adult, human adolescent, human juvenile, human child, human infant, human newborn, human fetus, or human embryo.
[0276] The subject can be, for example, a human infant less than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months of age.
[0277] The subject can be a human (e.g., a male human and / or a female human), e.g., at least 1 year old or older, e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 years old or older, and optionally less than 100, 90, 85, 80, 75, 65, 60, 55, 50, 45, or 40 years old.
[0278] The subject can be a human (e.g., a male human and / or a female human), e.g., at least 20 years of age or older, e.g., at least 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 80 years of age or older, and optionally under 100, 90, or 85 years of age.
[0279] A pathological subject can be a subject (such as a human subject) that has a pathological condition, has been diagnosed with a pathological condition, is suspected of having a pathological condition, is being treated for a pathological condition, has previously been treated for a pathological condition, and / or has previously had a pathological condition.
[0280] In the case of a histological specimen containing biological material from a pathological subject, the histological specimen may contain or be suspected of containing biological material containing a pathological condition present in the subject, including, but not limited to, those conditions discussed above in Section 2.3 of this application, and in a preferred embodiment, the pathological condition may be a form of cancer, such as a type of cancer discussed above in Section 2.3.1 of this application.
[0281] A subject may, for example, be one who has not previously been diagnosed with a pathological condition of interest (including, but not limited to, cancer), but for whom it is desirable to know, for example, whether they are at risk of developing or progressing to the pathological condition, and / or whether they are at risk.
[0282] The subject may be, for example, a subject having a pathological condition of interest (including, but not limited to, cancer), and for which it is desirable to obtain further information about the pathological condition, for example, to determine a prognosis for the pathological condition and / or to assist in determining the need for, nature of, and / or potential benefits of future treatment for the pathological condition.
[0283] The subject may be, for example, a subject who has been diagnosed with a pathological condition of interest (including, but not limited to, cancer) and who has already received one or more treatments for that pathological condition (e.g., as discussed above in Section 2.3.2 of this application), and for whom it is desirable to know whether the subject still has the pathological condition and / or to obtain further information regarding the biological material in the area of the subject's body where the pathological condition was previously treated or in other areas of the subject's body to help make a decision regarding the need for, nature of, and / or potential benefit of, future treatment for the pathological condition.
[0284] For example, if the pathological condition is cancer, samples may be taken from: Subjects who have not been previously diagnosed with cancer but would like to know if they have cancer, for example. subjects who have been diagnosed with cancer and who would like to obtain further information about the cancer, for example, to determine the prognosis of the cancer and / or to assist in making decisions about the need for, the nature and / or potential benefits of future treatment for the cancer; and / or A subject who has been diagnosed with cancer and has already received one or more treatments for that cancer, and who would like to know if the subject still has cancer and / or obtain further information about biological material at sites in the subject's body where the cancer has been previously treated or at other sites in the subject's body, e.g., to determine the prognosis of the cancer and / or to assist in decisions about the need for, the nature and / or potential benefits of future treatments for the cancer.
[0285] The subject (such as a human subject) that serves as the source of biological material present in the histological specimen used in accordance with the present invention may be a subject that is at risk (e.g., a higher than average risk) and / or has been determined to be at risk for developing a pathological condition (including, but not limited to, cancer).
[0286] Known risk factors associated with cancer include, but are not limited to, age, alcohol intake, exposure to cancer-causing agents, chronic inflammation, diet, hormones, immunosuppression, infectious agents, obesity, radiation, sunlight, and tobacco. The subject (such as a human subject) that serves as the source of the biological material present in the histological specimen used in accordance with the present invention may be at risk and / or determined to be at risk for developing cancer based on one or more of the aforementioned risk factors.
[0287] Increasing age is one of the most important risk factors for cancer overall and for many individual cancer types. According to the most recent statistical data from NCI's Surveillance, Epidemiology, and End Results program, the median age at cancer diagnosis is 66 years. This means that half of cancer cases occur in people younger than this age and half occur in people older than this age. One-quarter of new cancer cases are diagnosed in people aged 65 to 74. Similar patterns are seen for many common cancer types. For example, the median age at diagnosis is 61 years for breast cancer, 68 years for colorectal cancer, 70 years for lung cancer, and 66 years for prostate cancer. Thus, in one embodiment of the disclosure, the pathological condition is cancer and the subject is a human (e.g., a male human and / or a female human) at least 50, 55, 60, 65, 70, 75, or 80 years of age or older, and optionally less than 100, 90, or 85 years of age, e.g., 50-85 years, 60-85 years, 65-85 years, or 65-75 years of age.
[0288] However, cancer can occur at any age. For example, osteosarcoma is most frequently diagnosed among people under the age of 20, with over a quarter of cases occurring in this age group. Also, 10% of leukemias are diagnosed in children and adolescents under the age of 20, but only 1% of cancers overall are diagnosed in that age group. Some types of cancer, such as neuroblastoma, are more common in children or adolescents than in adults.
[0289] Smoking is a leading cause of cancer and cancer death. People who use tobacco products or are regularly exposed to environmental tobacco smoke (also known as secondhand smoke) are at increased risk of cancer because tobacco products and secondhand smoke contain many DNA-damaging chemicals. Smoking causes many types of cancer, including cancer of the lung, larynx (larynx), mouth, esophagus, throat, bladder, kidney, liver, stomach, pancreas, colon and rectum, and cervix, as well as acute myeloid leukemia. People who use smokeless tobacco (snuff or chewing tobacco) are at increased risk of cancer of the mouth, esophagus, and pancreas. Thus, in one embodiment of the present disclosure, the pathological condition is cancer, and the subject is a human (e.g., a male human and / or a female human) with a smoking history (including tobacco exposure, such as secondhand smoke).
[0290] Drinking alcohol can increase the risk of cancer, such as cancer of the mouth, throat, esophagus, larynx (voice box), liver, and / or breast. The more alcohol a subject drinks, the higher the risk. For people who drink alcohol and also smoke, the risk of cancer is much higher. Thus, in one embodiment of the present disclosure, the pathological condition is cancer, and the subject is a human (e.g., a male human and / or a female human) with a history of alcohol consumption (typically higher than average alcohol consumption) and / or smoking.
[0291] Hormone levels can affect cancer risk. For example, estrogen, a group of female hormones, is a known human carcinogen. While these hormones play essential physiological roles in both women and men, they are also associated with an increased risk of certain cancers. For example, combined menopausal hormone therapy (estrogen and progestin, a synthetic version of the female hormone progesterone) may increase a woman's risk of breast cancer. Menopausal hormone therapy with estrogen alone increases the risk of endometrial cancer and is only used in women who have had a hysterectomy. Research has shown that a woman's risk of breast cancer is related to the estrogen and progesterone produced in her ovaries (known as endogenous estrogen and progesterone). Prolonged and / or high exposure to these hormones is associated with an increased risk of breast cancer. Increased exposure may be caused by early onset of menstruation, late menopause, being older at the time of first pregnancy, and never having given birth. Conversely, having given birth may be a protective factor for breast cancer.
[0292] Diethylstilbestrol (DES) is a type of estrogen that was given to some pregnant women in the United States between 1940 and 1971 to prevent miscarriage, premature birth, and related pregnancy problems. Women who took DES during pregnancy have an increased risk of breast cancer. Their daughters have an increased risk of vaginal or cervical cancer. The possible effects on the sons and grandchildren of women who took DES during pregnancy are being studied.
[0293] Thus, in one embodiment of the present disclosure, the pathological condition is cancer and the subject is a human (eg, a male human and / or a female human) with associated hormonal risk factors.
[0294] Subjects receiving immunosuppression may also be at increased risk for cancer. For example, organ transplant recipients are typically given immunosuppressant drugs, which reduce the immune system's ability to detect and destroy cancer cells or fight cancer-causing infections. Infection with HIV and other immunosuppressive pathogens can also weaken the immune system and increase the risk of certain cancers.
[0295] The four most common cancers among transplant recipients, occurring more commonly in these individuals than in the general population, are non-Hodgkin lymphoma (NHL), and cancers of the lung, kidney, and liver. NHL can be caused by Epstein-Barr virus (EBV) infection, and liver cancer can be caused by chronic infection with hepatitis B (HBV) and hepatitis C (HCV) viruses. Lung and kidney cancers are not generally considered to be associated with infection.
[0296] People with HIV / AIDS are also at increased risk of cancers caused by infectious agents, including EBV, human herpesvirus 8 or Kaposi's sarcoma-associated virus, HBV and HCV, which cause liver cancer, and human papillomavirus, which causes cervical, anal, oropharyngeal, and other cancers. HIV infection is also associated with an increased risk of cancers, such as lung cancer, that are not thought to be caused by infectious agents.
[0297] Thus, in one embodiment of the present disclosure, the pathological condition is cancer and the subject is an immunosuppressed human (eg, a male human and / or a female human).
[0298] Exposure to certain infectious agents, including viruses, bacteria, and parasites, can cause cancer or increase the risk of cancer formation. Some viruses can disrupt signaling that normally keeps cells from continuing to grow and multiply. Some infections also weaken the immune system, making it impossible for the body to fight other cancer-causing infections. Some viruses, bacteria, and parasites also cause chronic inflammation, which can lead to cancer. Most viruses associated with an increased risk of cancer can be transmitted from person to person through blood and / or other bodily fluids. Exemplary infectious agents that may cause cancer or increase the risk of cancer include Epstein-Barr virus (EBV), hepatitis B virus and hepatitis C virus (HBV and HCV), human immunodeficiency virus (HIV), human papillomavirus (HPV), human T-cell leukemia / lymphoma virus type 1 (HTLV-1), Kaposi's sarcoma-associated herpesvirus (KSHV), Merkel cell polyomavirus (MCPyV), Helicobacter pylori (H. pylori), Opisthorchis viverrini, and Schistosoma hematobium.
[0299] Thus, in one embodiment of the present disclosure, the pathological condition is cancer, and the subject is a human (e.g., a male and / or female human), who has, is suspected of having, is being treated for, has previously been treated for, and / or has previously had an infection (whether or not actually diagnosed) with an infectious agent that may cause cancer or increase the risk of cancer formation.
[0300] Obese people may be at increased risk for several types of cancer, including breast cancer (in menopausal women), colon cancer, rectal cancer, endometrial (the lining of the uterus), esophageal cancer, kidney cancer, pancreatic cancer, and gallbladder cancer. Thus, in one embodiment of the present disclosure, the pathological condition is cancer and the subject is an obese human (e.g., a male human and / or a female human).
[0301] In a further embodiment of the present disclosure, the pathological condition is cancer, and the subject is a person (for example, a male person and / or a female person) who has a genetic risk factor for cancer, and optionally has been determined to have a genetic risk factor for cancer.Genetic risk factors can be, for example, inherited genetic traits or acquired genetic traits.Many types of genetic risk factors are well known in the art, and include, but are not limited to, Lynch syndrome, BRCA gene, retinoblastoma gene, etc.
[0302] 2.4 Preparation of histological specimens Those skilled in the art will be familiar with the many techniques known in the art for preparing histological specimens from biological material obtained from biological sources. Any suitable means for preparing a histological specimen can be used in the context of the present invention.
[0303] For example, but not by way of limitation, for specimens to be examined (eg, by light microscopy and / or WSI imaging), three techniques are commonly used: paraffin sectioning, frozen sectioning, and semi-thin sectioning.
[0304] The paraffin method is the most commonly used. In this technique, tissue is fixed and embedded in wax. This makes the tissue hard and makes it much easier to cut sections from it. The sections are then stained to help distinguish the tissue components.
[0305] Fixation involves chemically immobilizing biological samples. Added chemicals bind to and cross-link some proteins and denature others through dehydration. This hardens the tissue and inactivates enzymes that can degrade it. Fixation also kills bacteria and other pathogens, which can facilitate tissue staining. A common fixative is a 4% aqueous solution of formaldehyde at a neutral pH. Another common option is formalin-fixed, paraffin-embedded (FFPE) tissue specimens, which have been a staple in research and therapeutic applications for decades.
[0306] After fixation, samples typically undergo the steps of dehydration and clearing, embedding, sectioning, staining, and mounting.
[0307] Dehydration and Clearing: To cut sections, it may be desirable to embed fixed biological samples in paraffin wax. However, wax is not soluble in water or alcohol. However, it is soluble in solvents such as xylene. Therefore, the water within the tissue can be replaced with a solvent. To do this, the tissue is first dehydrated by gradually replacing the water within the sample with alcohol. This can be achieved, for example, by passing the tissue through increasing concentrations of ethyl alcohol (from 0% to 100%). Finally, once the water has been replaced with 100% alcohol, the alcohol is replaced with an alcohol-miscible solvent (e.g., xylene). This final step is called "clearing."
[0308] Following the dehydration and clearing steps, the sample is typically subjected to an embedding step. For example, the tissue can be placed in warm paraffin wax and the melted wax can fill the spaces previously occupied by water. After cooling, the tissue hardens and can be used for cutting (sectioned) slices.
[0309] In the sectioning step, tissue is trimmed and mounted on a cutting device (e.g., a microtome or ultramicrotome). Thin sections can be cut, stained, and mounted on microscope slides. In this context, "thin" sections can include sections approximately one cell deep or less. Typically, cells in samples taken from animal or human biological sources can be, for example, 10-15 µm in diameter, and thin sections can be equivalent or smaller, such as 5-10 µm thick.
[0310] In one option, it may be desirable to cut multiple physically separate (typically contiguous parallel or essentially parallel) slices of biological material for the purpose of obtaining a collection of source histological images that can together form a single three-dimensional source histological image.
[0311] In another option, it may be desirable to cut thick sections that can be stained and mounted on microscope slides, and multiple images (e.g., successive parallel or essentially parallel images) can be obtained by selective focusing techniques at planes within the multiple thick sections. Selective focusing techniques are known in the art, e.g., Mir et al., 2014: "An extensive empirical evaluation of focus measures for digital photography," in Proceedings of SPIE—The International Society for Optical Engineering 9023 (available online at https: / / cs.uwaterloo.ca / ~vanbeek / Publications / spie2014.pdf), and Hosseini et al., "Focus Quality Assessment of High-Throughput Whole Slide Imaging in Digital Pathology," submitted for publication to The IEEE in 2018 (available online at https: / / arxiv.org / pdf / 1811.06038.pdf), the contents of both of which are incorporated herein by reference.
[0312] In this context, a "thick" section can include a section having a thickness greater than one cell deep, e.g., up to 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more cells deep. As noted above, cells in samples taken from animal or human biological sources typically have a diameter of, for example, 10-15 μm. Thus, a thick section in this context can have a thickness of 10 μm, 15 μm, 20 μm, 25 μm, 30 μm, 40 μm, 50 μm, 60 μm, 70 μm, 80 μm, 90 μm, 100 μm, 150 μm, 200 μm, 250 μm, 300 μm, 400 μm, 500 μm, or more. For example, if multiple successive parallel or essentially parallel images are acquired by selective focusing techniques at multiple planes within a thick section, this is a further method that can provide an assembly of source histological images, which together can form a three-dimensional source histological image.
[0313] Staining and Mounting: Unfortunately, most staining solutions are water-based, so staining sections may require dissolving the wax and replacing it with water (rehydration). This is essentially the reverse of the dehydration and clearing steps. Sections are passed through wax solvents (e.g., ethylenediamine), then alcohols (100% to 0%), and finally water at decreasing strengths. Once stained, sections are dehydrated once more and placed in solvent (e.g., ethylenediamine). Sections can then be mounted on microscope slides in a mounting medium dissolved in the solvent. A coverslip can be placed on top to protect the sample. Evaporation of the solvent (e.g., ethylenediamine) around the edges of the coverslip can be allowed to dry the mounting medium and firmly adhere the coverslip to the slide.
[0314] Alternative means for preparing histological specimens include, but are not limited to, preparation of frozen sections and semi-thin sections.
[0315] In cryosectioning, tissue is rapidly frozen (e.g., with liquid nitrogen), cut while still cool (e.g., with a cold knife in a refrigerated cabinet (cryostat)), and then stained for observation. This procedure is faster and preserves tissue details that may be lost with paraffin techniques. Frozen sections are typically 5–10 μm thick, but a suitable thickness, including any of the thicknesses discussed above, can be selected depending on the desired purpose.
[0316] Semi-thin sections can be useful in situations where it is difficult to see the details of thicker sections. To circumvent this, sections can be embedded in epoxy or acrylic resin, which allows thinner sections (e.g., less than 2 μm) to be cut.
[0317] 2.4.1 Staining of histological specimens Histological or histopathological images can be prepared by appropriately staining histological or histopathological specimens, as already discussed above. At the microscopic scale, many interesting cellular features are transparent and colorless, making them essentially invisible. To reveal these features, specimens are usually stained with markers before being imaged under a microscope. Markers include one or more staining agents (dyes or pigments) designed to specifically bind to specific components of cellular structure, thereby revealing the histological features of interest. Those skilled in the art are familiar with numerous staining techniques well known in the art. Any suitable means for staining histological specimens can be used in the context of the present invention. Non-limiting examples of such staining techniques that may be suitable for use in the present invention are discussed further below.
[0318] The techniques used can be nonspecific, staining most cells in much the same way, or specific, selectively staining particular chemical groups or molecules within cells or tissues. Staining usually works by using a dye that stains some of the cellular components with a first color, together with one or more counterstains that stain the rest of the cell with one or more different colors.
[0319] For example, staining techniques can use basophilic and acidophilic stains.
[0320] Acid dyes (e.g., aegeosin) react with cationic or basic components within cells. Most proteins and many other components of the cytoplasm are basic and will bind to acid dyes. This includes, for example, the cytoplasmic filaments of muscle cells, intracellular membranes, and extracellular fibers.
[0321] Basic dyes (e.g., hematoxylin) react with anionic or acidic components within cells. Nucleic acids are acidic and therefore bind to basic dyes. For example, DNA in the nucleus (heterochromatin and nucleoli) and RNA in ribosomes and the rough endoplasmic reticulum are both acidic, so hematoxylin binds to them and stains them purple. Some extracellular substances (e.g., carbohydrates in cartilage) also bind to basic dyes.
[0322] For purposes of the present invention, the stain used in one preferred embodiment is a staining system called H&E (hemotoxylin and eosin). H&E contains two dyes: hemotoxylin and eosin. Eosin is an acidic dye, negatively charged, and stains basic (i.e., acidophilic) structures red or pink. This is sometimes referred to as "eosinophilic." Hematoxylin can be considered a basic dye and is used to stain acidic (i.e., basophilic) structures a purplish-blue color. Thus, when a histological sample is stained with the H&E system, the nucleus and the portion of the cytoplasm containing RNA typically stain one color (purple), while the remaining cytoplasm typically stains a different color (pink).
[0323] However, it will be understood that the present invention is not limited with respect to the use of any particular dyeing technique. For example, many other acid and basic dyes are known in the art. Examples are provided in the following table. Typical histological stains other than H&E may include: [Table 1] [Table 2]
[0324] In the case of basic dyes, reaction with cellular anionic groups (which include phosphate groups of nucleic acids, sulfate groups of glycosoaminoglycans, and carboxyl groups of proteins) may depend on the pH used.
[0325] In the case of acid dyes, the dye in question can often be even more selective for certain acidophilic components. For example, the well-known Mallory staining technique uses three acid dyes, namely aniline blue, acid fuscin, and orange G, which selectively stain collagen, cytoplasm, and red blood cells, respectively.
[0326] Additional staining techniques contemplated herein for use in the present invention include the following: The Periodic Acid-Schiff reaction (PAS) is a bleached basic fuchsin in which Schiff's reagent reacts with aldehyde groups. This reaction produces a deep red color in the sections. This is the basis of PAS staining. PAS stains carbohydrates and carbohydrate-rich macromolecules deep red (magenta). Thus, PAS stains glycogen, the intracellular storage form of carbohydrates in cells, mucus in cells and tissues, basement membranes and the reticular fibers (i.e., collagen) of the brush border connective tissue and cartilage of the renal tubules and small and large intestines.
[0327] Masson's trichrome is a commonly used stain for connective tissue. The term "trichrome" refers to the fact that this technique produces three colors. Nuclei and other basophilic structures stain blue, while cytoplasm, muscle, red blood cells, and keratin stain bright red. Collagen may stain green or blue, depending on the variation in the technique used.
[0328] Alcian blue is a mucin stain that stains certain types of mucin blue. It also stains cartilage blue. It can be used with other staining systems such as H&E and van Gieson stains.
[0329] The Van Gieson technique stains collagen red, nuclei blue, and red blood cells and cytoplasm yellow. It can also be combined with an elastin stain, which stains elastin blue / black. It is commonly used on blood vessels and skin.
[0330] The reticulin staining technique stains reticulin fibers blue / black. It can be used with other staining techniques, e.g., H&E.
[0331] Azan staining stains nuclei bright red, collagen, basement membranes and mucins blue, and muscle and red blood cells orange to red. This technique may be particularly suitable for staining connective tissue and epithelium.
[0332] The Giemsa staining technique is commonly used to stain blood and bone marrow smears. Nuclei are stained dark blue to purple, cytoplasm pale blue, and red blood cells pale pink.
[0333] Toluidine blue is a basic stain that stains acidic components in various shades of blue. It is typically used on thin acrylic or epoxy sections.
[0334] The silver and gold method can be used to demonstrate fine structures such as the cellular processes of neurons. This technique produces black, brown, or gold staining.
[0335] Chrome alum / hemotoxylin is a less commonly used technique that stains nuclei blue and cytoplasm red, which can be particularly useful in pancreatic samples, where glucagon-secreting cells stain pink and insulin-secreting cells stain blue.
[0336] Isamine blue / eosin is a staining technique similar to H&E, but with a darker blue color.
[0337] Nissl and methylene blue are staining techniques that can be used as basic dyes to stain, for example, the rough endoplasmic reticulum of neurons.
[0338] Sudan black and osmium are dyes that stain lipid-containing structures such as myelin brownish black.
[0339] Immunohistochemical (IHC) techniques can also be used (alone or in combination with any one or more of the aforementioned techniques). Immunohistochemical techniques are well known in the art. Typically, a primary antibody that specifically labels a protein (or other specific target) is used, and then a labeled (e.g., fluorescently labeled) secondary antibody is used to bind to the primary antibody and indicate where the first (primary) antibody bound. The label can be detected by any suitable means. For example, in methods using fluorescently labeled antibodies, a light microscope (or equivalent imaging device) equipped with fluorescence can be used to visualize the staining. Fluorescent antibodies are excited with light of one wavelength and then emit light of a different wavelength. By using the correct combination of filters, the staining pattern generated by the emitted fluorescence can be observed.
[0340] IHC techniques may become increasingly useful in the context of imaging samples containing or suspected of containing cancer. In such techniques, primary antibodies can target, for example, tumor markers. Tumor markers are molecules whose levels are considered signals, symbols, or representations of tumor cells and may increase in cancerous conditions. Tumor markers include, but are not limited to, proteins, gene expression patterns, and DNA alterations. Tumor markers visualized by IHC may include, among others, enzymes, oncogenes, tumor-specific antigens, tumor suppressor genes, and tumor proliferation markers.
[0341] 3. Ground Truth As discussed in more detail in Section 1 of this application, and as further defined by the claims, this application describes a computer-implemented system (100, 200, 300) for determining a classifier (118, 318), or an overall classifier (232, 332), for a source histological image (102, 202, 302), which may be a source histopathological image.
[0342] Also, as discussed in more detail in Section 2 of this application, during a training phase in which the systems (100, 200, 300) may be used to train machine learning algorithms within the systems (100, 200, 300), the source histological images used for training are images of histological specimens acquired from sources for which the ground truth is known. Each of the source histological images is then labeled with associated truth data that represents the ground truth for each source from which each histological specimen was acquired.
[0343] In this context, "ground truth" may be any information of interest associated with the or each source histological image. Any useful measurement that can be associated with each of the images may be applied as "ground truth." It will be appreciated that the selection of ground truth is an important feature for determining the functionality and usefulness of subsequently trained machine learning algorithms.
[0344] For example, an algorithm trained using source histopathological images of histopathological specimens, each image having associated ground truth, is a known diagnosis (e.g., presence or absence) of a particular pathological condition, and the trained algorithm is then suitable for use in diagnosing that pathological condition (and / or the risk of developing that pathological condition) in microscopic images of histological specimens taken from subjects ("test subjects") of unknown diagnosis, and therefore is also suitable for use in diagnosing the pathological condition in the subject.
[0345] In another example, an algorithm trained using source histopathological images of histopathological specimens with associated ground truth for each image may be of known prognosis for a particular pathological condition, and the trained algorithm is then suitable for use in prognosis of that pathological condition in microscopic images of histological specimens taken from subjects of unknown prognosis ("test subjects"), and therefore also suitable for use in prognosis of the subjects.
[0346] In the examples described herein, histological specimens were obtained from cancer patients and ground truth related to the classification of each patient into prognostic groups.
[0347] In one embodiment described herein, histological specimens may be obtained from patients with one or more types of cancer and ground truth related to the classification of the cancer present within each patient into stages and / or grades, particularly ground truth where the stages and / or grades are correlated with defined prognoses.
[0348] For example, the subsequently trained algorithm may be suitable for use to automatically identify the stage and / or grade of cancer present in microscopic images of histological specimens taken from one or more subjects ("test subjects") of undetermined stage and / or grade of cancer and / or unknown prognosis, and thus, optionally, also suitable for use in prognosing cancer in the test subjects.
[0349] In an additional or alternative option, the trained algorithm is then suitable for use in automatically and directly predicting a subject's cancer-specific prognosis (e.g., survival rate), and optionally used to make one or more further treatment decisions (such as a decision to engage in further treatment and / or a selection of the type of treatment) by identifying low-risk subjects who may have a more favorable prognosis and therefore less likely to benefit from further treatment, and / or by identifying high-risk subjects who are much more likely to benefit from further treatment, such as a more intensive treatment regimen.
[0350] Optionally, the one or more histological specimens may be from one or more test subjects who have undergone surgical treatment for cancer (e.g., tumor resection), and the trained algorithm may be used to automatically and directly predict the subject's cancer-specific prognosis (e.g., survival), and optionally make one or more further adjuvant treatment decisions.
[0351] Although any of the cancers described herein may be suitably evaluated, a cancer of particular interest is colon cancer, as exemplified herein.
[0352] More generally, by way of example, the training phase may involve the use of source histological images of histological specimens comprising biological material obtained from healthy subjects and / or diseased subjects (which may include, but are not limited to, the disease states discussed above in Section 2.3 of this application, and in embodiments of particular interest, may be forms of cancer as discussed above in Section 2.3.1 of this application), where the ground truth associated with each of the source histological images is known and can be provided to the system (100, 200, 300) during the training phase to be used to train machine learning algorithms within the system (100, 200, 300).
[0353] The ground truth associated with each of the source histological images is entirely at the discretion of the user and is not limited by the present invention.
[0354] However, in the context of a training phase involving, for example, source histopathological images derived from pathological subjects, the ground truth may optionally be selected from one or more of the following: (a) the presence or absence of a pathological condition in the subject; (b) the type, grade, and / or stage of the pathological condition of interest (e.g., distinguishing between different stages and / or grades of cancer); (c) the progression (or lack thereof) of the subject's pathological condition over a period of time following a defined event; (d) the subject's survival time following a defined event (typically excluding deaths not related to a specific pathological condition), and / or (e) the recurrence (or absence of recurrence) of the subject's pathological condition after previous treatment (e.g., surgery or non-surgical therapy) for that condition; Here, a "defined event" may be, for example, the time of collection of biological material from the subject from which the source histological image was generated, or the time of previous treatment for a pathological condition (e.g., surgical or non-surgical therapy, such as those discussed above in Section 2.3.2 of this application).
[0355] Optionally, the pathological condition can be a condition as described in Section 2.2 of the present application, for example, in a tissue, organ, or other body part as described in Section 2.2 of the present application. In an embodiment of particular interest, the pathological condition can be a cancer, for example, a solid cancer (e.g., carcinoma), as described in Section 2.3.1 above of the present application, a representative example of which is CRC, as shown in the Examples below. Optionally, if the pathological condition is cancer, the "defined event" can be the time of any previous treatment, for example, surgery (e.g., surgical removal of a tumor), as defined in Section 2.3.2 of the present application.
[0356] The time period following the defined event can be any time period of interest. In one embodiment, the time period can be 0 to 24 hours, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 hours. In another embodiment, the time period can be 0 to 28 days, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, or 28 days. In another embodiment, the time period can be 0 to 12 months, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months. In another embodiment, the period may be 1 year or more, and / or 0-10 years, 0-9 years, 0-8 years, 0-7 years, 0-6 years, 0-5 years, 0-4 years, 0-3 years, or 0-2 years, e.g., 0-20 years, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 years.
[0357] Thus, in one embodiment, in the context of a training phase involving source histopathological images derived from a pathological subject, the ground truth is the progression (or lack thereof) of a pathological condition (such as cancer) in the subject over a period following a defined event selected from the time of collection of biological material from the subject from which the source histopathological image was created, or the time of any previous treatment for the pathological condition, the period being as indicated above.
[0358] In another embodiment, in the context of a training phase involving source histopathological images derived from a pathological subject (such as a subject with cancer), the ground truth is the subject's survival time after a defined event, where the defined event is the time of harvesting biological material from the subject from which the source histological image was created, or the time of any prior treatment for the pathological condition (such as cancer), the time period being as indicated above.
[0359] If the algorithm is trained using source histopathological images of histopathological specimens where the ground truth associated with each image is a known prognosis for a particular pathological condition (such as cancer), the ground truth prognosis may take into account, for example: (i) condition-specific (e.g., cancer-specific) deaths within a defined period, and / or (ii) recurrence of a pathological condition (e.g., cancer) in a subject within a defined period of time.
[0360] The defined time periods for (i) and (ii) may be the same or different.
[0361] Truth data associated with a ground truth may be suitable for expressing the "category" of the ground truth, for example: Any ground truth prognostic information that passes a first predefined threshold (e.g., pathology-specific survival longer than a first predefined period and / or non-recurrence of the pathology for a second predefined period) may be considered to be associated with a "good prognosis" as the truth data for that sample. Any ground truth prognostic information that does not pass a second predefined threshold (e.g., pathology-specific survival shorter than a third predefined period, and / or pathology recurrence within a fourth predefined period) may be considered to be associated with a "poor prognosis" as the truth data for that sample. -And further, the first and second predefined thresholds may be the same or different; if they are different, a third category of truth data is possible, in which case the ground truth prognosis information passes the second predefined threshold but fails to pass the first predefined threshold, which may be considered to be associated with "uncertain prognosis" as the truth data for that sample.
[0362] In the training example described herein, histological specimens were obtained from cancer patients and ground truth data related to classifying each patient's known outcomes into prognostic groups. Patients in the clear prognosis group included patients with a good prognosis, a poor prognosis, or an unclear prognosis. Patients were defined as having a good prognosis if they were under 85 years of age at the time of surgery, had more than 6 years of follow-up after surgery, and had no documented cancer-specific deaths or documented recurrences. Patients were defined as having a poor prognosis if they were under 85 years of age at the time of surgery and experienced a cancer-specific death between 100 days (inclusive) and 2.5 years (exclusive) after surgery. Patients who did not meet any of these criteria were classified as the unclear prognosis group.
[0363] 4. Application of the system of the present invention As also discussed above, the computer-implemented systems (100, 200, 300) of the present invention can be used to process one or more source histological images (102, 202, 302), for example, one or more source histopathological images for which the outcome ("ground truth") is not known.
[0364] In this case, the system may apply a machine learning algorithm configured using training histological images (e.g., training histopathological images) such that the output of the system is a classifier (118, 318) and / or an overall classifier (232, 332) for the received source histological images 102.
[0365] As discussed further in Section 3 of this application, it will be appreciated that the nature of the classifier and / or the overall classifier will depend on the nature of the ground truth used during training.
[0366] For example, the classifier and / or the overall classifier can provide a diagnostic or prognostic decision for the subject for which the histological image is obtained.
[0367] Accordingly, the present invention also provides a computer-implemented method of processing one or more histological images according to the method of the present invention as further described above, said method being a method of generating a diagnosis and / or prognosis for a subject, said method comprising receiving one or more source histological images (202, 302) obtained from one or more histological samples obtained from a subject, said method comprising: determining an overall classifier (118, 318) for one or more source histological images (102, 302) according to the invention as further described above; and determining an overall classifier (232, 332) for one or more source histological images (202, 302) according to the invention as further described above; Optionally, attributing a diagnosis and / or prognosis assessment to the classifier and / or the overall classifier.
[0368] The subject can be any biological source, such as those described in Section 2.1 of this application, and in one preferred embodiment is a human.
[0369] One or more histological samples may be obtained from a subject having, suspected of having, being treated for, and / or previously having a pathological condition, such as a pathological condition described in Section 2.3 of the present application. In one embodiment of particular interest, the pathological condition is a cancer, such as those further described in Section 2.3.1 of the present application.
[0370] The one or more histological samples can be obtained from any tissue type, organ or other structure of interest of the subject, such as those described in Section 2.2 of this application.
[0371] Optionally, the or each histological sample obtained from the subject is in or has been obtained from a part of the body of a subject who has, is suspected of having, is being treated for, has been treated for, and / or previously had a pathological condition, e.g., a pathological condition described in Section 2.3 of the present application, more particularly, a cancer, such as a cancer further described in Section 2.3.1 of the present application.
[0372] In one preferred option, the method includes evaluating a plurality of source histological images (202, 302) obtained from a plurality of histological samples obtained from a subject. Optionally, the subject has, is suspected of having, is being treated for, has been treated for, and / or previously had a pathological condition such as cancer, and further optionally, the plurality of histological samples are samples obtained from a plurality of locations on the subject's body, the plurality of locations having, is suspected of having, and / or previously had biological material comprising the pathological condition, and / or is being treated for and / or previously treated for the pathological condition such as cancer, e.g., the pathological condition is cancer, and the plurality of histological samples may comprise or consist of multiple samples taken from the same tumor in the subject, in which case the method optionally allows for assessment of tumor heterogeneity.
[0373] In one embodiment, the method includes combining the determined classifier (118, 318) and / or the overall classifier (232, 332) with one or more additional diagnostic and / or prognostic markers for the pathological condition of interest.
[0374] For example, the pathological condition may be cancer, and the method may include combining the determined classifier (118, 318) and / or the overall classifier (232, 332) with one or more additional diagnostic and / or prognostic markers for the cancer. For example, if the cancer is CRC, one or more additional markers discussed in the examples of the present application may be used.
[0375] In the case where one or more such additional diagnostic and / or prognostic markers are evaluated, the step of attributing a diagnostic and / or prognostic assessment to the classifier (118, 318) and / or overall classifier (232, 332) may include evaluation of the results of the evaluation of the or each additional diagnostic and / or prognostic marker.
[0376] These methods can be used, for example, in methods for monitoring the progress or effectiveness of ongoing and / or previous treatments (such as surgery and / or therapy) and / or in methods for treating evaluated subjects, to make diagnoses, to generate prognoses, to stratify patient groups, and to assist in making subject treatment decisions based on diagnostic and / or prognostic assessments of stratified patient groups.
[0377] 5.Treatment method The present invention further provides a method of treatment in a subject in need thereof, wherein a diagnostic and / or prognostic assessment has been imparted to the subject by a method according to the present invention, said method comprising treating the subject by surgery and / or therapy.
[0378] Stated another way, the present invention provides surgical and / or therapeutic treatments (such as pharmaceutically acceptable compositions comprising one or more therapeutic agents) for use in methods of treating a subject in need thereof, wherein a diagnostic and / or prognostic assessment is imparted to the subject by a method according to the present invention.
[0379] Such methods of treatment can include, for example, therapeutic and / or prophylactic methods. Such types of treatment can include, for example, one or more forms of treatment discussed above in Section 2.3.2 of this application.
[0380] Preferably, the diagnostic and / or prognostic assessment attributed to the subject relates to a pathological condition, such as a pathological condition described in Section 2.3 of the present application. In one embodiment of particular interest, the pathological condition is a cancer, such as those further described in Section 2.3.1 of the present application.
[0381] The subject can be any organism, such as those described in Section 2.1 of this application, and in one preferred embodiment is a human.
[0382] The subject may have, be suspected of having, be being treated for, and / or have previously had a pathological condition, such as a pathological condition described in Section 2.3 of this application. In one embodiment of particular interest, the pathological condition is a cancer, such as those further described in Section 2.3.1 of this application.
[0383] Thus, the method of treatment can be a method of treating a pathological condition described in Section 2.3 of this application. In one embodiment of particular interest, the pathological condition is a cancer, such as those further described in Section 2.3.1 of this application.
[0384] In particularly interesting embodiments, the pathological condition is cancer, and the subject being treated has previously been treated for cancer (e.g., by surgery and / or non-surgical therapy) before the diagnostic and / or prognostic assessment attributed to the subject by the method according to the present invention. In this way, for example, the subject can be stratified, and the possible benefit of further treatment (e.g., adjuvant therapy, and / or further surgical and / or non-surgical therapy) can be determined before adopting a therapy for the subject.
[0385] The method of treatment can optionally include adapting one or more parameters of the surgical and / or non-surgical therapy in view of the diagnostic and / or prognostic assessment attributed to the subject. Optionally, the one or more parameters of the surgical and / or non-surgical therapy are selected from the group consisting of: the nature of the surgical and / or non-surgical therapy, the timing of the surgical and / or non-surgical therapy, the duration of the surgical and / or non-surgical therapy, the dosage of the non-surgical therapy, the route of administration of the non-surgical therapy, and the site in the body targeted by the surgical and / or non-surgical therapy.
[0386] In one option, the diagnostic and / or prognostic assessment of the subject includes evaluating the effect of previous or ongoing treatment with, for example, such therapy and / or surgery on the subject to monitor progression and / or effectiveness of the treatment, and further optionally, the method includes making a further treatment decision, such as discontinuing, continuing, repeating or modifying the previous or ongoing treatment and / or implementation of a different treatment modality, and further optionally, implementing the treatment decision for the patient. *********************************************
[0387] This invention is further illustrated by the following examples, which should not be construed as limiting the scope of the invention.
[0388] Example 1 method Training and Calibration Cohorts This study included (i) 160 patients with stage I, II, and III colon cancer treated at Akershus University Hospital, Norway, between 1988 and 2000, as described by Bondi et al., J Clin Pathol, 2005;58:509-14; (ii) 576 patients with stage I, II, and III CRC treated at Aker University Hospital, Norway, between 1993 and 2003, as described by Danielsen et al., Ann Oncol 2018;29:616-23; and (iii) 576 patients with stage I, II, and III CRC treated at Gloucester Colorectal Cancer Center, Norway, between 1988 and 1996, as described by Petersen et al., Gut, 2002;51:65-9 and Mitchard et al., Histopathology, 2010;57:671-9. The study utilized four training cohorts consisting of (i) 970 patients with stage I, II, and III CRC treated in the VICTOR trial, UK, from 2002 to 2004, as described in Kerr et al., N Engl J Med, 2007;357:360-9 and Midgley et al., J Clin Oncol, 2010;28:4575-80. These cohorts are further described in Section 1 of this application.
[0389] Patients were classified into a definite or unclear prognosis group according to their age at surgery and follow-up data. Patients in the definite prognosis group included those with a good or poor prognosis. Patients were defined as having a good prognosis if they were younger than 85 years at surgery, had more than 6 years of follow-up after surgery, and had no documented cancer-specific death or recurrence. Patients were defined as having a poor prognosis if they were younger than 85 years at surgery and had a cancer-specific death between 100 days (inclusive) and 2.5 years (exclusive) after surgery. Patients who did not meet any of these criteria were classified as the unclear prognosis group.
[0390] The training cohort consisted of 1652 WSIs from 828 patients with a defined prognosis across the four cohorts, and the calibration cohort consisted of 3280 WSIs from 1645 patients with an unclear prognosis. WSIs were prepared by laboratory technicians at Cancer Genetics and Informatics (ICGI), Norway. The demographics of the training and calibration patients are summarized in Table 1 below.
[0391] Study cohort The study cohort comprised 1824 WSIs prepared from 920 patients in the Gloucester Colorectal Cancer Study in Cheltenham, UK. WSIs were obtained from formalin-fixed, paraffin-embedded (FFPE) tumor tissue blocks different from those used in the training and calibration cohorts. The demographics of the study patients are summarized in Table 1 below.
[0392] Validation cohort The validation cohort consisted of 2,234 WSIs prepared by ICGI from 1,122 patients recruited in the QUASAR 2 trial (Kerr et al, Lancet Oncol, 2016;17:1543-57).
[0393] The open-label, randomized, controlled QUASAR 2 trial (ISRCTN registration number ISRCTN45133151) enrolled 1,952 patients with histologically proven stage III or high-risk stage II colorectal cancer from 170 hospitals in seven countries (Australia, Austria, Czech Republic, New Zealand, Serbia, Slovenia, and the UK) between April 2005 and October 2010, of whom 1,941 had evaluable data (Kerr et al., 2016, supra).
[0394] This trial was designed to investigate whether bevacizumab improved disease-free survival after potentially curative surgery for the primary tumor. All patients received adjuvant chemotherapy in the form of capecitabine, but none received neoadjuvant treatment. No significant differences were observed between treatment groups, and the researchers concluded that the addition of bevacizumab to capecitabine should not be used in this adjuvant setting (Kerr et al., 2016, supra).
[0395] FFPE tissue blocks were collected from 1,251 QUASAR 2 patients with stage II or III colorectal cancer, with recommended but not required blood and tumor samples from primary resection. These patients were representative of the entire study population in terms of clinical and pathological characteristics (Kerr et al., 2016, supra). Pathological assessment was performed by pathologists at participating hospitals. All patients provided written informed consent for treatment and use of tissue samples. The West Midlands Research Ethics Committee (no. 04 / MRE / 11 / 18) and the Norway Regional Committees for Medical and Health Research Ethics (REK) (no. 2015 / 1607) approved this study.
[0396] Tissue blocks from 1,140 patients were received, sectioned, and prepared as 3-μm H&E-stained tissue slides by ICGI laboratory technicians (Figure 18). A local pathologist blinded to the clinical results confirmed the presence of tumor in each tissue section. Digital images of 1,132 tumor-bearing sections were acquired using the same two scanners as the training cohort: the Aperio AT2 and the NanoZoomer XR. A previously developed segmentation model was blindly applied to automatically identify tumor-bearing regions, resulting in 1,113 patients with Aperio AT2 segmentation and 1,121 patients with NanoZoomer XR segmentation (Figure 18). Although the slide images are tiled as in the training cohort, no 10x tiles could be fitted within the automated tumor segmentation for three Aperio AT2 segmentations and two NanoZoomer XR segmentations (Figure 18). The QUASAR 2 cohort was defined as 40x tiles from Aperio AT2 slide images (available for 1,113 patients), 10x tiles from Aperio AT2 slide images (available for 1,110 patients), 40x tiles from NanoZoomer XR slide images (1,121 patients), and 10x tiles forming NanoZoomer XR slide images (available for 1,119 patients).
[0397] The QUASAR 2 cohort represents patients eligible for the QUASAR 2 trial. Eligible patients had to meet all of the following inclusion criteria (originally described in Kerr et al., 2016, supra): -18 years of age or older. - Colorectal adenocarcinoma. Histologically proven R0 M0 stage III or high-risk stage II colorectal cancer, defined as the presence of one or more of the following poor prognostic features: T4 stage, lymphatic invasion, vascular invasion, peritoneal involvement, poor differentiation, and preoperative obstruction or perforation of the primary tumor. Primary resection 4 to 10 weeks prior to randomization. - World Health Organization (WHO) Performance Status of 0 or 1. - A life expectancy of at least 5 years, taking into account comorbidities, excluding cancer risk.
[0398] Additionally, eligible patients could not meet any of the following exclusion criteria (originally described in Kerr et al., 2016, supra): - History of cancer other than treatment for carcinoma in situ, basal cell carcinoma, or squamous cell carcinoma of the cervix, or disease-free interval greater than 10 years after previous cancer. - Inflammatory bowel disease and / or active peptic ulcer requiring treatment within the past 2 years. - Lack of physical integrity of the upper gastrointestinal tract, malabsorption syndrome, or inability to take oral medications. - Moderate or severe renal impairment (creatinine clearance <30 mL / min). Any of the following hematological abnormalities: Absolute neutrophil count <1.5 x 109 / L. ○Platelet count <100×109 / L. o Total bilirubin concentration > 1.5 times the upper limit of normal (ULN). o Alanine aminotransferase, aspartate aminotransferase, or alkaline phosphatase concentrations >2.5 times the upper limit of normal (ULN). - Proteinuria > 500 mg per 24 hours. - Patients who have received previous chemotherapy, immunotherapy, or subdiaphragmatic radiation therapy (including neoadjuvant therapy to the rectum), or who are expected to require radiation therapy to these sites within the next 12 months. Use of any investigational drug or medication / procedure within 4 weeks of randomization. - Chronic use of full-dose anticoagulants, high-dose aspirin (>325 mg / day), antiplatelet agents, or known bleeding diathesis (low-dose aspirin is permitted). - Concomitant treatment with sorivudine or its chemically related analogues. - History of uncontrolled seizures, central nervous system disorders, or psychiatric history that would interfere with informed consent or compliance with oral medication intake. Clinically significant cardiovascular disease, i.e., active or less than 12 months since, for example, cerebrovascular accident, myocardial infarction, unstable angina, New York Heart Association (NYHA) grade II or greater congestive heart failure, serious cardiac arrhythmia requiring medication, or uncontrolled hypertension. - Known coagulation disorder. Known allergy to Chinese hamster ovary cell proteins or other recombinant human or humanized antibodies, or to any excipients in bevacizumab formulations. - Pregnant or lactating women, or premenopausal women not using contraception.
[0399] The demographics of the study patients are summarized in Table 1 below. [Table 3]
[0400] Sample preparation 3 μm FFPE tissue block sections were stained with hematoxylin and eosin (H&E) and confirmed to contain tumor by a pathologist (MP).
[0401] WSIs were acquired with two scanners, an Aperio AT2 (Leica Biosystems, Germany) and a NanoZoomer XR (Hamamatsu Photonics, Japan), at the highest available resolution (designated 40×).
[0402] Regions of high tumor content were identified by automated segmentation methods (described in Section 1 of this application). Typically, a 40x resolution WSI contains on the order of 100,000x100,000 pixels, which is several orders of magnitude larger than images currently achievable for classification by deep learning methods. To preserve the prognostic information contained at high resolution, the WSI was divided into multiple non-overlapping image regions, called tiles, at 10x and 40x resolution, with each 40x pixel measuring approximately 0.24x0.24 μm. 2 Represents the physical size of
[0403] classification In the training cohort, five convolutional neural networks were trained with 10x tiles of 634,564 and five with 40x tiles of 11,591,555. All networks were DoMorev1 networks, a specialized network for classifying very large, heterogeneous images, and consisted of a MobileNetV2 (Sandler et al., 2018) representation network, a Noisy-AND pooling function (Krauss et al., 2016), and a fully connected classification network (Figure 3).
[0404] Traditional classification networks are trained on images, each of which has an associated label. One approach is to have each tile inherit the label of its WSI, but due to spatial heterogeneity, tiles do not necessarily contain predictive information that reflects the WSI. Using ideas from multiple instance learning, we instead train on ensembles of tiles from WSIs, each of which is labeled with the WSI. By using a novel gradient approximation method, we can train the network end-to-end with an increasing number of tiles representing each WSI during training.
[0405] The network was trained beyond convergence and evaluated at 21 equidistant positions in the training progression, resulting in 21 models per training run. Model performance on the calibration cohort was used to select one model from each training run, resulting in five models per resolution.
[0406] To evaluate the WSI, each of the five selected models provided a prediction of the probability of poor prognosis, and the average predicted probability was defined as the predicted probability of the ensemble model. The appropriate thresholds for the predicted probabilities of two ensemble models (one 10x and one 40x) were determined by evaluating the training cohort, whereby the two ensemble markers predicted either good or poor prognosis.
[0407] The performance of these two ensemble markers and several other candidate markers was evaluated in the test cohort, and a combination of two ensemble markers was selected for the primary analysis of the validation cohort. This combination, called DoMore-v1-CRCmarker, predicted a favorable prognosis when both ensemble markers predicted a favorable prognosis, an uncertain prognosis when the ensemble markers made different predictions, a poor prognosis when both ensemble markers predicted a poor prognosis, and no prediction was defined when the ensemble markers could not be evaluated due to missing tiles.
[0408] Primary analysis We predefined the primary analysis of the DoMore-v1-CRC marker (embodied as the overall classifier (232, 332) in Figures 2 and 3) for both scanners in the validation cohort. The metric chosen to measure model performance was the hazard ratio (HR) with 95% confidence interval (CI) for patients predicted to have a poor prognosis compared with patients predicted to have an uncertain prognosis and those predicted to have a good prognosis. The two HRs were calculated by analyzing a Cox proportional hazards model (DoMore-v1-CRC markers were included as categorical variables, i.e., the model consisted of two indicator variables: uncertain prognosis and poor prognosis) with the DoMore-v1-CRC marker as the only variable, and cancer-specific survival (CSS) as the endpoint (using Efron's method for related events).
[0409] The study selected to evaluate whether the DoMore-v1-CRC marker predicts CSS was a two-sided Mantel-Cox log-rank test with a significance level of 0.05. Time to CSS was calculated from the date of randomization to the date of cancer-specific death or loss to follow-up. The primary analysis was an unbiased evaluation of the ability of the DoMore-v1-CRC marker to predict CSS in the target population of patients who received adjuvant chemotherapy (specifically capecitabine) and met the eligibility criteria for the QUASAR 2 trial (as described above).
[0410] statistical analysis Primary and secondary analyses were planned prior to evaluation in the validation cohort and described in the protocol. Markers were included in multivariable models if available at the time of analysis and were significant in the univariate analysis of CSS. All reported CIs have a confidence level of 95%. A two-sided P < 0.05 was considered statistically significant. Survival analyses were performed in Stata / SE15.1 (StataCorp, TX).
[0411] result DoMore-v1-CRC was a statistically significant marker of CSS in the primary analysis of the validation cohort for both the Aperio AT2 scanner (HR for patients with an uncertain prognosis: 1.89, CI: 1.14–3.15; HR for patients with a poor prognosis: 3.84, CI: 2.72–5.43; p<0.0001 (Figure 16A)) and the NanoZoomer XR scanner (HR for patients with an uncertain prognosis: 2.42, CI: 1.45–4.03; HR for patients with a poor prognosis: 3.39, CI: 2.36–4.87; p<0.0001 (Figure 16B)). Results for the Aperio AT2 scanner are shown below. Corresponding analyses based on the NanoZoomer XR scanner, as well as results based on the 10x and 40x resolution levels of the two scanners, were also significant (data not shown).
[0412] DoMore-v1-CRC significantly predicted CSS after adjusting for the covariates pN stage, pT stage, lymphatic invasion, and venous vascular invasion in multivariate analysis (HR for poor vs. good prognosis: 3.04, CI: 2.07-4.47 (Table 2)). DoMore-v1-CRC correlated with many established prognostic factors, such as age, pN stage, pT stage, histological grade, location, laterality, BRAF mutation, and microsatellite instability, but was not associated with sex, vascular invasion, venous vascular invasion, or KRAS mutation (Table 3). [Table 4] [Table 5]
[0413] CSS was significantly predictive of DoMore-v1-CRC in stage II (HR for poor prognosis was 2.71, CI 1.25-5.86 (Figure 16C)) and stage III (HR for poor prognosis was 4.09, CI 2.77-6.03 (Figure 16D)).
[0414] Additionally, the binary DoMore-v1-CRC marker significantly identified patients at high risk of cancer death in stages IIIA, IIIB, and IIIC (data not shown), as well as in pN (Figure 16E and additional data not shown) or pT (pT1-3 vs. pT4 (Figure 16F and additional data not shown)) stages. The binary DoMore-v1-CRC marker provided a similar HR for poor prognosis prediction as the standard DoMore-v1-CRC marker (data not shown).
[0415] Inception v3, a state-of-the-art convolutional neural network, was trained, tuned, and evaluated in the same research setup as DoMore-v1-CRC (see discussion in Section 1 of this application) and provided a statistically significant marker of CSS, performing slightly worse than DoMore-v1-CRC (data not shown).
[0416] In a test cohort using sections from fresh tumor blocks prepared at another hospital, DoMore-v1-CRC significantly identified patients at increased risk of cancer death (HR for poor vs. favorable prognosis: 4.83, CI: 3.27–7.12 (Figure 17A)).
[0417] Robustness to laboratory preparation was also evident when analyzing DoMore-v1-CRC in stages II and III (Figure 17C-D). The dichotomous DoMore-v1-CRC also provided a significant predictor of CSS (Figure 17B).
[0418] It is important to remember that all patients in the QUASAR 2 validation cohort received adjuvant chemotherapy with capecitabine (the addition of bevacizumab did not affect disease-free or overall survival), which explains the observation that survival curves are generally better, as only a small number of patients received chemotherapy in the test cohort.
[0419] Consideration Building on recent developments in machine learning (LeCun et al., Nature, 2015;521:436-44), we developed a fully automated system for predicting patient outcomes, such as those of patients with CRC, using standard laboratory H&E-stained tissue sections. Our method first outlines pathological (e.g., cancerous) tissue within the image and then stratifies patients into prognostic categories, which, in validation, differed by 3-4 fold in HRs of disease-specific mortality.
[0420] Deep learning has already been shown to be suitable for the detection and delineation of several tumor types (Ehteshami Bejnordi et al. JAMA, 2017;318:2199-210) and various cancer classifications have been reported (Coudray et al. Nat Med, 2018;24:1559-67). However, we have not yet seen a validated system for directly predicting patient outcomes based on histological images.
[0421] Automated prognostic procedures have the potential to reduce human intervention and increase the objectivity and reproducibility of prognoses. Furthermore, with the increasing robotization of wet-lab procedures, analytical throughput increases, allowing decisions to be based on multiple samples from tumors. This may mitigate the problem of tumor heterogeneity, which may be key to improving prognostic accuracy.
[0422] MMR status (Sinicrope, Nat Rev Clin Oncol,2010;7:174-7, Mouradov et al.,Am J Gastroenterol,2013;108:1785-93), stromal estimation (Danielsen et al.,Ann Oncol,2018;29:616-23), lymphatic invasion (Akagi et al.,Anticancer Res,2013;33:2965-70), RNA profile (Salazar et al.,J Clin Oncol,2011;29:17-24, Gray et al.,J Clin Oncol,2011;29:4611-9), mutation burden (Mouradov et al.,Am J Gastroenterol, 2013;108:1785-93), while they may be biologically plausible, do not perform as well as DoMore-v1-CRC, which outperforms these markers in terms of HR and provides more clinically useful stratification and distribution of patients among risk groups.
[0423] DoMore-v1-CRC is technically easy to apply and can be delivered in any standard pathology laboratory. While training the network requires resources, applying the DoMore-v1-CRC marker to new patient slides is a much smaller computational task and can be performed in clinical practice within 10 minutes using consumer hardware. The clinical utility of the marker is its ability to guide discussions with patients about the pros and cons of various treatment options (chemotherapy doses / schedules). While the number of drugs used in adjuvant therapy is currently limited to fluoropyrimidines ± oxaliplatin, recent data suggest that 3 months of treatment provides survival outcomes similar to 6 months for the majority of patients with stage III disease, and that high-risk patients (pT4 and pN2) may benefit from longer-term treatment (Grothey et al., N Engl J Med, 2018;378:1177-88; Iveson et al., Lancet Oncol, 2018;19:562-7).
[0424] The proportional reduction in HR for CRC recurrence and death after adjuvant therapy is remarkably consistent at 20% across most well-designed clinical trials. However, this translates into very different absolute survival improvements for low-risk and high-risk subgroups. Although there are no prospective adjuvant trials for these risk-stratified groups, this does not prevent clinicians from interpreting existing trial data and applying them to individuals proven to be at low or high risk of recurrence.
[0425] Taking Figure 16C as an example, these data can be interpreted to suggest that stage II individuals in the poor-risk group (approximately 20%) would benefit from single-agent fluoropyrimidines, such as capecitabine, whereas the favorable-risk group is more likely to be cured by surgery alone. Patients with stage III disease, pN2 and pT4, exhibit more diverse survival curves (Figures 16D-F). For patients who are poorly suited for favorable and possibly uncertain combination chemotherapy, the surviving group exhibits very reasonable survival with single-agent capecitabine, whereas the poor-risk group may benefit more from 3 or 6 months of combination chemotherapy (absolute survival benefit of approximately 8-10%). Clearly, these survival curves are indicative and can be used by clinicians and patients to collaboratively make more informed decisions regarding the choice of adjuvant chemotherapy.
[0426] In summary, using deep learning techniques associated with digital scanning of conventional H&E-stained FFPE tumor tissue sections, it was possible to develop a clinically useful prognostic marker. This assay has been extensively evaluated in large, independent patient populations, correlates with and outperforms existing molecular and morphological prognostic markers, provides consistent results across tumor and lymph node stages, and can be used by clinicians to support decision-making regarding adjuvant treatment options.
[0427] We also demonstrated that the deep learning approach described herein can be used to predict cancer-specific survival (CSS) fully automatically and directly from scanned histopathological images (illustrated using hematoxylin and eosin-stained, formalin-fixed, paraffin-embedded tumor tissue sections). Independent validation of the classifier showed that the demonstrated system stratifies stage II and III colon cancer patients into distinct prognostic groups, complementing established prognostic markers and outperforming most existing markers in terms of hazard ratios. This system can be used to improve adjuvant therapy selection after resection of cancerous tissue (e.g., colon cancer) by identifying very low-risk patients who may be cured by surgery alone and high-risk patients who are much more likely to benefit from a more intensive regimen.
[0428] Therefore, this application describes an approach that provides a fully automated deep learning system for predicting patient outcomes from conventional histopathological images, allowing stratification of those individuals and making enhanced treatment decisions.
[0429] Example 2 This example provides additional supporting data showing preliminary results of applying the methods of the present invention to the assessment of lung cancer.
[0430] Experiments similar to those described in Section 1.3 and Example 1 were performed using whole slide images of lung tissue from lung cancer patients collected at Oslo University Hospital. For the colon cancer example, one tissue block per patient was used, from which one section was scanned with two scanners (Aperio AT2 and NanoZoomer XR) to generate two whole slide images per patient.
[0431] The methods and experiments were as described in Section 1.3 and Example 1, with the following exceptions. - All patients are from the same cohort - The study cohort will consist only of patients with either a clear favorable outcome or a clear unfavorable outcome Manual tumor segmentation was used to define the tiling area. [Table 6]
[0432] A survival analysis of the outcomes in the test cohort is shown in Figure 19. This is one Kaplan-Meier plot for each of the two scanners included in the experiment. This is the same format as the analysis shown in Figure 16 for the colon cancer experiment. The plot shows cancer-specific survival, i.e., the survival rate free of lung cancer death, for the three predicted outcome groups. The estimated survival probability can be read off the y-axis, while the x-axis represents time since surgery. As can be seen from the plot, there is a large difference in cancer-specific survival when comparing the predicted good-prognosis and poor-prognosis groups. The uncertain group has a cancer-specific survival rate between the good-prognosis and poor-prognosis groups.
[0433] Note that the number of test patients is smaller than in the experiment in Example 1 (see Table 4). Also note that, unlike in Example 1, the test patients are from the same cohort as the training patients, so any weight that may be attributed to these results may be further confirmed in future evaluations by selecting test patients and training training patients from different cohorts. On the other hand, compared to Example 1, although fewer patients were used for training and all came from the same cohort, the results could be further improved by using a larger, more diverse training cohort.
[0434] Therefore, although further steps to validate in an independent cohort would be optimal to critically evaluate the performance of the method on lung cancer samples, the preliminary results provided here clearly demonstrate that the method can predict outcomes in patients with forms of cancer other than CRC, including lung cancer. *********************************************
[0435] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. Any reference signs in the claims should not be construed as limiting the scope. [Explanation of symbols]
[0436] 100 Computer-implemented systems 102 Source Histological Images 104 Tile Generator 106 tiles 108 The First Neural Network 110 Tile Features 111 Network Architecture 112 Pooling Functions 113 Processing Block 114 Bag Features 116 The Second Neural Network 118 Classifier 120 Truth Data 126 Loss Function 200 Computer-implemented systems 202 Source Histological Images 204 First Tile Generator 205 Second Tile Generator 206 First Tile 207 Second Tile 211 Machine Learning Network 215 Machine Learning Networks 218 First Classifier 219 Second Classifier 222 Segmentation Block 232 Global classifier 300 System 302 Source Histological Images 306 First Tile 307 Second Tile 308 Representation Network (First Neural Network) 310 Tile Features 311 Machine Learning Network 312 Pooling Functions 316 Classification Network (Second Neural Network) 318 First Classifier 319 Second Classifier 320 Truth Data 324 WSI histopathological images 330 Classifier synthesizer 332 Global classifier 340 Averaged First Classifier 341 Averaged Second Classifier 342 First Thresholding Function 343 Second Thresholding Function 608 Expression Network 614 tiles 618 Classifier
Claims
1. 1. A computer-implemented system (300) for determining a global classifier (332) of one or more source histopathological images (302), comprising: the or each source histopathological image (302) is obtained from one or more histopathological samples obtained from one or more subjects, the or each subject having cancer, being diagnosed with cancer, being suspected of having cancer, being treated for cancer, having previously been treated for cancer, and / or having previously had cancer; The system comprises: a first tile generator configured to generate a plurality of first tiles (306) from the one or more source histopathological images (302), each of the plurality of first tiles (306) including a plurality of pixels representing a region of the one or more source histopathological images having a first area and a first resolution; a second tile generator configured to generate a plurality of second tiles (307) from the one or more source histopathological images (302), each of the plurality of second tiles (307) comprising a plurality of pixels representing a region of the one or more source histopathological images having a second area and a second resolution; the first area of the first tile (306) is greater than the second area of the second tile (307); a second tile generator, wherein the second resolution of the second tiles (307) is higher than the first resolution of the first tiles (306); a machine learning network (211, 311) configured to process the plurality of first tiles (306) to determine a first classifier (318) for the one or more source histopathological images (302), the machine learning network (311) comprising: a first neural network (308) configured to process the plurality of first tiles (306) to determine tile features (310) for each of the plurality of first tiles (306); a pooling function (312) configured to combine subsets of the tile features to generate a bag feature (314) for each of the subsets; a machine learning network (211, 311) comprising a second neural network (316) configured to process the bag features (314) to determine a first classifier (318) for the one or more source histopathological images (302), the second neural network (316) being a classification network; a machine learning network (215, 311) configured to process the plurality of second tiles (307) to determine a second classifier (319) for the one or more source histopathological images (302), the machine learning network (215, 311) comprising: a first neural network (308) configured to process the plurality of second tiles (307) to determine tile features (310) for each of the plurality of second tiles (307); a pooling function (312) configured to combine subsets of the tile features to generate a bag feature (314) for each of the subsets; a machine learning network (215, 311) comprising a second neural network (316) configured to process the bag features (314) to determine a second classifier (319) for the one or more source histopathological images (302), the second neural network (316) being a classification network; a classifier combiner configured to combine the first classifier (318) and the second classifier (319) to determine an overall classifier (332) for the one or more source histopathological images (302).
2. The classifier combiner (230, 330) applying a thresholding function to the first classifier (218, 318) to determine a thresholded first classifier; applying a thresholding function to the second classifier (219, 319) to determine a thresholded second classifier; and combining the thresholded first classifier and the thresholded second classifier to determine an overall classifier (232, 332).
3. the machine learning network (211, 311) is configured to process the plurality of first tiles (206, 306) to determine a plurality of first classifiers (218, 318) for the one or more source histopathological images (202, 302); the machine learning network (215, 311) is configured to process the plurality of second tiles (207, 307) to determine a plurality of second classifiers (219, 319) for the one or more source histopathological images (202, 302); The classifier combiner (230, 330) applying a statistical function to the plurality of first classifiers (218, 318) to determine a combined first classifier (340); applying a statistical function to the plurality of second classifiers (219, 319) to determine a combined first classifier (341); and combining the combined first classifier (340) and the combined second classifier (341) to determine the overall classifier (232, 332).
4. 2. The system of claim 1, wherein the classifier combiner (330) is configured to perform a logical combination of the first classifier (218, 318) and the second classifier (219, 319) to determine the overall classifier (232, 332) for the one or more source histopathological images (202, 302).
5. comparing the classifier (118, 318) determined by the second neural network (116, 316) with ground truth represented by truth data (120, 320); 10. The system of claim 1, further comprising: a loss function configured to: set trainable parameters for the first neural network, the pooling function, and the second neural network based on a result of the comparison.
6. The system of any one of claims 1 to 5, further comprising a segmentation block (122) configured to apply an image segmentation method to the whole slide image histopathology image (124, 324) to provide a source histopathology image (102, 302).
7. The system of claim 1 , wherein the first tile generator (204) and the second tile generator (205) are configured to generate their respective tiles independently of each other.
8. the machine learning network (211, 311) is trained using training histopathological images and associated ground truth; 2. The system of claim 1, wherein the or each training histopathological image (302) is obtained from one or more histopathological samples obtained from one or more subjects, and the or each subject has cancer, has been diagnosed with cancer, is suspected of having cancer, is being treated for cancer, has previously been treated for cancer, and / or has previously had cancer.
9. 1. A computer-implemented method for determining a global classifier (332) for one or more source histopathological images (302), comprising: the or each source histopathological image (302) is obtained from one or more histopathological samples obtained from one or more subjects, the or each subject having cancer, being diagnosed with cancer, being suspected of having cancer, being treated for cancer, having previously been treated for cancer, and / or having previously had cancer; The method comprises: generating a plurality of first tiles (306) from the one or more source histopathological images (302), each of the plurality of first tiles (306) including a plurality of pixels representing a region of the one or more source histopathological images having a first area and a first resolution; generating a plurality of second tiles (307) from the one or more source histopathological images (302), each of the plurality of second tiles (307) including a plurality of pixels representing a region of the one or more source histopathological images having a second area and a second resolution; the first area of the first tile (306) is greater than the second area of the second tile (307); generating a second resolution of the second tile (307) that is higher than the first resolution of the first tile (306); applying a machine learning network (211, 311) to the plurality of first tiles (306) to determine a first classifier (318) for the one or more source histopathological images (302), wherein applying the machine learning network (211, 311) includes: applying a first neural network (308) to the plurality of first tiles (306) to determine tile features (310) for each of the plurality of first tiles (306); combining the subsets of tile features to generate a bag feature (314) for each of the subsets; applying a second neural network (316), which is a classification network, to the bag features (314) to determine a first classifier (318) for the one or more source histopathological images (302); applying a machine learning network (215, 311) to the plurality of second tiles (307) to determine a second classifier (319) for the one or more source histopathological images (302), wherein applying the machine learning network (311) includes: applying a first neural network (308) to the plurality of second tiles (307) to determine tile features (310) for each of the plurality of second tiles (307); combining the subsets of tile features to generate a bag feature (314) for each of the subsets; applying a second neural network (316), which is a classification network, to the bag features (314) to determine a second classifier (319) for the one or more source histopathological images (302); combining the first classifier (318) and the second classifier (319) to determine the overall classifier (332) for the one or more source histopathological images (302).
10. 10. The computer-implemented method of claim 9 for processing one or more histopathological images, comprising: wherein the method is a method of making a diagnostic and / or prognostic determination for a subject who has cancer, has been diagnosed with cancer, is suspected of having cancer, is being treated for cancer, has previously been treated for cancer, and / or has previously had cancer; The method includes receiving one or more source histopathological images (102, 202, 302) obtained from one or more histopathological samples (typically ex vivo histopathological samples) obtained from the subject, the method comprising: Determining the overall classifier (232, 332) for the one or more source histopathological images (202, 302) according to the method of claim 9; and attributing a diagnosis and / or prognosis assessment to said classifier (118) and / or said overall classifier (232, 332).
11. The method of claim 10 , wherein the subject is a human.
12. 12. The method of claim 10 or 11, wherein the or each histopathological sample obtained from the subject is obtained from a part of the body of the subject that has cancer, is suspected of having cancer, is being treated for cancer, has previously been treated for cancer, and / or has previously had cancer.
13. The system of any one of claims 1 to 8 or the method of any one of claims 9 to 12, wherein the cancer is selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed cancers.
14. 14. The system or method of claim 13, wherein the cancer is colon cancer or lung cancer.
15. evaluating a plurality of source histopathological images (102, 202, 302) obtained from a plurality of histopathological samples obtained from the subject to determine a plurality of classifiers and / or an overall classifier; Optionally, attributing said diagnosis and / or prognosis assessment to said plurality of classifiers and / or an overall classifier.
16. the method comprises assessing one or more additional diagnostic and / or prognostic markers for the cancer; 16. The method according to any one of claims 10 to 15, wherein the step of attributing a diagnostic and / or prognostic assessment to said classifier and / or said overall classifier comprises evaluation of the result of the or each evaluation of the or each further diagnostic and / or prognostic marker.
17. the method further comprising making a treatment decision for the subject based on the diagnostic and / or prognostic assessment; 17. The method of any one of claims 10 to 16, optionally wherein the treatment decision relates to a diagnosed or prognosticated cancer condition, such as a cancer selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed cancers, and optionally wherein the cancer is colon cancer or lung cancer.
18. 18. A method of treating in a subject in need thereof, wherein a diagnosis and / or prognosis assessment has been imparted to said subject by the method of any one of claims 10 to 17, said method comprising treating said subject by surgery and / or non-surgical therapy, The method, wherein the treatment of the diagnosed or prognosticated pathological condition is, for example, treatment of a cancer selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed cancers, and optionally, the cancer is colon cancer or lung cancer.
19. 19. The method of claim 18, wherein the subject is a human.
20. The object is (a) has cancer, has been diagnosed with cancer, is suspected of having cancer, is being treated for cancer, has previously been treated for cancer, and / or has previously had cancer; and / or (b) a diagnosis and / or prognosis assessment of a cancer condition is imputed to said subject by a method according to any one of claims 10 to 17.
21. 21. The method of claim 20, wherein the pathological condition is a cancer selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed cancers, and optionally, the cancer is colon cancer or lung cancer.
22. The method comprises adapting one or more parameters of the surgical and / or non-surgical therapy taking into account the diagnostic and / or prognostic assessment attributed to the subject by the method of any one of claims 10 to 17, and optionally 22. The method of any one of claims 18 to 21, wherein the one or more parameters of the surgical and / or non-surgical therapy are selected from the group consisting of: the nature of the surgical and / or non-surgical therapy, the timing of the surgical and / or non-surgical therapy, the duration of the surgical and / or non-surgical therapy, the dosage of the therapy, the route of administration of the non-surgical therapy, and the site in the body targeted by the surgical and / or non-surgical therapy.
23. said diagnostic and / or prognostic assessment of said subject comprises evaluating the effect of previous or ongoing treatment with surgical and / or non-surgical therapy on said subject; For example, to monitor the progress and / or effectiveness of such treatment, and further optionally: Optionally, the method further comprises making a further treatment decision, such as discontinuing, continuing, repeating or modifying a previous or ongoing treatment and / or the implementation of a different treatment modality; implementing the further treatment decision for the subject; 23. The method of any one of claims 17 to 22, wherein said diagnostic and / or prognostic assessment, said treatment and / or said treatment decision relates to a cancerous condition, such as a cancer selected from the group consisting of carcinoma, sarcoma, myeloma, leukemia, lymphoma, and mixed cancers, and optionally said cancer is colorectal cancer.
Citation Information
Patent Citations
Method and Apparatus for Automatically Establishing High Performance Classifiers for Generating Medically Meaningful Descriptions in Medical Diagnostic Imaging
JP2008523876A
Method for image analysis, image analyzer, program, method for manufacturing learned deep learning algorithm, and learned deep learning algorithm
JP2019148473A
Scaling up convolutional networks
US20170161891A1