Artificial Intelligence Architecture for Predicting Cancer Biomarkers
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-07
- Publication Date
- 2026-03-16
AI Technical Summary
Conventional methods for detecting clinically actionable biomarkers for personalized cancer treatment rely heavily on genomic sequencing, which is costly, time-consuming, and not universally accessible.
A deep learning architecture that uses multi-resolution convolutional neural networks to predict cancer biomarkers directly from digital images of hematoxylin and eosin (H&E) stained tissue slides, eliminating the need for DNA sequencing.
This approach enables accurate prediction of biomarkers with at least 80% accuracy compared to genomic sequencing, reducing the need for costly and time-consuming sequencing processes and making biomarker detection more accessible.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This patent document claims the priority and benefit of U.S. Provisional Patent Application No. 63 / 269,033, “ARTIFICIAL INTELLIGENCE ARCHITECTURE FOR PREDICTING CANCER BIOMARKERS,” filed on March 8, 2022, and U.S. Provisional Patent Application No. 63 / 483,237, “ARTIFICIAL INTELLIGENCE ARCHITECTURE FOR PREDICTING CANCER BIOMARKERS,” filed on February 3, 2023. The entire contents of the foregoing patent applications are incorporated by reference as part of the disclosure of this patent document.
[0002] The disclosed technology relates to systems and methods for detecting clinically actionable biomarkers.
Background Art
[0003] Conventional methods for detecting clinically actionable biomarkers for personalized cancer treatment rely largely on genomic sequencing or genotyping platforms (e.g., microarrays, targeted panel sequencing, whole exome sequencing, or whole genome sequencing). Researchers are conducting studies to avoid conventional sequencing techniques.
Summary of the Invention
[0004] Disclosed are materials, systems, devices, and methods for predicting cancer biomarkers using deep learning architectures.
[0005] In some embodiments of the disclosed technology, a method for determining the presence of a biomarker in a biological sample includes obtaining a section of the biological sample treated with a stain, imaging one or more regions of the stained section of the biological sample at a first resolution and a second resolution to generate first and second pluralities of image data, reducing the parameter spaces of the first and second pluralities of image data to generate reduced first and second pluralities of image data, providing the first and second pluralities of image data to a trained prediction neural network, and determining the presence of the biomarker in the biological sample as an output of the trained prediction neural network.
[0006] In some embodiments of the disclosed technology, a method for generating a trained prediction model configured to determine the presence of a biomarker in a biological sample includes generating stained sections of one or more biological samples and corresponding biomarker labels, imaging one or more regions of the stained sections of the one or more biological samples at a first resolution and a second resolution to generate first and second pluralities of image data, reducing the parameter spaces of the first and second pluralities of image data to generate reduced first and second pluralities of image data, and generating a trained prediction model including a first prediction model trained with the reduced first plurality of image data and the corresponding biomarker labels and a second prediction model trained with the reduced second plurality of image data and the corresponding biomarker labels.
[0007] In some embodiments of the disclosed technology, a computer system configured to determine the presence of biomarkers in a biological sample includes one or more processors and, as a result of execution, images one or more regions of a stained section of the biological sample at a first resolution and a second resolution to generate first and second pluralities of image data, reduces the parameter spaces of the first and second pluralities of image data to generate reduced first and second pluralities of image data, provides the first and second pluralities of image data to a trained prediction model, and determines the presence of biomarkers in the biological sample as an output of the trained prediction model. A non-transitory computer-readable storage medium storing software including executable instructions that cause the one or more processors of the computer system to perform the foregoing is provided.
[0008] In some embodiments of the disclosed technology, a method of determining the presence of biomarkers in a biological sample includes obtaining a stained section of the biological sample, imaging one or more regions of the stained section of the biological sample to generate a plurality of images of the stained section, providing the plurality of images of the stained section as an input to a trained prediction model, and determining the presence of biomarkers in the biological sample as an output of the trained prediction model, wherein the trained prediction model is configured such that a preset accuracy for determining the presence of biomarkers is set to at least 80% of genome sequencing.
[0009] In some embodiments of the disclosed technology, a treatment method for treating a target cancer as needed includes obtaining a stained section of a biological sample, imaging one or more regions of the stained section of the biological sample to generate a plurality of images of the stained section, providing the plurality of images of the stained section to a trained prediction model, and determining the presence of a biomarker in the biological sample as an output of the trained prediction model, wherein the trained prediction model is configured such that a preset accuracy for determining the presence of the biomarker is set to at least 80% of genomic sequencing, and treating the patient based on the presence of the biomarker.
[0010] The above and other aspects and embodiments of the disclosed technology will be described in more detail in the drawings, this specification, and the claims.
Brief Description of the Drawings
[0011]
Fig. 1A
Fig. 1B
Fig. 2A
Fig. 2B
Fig. 2C
Fig. 2D
Fig. 3A
Fig. 3B
Fig. 3C
Fig. 4A
Fig. 4B
Fig. 4C
Fig. 5A
Fig. 5B
Fig. 6
Fig. 7
Fig. 8
Fig. 9
Fig. 10
[0012] In some embodiments, the implementation of the disclosed technology can provide an artificial intelligence (AI) architecture platform for predicting cancer biomarkers and a treatment method based on the biomarkers identified by the AI architecture platform.
[0013] In some embodiments, the implementation of the disclosed technology can provide a method and system for detecting clinically actionable cancer biomarkers and mutation signatures directly from digital hematoxylin and eosin (H&E) slides without sequencing.
[0014] Also, in some embodiments, the implementation of the disclosed technology can provide a novel deep learning architecture that can be trained to predict clinically actionable molecular cancer biomarkers directly from digital images based on scans of slides stained with hematoxylin and eosin (H&E) with little or no customization. The present invention enables the omission of DNA sequencing and can directly predict these biomarkers from the scanned slides.
[0015] Conventional methods for detecting clinically acting biomarkers for cancer treatment individualization rely largely on genomic sequencing or genotyping platforms (e.g., microarrays, targeted panel sequencing, whole exome sequencing, or whole genome sequencing). The developed convolutional neural network architecture completely avoids conventional sequencing methods by directly predicting clinically acting biomarkers using digital images from hematoxylin and eosin (H&E) tissue images sampled from individual patients. From the perspective of implementing a machine learning architecture, this method uses a semi-supervised convolutional neural network to perform segment predictions in the grid space over whole slide H&E images composed of three color channels. In the machine learning method, these predictions are aggregated to identify regions of interest and to perform a final action status prediction for each patient. Furthermore, this model provides a reliability measure for refining the final prediction value by introducing a multi-resolution approach that captures morphological patterns at various zoom magnifications and implements Monte Carlo dropout. Also, regions of interest are selected by using an unsupervised machine learning module composed of dimensionality reduction using principal component analysis and a custom k-means clustering algorithm for the feature vectors extracted for each component of the grid space. From the perspective of application, this method needs to be trained before it can be applied to specific clinical biomarkers. Specifically, this method requires at least 1,000 patients with digital H&E slides and known molecular biomarkers to generate a cancer-specific / biomarker-specific prediction model. The model can predict whether an individual patient's cancer has a biomarker by applying it to individual patients after training.
[0016] Hematoxylin and eosin (H&E) stain is one of the major tissue stains used in histology. H&E slides are routinely generated by pathologists for cancer diagnosis. However, in most cases, these slides do not enable pathologists to detect clinically actionable molecular biomarkers, nor do they provide guidance for individualized treatment. For this reason, in most cases, cancer samples are later sent for DNA and / or RNA sequencing to detect individual and / or multiple sets of biomarkers. The novel deep learning architecture implemented based on some embodiments of the disclosed technology enables the training of an AI model that can directly predict biomarkers incorporated into treatment regimens, thereby improving patient response to treatment and post-treatment survival of patients targeted by the detected biomarkers. The methods described herein eliminate the need for shipping and sequencing of biological samples by external providers, as they enable the identification of biomarkers directly from digital H&E slides (see FIGS. 4A-4C). For example, instead of mailing a biological sample to an external CLIA lab and waiting 14 days for sequencing results, this AI approach can directly detect these biomarkers in digital slides within fractions of a second.
[0017] Also, in some embodiments, through the implementation of the disclosed technology, a novel deep learning architecture can be provided that can be trained to predict clinically actionable molecular signatures and / or epidemiologically relevant molecular signatures directly from digital images based on scans of slides stained with hematoxylin and eosin (H&E) with little or no customization. The present invention enables the omission of DNA sequencing and can directly predict these signatures from digital images of scanned slides. Conventional methods for detecting clinically actionable biomarkers for personalized cancer treatment or epidemiologically relevant biomarkers for large-scale genetic epidemiology studies have relied largely on DNA sequencing or genotyping platforms (e.g., microarrays, targeted panel sequencing, whole exome sequencing, and / or whole genome sequencing). The developed convolutional neural network architecture completely avoids conventional sequencing techniques by accurately assessing the presence of molecular signatures using digital images from hematoxylin and eosin stained tissue images sampled from individual patients. The developed architecture requires at least 1,000 digital images of whole slides for the training of models of specific molecular signatures in specific cancer types. Nevertheless, accurate predictions can be made for digital images from a single cancer patient after the training of the model.
[0018] From a machine learning and implementation perspective, the developed architecture utilizes a semi-supervised convolutional neural network to perform segment predictions in the grid space over whole-slide H&E images composed of 3 color channels. The overall architecture represents a multi-resolution model that captures morphological patterns at two zoom magnifications, each magnification reflecting a separate convolutional neural network (CNN). Generally, the model makes an initial prediction at a low resolution (i.e., typically 5x magnification) and automatically locates the regions of interest by identifying the regions with the highest predictive power. Subsequently, the model makes a secondary prediction at a higher resolution scale (i.e., typically 20x magnification) over the entire selected region of interest. The resulting predictions are used in a final module that aggregates the scores of the entire multi-resolution model to provide the final prediction of a given molecular signature in a specific cancer type. In this architecture, generally little customization is done related to the adjustment of one of the two zoom levels (e.g., using 25x magnification to generate a better model for a specific molecular signature).
[0019] The developed convolutional neural network is based on and uses a convolutional neural network architecture (e.g., ResNet) with multiple important and significant changes. In contrast, the newly developed convolutional neural network is trained using a binary cross-entropy loss function based on the most predictive tiles derived from each sample at a given magnification level. For example, a digital image of a whole slide is tiled at 5x magnification and all sub-tiles are evaluated at the inference stage. And in a single training pass of the model, the sub-tile with the highest prediction probability for each whole slide is used. This process is repeated for each epoch throughout the training of the model. First, in this method, segment predictions are aggregated at low magnification to identify regions of interest. The automatic selection of regions of interest is performed using an unsupervised machine learning module. The unsupervised machine learning module includes a dimensionality reduction algorithm based on principal component analysis of the feature vectors extracted for each component of the grid space. The feature vectors are collected from the second-to-last layer of the trained CNN and then reduced to two principal components that contribute the most to the maximum variance across the entire collection of vectors. A custom k-means clustering module determines the optimal number of clusters per sample by selecting the solution with the maximum silhouette coefficient over all iterations utilized. The final region of interest includes the tile with the highest predicted value and uses the clusters that contain all other instances with silhouette coefficients within the top 50 percentile of the selected clusters. Subsequently, the complete set of selected tiles is used for the training of a second CNN model. Here, the CNN has an extended resolution (usually 20x). The second CNN at the extended resolution is based on the same architecture as the first CNN and is trained only on the regions of interest selected in the first stage of the model. Also, by resampling at the magnification of each region of interest from the original whole slide image, more details are captured at the cellular level. The proposed model was generally trained and tested after resampling at 20x magnification. However, this model can be used at any zoom setting.During inference, the final prediction of a specific molecular signature across all regions of interest in a given sample is made by using the tile with the highest predicted value. This extended prediction score is averaged with the prediction score from the first model to arrive at the final action status prediction for each patient.
[0020] In both CNN models, the robustness of this approach is improved by incorporating random dropout of nodes within the fully connected layers. Dropout is utilized both during model training to overcome the problem of overfitting and during the application of the model for inference as a way to quantify the level of confidence for each patient-level prediction. Specifically, inference is performed with multiple iterations (at least 100 iterations, usually less than 1,000 iterations) on a single whole-slide image by a new set of randomly dropped nodes in each path of the model. This method acts as a Bayesian approximation of the distribution underlying the potential models of the developed architecture. Each path of the model presents a single prediction score, but the analysis of the resulting distribution of scores across all iterations for a single patient can determine the level of certainty in the final prediction. For example, a certain prediction is one with a low variance from the average prediction score, while an uncertain prediction tends to have a high variance from the average prediction score. In the case of application to a single image of a whole-slide for an individual patient, the developed architecture will provide a normalized prediction score between 1 (low) and 100 (high) and a confidence interval for the score.
[0021] Figures 1A and 1B show an example of a multi - resolution convolutional neural network architecture for detecting molecular biomarkers from histopathological tissue slides based on some embodiments of the disclosed technology. As shown in FIGS. 1A and 1B, the multi - resolution convolutional neural network architecture implemented based on some embodiments can detect homologous recombination repair defects from histopathological tissue slides. FIG. 1A shows the training of a neural network (e.g., DeepHRD) model for detecting homologous recombination repair defects (HRD) from whole - slide images (WSIs). For each WSI, a single prediction score is estimated based on the detection of HRD. Specifically, at 101, pre - processing and quality control of each WSI are performed. This module consists of tissue segmentation, filtering of out - of - focus tissue, and final tiling at 5x magnification of the area containing the tissue. At 102, all tiles of a single image are processed by the first multi - instance learning (MIL) ResNet18 convolutional neural network. In this architecture, the average of the top 25 predicted tile scores is used as the WSI prediction score. In the feature extraction module, dropout is incorporated into the fully - connected layer to suppress over - fitting during training. During inference, the same dropout technique is incorporated to simulate Monte Carlo dropout, which is used to calculate the confidence interval of the final WSI prediction. At 103, by using the tile feature vectors from the second - last layer of feature extraction, regions of interest (ROIs) for additional evaluation are automatically selected from the original WSI. The feature vectors are dimensionally reduced by principal component analysis, and the optimal number of clusters per sample is determined by using a custom k - means clustering module. Then, at 104, the selected tiles are resampled at 20x magnification. At 105, the second MIL - ResNet18 model is trained by using these multiple sets of tiles by using the same architecture as that used at 102. At 106, for a single WSI, the average predictions across both models are aggregated.Then, by using the obtained score distribution, a confidence interval is calculated and a threshold for the reliability of the final prediction is established. Figure 1B shows a trained neural network (e.g., DeepHRD) model for HRD prediction from a single whole-slide image. The trained neural network (e.g., DeepHRD) model generates a final prediction score for the biopsy of an individual patient in an algorithm-based diagnosis for subsequent clinical actions.
[0022] Figures 2A-2D are diagrams showing a neural network (e.g., DeepHRD) for detecting homologous recombination repair deficiency and predicting the response to treatment of primary and metastatic breast cancers. Figure 2A shows a receiver operating characteristic curve (ROC) for classifying homologous recombination repair deficiency (HRD) in the TCGA presentation set (202) and independent set (204) of primary breast cancers, including the independent CPTAC and METABRIC primary breast cancer cohorts. Figure 2B shows representative TCGA tissue slides for both HRD and homologous recombination proficiency (HRP) samples across multiple breast cancer subtypes, along with the predictions obtained for each segment tile at 5x and 20x resolutions. Figure 2C shows the ROC for the formalin-fixed paraffin-embedded (FFPE) diagnostic model in the TCGA presentation set (212) and the ROC for classifying metastatic breast cancer (MBC) patients who are complete responders to platinum treatment. Figure 2D shows the Kaplan-Meier survival curves of MBC patients treated with platinum chemotherapy, separated by DeepHRD model prediction (220), BRCA1 / 2 mutation status (230), and SBS3 activity predicted by SigMA (240). The Q value is corrected by considering the breast cancer subtype, age at diagnosis, and the binary HRD classification score of the standard treatment ≧42 (i.e., HRD score). Also, the Cox regression showing the log 10 transformed hazard ratio is shown together with the respective 95% confidence intervals (lower part of 220, 230, 240). An asterisk (*) annotation is attached to Q values of 0.05 or less, while an n.s. (i.e., non-significant) annotation is attached to q values greater than 0.05.
[0023] Figures 3A to 3C show transfer learning (e.g., DeepHRD transfer learning) in ovarian cancer for predicting response to platinum treatment. Figure 3A is a schematic diagram showing a transfer learning method for training an ovarian homologous recombination repair deficiency (HRD) model from whole slide H&E images (WSIs) using a pre-trained breast DeepHRD model. By using the pre-trained frozen breast model, the weights and biases of all parameters in the ovarian model are brought about. Also, an HRD score is calculated from the SNP6 genotyping microarray by deriving loss of heterozygosity (LOH), large-scale transition (LST), and telomere allelic imbalance (TAI). Figure 3B shows Kaplan-Meier survival curves comparing the outcomes of patients treated with platinum chemotherapy, divided by the prediction of the DeepHRD transfer learning model. Figure 3C shows Kaplan-Meier survival curves comparing the outcomes of platinum-treated patients, divided by the reference model prediction (310) without the application of transfer learning, the BRCA1 / 2 mutation status (320), and the SBS3 activity predicted by SigMA (330). The Q value is corrected by considering the stage of ovarian cancer, age at diagnosis, and the binary HRD classification score ≥ 63 (i.e., the HRD score) of the standard treatment. Also shown is Cox regression showing the log10-transformed hazard ratio along with the respective 95% confidence intervals (lower part of 310, 320, 330). An asterisk annotation is attached to Q values of 0.05 or less, while an n.s. (i.e., non-significant) annotation is attached to q values greater than 0.05.
[0024] Figures 4A-4C are diagrams showing a workflow for independently training a neural network (e.g., DeepHRD) model on digitized rapid frozen·FFPE breast cancer slides. Figure 4A shows that, prior to training, the numbers of HRD and HRP samples within each breast cancer subtype were balanced by using all available PAM50 annotations. Figures 4B and 4C show collections of rapid frozen (Figure 4B) slides and formalin-fixed paraffin-embedded (FFPE) (Figure 4C) slides of the TCGA breast cancer cohort used for training two independent DeepHRD models. Prior to training, the numbers of HRD and HRP samples within each breast cancer subtype were balanced. Also, all downsampled individuals were added to an internal presentation test set. A validation set was used for optimizing the classification threshold.
[0025] Figures 5A and 5B are diagrams showing a workflow for testing the performance of a neural network (e.g., DeepHRD) model on digitized rapid frozen·FFPE breast cancer slides. Figure 5A shows a collection of breast cancers from CPTAC and METABRIC used for independent validation of the rapid frozen breast cancer model. The DeepHRD prediction scores were averaged for samples over multiple images. Figure 5B shows an independent collection of metastatic breast cancers treated with platinum chemotherapy, which was used for validation of a formalin-fixed paraffin-embedded (FFPE) breast cancer model based on the response of individual patients to the treatment.
[0026] In some embodiments, implementation of the disclosed technology can provide a deep learning artificial intelligence architecture for predicting genomic homologous recombination repair deficiency and platinum response from routine tissue slides in breast and ovarian cancers.
[0027] For cancers with homologous recombination deficiency (HRD), platinum-based chemotherapy and PARP inhibitors can be beneficial. Current standard diagnostic tests for detecting HRD in breast and ovarian cancers require genotyping or sequencing-based assays, which are not available everywhere.
[0028] In some embodiments, the implementation of the disclosed technology can provide a novel multi-resolution deep learning approach that enables the training of a robust model for directly detecting genomic biomarkers from digitalized images of hematoxylin and eosin (H&E)-stained optical microscopy histopathological slides. In some embodiments, a model for predicting a genome-derived HRD score can be trained using many primary breast cancers (e.g., 1,008 primary breast cancers from the The Cancer Genome Atlas (TCGA) project). In addition to a set of TCGA-presented samples, the trained breast cancer model was externally validated for 535 primary breast cancers and 77 platinum-treated metastatic breast cancers from two independent study cohorts. The applicability to 589 TCGA ovarian tumors was also demonstrated by training and validating a model using transfer learning to predict platinum response.
[0029] Across the entire validation cohort of TCGA breast cancer, a trained deep learning model for primary breast cancer implemented based on some embodiments of the disclosed technology can predict genomic HRD scores from digital H&E slides with an AUC of 0.81 ([0.77-0.85] 95% confidence interval (CI)). This performance was confirmed in two independent primary breast cancer cohorts (AUC = 0.76, [0.71-0.82] 95% CI). In an external clinical cohort of platinum-treated metastatic breast cancer, samples predicted as HRD by the deep learning model had a high complete response rate (AUC = 0.76, [0.54-0.93] 95% CI) and an average progression-free survival 3.7 times longer (hazard ratio = 0.47, q = 0.0087). Notably, deep learning techniques based on some embodiments of the disclosed technology can identify platinum-sensitive BRCA1 / 2 wild-type tumors as HRD positive. Also, by applying transfer learning, techniques based on some embodiments of the disclosed technology can predict overall survival after first-line platinum treatment in advanced high-grade serous ovarian cancer (hazard ratio = 0.45, q = 0.024). The deep learning model implemented based on some embodiments of the disclosed technology can outperform multiple existing genomic HRD biomarkers within each cohort.
[0030] The deep learning model applied to digitized H&E histopathological slides from breast and ovarian cancer detected genomic HRD and predicted a direct clinical benefit to standard-of-care platinum-based treatment. This approach outperformed existing gene biomarkers across multiple cohorts, slide scanners, and tissue fixation techniques. These results have important implications for the fair and efficient clinical management of cancer patients sensitive to DNA damage response targeted therapies.
[0031] Precision oncology aims to individualize cancer treatment by first identifying the molecular defects in each individual's tumor and then targeting them. Many cancers have defects in specific DNA repair pathways, and exploiting synthetic lethal relationships between peripheral pathways has proven to be an effective treatment approach. Specifically, in cells that originally have DNA repair deficiencies, using treatments that result in increased DNA damage and / or inhibition of additional DNA repair pathways may lead to selective cancer cell death. Past mechanistic studies and clinical trials have shown that breast and ovarian cancers lacking the ability to repair DNA double-strand breaks by homologous recombination are highly sensitive to DNA damage response targeted therapies such as platinum treatment and poly(ADP-ribose) polymerase (PARP) inhibitors.
[0032] Historically, homologous recombination repair deficiency (HRD) has been associated with germline mutations in specific genes that lead to an increased risk of developing breast and ovarian cancers, where the most notable susceptibility genes are BRCA1 and BRCA2. In addition to germline mutations, somatic mutations and epigenetic dysregulation in breast and ovarian cancers have also been shown to lead to HRD. Importantly, cancers with defective homologous recombination exhibit genomic instability with characteristic patterns of somatic mutations and gene expression. Some of these patterns have also been used to detect HRD in the absence of standard germline or somatic defects within HRD-related genes. For example, the pattern of single-base substitution signature 3 (SBS3), which is part of the Catalogue of Somatic Mutations in Cancer (COSMIC) mutation signature, is caused by HRD independent of the molecular mechanism that renders DNA repair by homologous recombination ineffective. Importantly, SBS3 has been used in the past as a clinical biomarker for detecting HRD in breast and ovarian cancers.
[0033] In the United States, two commercially available HRD Companion Diagnostic (CDx) tests have been approved by the US Food and Drug Administration (FDA) for patients with ovarian cancer and metastatic breast cancer. Both Myriad myChoice® CDx and FoundationOne® CDx determine HRD by quantifying genomic instability in combination with the status of BRCA1 and BRCA2. In addition, multiple studies and CLIA-certified diagnostic tests have been developed to detect HRD by examining germline variants, somatic mutations, mutation patterns, changes in gene expression, and / or epigenetic modifications. Detection of homologous recombination repair deficiency can be performed using a number of different methods, but all of these methods essentially rely on assays that profile DNA and / or RNA, resulting in bottlenecks that are a major cause of the solubility of molecular tests, time to decision-making, and overall cost. As a result, the widespread use of companion and complementary diagnostic biomarkers in CLIA-certified research trials for standard treatment and clinical trials is hampered. For example, the cost of a CLIA-certified companion for a complementary HRD test is in the thousands of dollars, which is not affordable for many patients in the United States and most countries around the world. Furthermore, the results of sequencing-based diagnostics can take 3 to 6 weeks, significantly delaying the clinical management of many lethal and rapidly progressing solid tumors. Finally, recent reports have shown that only a small percentage of patients worldwide have access to sequencing-based diagnostic tests, including FDA-approved companion diagnostic tests, and the testing rate is even lower in various populations that are not receiving adequate services. This large "gap" in cancer genomics testing is an important issue in providing fair and efficient clinical management to all cancer patients worldwide who need to identify low-cost and scalable biomarkers for clinical oncology.
[0034] Although access to and uptake of array-based diagnostics are limited, tissue biopsies are routinely sampled and processed by hematoxylin and eosin staining (H&E) for the diagnosis of solid tumors in most patients worldwide. In combination with recent advances in computer vision and computational pathology to detect recurrence patterns in complex, data-rich whole-slide H&E images, deep learning artificial intelligence (AI)-based models enable both prognostic and diagnostic prediction using only histopathological tissue slides. Here, we introduce DeepHRD, a multi-resolution convolutional neural network architecture with weak teachers for directly detecting HRD from digitized H&E tissue slides. We train and validate the DeepHRD model with data from The Cancer Genome Atlas (TCGA) project and demonstrate the ability to detect HRD using data from two external research consortia. Importantly, using independent clinical samples, we show that DeepHRD outperforms existing genomic biomarkers in predicting patient response to platinum-based therapy. By circumventing current bottlenecks in genomic testing, methods based on some embodiments of the disclosed technology are directly relevant to addressing the global socioeconomic disparities in the diagnosis and treatment of breast and ovarian cancers.
[0035] To train a downstream model that can directly predict the HRD status from digitized tissue slides, we implemented DeepHRD, a convolutional neural network architecture with weak teachers based on the basic premise of multi-instance learning (MIL (Figure 1)). We also trained the model using digitized H&E images from the TCGA breast cancer cohort, which included 1,008 samples with snap-frozen slides and 1,055 samples with formalin-fixed paraffin-embedded (FFPE) slides, and performed internal validation (Figures 4A–4C). Additionally, we trained a separate model using the TCGA ovarian cancer cohort, which included 589 samples with snap-frozen slides. All samples of interest were accompanied by whole-exome sequencing data and microarray genotyping data for calculating the genomic HRD score (Figures 4A–4C). By separately training the DeepHRD method to predict the HRD status from FFPE and snap-frozen tissues of breast cancer and snap-frozen tissues of ovarian cancer, we obtained a total of three independent trained DeepHRD models.
[0036] (i) The primary breast cancers from the CPTAC (Clinical Proteomic Tumor Analysis Consortium) consisting of 116 samples associated with whole exome array determination and (ii) the METABRIC (Molecular Taxonomy of Breast Cancer International Consortium) consisting of 419 samples associated with microarray genotype classification data were used to perform external validation of the rapid frozen breast cancer model (Figure 5A). Also, with the available metastatic biopsies and associated genomic and clinical annotations, external validation of the FFPE breast cancer model was performed in an independent clinical cohort of metastatic breast cancer consisting of 77 patients treated with platinum-based chemotherapy. Also, the RECIST 1.1 (Response Evaluation Criteria in Solid Tumors, version 1.1) (Figure 5B) was used to evaluate the clinical response and progression-free survival (PFS) to platinum-based treatment. Furthermore, the Aperio ScanScope system was used to digitize a collection of fixed tissue slides from the TCGA breast cancer, TCGA ovarian cancer, CPTAC cohort, and METABRIC cohort. Also, the Hamamatsu Photonics Nanozoomer system was used to digitize the clinical cohort of metastatic breast cancer.
[0037] Definition of genomic HRD score and related gene markers The HRD score was calculated from sequencing or genotyping data reported previously using scarHRD. Briefly, scarHRD takes into account an aggregate score of telomeric allelic imbalance score, loss of heterozygosity score, and large-scale transition score calculated for each patient using copy number calls derived from ASCAT from SNP6 genotyping microarrays. Conventionally, in triple-negative breast cancer, an HRD score above 42 has been used to determine eligibility for platinum-based chemotherapy or treatment with PARP inhibitors, while in ovarian cancer, an HRD score above 63 has been utilized. For Grand Turismo, soft binning was incorporated during training to place the HRD score cutoff at the median score of all breast cancer samples (i.e., set the HRD score cutoff at 30) so that the model does not become overconfident with a single image. All breast cancer samples with an HRD score above 50 are considered deficient, while all samples with an HRD score below 10 are considered proficient. For other breast cancers with an HRD score between 10 and 50, they were modeled using soft binning centered around 30. In this case, the symmetric probability that the sample is deficient or proficient is equal. In the case of ovarian cancer, all ovarian samples with an HRD score above 73 are considered deficient, while all samples with an HRD score below 53 are considered proficient. Similar to breast cancer, for other ovarian cancers with an HRD score between 53 and 73, they were modeled using soft binning centered around 63. In this case, the symmetric probability that the sample is deficient or proficient is equal.
[0038] For the pathogenicity of mutations found within BRCA1 and BRCA2, it was determined using InterVar as described above for the TCGA ovarian cohort. All variants predicted to be pathogenic were also considered as defects. For the mutation status of BRCA1 and BRCA2 within the metastatic breast cancer cohort, it was determined by screening variants across all existing database annotations including Clinvar, Swissprot, LOVD (Leiden Open Variation Database), and UMD (Universal Mutations Database), as per past reports.
[0039] The activity of the COSMIC mutation signature SBS3 has been shown to be abundant in HRD samples. The presence of SBS3 reflects many potential defects that can occur across the HR machinery independent of the underlying mechanism of inactivation and has recently been used for the detection of HRD from clinical sequencing data. To determine the presence of SBS3 activity within the TCGA ovarian cohort and the metastatic breast cancer cohort, reliable predictions generated by the machine learning tool SigMA were used. SigMA classifies samples as either SBS3 positive or SBS3 negative with an estimated false positive rate (FPR) of 1% and a sensitivity of 50%.
[0040] Deep learning architectures implemented based on some embodiments of the disclosed technology can be constructed based on the concept of multi-instance learning with weak teachers (Figure 1A). Specifically, for all segmented regions of slides or tiles within whole-slide H&E images (WSIs), weak labels based on slide-level classification of each sample are assigned. While all tiles within negatively labeled slides have homologous recombination proficiency (HRP), within positively labeled slides, it is thought that at least one tile must exhibit the HRD phenotype. These assumptions enable the training of models that use only a single classification label for the entire image without requiring detailed manual annotation by pathologists (which does not exist for HRD characterization).
[0041] This model is based on multi - resolution determination that makes an initial prediction at a low magnification (i.e., 5x magnification), then automatically selects a region of interest (ROI), and makes a secondary prediction at an enlarged magnification (i.e., 20x magnification) within the selected ROI (Figure 1A). Here, after selecting ROIs at low magnification across the entire tissue slide, a DeepHRD multi - resolution architecture framework was designed to mimic the standard diagnostic protocol used by pathologists to examine H&E images by refining the characteristics and subtypes of specific tumors captured at high resolution. Further, DeepHRD maps individual tile predictions back to the original WSI, which enables visualization of the relative contribution and importance of specific tissue regions to the model's prediction without using pixel - level annotations (Figure 1B). The final model consists of a collection of five identical architectures, each of which generates a multi - resolution prediction score. The average of these scores was used to make the final prediction for each tissue slide. Due to the computational cost associated with processing the entire WSI, each slide was first segmented into small tiles for each resolution used. In the first stage of the model, each slide was tiled at 5x magnification with tiles of 256x256 pixels per tile and approximately 2μm of tissue per pixel. Tiles with less than 80% of pixels containing unclear tiles and tissue were removed (Figure 1A). In both stages of the model, a ResNet18 convolutional neural network was trained to extract features from the set of tiles that make up a single WSI. Also, using the encoded features obtained from the second - last fully - connected layer, ROIs were automatically selected at 5x resolution. Specifically, principal component analysis (PCA) was used to project the encoded features into a latent space containing the maximum variance. Then, the k - means clustering was used to group each tile representation. Also, the total number of clusters per sample was determined based on the value of k that provides the maximum silhouette coefficient across all tile representations. And the cluster containing the tile with the maximum prediction probability was selected, together with all tiles within the same cluster that have a silhouette score exceeding the 95th percentile of all silhouette scores for the entire WSI.The final ROI was tiled at a magnification of 20× (0.5 μm per pixel), and this was used to train and test the second model. The top 25 tiles were averaged to calculate the final prediction score at a given resolution during the inference pass of the WSI. Importantly, during training, overfitting of the training dataset was prevented by incorporating random dropout of nodes within the fully connected layers of the ResNet architecture. This same dropout technique, known as Monte Carlo dropout, was also applied during the inference of each WSI to estimate the model uncertainty by running multiple inference passes for a single WSI. The distribution of the obtained predictions was averaged to calculate the final score including any epistemic uncertainty, and this was used to calculate the confidence threshold for a given sample (Figure 1A).
[0042] Once the training and calibration of the final HRD model are complete, predictions for individual patients can be made using only digitized cancer biopsies by using DeepHRD (Figure 1B). By using a single WSI as input, DeepHRD will generate a prediction score with a confidence interval for use in computer-aided diagnosis recommendations. The purpose of using this method is to provide computer-aided diagnosis for subsequent clinical actions. Specifically, individuals with high prediction reliability are labeled as either HRD or HRP (Figure 1B).
[0043] To characterize the classification performance of the proposed method, the area under the receiver operating characteristic curve (AUROC or AUC) was calculated. Confidence intervals were also calculated using non-parametric resampling. Comparisons between survival curves were calculated using the log-rank test. Additionally, multivariate analysis was performed to calculate the hazard ratio using Cox regression. Furthermore, the Benjamini Hochberg method was used to perform correction for multiple hypothesis testing.
[0044] The DeepHRD model was trained to detect HRD samples using a subset of frozen tissue slides from the TCGA breast cancer cohort (Figure 1A). The TCGA breast cancer cohort was separated into (i) 70% of the samples used for training, (ii) 15% of the samples used for training validation and adjustment of prediction parameters, and (iii) 15% of the samples used for testing the final trained model presented (n = 1,008 total breast cancer samples (Figures 4A - 4C)). Prior to training, the number of HRD and HRP samples for each breast cancer subtype was balanced to prevent the model from learning features specific to the tissue of individual subtypes rather than features directly associated with HRD (Figures 4A and 4B). Subsequently, the trained DeepHRD model can be applied to individual digital slides to provide patient-level predictions that clarify whether breast cancer is HRD or HRP (Figure 1B). Importantly, DeepHRD can overlay an HRD probability mask on each digital slide, which can be used for subsequent investigation of the pathological characteristics of each breast cancer patient (Figure 1B).
[0045] Subsequently, the final trained DeepHRD breast cancer model was tested against the TCGA presented test set to evaluate overall performance, and the AUC was 0.81 ([0.77 - 0.85] 95% confidence interval (CI)) (Figure 2A). Also, to evaluate the generalizability of this model, external validation was performed using a set of breast cancer slides from CPTAC and METABRIC (Figure 5A), and the AUC was 0.76 ([0.71 - 0.82] 95% CI) (Figure 2A). Importantly, HRD is elevated in luminal B, basal-like, and Her2-enriched breast cancers (Figure 4A), but the models implemented based on some embodiments of the disclosed technology can distinguish HR deficiency and proficiency across all subtypes (Figure 2B).
[0046] Snap-frozen tissue slides are generally used for downstream molecular analysis, while FFPE tissue slides are the standard in the clinical setting. Therefore, an independent model for directly classifying HRD from FFPE slides from the TCGA breast cancer cohort was trained following the same procedure described above for training on snap-frozen tissue images (Figures 4A - 4C). The final TCGA FFPE model showed an AUC of 0.81 ([0.77 - 0.86] 95% CI (Figure 2C)), which was identical to the TCGA snap-frozen model. These results indicate that differences in fixation procedures and staining have a minimal impact on the performance to predict HRD status directly from breast cancer tissue slides.
[0047] Importantly, the FFPE model was able to distinguish metastatic breast cancer (MBC) samples (part of an independent clinical cohort) (n = 9) with a complete response to platinum chemotherapy from samples that showed only a partial response to treatment or no response (n = 68), resulting in an AUC of 0.76 ([0.54 - 0.93] 95% CI) (Figure 2C, Figure 5B). Separating MBC samples treated with platinum based on DeepHRD predictions revealed different clinical benefits between HRD and HRP predicted samples, with a median progression-free survival of 14.4 months in HRD patients and 3.9 months in HRP patients (p-value = 0.0019, log-rank test). The predictive value of this model was consistent even after adjusting for breast cancer subtype, age at diagnosis, and genomic HRD score, with a hazard ratio of 0.47 ([0.27 - 0.83] 95% CI, q-value = 0.0087, Cox proportional hazards regression) (Figure 2D). Furthermore, DeepHRD captured 7 out of 9 samples with a complete response to platinum treatment. In comparison, separating samples based on the genomic status of BRCA1 / BRCA2 or using the elevated activity of the mutation signature associated with HRD (SBS3 predicted by SigMA) did not result in a significant difference in progression-free survival (q-value = 0.13 and q = 0.34, Cox proportional hazards regression) (Figure 2D). The small sample size of BRCA1 / 2 mutant tumors (about 8% of the MBC samples examined) affected the significance level compared to wild-type tumors, but DeepHRD predictions captured more than 4 times as many platinum-sensitive samples than when using only BRCA1 / 2 status. Finally, tissue slides from the MBC cohort were digitized using the Hamamatsu Nanozoomer system, while TCGA breast cancer was digitized using the Aperio ScanScope system, demonstrating the generalizability of DeepHRD across different scanning protocols.
[0048] To further evaluate the generalizability of the proposed method in the application to other cancers, we performed transfer learning on the TCGA ovarian cancer cohort (Figure 3A). Individuals with ovarian cancer have conventionally received platinum chemotherapy as the first-choice standard treatment, and this cohort is ideal for verifying whether HRD prediction from tissue slides can have direct clinical benefits. Specifically, we trained an independent model to predict HRD status from snap-frozen slides using the TCGA ovarian cohort. Due to the cohort size being twice as small (n = 589 ovarian cancers), we started the ovarian model using pre-trained weights and biases generated from the snap-frozen breast cancer model with convolutional weights and biases frozen during training (Figure 3A). We also applied the final model to the presentation test set of TCGA ovarian cancer to evaluate the model's ability to separate individuals who would benefit from treatment with platinum chemotherapy (Figure 3A). Of the 117 patients in the presentation set, 66 received first-choice platinum chemotherapy for advanced high-grade serous ovarian cancer. As a result of separating these individuals by their respective DeepHRD predictions, a difference in median survival time occurred between the predicted patients with HRD and HRP (Figure 3B). Specifically, the median survival time of patients predicted to have HRD was 4.6 years, while the median survival time of patients predicted to have HRP was 3.2 years (q-value = 0.024), and the hazard ratio after correcting for cancer stage, age, and genomic HRD score was 0.45 ([0.22 - 0.90] 95% CI, Cox proportional hazard ratio) (Figure 3B). For comparison, in the case of using a reference model without transfer learning, the separation of samples deteriorated (q-value = 0.076), and the hazard ratio after correction was 0.53 ([0.26 - 1.07] 95% CI, Cox proportional hazard ratio) (Figure 3C). This suggests that transfer learning is beneficial when attempting to train AI-based methods for small datasets. Consistent with the breast cancer cohort, no significant difference in survival time occurred when separating samples based on the mutation status of BRCA1 / BRCA2 or based on the increased activity of SBS3 (q-value = 0.47 and q-value = 0.32, Cox proportional hazard regression) (Figure 3C).
[0049] The development of the DeepHRD prediction model for breast and ovarian cancers demonstrates the practicality of adopting AI-based guidance in clinical diagnosis and precision medicine workflows. Results across multiple publicly available cohorts (TCGA, CPTAC, and METABRIC) as well as additional external cohorts show that the model is applicable to routinely sampled tissue blocks and generalizable across diverse cancers, histological and molecular subtypes, digital scanning systems, and tissue fixation procedures, including variability in H&E tissue staining. The performance of DeepHRD is consistent across primary and metastatic breast cancers, and by incorporating transfer learning, the model has also become applicable to serous ovarian cancers. Most importantly, based on the RECIST 1.1 criteria, DeepHRD predicted clinical response and progression-free survival to DNA damage response-targeted therapy, i.e., platinum therapy, outperforming existing genome-based diagnostic biomarkers (i.e., BRCA1 / 2 status and signature SBS3). Furthermore, in line with prior breast cancer genomic studies, DeepHRD captured patients with BRCA1 / 2 wild-type tumors who responded to platinum therapy and identified patients with four-fold higher reactivity than BRCA1 / 2 mutation testing alone (Figure 2D). These results indicate that an AI-based deep learning model capable of rapidly predicting clinical response from routine digitized diagnostic histopathological slides can substitute for and / or complement genotype-based and sequencing-based assays conventionally used for HRD assessment in the clinic. Avoiding reliance on genomic profiling provides a more readily clinically deployable solution while enhancing access to cutting-edge diagnostics for a larger percentage of the population across diverse socioeconomic groups.
[0050] Historically, the diagnosis of solid tumors has evolved from microscopic and morphological evaluation of H&E slides to genomic biomarker testing. Genomic testing is essential for the clinical management of specific cancers, but in addition to repeat biopsies to obtain sufficient tumor tissue for molecular assays, extensive analysis is often required to analyze the large-scale data generated by these molecular assays, substantially complicating the daily clinical oncology workflow. Recent deep learning AI techniques have demonstrated the ability to directly detect genomic biomarkers from H&E images, including those indirectly related to treatment outcomes (e.g., detection of microsatellite instability that can predict response to immunotherapy). In some embodiments, there is no prior study showing the direct clinical significance of an AI-based model for HRD detection by predicting treatment benefit with external validation. Since HRD is a complementary biomarker that helps guide FDA-approved companion diagnostic tests for the use of platinum therapy and the use of PARP inhibitors, the performance of neural networks (e.g., DeepHRD) implemented based on some embodiments of the disclosed technology has direct implications for predicting response to DNA damage response targeted therapies within breast, ovarian, and other cancer types with known HR deficiencies. One such example is pancreatic ductal adenocarcinoma (PDAC), where patients in whom HRD was detected using an FDA-approved targeted sequencing assay had improved clinical outcomes with standard first-line platinum-based therapy. As a strict limitation to the use of genomic detection of HRD in the standard clinical care of PDAC, in addition to the issue of tissue procurement for genomic studies, in some PDAC patients, the median survival is only about 3 to 6 months due to rapid progression, so a 3- to 6-week turnaround for molecular profiling was not suitable for the first-line treatment of advanced disease. In the case of such cancer types, AI-based detection of HRD from routinely generated H&E slides may provide a superior rapid diagnostic option.
[0051] While deep learning and computer vision-based approaches in digital pathology are increasing explosively, due to the necessary infrastructure and high overhead associated with the rationalization of digital pathology, in developing countries and resource-constrained communities, the lack of global accessibility restricts the immediate conversion to clinical practice. Nevertheless, recent research has shown the possibility of directly deploying deep learning models trained on whole-slide images to handheld photographs of core needle biopsy tissues taken from microscopic fields of view. These types of digitized tissue images are smaller in size and require only one-tenth of the computational resources ultimately needed for downstream diagnosis, making them easily deployable on local smartphone devices equipped with standard high-resolution cameras attached to the eyepieces of conventional optical microscopes. This approach guarantees an inexpensive, efficient, and accurate deep learning readout within seconds from the preparation of H&E slides. In conjunction with the development of lightweight deep learning architectures, there is an opportunity to deploy diagnostic applications that were conventionally computationally costly into manageable packages without a significant loss of predictive power. This shift will provide AI-based diagnostic solutions for fair and efficient clinical management of all cancer patients worldwide by relying on smartphone microscopic images.
[0052] Additional method Data source For the collection of frozen and formalin-fixed paraffin-embedded (FFPE) slides from TCGA, all were downloaded from the Genomic Data Commons (GDC) (https: / / gdc.cancer.gov / ) along with all clinical characteristics. For the collection of frozen slides from CPTAC, they were downloaded from The Cancer Imaging Archive (TCIA), and the genomic data were downloaded from GDC. For the collection of images and related SNP6 genotyping microarray data from METABRIC, they were downloaded from EGA (accession numbers: EGAD00010000270 and EGAD00010000266). For the predicted cancer subtypes for a subset of the TCGA breast cancer cohort, they were obtained from past studies using the 50-gene PAM50 model. For the HRD scores of TCGA breast and ovarian cancers, they were obtained from past studies. Between June 2018 and March 2020, 77 patients were enrolled from a clinical cohort of metastatic breast cancer with whole-exome sequencing, and all had received at least one cycle of platinum chemotherapy. Also, as reported in the past, all clinical evaluations were determined locally at the Georges Francois Leclerc Cancer Center.
[0053] Data preprocessing Each whole-slide image (WSI) was segmented into 256×256 tiles at magnifications of 5× and 20×, containing 2μm and 0.5μm per pixel, respectively. Tiles with less than 80% pixels representing unclear tiles and tissues were removed from all training and test cohorts. To filter out unclear tiles, a Laplacian filter was applied to each tile using a 3×3 kernel, and all tiles with a variance less than 0.02 were excluded from further analysis. For all green, red, and blue pen marks and other annotation artifacts, they were removed by thresholding in the RGB color channels within each pixel.
[0054] Calculation of HRD score As previously reported using scarHRD, the HRD score was calculated. Specifically, the HRD score was the sum of the telomeric allelic imbalance score, loss of heterozygosity score, and large-scale transition score calculated for each patient using copy number calls derived from ASCAT from the SNP6 genotyping microarray. The HRD score for CPTAC breast cancer samples was calculated based on copy number calls derived from whole exome sequencing using Sequenza, which has been shown to have a similar HRD score distribution to that calculated using copy number calls derived from ASCAT from the SNP6 genotyping microarray.
[0055] For both the TCGA breast and ovarian cancer cohorts, soft labeling was applied to the HRD score using the thresholds designated for samples reliably labeled as HRD or HRP. These thresholds were determined by first dividing the samples within each cancer type into HRD and HRP categories using a single cutoff (HRD ≥ 30 for breast cancer and HRD ≥ 63 for ovarian cancer). Also, the median values of the two sample categories obtained for each cancer type were used to set the range of reliable thresholds for HRD and HRP. By using a quadratic function, all intermediate HRD scores were modeled as probabilities (Equation 1).
[0056]
Number
[0057] Specifically, in the breast cancer cohort, an HRD score greater than 50 was considered HR deficient, and a score less than 10 was considered proficient. Also, all intermediate scores were modeled as the probability of being deficient or proficient, with an equal probability of both states when the HRD score was 30 (Equation 2).
[0058]
Number
[0059] In the TCGA ovarian cohort, an HRD score above 73 was considered deficient, a score below 53 was considered proficient, and the intermediate probability was centered around 63 (Equation 3).
[0060]
Number
[0061] Training and Testing of the Model Prior to training, the use of PAM50 model classification was used to balance the number of HRD and HRP samples across all breast cancer subtypes and to normalize the increase or decrease in HRD samples for specific breast cancer subtypes. All samples without annotation of PAM50 subtype labels were considered unknown, and the number of cases of HRD and HRP was also balanced (Figure 4A). Also, to prevent overfitting during training and to classify true HRD samples considering the ambiguity in the proficiency of the HRD score, soft labeling was incorporated. All training and testing were performed using the machine learning Python framework Pytorch (v.1.5.0). For both resolution models, a learning rate of 10 -3 a weight decay of 10 -4 and the Adam optimizer were used for training in mini - batches consisting of 64 tiles. Also, each model was initialized using the ResNet18 architecture pre - trained on the ImageNet (http: / / www.image - net.org / ) database and trained over 200 epochs. During training, all convolutional weights were frozen. Also, early stopping was incorporated to prevent overfitting.
[0062] After training the 5x resolution model, the final inference pass is executed for all slides. From the second-to-last layer of the feature extractor, all features of a single WSI were selected and projected into a low-dimensional latent space using principal component analysis. Also, k-means clustering was used to automatically select regions of interest (ROIs) to be re-tiled at 20x magnification. Furthermore, the number of clusters was determined by selecting the solution with the maximum silhouette coefficient. And the cluster containing the tile with the highest prediction probability was used for ROI selection. All tiles belonging to this cluster and having a silhouette score exceeding the 95th percentile of all silhouette scores of a given WSI were selected as the final ROIs. Subsequently, each ROI was tiled at 20x magnification into 256×256 pixel sub-tiles. As a result, for each ROI at 5x magnification, 16 tiles at 20x magnification were obtained. To execute the inference pass of the model, a single WSI image was processed over 10 iterations with a random dropout probability of 0.20 for all nodes in the fully connected layer.
[0063] Transfer learning For the initiation of the model weights of the ovarian model, known as transfer learning, weights collected from the final model trained to detect HRD from rapidly frozen breast slides were used. Also, a presentation internal validation set was used to perform survival analysis based on pre-treatment with platinum chemotherapy. For the training and testing of the DeepHRD model on FFPE ovarian cancer samples, there were not sufficient FFPE slides in the ovarian cohort.
[0064] Visualization of DeepHRD predictions If the training is successful, use DeepHRD to make predictions for individual whole-slide images. When performing multiresolution inference, DeepHRD generates HRD probabilities at 5x magnification for each tile and at 20x magnification for each tile within the automatically selected region of interest. By mapping the probabilities to their original positions within the whole-slide image using the positions of the original tiles, the regional patterns that influence the final model predictions can be visualized.
[0065] Survival analysis Survival analysis was performed using the Lifelines Python package (v.0.24.4). Samples were split based on the predictions from each DeepHRD model for both the metastatic breast cancer (MBC) and TCGA ovarian cohorts. Only samples treated with platinum chemotherapy were considered in the survival comparisons. Also, the log-rank test was used to compare the survival curves. Hazard ratios were calculated by Cox regression after adjusting for age at diagnosis, primary breast cancer subtype, and genomic HRD score within the MBC cohort, and age at diagnosis, ovarian cancer stage, and genomic HRD score within the TCGA ovarian cohort. The median survival time was calculated as the point at which there is a 50% probability of surviving beyond it.
[0066] Statistics All performance metrics were calculated using the scikit-learn Python package (v.0.22.1). Confidence intervals were calculated using non-parametric resampling. Additionally, standard error bars were calculated using the NumPy Python package (v.1.18.1).
[0067] Figure 6 shows an exemplary method 600 for determining the presence of biomarkers in a biological sample based on some embodiments of the disclosed technology.
[0068] In some embodiments of the disclosed technology, method 600 may include, at 610, obtaining a section of a biological sample treated with a stain; at 620, imaging one or more regions of the stained section of the biological sample at a first resolution and a second resolution to generate first and second pluralities of image data; at 630, reducing a parameter space of the first and second pluralities of image data to generate reduced first and second pluralities of image data; and at 640, providing the first and second pluralities of image data to a trained prediction neural network and determining the presence of a biomarker in the biological sample as an output of the trained prediction neural network.
[0069] Figure 7 shows an exemplary method 700 for generating a trained prediction model configured to determine the presence of a biomarker in a biological sample based on some embodiments of the disclosed technology.
[0070] In some embodiments of the disclosed technology, method 700 may include, at 710, generating stained sections of one or more biological samples and corresponding biomarker labels; at 720, imaging one or more regions of the stained sections of the one or more biological samples at a first resolution and a second resolution to generate first and second pluralities of image data; at 730, reducing a parameter space of the first and second pluralities of image data to generate reduced first and second pluralities of image data; and at 740, generating a trained prediction model including a first prediction model trained with the reduced first plurality of image data and the corresponding biomarker labels and a second prediction model trained with the reduced second plurality of image data and the corresponding biomarker labels.
[0071] Figure 8 shows another exemplary method 800 for determining the presence of a biomarker in a biological sample based on some embodiments of the disclosed technology.
[0072] In some embodiments of the disclosed technology, method 800 includes, at 810, obtaining a stained section of a biological sample; at 820, imaging one or more regions of the stained section of the biological sample to generate a plurality of images of the stained section; and at 830, providing the plurality of images of the stained section as input to a trained prediction model and, as an output of the trained prediction model, determining the presence of a biomarker in the biological sample, wherein the trained prediction model is configured such that a preset accuracy for determining the presence of the biomarker is set to at least 80% of genomic sequencing.
[0073] FIG. 9 shows a treatment method 900 for treating a target cancer as needed based on some embodiments of the disclosed technology.
[0074] In some embodiments of the disclosed technology, method 900 includes, at 910, obtaining a stained section of a biological sample; at 920, imaging one or more regions of the stained section of the biological sample to generate a plurality of images of the stained section; at 930, providing the plurality of images of the stained section to a trained prediction model and, as an output of the trained prediction model, determining the presence of a biomarker in the biological sample, wherein the trained prediction model is configured such that a preset accuracy for determining the presence of the biomarker is set to at least 80% of genomic sequencing; and at 940, administering treatment to a patient based on the presence of the biomarker.
[0075] FIG. 10 shows an example of a computer system 1000 configured to determine the presence of a biomarker in a biological sample based on some embodiments of the disclosed technology.
[0076] In some embodiments of the disclosed technology, system 1000 comprises processor 1010 and memory or storage medium 1020. Processor 1010 reads code from memory 1020 and executes the methods discussed in this patent document.
[0077] Accordingly, based on the above disclosure, various embodiments of the features of the disclosed technology are possible, including the examples listed below.
[0078] A method for determining the presence of a biomarker in a section using a trained prediction model Example 1. A method for determining the presence of a biomarker in a biological sample, comprising: (a) preparing a section of a biological sample treated with a stain; (b) generating first and second pluralities of image data by imaging one or more regions of the stained section of the biological sample at a first resolution and a second resolution; (c) generating reduced first and second pluralities of image data by reducing the parameter spaces of the first and second pluralities of image data; and (d) determining the presence of the biomarker in the biological sample as an output of a trained prediction model when the reduced first and second pluralities of image data are provided as inputs to the trained prediction model.
[0079] Example 2. The method according to Example 1, wherein the trained prediction model is configured to determine the presence of the biomarker with at least 80% accuracy as compared to genomic sequencing.
[0080] Example 3. The method according to Example 2, wherein the accuracy includes at least 85%, at least 92%, at least 95%, at least 97%, or at least 99% as compared to genomic sequencing.
[0081] Example 4. The method according to Example 1, wherein the trained prediction model includes a first prediction model trained with respect to the first plurality of image data and a second prediction model trained with respect to the second plurality of image data.
[0082] Example 5. The method according to Example 1, wherein the biomarker comprises a deletion of chromosome 9p.
[0083] Example 6. The method according to Example 1, wherein the biomarker comprises the presence of clustered mutations in the gene TP53. In some embodiments, TP53 comprises a tumor suppressor gene. In one example, TP53 refers to tumor protein P53.
[0084] Example 7. The method according to Example 1, wherein the biomarker comprises the presence of clustered mutations in the gene EGFR (epidermal growth factor receptor).
[0085] Example 8. The method according to Example 1, wherein the biomarker comprises the presence of clustered mutations in the gene BRAF. In some embodiments, BRAF comprises a human gene encoding a protein called B-Raf. In one example, BRAF refers to v-raf murine sarcoma viral oncogene homolog B1.
[0086] Example 9. The method according to Example 1, wherein the biomarker comprises the presence of clustered mutations in the gene KIT.
[0087] Example 10. The method according to Example 1, wherein the biomarker comprises the presence of MSI (versus MSS) and / or MMR gene (e.g., POLE, MLH1, MLH3, MGMT, MSH6, MSH3, MSH2, PMS1, or PMS2) deficiency. In some embodiments, MSI gene deficiency indicates microsatellite instability (MSI) gene deficiency, and MMR gene deficiency indicates mismatch repair gene deficiency. In some embodiments, MSS indicates microsatellite stability.
[0088] Example 11. The method according to Example 1, wherein the biomarker comprises the presence of a high amount of tumor mutations.
[0089] Example 12. The method according to Example 1, wherein the biomarker comprises the presence of a hypermutator mutation signature selected from POLE, MSI-COSMIC15 which is a combination of POLE and MSI composed of POLE and MSI-COSMIC14 (POLE+MSI), MSI-COSMIC20 (POLD+MSI), MSI-COSMIC21, MSI-COSMIC26, and MSI-COSMIC6.
[0090] Example 13. The method according to Example 1, wherein the biomarker comprises the presence of apolipoprotein B mRNA editing enzyme, catalytic polypeptide (APOBEC) changes, and a mutation signature. In some embodiments, APOBEC represents a family of evolutionarily conserved cytidine deaminases.
[0091] Example 14. The method according to Example 1, wherein the biomarker comprises the presence of homologous recombination repair deficiency (HRD). To determine HRD by quantifying genomic instability in combination with BRCA1 and BRCA2 status, two commercially available HRD companion diagnostic (CDx) tests, Myriad myChoice® CDx and FoundationOne® CDx, are FDA-approved, and there are at least three academic HRD detection methods (SigMA, HRDetect, and CHORD). In some embodiments, BRCA represents breast cancer genes.
[0092] Example 15. The method according to Example 1, wherein the biomarker comprises the presence of HRD negative (or homologous recombination proficiency (HRP)) or HRD positive, for example, using the genomic test of Example 14.
[0093] Example 16. The method according to Example 1, wherein the biomarker comprises the presence of BRCA1 / 2 mutations.
[0094] Example 17. The method according to Example 1, wherein the biomarker comprises the presence of the "COSMIC3-BRCA" mutation signature, which includes a specific pattern of genome-wide somatic single nucleotide variations (SNVs) defined as "mutation signature 3" (Sig3) in the COSMIC signature catalog, or the presence of a genomic "scar" signature.
[0095] Example 18. The method according to Example 1, wherein the biomarker comprises the presence of a genomic instability score (GIS) composed of a pattern (or signature) of loss of heterozygosity (LOH), the number of telomeric imbalances (telomeric allelic imbalances or TAI), which is the number of regions with allelic imbalance extending to the subtelomere but not exceeding the centromere, and large-scale state transitions (LSTs) that are chromosomal breaks (deletions, translocations, and inversions).
[0096] Example 19. The method according to Example 1, wherein the biomarker comprises the presence of a set of homologous recombination features including the total number and percentage of deletions in the microhomology features of the sequencing data, the total number and percentage of genomic segments with deletions in the heterozygosity features of the sequencing data, the total number and percentage of heterozygous genomic segment features of the sequencing data, the total number and percentage of C:G>T:A single base substitutions in the 5'-NpCpG-3' context features of the sequencing data, or any combination thereof.
[0097] Example 20. The method according to Example 1, wherein the biomarker comprises the presence of genomic alterations in one or more of homologous recombination repair (HRR)-related or associated genes beyond BRCA1 and BRCA2 (also referred to as "BRCA-ness"), and the genomic alterations comprise alterations in PALB2, ATM, ATR, CHEK1 / 2, FANC genes (FANCA / C / D2 / E / F / G / I / L / M / 1), RAD50, RAD51 genes (RAD51B / C / D / L1 / 3), RAD52, RAD54L / C / D / B, ATRX, BAP1, BARD1, BRIP1, CDK12, PPP2R2A, MRE11, MRE11A, NBN, TP53, NCOR1, PTK2, ARID1A, BLM, WRN, CDK12, RPA1, EMSY, CCNE1, ERCC3, TAD54, XRCC2 / 3, HDAC2.
[0098] Example 21. The method according to Example 1, wherein the biomarker comprises the presence of genomic alterations that may potentially act in one or more of the genes such as ABL1, AKT1, ALK, APC, ATM, BRAF, RET, ROS, KRAS, NRAS, HRAS, RAF1, IDH1, IDH2, JAK1, JAK2, JAK3, KDR, KIT, MAP2K1, MET, NTRK, NTRK1, CCNE, CCNE1, CDK4 / 6, CCND1 / 2, AR, PDGFRA, PIK3CA, PTEN, CDH1, CDKN2A, CSF1R, CTNNB1, DDR2, DNMT3A, EGFR, ERBB2, ERBB3, ERBB4, HER2 / NEU, EZH2, FBXW7, FGF, FGFR, FGFR1, FGFR2, FGFR3, FLT3, FOXL2, GNA11, GNAQ, GNAS, HNF1A, MLH1, MPL, MSH6, NOTCH1, VEGFA, HGF, NPM1, PTPN11, RB1, SMAD4, SMARCB1, SMO, SRC, STK11, TP53, TSC1, VHL, ESR1, MAPK3K1, GATA3, CDH1, FBXW7, NF1, KMT2C, CTNNB1, GNA13, GNAQ, GNA11, RRAS2, KIF1A, KIF5B.
[0099] Example 22. The method according to Example 1, wherein the biomarker indicates a copy number change, deletion, amplification, fusion, mutation cluster, mutation signature, or any combination thereof of the genome of a biological sample.
[0100] Example 23. The method according to Example 1, wherein the section of the biological sample includes a paraffin-embedded section, a formalin-fixed section, a frozen section, a fresh section, or a combination of these sections.
[0101] Example 24. The method according to Example 1, wherein the trained prediction model includes a convolutional neural network.
[0102] Example 25. The method according to Example 1, wherein the trained prediction model includes a neural network such as a ResNet model.
[0103] Example 26. The method according to Example 1, further comprising generating reduced first and second pluralities of image data by reducing the parameter spaces of the first and second pluralities of image data. In some embodiments, the parameter spaces of the first and second pluralities of image data represent tiles at a magnification of 5 times, and the parameter spaces of the first and second pluralities of image data are reduced to 25%, 10%, or 5% of the tiles carrying prediction information.
[0104] Example 27. The method according to Example 26, wherein the reduction is completed by principal component analysis.
[0105] Example 28. The method according to Example 1, wherein the biological sample includes a non-cancerous biological sample or a cancerous biological sample.
[0106] Example 29. The method according to Example 1, wherein the biological sample includes healthy tissue, unhealthy tissue, or any combination of these tissues.
[0107] Example 30. The method according to Example 29, wherein the unhealthy tissue includes virus-infected tissue.
[0108] Example 31. The method according to Example 30, wherein the virus-infected tissue comprises human papillomavirus (HPV)-positive tissue.
[0109] Example 32. The method according to Example 30, wherein the virus-infected tissue comprises Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human immunodeficiency virus (HIV), human herpesvirus 8 (HHV-8), and / or human T-lymphotropic virus type 1 (HTLV-1), also known as human T-cell leukemia virus type 1.
[0110] Example 33. The method according to Example 29, wherein the unhealthy tissue comprises a nuclear morphology different from that of healthy tissue.
[0111] Example 34. The method according to Example 33, wherein the unhealthy tissue comprises a pre-cancerous condition or pre-cancerous tissue.
[0112] Example 35. The method according to Example 1, wherein the stain comprises hematoxylin and eosin stain.
[0113] Example 36. The method according to Example 1, wherein the first resolution comprises a magnification of 5 times and the second resolution comprises a magnification of 20 times.
[0114] Example 37. The method according to Example 26, further comprising generating first and second clustered data sets by clustering a plurality of reduced first and second image data.
[0115] Example 38. The method according to Example 37, wherein the clustering is completed by k-means clustering.
[0116] Example 39. The method according to Example 37, wherein the trained prediction model is trained with a clustered data set representing the top 15% of the variance between the clustered data set of the first and second clustered data sets and the corresponding biomarker label.
[0117] Example 40. The method according to Example 37, wherein the trained prediction model is trained with first and second clustered data sets of biological samples and corresponding biomarker labels, and the first and second clustered data sets include clustered data sets having silhouette coefficients within the top 50 percentile across all clusters of the first and second clustered data sets.
[0118] Example 41. The method according to Example 40, wherein the corresponding biomarker labels of the biological samples are determined by genome sequencing.
[0119] Example 42. The method according to Example 1, wherein the output of the trained prediction model includes the average prediction probability scores of the first and second prediction models.
[0120] Example 43. The method according to Example 1, wherein the one or more regions include at least 100 regions.
[0121] Example 44. The method according to Example 1, wherein the one or more regions include at most 10,000 regions.
[0122] Example 45. The method according to Example 1, including removing one or more nodes of the trained prediction model when the reduced first and second plurality of image data are provided as input to the trained prediction model.
[0123] Method for training a prediction model A method for generating a trained prediction model configured to determine the presence of biomarkers in a biological sample, the method comprising: (a) providing stained sections of one or more biological samples and corresponding biomarker labels; (b) generating first and second pluralities of image data by imaging one or more regions of the stained sections of the one or more biological samples at a first resolution and a second resolution; (c) generating reduced first and second pluralities of image data by reducing the parameter spaces of the first and second pluralities of image data; and (d) generating a trained prediction model comprising a first prediction model trained with the reduced first plurality of image data and corresponding biomarker labels and a second prediction model trained with the reduced second plurality of image data and corresponding biomarker labels.
[0124] Example 47. The method according to Example 46, wherein the trained prediction model is configured to determine the presence of biomarkers with at least 80% accuracy as compared to genomic sequencing.
[0125] Example 48. The method according to Example 47, wherein the accuracy includes at least 85%, at least 92%, at least 95%, at least 97%, or at least 99% as compared to genomic sequencing.
[0126] Example 49. The method according to Example 46, wherein the biomarker label includes a deletion of chromosome 9p.
[0127] Example 50. The method according to Example 46, wherein the biomarker includes the presence of clustered mutations in the gene TP53.
[0128] Example 51. The method according to Example 46, wherein the biomarker includes the presence of clustered mutations in the gene EGFR.
[0129] Example 52. The method according to Example 46, wherein the biomarker includes the presence of clustered mutations in the gene BRAF.
[0130] Example 53. The method according to Example 46, wherein the biomarker comprises the presence of a clustering mutation in the gene KIT.
[0131] Example 54. The method according to Example 46, wherein the biomarker comprises the presence of MSI (vs. MSS) and / or the presence of MMR gene deficiency including one or more of POLE, MLH1, MLH3, MGMT, MSH6, MSH3, MSH2, PMS1, and PMS2.
[0132] Example 55. The method according to Example 46, wherein the biomarker comprises the presence of a hypermutator mutation signature selected from POLE, MSI-COSMIC14 (POLE+MSI), MSI-COSMIC15 which combines MSI, MSI-COSMIC20 (POLD+MSI), MSI-COSMIC21, MSI-COSMIC26, and MSI-COSMIC6.
[0133] Example 56. The method according to Example 46, wherein the biomarker comprises the presence of APOBEC changes and mutation signatures.
[0134] Example 57. The method according to Example 45, wherein the biomarker comprises the presence of a high tumor mutation burden.
[0135] Example 58. The method according to Example 46, wherein the biomarker comprises the presence of homologous recombination repair deficiency (HRD). To determine HRD by quantifying genome instability in combination with BRCA1 and BRCA2 status, two commercially available HRD companion diagnostic (CDx) tests, Myriad myChoice® CDx and FoundationOne® CDx, are FDA-approved, and there are at least three academic HRD detection methods, SigMA, HRDetect, and CHORD.
[0136] Example 59. The method according to Example 46, wherein the biomarker comprises the presence of HRD negative (or homologous recombination proficiency (HRP)) or HRD positive, for example, using the genomic test of Example 14.
[0137] Example 60. The method according to Example 46, wherein the biomarker comprises the presence of a BRCA1 / 2 mutation.
[0138] Example 61. The method according to Example 46, wherein the biomarker comprises the presence of the "COSMIC3 - BRCA" mutation signature, which includes a specific pattern of genome - wide somatic single - nucleotide variations (SNVs) that includes "mutation signature 3" (Sig3) in the COSMIC signature catalog, or the presence of a genomic "scar" signature.
[0139] Example 62. The method according to Example 46, wherein the biomarker comprises the presence of a genomic instability score (GIS) composed of a pattern (or signature) of loss of heterozygosity (LOH) in an intermediate - sized region (greater than 15 MB and less than the entire chromosome), the number of telomeric allelic imbalances (telomeric allelic imbalance (TAI)), which is the number of regions with allelic imbalance that extends to the subtelomere but does not cross the centromere, and large - scale state transitions (LSTs) that are chromosomal breaks (deletions, translocations, and inversions).
[0140] Example 63. The method according to Example 46, wherein the biomarker comprises the presence of a set of homologous recombination features including the total number and proportion of deletions in the micro - homology features of sequencing data, the total number and proportion of genomic segments with deletions in the heterozygosity features of sequencing data, the total number and proportion of heterozygous genomic segment features of sequencing data, the total number and proportion of C:G>T:A single - base substitutions in the 5'-NpCpG - 3' context features of sequencing data, or any combination thereof.
[0141] Example 64. The method according to Example 46, wherein the biomarker comprises the presence of genomic changes in one or more of homologous recombination repair (HRR)-related or associated genes beyond BRCA1 and BRCA2 (also referred to as "BRCA-ness"), and the genomic changes include changes in PALB2, ATM, ATR, CHEK1 / 2, FANC genes (FANCA / C / D2 / E / F / G / I / L / M / 1), RAD50, RAD51 genes (RAD51B / C / D / L1 / 3), RAD52, RAD54L / C / D / B, ATRX, BAP1, BARD1, BRIP1, CDK12, PPP2R2A, MRE11, MRE11A, NBN, TP53, NCOR1, PTK2, ARID1A, BLM, WRN, CDK12, RPA1, EMSY, CCNE1, ERCC3, TAD54, XRCC2 / 3, HDAC2.
[0142] Example 65. The method according to Example 46, wherein the biomarker comprises the presence of genomic changes that can potentially act in one or more of genes such as ABL1, AKT1, ALK, APC, ATM, BRAF, RET, ROS, KRAS, NRAS, HRAS, RAF1, IDH1, IDH2, JAK1, JAK2, JAK3, KDR, KIT, MAP2K1, MET, NTRK, NTRK1, CCNE, CCNE1, CDK4 / 6, CCND1 / 2, AR, PDGFRA, PIK3CA, PTEN, CDH1, CDKN2A, CSF1R, CTNNB1, DDR2, DNMT3A, EGFR, ERBB2, ERBB3, ERBB4, HER2 / NEU, EZH2, FBXW7, FGF, FGFR, FGFR1, FGFR2, FGFR3, FLT3, FOXL2, GNA11, GNAQ, GNAS, HNF1A, MLH1, MPL, MSH6, NOTCH1, VEGFA, HGF, NPM1, PTPN11, RB1, SMAD4, SMARCB1, SMO, SRC, STK11, TP53, TSC1, VHL, ESR1, MAPK3K1, GATA3, CDH1, FBXW7, NF1, KMT2C, CTNNB1, GNA13, GNAQ, GNA11, RRAS2, KIF1A, KIF5B / / SEE RK MSKCC LINK, DS / CARIS EMAIL.
[0143] Example 66. The method according to Example 46, wherein the stained sections of one or more biological samples include paraffin-embedded sections, formalin-fixed sections, frozen sections, fresh sections, or any combination of these sections.
[0144] Example 67. The method according to Example 46, wherein the trained prediction model includes a convolutional neural network.
[0145] Example 68. The method according to Example 46, wherein the trained prediction model includes a ResNet model.
[0146] Example 69. The method according to Example 46, wherein the reduction is completed by principal component analysis.
[0147] Example 70. The method according to Example 46, wherein the one or more biological samples include non-cancerous biological samples, cancerous biological samples, healthy tissues, unhealthy tissues, or any combination of healthy and unhealthy tissues.
[0148] Example 71. The method according to Example 70, wherein the unhealthy tissue includes virus-infected tissue.
[0149] Example 72. The method according to Example 71, wherein the virus-infected tissue includes human papillomavirus (HPV)-positive tissue.
[0150] Example 73. The method according to Example 46, wherein the virus-infected tissue includes Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human immunodeficiency virus (HIV), human herpesvirus 8 (HHV-8), and / or human T-cell leukemia virus type 1, also known as human T-lymphotropic virus (HTLV-1).
[0151] Example 74. The method according to Example 70, wherein the unhealthy tissue includes a nuclear morphology different from that of the healthy tissue.
[0152] Example 75. The method according to Example 74, wherein the unhealthy tissue includes a pre-cancerous state or pre-cancerous tissue.
[0153] Example 76. The method according to Example 46, wherein the stain comprises hematoxylin-eosin stain.
[0154] Example 77. The method according to Example 46, wherein the first resolution includes a magnification of 5 times and the second resolution includes a magnification of 20 times.
[0155] Example 78. The method according to Example 46, further comprising generating first and second clustered data sets by clustering a plurality of reduced first and second image data.
[0156] Example 79. The method according to Example 78, wherein the clustering is completed by k-means clustering.
[0157] Example 80. The method according to Example 78, wherein the trained prediction model is trained with a clustered data set representing the top 15% of the variance between the clustered data set among the first and second clustered data sets and the corresponding biomarker label.
[0158] Example 81. The first and second prediction models are trained with first and second clustered data sets of one or more biological samples and corresponding biomarker labels, and the first and second clustered data sets include a clustered data set having a silhouette coefficient within the top 50 percentile across all clusters of the first and second clustered data sets. The method according to Example 78.
[0159] Example 82. The method according to Example 46, wherein the corresponding biomarker label of one or more biological samples is determined by genome sequencing.
[0160] Example 83. The method according to Example 46, wherein the output of the trained prediction model includes the average prediction probability scores of the first and second prediction models.
[0161] Example 84. The method according to Example 46, wherein the one or more regions include at least 100 regions.
[0162] Example 85. The method according to Example 46, wherein one or more regions comprise up to 1,000 regions.
[0163] Example 86. The method according to Example 46, wherein generating a trained prediction model includes removing one or more nodes of the first and second prediction models during training.
[0164] A system for determining the presence of a biomarker in a biological sample using a trained prediction model Example 87. A computer system configured to determine the presence of a biomarker in a biological sample, comprising one or more processors and, as a result of execution, (i) receiving a section of a stained biological sample, (ii) generating first and second pluralities of image data by imaging one or more regions of the stained section at a first resolution and a second resolution, (iii) generating reduced first and second pluralities of image data by reducing a parameter space of the first and second pluralities of image data, and (iv) determining the presence of a biomarker in the biological sample as an output of the trained prediction model when the reduced first and second pluralities of image data are provided as input to the trained prediction model, the non-transitory computer-readable storage medium including software including executable instructions causing the one or more processors of the computer system to perform the above.
[0165] Example 88. The system according to Example 87, wherein the trained prediction model is configured to determine the presence of a biomarker with at least 80% accuracy compared to genomic sequencing.
[0166] Example 89. The system according to Example 88, wherein the accuracy includes at least 85%, at least 92%, at least 95%, at least 97%, or at least 99% compared to genomic sequencing.
[0167] Example 90. The system according to Example 87, wherein the trained prediction model includes a first prediction model trained with respect to a first plurality of image data and a second prediction model trained with respect to a second plurality of image data.
[0168] Example 91. The system according to Example 87, wherein the biomarker includes a deletion of chromosome 9p.
[0169] Example 92. The system according to Example 87, wherein the biomarker includes the presence of clustered mutations in the gene TP53.
[0170] Example 93. The system according to Example 87, wherein the biomarker includes the presence of clustered mutations in the gene EGFR.
[0171] Example 94. The system according to Example 87, wherein the biomarker includes the presence of clustered mutations in the gene BRAF.
[0172] Example 95. The system according to Example 87, wherein the section of the biological sample includes a paraffin-embedded section, a formalin-fixed section, a frozen section, a fresh section, or any combination of these sections.
[0173] Example 96. The system according to Example 87, wherein the trained prediction model includes a convolutional neural network.
[0174] Example 97. The system according to Example 87, wherein the trained prediction model includes a ResNet model.
[0175] Example 98. The system according to Example 87, wherein the reduction is completed by principal component analysis.
[0176] Example 99. The system according to Example 87, wherein the biological sample includes healthy tissue, unhealthy tissue, or any combination of these tissues.
[0177] Example 100. The system according to Example 99, wherein the unhealthy tissue includes virus-infected tissue.
[0178] Example 101. The system according to Example 100, wherein the virus-infected tissue comprises human papillomavirus (HPV)-positive tissue.
[0179] Example 102. The system according to Example 100, wherein the virus-infected tissue comprises Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human immunodeficiency virus (HIV), human herpesvirus 8 (HHV-8), and / or human T-lymphotropic virus type 1 (HTLV-1), also known as human T-cell leukemia virus type 1.
[0180] Example 103. The system according to Example 99, wherein the unhealthy tissue comprises a nuclear morphology different from that of healthy tissue.
[0181] Example 104. The system according to Example 99, wherein the unhealthy tissue comprises a pre-cancerous condition or pre-cancerous tissue.
[0182] Example 105. The system according to Example 87, wherein the biological sample comprises a non-cancerous biological sample or a cancerous biological sample.
[0183] Example 106. The system according to Example 87, wherein the stain comprises hematoxylin and eosin stain.
[0184] Example 107. The system according to Example 87, wherein the first resolution comprises a magnification of 5 times and the second resolution comprises a magnification of 20 times.
[0185] Example 108. The system according to Example 87, wherein the instruction further comprises generating first and second clustered data sets by clustering a plurality of reduced first and second image data.
[0186] Example 109. The system according to Example 108, wherein the instruction for clustering is completed by k-means clustering.
[0187] Example 110. The system according to Example 108, wherein the trained prediction model is trained with a clustered data set that represents the top 15% of the variance between the clustered data set among the first and second clustered data sets and the corresponding biomarker label.
[0188] Example 111. The system according to Example 108, wherein the trained prediction model is trained with first and second clustered data sets of biological samples and corresponding biomarker labels of the biological samples, and the first and second clustered data sets include a clustered data set having a silhouette coefficient within the top 50 percentile across all clusters of the first and second clustered data sets.
[0189] Example 112. The system according to Example 111, wherein the corresponding biomarker label of the biological sample is determined by genome sequencing.
[0190] Example 113. The system according to Example 87, wherein the output of the trained prediction model includes an average prediction probability score of the first and second prediction models.
[0191] Example 114. The system according to Example 87, wherein the one or more regions include at least 100 regions or at most 1,000 regions, or at least 100 regions and at most 1,000 regions.
[0192] Example 115. The system according to Example 87, wherein the one or more processors include one or more processors of a smartphone, tablet, laptop, desktop, server, cloud computer architecture, or any combination thereof.
[0193] General method for converting image data into biomarker indicators Example 116. A method for determining the presence of a biomarker in a biological sample, comprising: (a) preparing a section of a stained biological sample; (b) generating a plurality of images of the stained section by imaging one or more regions of the stained section of the biological sample; and (c) when the plurality of images of the stained section are provided as input to a trained prediction model, determining the presence of the biomarker in the biological sample as an output of the trained prediction model, wherein the trained prediction model provides at least 80% accuracy in determining the presence of the biomarker as compared to genomic sequencing.
[0194] Example 117. The method according to Example 116, wherein the accuracy is at least 85%, at least 92%, at least 95%, at least 97%, or at least 99% as compared to genomic sequencing.
[0195] Example 118. The method according to Example 116, wherein the trained prediction model includes a first prediction model trained with respect to a first plurality of images acquired at a first resolution and a second prediction model trained with respect to a second plurality of images acquired at a second resolution.
[0196] Example 119. The method according to Example 116, wherein the biomarker includes a deletion of chromosome 9p.
[0197] Example 120. The method according to Example 116, wherein the biomarker includes the presence of clustered mutations in the gene TP53.
[0198] Example 121. The method according to Example 116, wherein the biomarker includes the presence of clustered mutations in the gene EGFR.
[0199] Example 122. The method according to Example 116, wherein the biomarker includes the presence of clustered mutations in the gene BRAF.
[0200] Example 123. The method according to Example 116, wherein the biomarker includes the presence of clustered mutations in the gene KIT.
[0201] Example 124. The method according to Example 16, wherein the biomarker comprises the presence of MSI (vs. MSS) and / or MMR gene (e.g., POLE, MLH1, MLH3, MGMT, MSH6, MSH3, MSH2, PMS1, or PMS2) deficiency.
[0202] Example 125. The method according to Example 116, wherein the biomarker comprises the presence of a hypermutator mutation signature comprising "POLE" and POLE composed of SBS6, SBS14, SBS15, SBS20, SBS21, SBS2, SBS26, SBS44.
[0203] Example 126. The method according to Example 116, wherein the biomarker comprises the presence of a high tumor mutation burden.
[0204] Example 127. The method according to Example 116, wherein the biomarker comprises the presence of APOBEC alterations and mutation signatures.
[0205] Example 128. The method according to Example 116, wherein the biomarker comprises the presence of homologous recombination repair deficiency (HRD). To determine HRD by quantifying genomic instability in combination with BRCA1 and BRCA2 status, two commercially available HRD companion diagnostic (CDx) tests, Myriad myChoice® CDx and FoundationOne® CDx, are FDA-approved, and there are at least three academic HRD detection methods, SigMA, HRDetect, and CHORD.
[0206] Example 129. The method according to Example 116, wherein the biomarker comprises the presence of HRD negative (or homologous recombination proficiency (HRP)) or HRD positive, for example, using the genomic test of Example 14.
[0207] Example 130. The method according to Example 116, wherein the biomarker comprises the presence of BRCA-1 and / or BRCA-2 mutations.
[0208] Example 131. The method according to Example 116, wherein the biomarker comprises the presence of the "COSMIC3-BRCA" mutation signature, which includes a specific pattern of genome-wide somatic single nucleotide variations (SNVs) defined as "mutation signature 3" (Sig3) in the COSMIC signature catalog, or the presence of a genomic "scar" signature.
[0209] Example 132. The method according to Example 116, wherein the biomarker comprises a genomic instability score (GIS) composed of a pattern (or signature) of loss of heterozygosity (LOH) in an intermediate-sized region, the number of telomere imbalances (telomere allele imbalance or TAI), and the presence of large-scale state transitions (LSTs) that are chromosomal breaks (deletions, translocations, and inversions).
[0210] Example 133. The method according to Example 116, wherein the biomarker comprises the presence of a homologous recombination feature set that includes the total number and percentage of deletions in the microhomology features of the sequencing data, the total number and percentage of genomic segments with deletions in the heterozygosity features of the sequencing data, the total number and percentage of heterozygous genomic segment features of the sequencing data, the total number and percentage of C:G>T:A single base substitutions in the 5'-NpCpG-3' context features of the sequencing data, or any combination thereof.
[0211] Example 134. The biomarker comprises the presence of genomic changes in one or more of the homologous recombination repair (HRR)-related or associated genes beyond BRCA1 and BRCA2 (also referred to as "BRCA-ness"), and the genomic changes include changes in PALB2, BARD1, ATM, BRIP1, CHEK1 / 2, CDK12, ATR, ATRX, BAP1, ARID1A, FANC genes (FANCA / C / D2 / E / F / G / I / / L / M, FANC1), RAD50, RAD51 genes (RAD51B / C / D / L1 / 3), RAD52, RAD54L / C / D / B, and other non-common HRR gene changes in PPP2R2A, MRE11, MRE11A, NBN, TP53, NCOR1, PTK2, BLM, WRN, RPA1, EMSY, CCNE1, ERCC3, TAD54, XRCC2 / 3, HDAC2, NPM1, PTWN, H2AX, RPA, or at least one of PRK2, NF1, the method according to Example 116.
[0212] Example 135. The method according to Example 116, comprising the presence of genomic changes that may potentially act on one or more of the genes such as ABL1, AKT1, APC, ALK, APC, BRAF, RET, ROS, KRAS,NRAS, HRAS, RAF1, KDR, MET, NTRK, NTRK1 / 2 / 3, CCNE, CCNE1, CDK4 / 6, CCND1 / 2, AR, PDGFRA, PIK3CA, PTEN, CDH1, CDKN2A, CSF1R, CTNNB1, DDR2, DNMT3A, EGFR, ERBB2, ERBB3, ERBB4, HER2 / NEU, EZH2, FBXW7, FGF, FGFR, FGFR1, FGFR2, FGFR3, FLT3, FOXL2, GNA11, GNA13, GNAQ, GNAS, HNF1A, MLH1, MPL, MSH6, NOTCH1, VEGFA, HGF, NPM1, PTPN11, RB1, SMAD4, SMARCB1, SMO, SRC, STK11, TP53, TSC1, VHL, ESR1, MAPK3K1, GATA3, CDH1, FBXW7, NF1, KMT2C, CTNNB1, RRAS2, KIF1A, KIF5B, IDH1 / 2, JAK1 / 2 / 3, MAP2K1, MAP3K1, GATA3, PTPN11, SRC, SETBP1, FAT1, KEAP1, LRP1B, FAT3, NF1, RB.
[0213] Example 136. The method according to Example 116, wherein the section of the biological sample comprises a paraffin-embedded section, a formalin-fixed section, a frozen section, a fresh section, or any combination of these sections.
[0214] Example 137. The method according to Example 116, wherein the trained prediction model comprises a convolutional neural network.
[0215] Example 138. The method according to Example 116, wherein the trained prediction model comprises a ResNet model.
[0216] Example 139. The method according to Example 116, further comprising generating reduced first and second pluralities of image data by reducing the parameter spaces of the first and second pluralities of image data.
[0217] Example 140. The method according to Example 116, wherein the reduction is accomplished by principal component analysis.
[0218] Example 141. The method according to Example 116, wherein the biological sample comprises a non-cancerous biological sample or a cancerous biological sample.
[0219] Example 142. The method according to Example 116, wherein the biological sample comprises healthy tissue, unhealthy tissue, or any combination of these tissues.
[0220] Example 143. The method according to Example 142, wherein the unhealthy tissue comprises virus-infected tissue.
[0221] Example 144. The method according to Example 143, wherein the virus-infected tissue comprises human papillomavirus (HPV)-positive tissue.
[0222] Example 145. The method according to Example 144, wherein the virus-infected tissue comprises Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human immunodeficiency virus (HIV), human herpesvirus 8 (HHV-8), and / or human T-lymphotropic virus type 1 (HTLV-1), also known as human T-cell leukemia virus type 1.
[0223] Example 146. The method according to Example 143, wherein the unhealthy tissue comprises a nuclear morphology different from that of healthy tissue.
[0224] Example 147. The method according to Example 143, wherein the unhealthy tissue comprises a pre-cancerous state or pre-cancerous tissue.
[0225] Example 148. The method according to Example 116, wherein the stain comprises hematoxylin and eosin stain.
[0226] Example 149. The method according to Example 118, wherein the first resolution comprises a magnification of 5× and the second resolution comprises a magnification of 20×.
[0227] Example 150. The method according to Example 143, further comprising generating first and second clustered data sets by clustering a plurality of reduced first and second image data.
[0228] Example 151. The method according to Example 150, wherein the clustering is completed by k-means clustering.
[0229] Example 152. The method according to Example 150, wherein the trained prediction model is trained with first and second clustered data sets of biological samples and corresponding biomarker labels of the biological samples, and the first and second clustered data sets include clustered data sets having a silhouette coefficient within the top 50 percentile across all clusters of the first and second clustered data sets.
[0230] Example 153. The method according to Example 152, wherein the corresponding biomarker label of the biological sample is determined by genome sequencing.
[0231] Example 154. The method according to Example 116, wherein the output of the trained prediction model includes the average prediction probability scores of the first and second prediction models.
[0232] Example 155. The method according to Example 116, wherein the one or more regions include at least 100 regions.
[0233] Example 156. The method according to Example 116, wherein the one or more regions include a maximum of 1,000 regions.
[0234] Example 157. The method according to Example 125, further comprising removing one or more nodes of the trained prediction model when a plurality of reduced first and second image data are provided as input.
[0235] A treatment method for treating cancer of a subject Example 158. A treatment method for treating a target cancer as needed, comprising preparing a section of a stained biological sample, generating a plurality of images of the stained section by imaging one or more regions of the stained section of the biological sample, and when the plurality of images of the stained section are provided as input to a trained prediction model, determining the presence of a biomarker in the biological sample as the output of the trained prediction model, wherein the trained prediction model provides at least 80% as the accuracy of determining the presence of the biomarker as compared to genomic sequencing, and administering treatment to a patient based on the presence of the biomarker.
[0236] Example 159. The treatment method according to Example 158, wherein the biomarker comprises a deletion of chromosome 9p.
[0237] Example 160. The treatment method according to Example 158, wherein the biomarker comprises the presence of clustered mutations in the gene TP53.
[0238] Example 161. The treatment method according to Example 158, wherein the biomarker comprises the presence of clustered mutations in the gene EGFR.
[0239] Example 162. The treatment method according to Example 158, wherein the biomarker comprises the presence of clustered mutations in the gene BRAF.
[0240] Example 163. The treatment method according to Example 158, wherein the biomarker comprises the presence of clustered mutations in the gene KIT.
[0241] Example 164. The treatment method according to Example 158, wherein the biomarker comprises the presence of MSI (vs. MSS) and / or MMR gene (e.g., POLE, MLH1, MLH3, MGMT, MSH6, MSH3, MSH2, PMS1, or PMS2) deficiency.
[0242] Example 165. The treatment method according to Example 158, wherein the biomarker comprises the presence of a hypermutator mutation signature selected from "POLE" and the POLE composed of SBS6, SBS14, SBS15, SBS20, SBS21, SBS2, SBS26, SBS44.
[0243] Example 166. The treatment method according to Example 158, wherein the biomarker comprises the presence of a high tumor mutation burden.
[0244] Example 167. The treatment method according to Example 158, wherein the biomarker comprises the presence of APOBEC alterations and mutation signatures.
[0245] Example 168. The treatment method according to Example 158, wherein the biomarker comprises the presence of homologous recombination repair deficiency (HRD). To determine HRD by quantifying genome-wide instability in combination with BRCA1 and BRCA2 status, two commercially available HRD companion diagnostic (CDx) tests, Myriad myChoice® CDx and FoundationOne® CDx, are FDA-approved, and there are at least three academic HRD detection methods, SigMA, HRDetect, and CHORD.
[0246] Example 169. The treatment method according to Example 158, wherein the biomarker comprises the presence of HRD negative (or homologous recombination proficiency (HRP)) or HRD positive, for example, using the genomic test of Example 168.
[0247] Example 170. The treatment method according to Example 158, wherein the biomarker comprises the presence of BRCA-1 and / or BRCA-2 mutations.
[0248] Example 171. The treatment method according to Example 158, wherein the biomarker comprises the presence of the "COSMIC3-BRCA" mutation signature containing a specific pattern of genome-wide somatic single nucleotide variations (SNVs) defined as "mutation signature 3" (Sig3) in the COSMIC signature catalog, or the presence of a genomic "scar" signature.
[0249] Example 172. The treatment method according to Example 158, wherein the biomarker comprises a genomic instability score (GIS) constituted by a pattern (or signature) of loss of heterozygosity (LOH) in an intermediate-sized region, the number of telomere imbalances (telomere allele imbalance, i.e., TAI), and the presence of large-scale state transitions (LSTs) that are chromosomal breaks (deletions, translocations, and inversions).
[0250] Example 173. The treatment method according to Example 158, wherein the biomarker comprises the total number and proportion of deletions in the microhomology features of sequencing data, the total number and proportion of genomic segments with deletions in the heterozygosity features of sequencing data, the total number and proportion of heterozygous genomic segment features of sequencing data, the total number and proportion of C:G>T:A single nucleotide substitutions in the 5'-NpCpG-3' context features of sequencing data, or the presence of a homologous recombination feature set comprising any combination thereof.
[0251] Example 174. The treatment method according to Example 158, wherein the biomarker comprises the presence of genomic changes in one or more of the homologous recombination repair (HRR)-related or associated genes beyond BRCA1 and BRCA2 (also referred to as "BRCA-ness"), and the genomic changes include changes in PALB2, BARD1, ATM, BRIP1, CHEK1 / 2, CDK12, ATR, ATRX, BAP1, ARID1A, FANC genes (FANCA / C / D2 / E / F / G / I / / L / M, FANC1), RAD50, RAD51 genes (RAD51B / C / D / L1 / 3), RAD52, RAD54L / C / D / B, as well as other non-common HRR gene changes in PPP2R2A, MRE11, MRE11A, NBN, TP53, NCOR1, PTK2, BLM, WRN, RPA1, EMSY, CCNE1, ERCC3, TAD54, XRCC2 / 3, HDAC2, NPM1, PTWN, H2AX, RPA, or at least one of PRK2, NF1.
[0252] Example 175. The treatment method according to Example 158, which includes the presence of genomic changes that can act on one or more of the following genes: ABL1, AKT1, APC, ALK, APC, BRAF, RET, ROS, KRAS,NRAS, HRAS, RAF1, KDR, MET, NTRK, NTRK1 / 2 / 3, CCNE, CCNE1, CDK4 / 6, CCND1 / 2, AR, PDGFRA, PIK3CA, PTEN, CDH1, CDKN2A, CSF1R, CTNNB1, DDR2, DNMT3A, EGFR, ERBB2, ERBB3, ERBB4, HER2 / NEU, EZH2, FBXW7, FGF, FGFR, FGFR1, FGFR2, FGFR3, FLT3, FOXL2, GNA11, GNA13, GNAQ, GNAS, HNF1A, MLH1, MPL, MSH6, NOTCH1, VEGFA, HGF, NPM1, PTPN11, RB1, SMAD4, SMARCB1, SMO, SRC, STK11, TP53, TSC1, VHL, ESR1, MAPK3K1, GATA3, CDH1, FBXW7, NF1, KMT2C, CTNNB1, RRAS2, KIF1A, KIF5B, IDH1 / 2, JAK1 / 2 / 3, MAP2K1, MAP3K1, GATA3, PTPN11, SRC, SETBP1, FAT1, KEAP1, LRP1B, FAT3, NF1, RB.
[0253] Example 176. The treatment method according to Example 163, wherein the patient has GIST and the treatment method does not include administering the c-Kit inhibitor imatinib.
[0254] Example 177. The treatment method according to Example 163, wherein the patient has GIST or another solid tumor, and the treatment method does not include administering a c-Kit inhibitor including axitinib, dovitinib, dasatinib, motesanib diphosphate, pazopanib, sunitinib, masitinib, bafetinib, cabozantinib, tivozanib, amuvatinib, teratinib, pazopanib, regorafenib, ripretinib, and dovitinib, in addition to imatinib.
[0255] Example 178. Further comprising administering to the subject a treatment for cancer comprising a drug class such as immune checkpoint inhibitors (ICIs) and other immunotherapies when the MSI, MMR gene (e.g., POLE, MLH1, MLH3, MGMT, MSH6, MSH3, MSH2, PMS1, or PMS2) deficiency, hypermutator mutation signature (e.g., COSMIC14 / 15 / 21 / 26 / 6), and / or high TMB of the sample is detected, the treatment method according to any one of Examples 164 to 166.
[0256] Example 179. Further comprising administering to the subject a treatment for cancer comprising a drug class such as PD-1 inhibitors (e.g., pembrolizumab, nivolumab, semipilumab, pidilizumab, dostarlimab, larotrectinib), PD-L1 inhibitors (e.g., atezolizumab, avelumab, durvalumab), CTLA-4 inhibitors (e.g., ipilimumab and tremelimumab), LAG-3 inhibitors (e.g., tebotelimab, eftilagimod alpha, relatlimab), TIM-3 inhibitors (e.g., MBG453, Sym023, TSR-022), other immunomodulatory therapies alone or in combination with other ICIs or other drugs when the MSI, MMR gene (e.g., POLE, MLH1, MLH3, MGMT, MSH6, MSH3, MSH2, PMS1, or PMS2) deficiency, hypermutator mutation signature (e.g., COSMIC14 / 15 / 21 / 26 / 6), and / or high TMB of the sample is detected, the treatment method according to any one of Examples 164 to 166.
[0257] Example 180. When the HRD or surrogate gene or signature is detected, further comprising treating the subject with cancer treatment comprising drug classes such as platinum agents, poly(ADP-ribose) polymerase (PARP) inhibitors, and / or novel drugs such as ATR, Wee1 or CHK, Pol-theta or RAD52 inhibitors, wherein the cancer comprises breast cancer, ovarian cancer, pancreatic adenocarcinoma, prostate cancer, sarcoma, or any solid tumor, or a combination of these cancers, the treatment method according to any one of Examples 168 to 174. In some embodiments, the use of weights collected from the final DeepHRD model trained to detect HRD in one tissue type or modality can result in model weights for another tissue type or modality. Since all other training procedures remain the same, knowledge transfer from the training of one tissue type or modality to another tissue type or modality is possible. In this treatment method, this approach can be utilized to train an ovarian cancer model by leveraging prior knowledge from a breast cancer model based on some embodiments of the disclosed technology. According to this AI algorithm, this deep learning AI technology can be applied to other genomic changes and other cancer types.
[0258] Example 181. When the HRD or surrogate gene or its signature of the sample is detected, further comprising treating the subject with cancer treatment comprising platinum agents including cisplatin, carboplatin, oxaliplatin, nedaplatin, lobaplatin, heptaplatin, or satraplatin alone or in combination with other drugs (e.g., FOLFOX), the treatment method according to any one of Examples 168 to 174. In some embodiments, this cancer treatment results in an interstrand break of the genomic molecules of the subject's cells, leading to p53-initiated apoptosis.
[0259] Example 182. When the HRD or surrogate gene or its signature of the sample is detected, in addition to olaparib (Lynparza), niraparib (Zejula), rucaparib (Rubraca), and talazoparib (Talzenna), which are four major poly ADP-ribose polymerase (PARP) inhibitors, further comprising administering to the subject a cancer treatment comprising a PARP inhibitor including other PARP inhibitors, the treatment method according to any one of Examples 168 to 174.
[0260] Example 183. The treatment method according to Example 158, comprising not administering an immune checkpoint inhibitor (ICI) and other immunotherapies to the subject when the 9p deletion of the sample is detected.
[0261] Example 184. The treatment method according to Example 158, wherein the biomarker comprises the presence of an EGFR / ErbB1 mutation including one or more of L858R, exon19del, and exon20 changes.
[0262] Example 185. The treatment method according to Example 184, further comprising administering afatinib, dacomitinib, erlotinib, gefitinib, osimertinib [T790], or amivantamab to the patient.
[0263] Example 186. The treatment method according to Example 158, wherein the biomarker comprises the presence of HER2 / ErbB2 amplification.
[0264] Example 187. The treatment method according to Example 186, further comprising administering trastuzumab, ado-trastuzumab emtansine, lapatinib, margetuximab, neratinib, pertuzumab, tucatinib, deruxtecan, trastuzumab deruxtecan, or neratinib to the patient.
[0265] Example 188. The treatment method according to Example 158, wherein the biomarker comprises the presence of a BRAF mutation.
[0266] Example 189. The treatment method according to Example 188, further comprising administering to a patient encorafenib, bemrafenib, dabrafenib, trametinib, or cobimetinib.
[0267] Example 190. The treatment method according to Example 158, wherein the biomarker comprises the presence of an FGFR1 / 2 / 3 fusion.
[0268] Example 191. The treatment method according to Example 190, further comprising administering to a patient erdafitinib, futibatinib, infliximab, pemigatinib, dovitinib, lenvatinib, pazopanib, ponatinib, or regorafenib.
[0269] Example 192. The treatment method according to Example 158, wherein the biomarker comprises the presence of a PDGFRA exon18 mutation.
[0270] Example 193. The treatment method according to Example 192, further comprising administering to a patient avapritinib or dasatinib.
[0271] Example 194. The treatment method according to Example 158, wherein the biomarker comprises the presence of a KIT mutation in GIST.
[0272] Example 195. The treatment method according to Example 194, further comprising administering to a patient imatinib, axitinib, dovitinib, dasatinib, motesanib diphosphate, pazopanib, sunitinib, masitinib, batatinib, cabozantinib, tiboxanib, amuvatinib, teratinib, pazopanib, regorafenib, ripretinib and vitinib, or sorafenib.
[0273] Example 196. The treatment method according to Example 158, wherein the biomarker comprises the presence of an NRG1 fusion.
[0274] Example 197. The treatment method according to Example 196, further comprising administering to a patient zilucoplan or seribantumab.
[0275] Example 198. The treatment method according to Example 158, wherein the biomarker comprises the presence of a RET fusion.
[0276] Example 199. The treatment method according to Example 198, further comprising administering to a patient pralsetinib, selpercatinib, crizotinib, ceritinib, cabozantinib, or vandetanib.
[0277] Example 200. The treatment method according to Example 158, wherein the biomarker comprises the presence of a ROS1 fusion.
[0278] Example 201. The treatment method according to Example 200, further comprising administering to a patient crizotinib or entrectinib.
[0279] Example 202. The treatment method according to Example 158, wherein the biomarker comprises the presence of an NTRK1 / 2 or 3 fusion.
[0280] Example 203. The treatment method according to Example 202, further comprising administering to a patient entrectinib, larotrectinib, or repotrectinib.
[0281] Example 204. The treatment method according to Example 158, wherein the biomarker comprises the presence of an ALK fusion.
[0282] Example 205. The treatment method according to Example 204, further comprising administering to a patient crizotinib, alectinib, brigatinib, ceritinib, or lorlatinib.
[0283] Example 206. The treatment method according to Example 158, wherein the biomarker comprises the presence of a PIK3CA alteration.
[0284] Example 207. The treatment method according to Example 206, further comprising administering to a patient alpelisib, temsirolimus, or everolimus.
[0285] Example 208. The treatment method according to Example 158, wherein the biomarker comprises the presence of an Mtor or TSC1 / 2 mutation.
[0286] The treatment method according to Example 208, further comprising administering temsirolimus or everolimus to a patient.
[0287] The treatment method according to Example 158, wherein the biomarker comprises the presence of Akt or PTEN alteration.
[0288] The treatment method according to Example 210, further comprising administering capivasertib to a patient.
[0289] The treatment method according to Example 158, wherein the biomarker comprises the presence of MET amplification or mutation.
[0290] The treatment method according to Example 212, further comprising administering crizotinib, tepotinib, capmatinib, terizotuzumab, tepotinib, or savolitinib to a patient.
[0291] The treatment method according to Example 158, wherein the biomarker comprises the presence of MEK mutation.
[0292] The treatment method according to Example 214, further comprising administering trametinib, cobimetinib, or selumetinib to a patient.
[0293] The treatment method according to Example 158, wherein the biomarker comprises the presence of NF1 / 2 alteration.
[0294] The treatment method according to Example 216, further comprising administering trametinib, temsirolimus, everolimus, or selumetinib to a patient.
[0295] The treatment method according to Example 158, wherein the biomarker comprises the presence of STK11 alteration.
[0296] The treatment method according to Example 218, further comprising administering dasatinib, everolimus, temsirolimus, or bosutinib to a patient.
[0297] Example 220. The treatment method according to Example 158, wherein the biomarker comprises the presence of a KDR change.
[0298] Example 221. The treatment method according to Example 220, further comprising administering pazopanib, regorafenib, or vandetanib to a patient.
[0299] Example 222. The treatment method according to Example 158, wherein the biomarker comprises microsatellite stability (MS) with a DNA polymerase-ε (POLE) mutation, CD274 amplification, or the presence of a 9p24.1 amplicon.
[0300] Example 223. The treatment method according to Example 222, further comprising administering an ICI to a patient.
[0301] Example 224. The treatment method according to Example 158, wherein the biomarker comprises the presence of a MAP2K change.
[0302] Example 225. The treatment method according to Example 224, further comprising administering trametinib to a patient.
[0303] Example 226. The treatment method according to Example 158, wherein the biomarker comprises the presence of a change in CCND2, CDK4, or CDKN2A / B.
[0304] Example 227. The treatment method according to Example 226, further comprising administering palbociclib to a patient.
[0305] Example 228. The treatment method according to Example 158, wherein the biomarker comprises the presence of an IDH1 mutation.
[0306] Example 229. The treatment method according to Example 228, further comprising administering ivosidenib to a patient.
[0307] Example 230. The treatment method according to Example 158, wherein the biomarker comprises the presence of a truncated or oncogenic mutation in B2M, PTEN, JAK1, JAK2, STK11, and EGFR, and / or a deletion in the 9p21 or 9p arm / gene region.
[0308] The treatment method according to Example 230, further comprising not administering an immune checkpoint inhibitor to a patient.
[0309] The treatment method according to Example 158, wherein the biomarker comprises the presence of mutations in the RAS genes KRAS and NRAS.
[0310] The treatment method according to Example 232, further comprising not administering an epidermal growth factor receptor (EGFR) therapy such as cetuximab and panitumumab to a patient with colorectal cancer and not administering an EGFR tyrosine kinase inhibitor such as erlotinib to a patient with lung cancer.
[0311] In some embodiments of the disclosed technology, genome sequencing encompasses any type of genomic profiling in which DNA and RNA are subjected to genotyping via next-generation ultra-parallel sequencing protocols or microarray hybridization.
[0312] In some embodiments of the disclosed technology, accuracy encompasses mathematical terms such as sensitivity, specificity, accuracy, negative predictive value, precision, and balanced accuracy, or any combination of these mathematical terms.
[0313] The subject matter and embodiments of the functional operations described in this patent document can be implemented in various systems, digital electronic circuits, or computer software, firmware, or hardware (including the structures disclosed herein and their respective structural equivalents), or in combinations of one or more of these. Embodiments of the subject matter described herein can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition that affects a machine-readable propagated signal, or a combination of one or more of these. The term "data processing unit" or "data processing apparatus" encompasses all devices, devices, and machines that process data, and by way of example, includes programmable processors, computers, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question (e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these).
[0314] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, such as a compiler-type or interpreter-type language, and can be deployed in any form, such as a single program or a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program in question, or multiple coordinated files (e.g., files that hold one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers located at one site or on multiple computers distributed across multiple sites and interconnected by a communication network.
[0315] The processes and logic flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operations on input data and generation of output. Also, these processes and logic flows can be executed by dedicated logic circuitry (e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit)). Also, the apparatus can be implemented as dedicated logic circuitry (e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit)).
[0316] Examples of processors suitable for the execution of a computer program include both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, the processor will receive instructions and data from read-only memory, random access memory, or both. Essential elements of a computer are a processor that executes instructions and one or more memory devices that store instructions and data. Also, generally, a computer will include one or more mass storage devices (such as magnetic, magneto-optical disks, or optical disks) for storing data, or will be operatively coupled to receive, transmit, or both, data to one or more mass storage devices. However, a computer does not necessarily have such devices. Examples of computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices). The processor and memory may be supplemented by, or incorporated in, dedicated logic circuitry.
[0317] This specification, together with the drawings, is intended to be considered as illustrative only, where "illustrative" here means by way of example. As used in this specification, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, the use of "or" is intended to include "and / or" as well, unless the context clearly dictates otherwise.
[0318] This patent document contains many details, but these should not be construed as limiting the scope of any invention or the claims, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Also, features described in this patent document in the context of separate embodiments may be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although a feature may be described as acting in a certain combination or even claimed as such from the start, in some cases it may be possible to remove one or more features from the combination according to the claims, and the combination according to the claims may also be directed to a sub-combination or a variation of a sub-combination.
[0319] Similarly, in the drawings, the operations are shown in a specific order, but it is understood that this does not require that such operations be performed in the specific order shown or the order of occurrence in order to achieve the desired result, nor does it require that all the operations shown be performed. Furthermore, it is understood that the separation of the various system components in the embodiments described in this patent document is not required in all embodiments.
[0320] Only a few embodiments and examples are described, and other embodiments, extensions, and variations based on what is described and illustrated in this patent document are also possible.
Claims
1. A method for determining the presence of biomarkers in a biological sample, Obtaining sections of biological samples treated with stain, To generate a plurality of first and second image data by imaging one or more regions of the stained section of the biological sample at a first resolution and a second resolution, Reducing the parameter space of the first and second plurality of image data to generate the reduced first and second plurality of image data, The first and second plurality of image data are fed to a trained predictive neural network, and the presence of the biomarker in the biological sample is determined as the output of the trained predictive neural network. Methods that include...
2. The method according to claim 1, wherein the trained predictive neural network is configured to determine the presence of the biomarker with a pre-set accuracy that is at least 80% of the genome sequencing accuracy.
3. The method according to claim 1, wherein the trained predictive neural network includes a first predictive model trained on the first plurality of image data and a second predictive model network trained on the second plurality of image data.
4. The method according to claim 1, wherein the biomarker includes a deletion of chromosome 9p.
5. The method according to claim 1, wherein the biomarker comprises the presence of at least one of microsatellite instability (MSI) deficiency or mismatch repair (MMR) gene deficiency, and the MMR gene deficiency comprises at least one of POLE, MLH1, MLH3, MGMT, MSH6, MSH3, MSH2, PMS1, or PMS2.
6. The method according to claim 1, wherein the biomarker includes the presence of a high tumor mutational load.
7. The method according to claim 1, wherein the biomarker includes the presence of a hypermutator mutation signature selected from MSI-COSMIC15, MSI-COSMIC20, MSI-COSMIC21, MSI-COSMIC26, and MSI-COSMIC6, which are combinations of POLE and MSI, including MSI-COSMIC14.
8. The method according to claim 1, wherein the biomarker includes the presence of homologous recombination repair defects (HRD).
9. The method according to claim 1, wherein the biomarker is HRD-negative, homologous recombination (HRP)-positive, or HRD-positive.
10. The method according to claim 1, wherein the biomarker includes the presence of at least one of a breast cancer gene (BRCA)-1 mutation or a BRCA-2 mutation.
11. The method according to claim 1, wherein the biomarker includes the presence of a genome instability score (GIS) that includes one or more of the following: a pattern or signature of loss of heterozygosity (LOH), a number of telomere imbalances corresponding to a number of allele-unbalanced regions that extend to the subtelomere but not beyond the centromere, or large-scale state transitions (LSTs) corresponding to chromosome breaks, wherein the telomere imbalance includes telomere-allelic imbalance (TAI), and the chromosome breaks include deletions, translocations, and inversions.
12. The method according to claim 1, wherein the biomarker comprises the presence of a potentially acting genomic alteration in at least one of ALK, BRAF, RET, ROS, KRAS, NRAS, HRAS, JAK1 / 2 / 3, KDR, KIT, MAPK, MTAP, MET, NTRK, NTRK1, PDGFRA, PIK3CA, EGFR, ERBB2, ERBB3, ERBB4, HER2 / NEU, FGFR, FGFR1, FGFR2, FGFR3, FLT3, ESR1, or NRG1.
13. The method according to claim 1, wherein the biomarker indicates at least one of immunohistochemical changes, copy number changes, deletions, amplifications, fusions, mutation clusters, mutation signatures, or any combination thereof of the genome of the biological sample.
14. The method according to claim 1, wherein the section of the biological sample includes a paraffin-embedded section, a formalin-fixed section, a frozen section, a fresh section, or a combination thereof.
15. The method according to claim 1, wherein the trained predictive neural network includes a convolutional neural network.
16. The method according to claim 1, wherein the parameter space of the first and second plurality of image data represents tiles at a magnification of 5x, and the parameter space of the first and second plurality of image data is reduced to 25%, 10%, or 5% of the tiles carrying prediction information.
17. The method according to claim 16, wherein the reduction is completed by principal component analysis.
18. The method according to claim 1, wherein the biological sample comprises healthy tissue, unhealthy tissue, or any combination thereof.
19. The method according to claim 18, wherein the unhealthy tissue includes virus-infected tissue containing one or more human T-cell leukemia virus types corresponding to Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human immunodeficiency virus (HIV), human herpesvirus 8 (HHV-8), and / or human T-lymphotropic virus (HTLV-1).
20. The method according to claim 1, wherein the stain comprises hematoxylin eosin stain.
21. The method according to claim 1, wherein the first resolution includes a low magnification of 5x or 10x, and the second resolution includes a high magnification of 20x or 40x.
22. The method according to claim 16, further comprising performing training to generate the trained predictive neural network by clustering the reduced first and second plurality of image data to generate first and second clustered datasets.
23. The method according to claim 22, wherein clustering is completed by k-means clustering.
24. The method according to claim 22, wherein the trained predictive neural network is trained on a clustered dataset representing the top 15% of the variance between the clustered dataset and the corresponding biomarker labels among the first and second clustered datasets.
25. The method according to claim 22, wherein the trained predictive neural network is trained on the first and second clustered datasets of the biological sample and corresponding biomarker labels, the first and second clustered datasets comprising a clustered dataset having silhouette coefficients within the top 50 percentile across all clusters of the first and second clustered datasets.
26. The method according to claim 25, wherein the corresponding biomarker label of the biological sample is determined by genome sequencing.
27. The method according to claim 1, wherein the output of the trained predictive neural network includes the average predicted probability scores of the first and second predictive neural networks.
28. The method according to claim 1, wherein the one or more regions include at least 100 regions and a maximum of 10,000 regions.
29. The method according to claim 1, comprising removing one or more nodes of the trained predictive neural network when the reduced first and second plurality of image data are given to the trained predictive neural network as input.
30. A computer system configured to determine the presence of biomarkers in a biological sample, One or more processors, As a result of execution, Acquiring a first set of multiple image data and a second set of multiple image data corresponding to a first set of images and a second set of images, by imaging one or more regions of a stained section of a biological sample at a first resolution and a second resolution, Reducing the parameter spaces of the first and second sets of image data to generate the reduced first and second sets of image data, The first and second sets of image data are provided to the trained predictive model. The output of the trained predictive model is to determine the presence of the biomarker in the biological sample, A non-temporary computer-readable storage medium storing software that includes executable instructions causing one or more processors of the computer system to perform the following: A system equipped with these features.
31. The system according to claim 30, wherein the trained predictive model is configured to determine the presence of the biomarker with a preliminary set accuracy of at least 80% of the genome sequencing accuracy.
32. The system according to claim 30, wherein the trained predictive model includes a first predictive model trained on a first plurality of image data and a second predictive model trained on a second plurality of image data.
33. The system according to claim 30, wherein the trained predictive model includes a convolutional neural network.
34. The system according to claim 30, wherein the computer system is configured to reduce the parameter space by principal component analysis.
35. The system according to claim 30, wherein the first resolution includes a magnification of 5x, and the second resolution includes a magnification of 20x.
36. The system according to claim 30, wherein the software further includes an instruction to one or more processors of the computer system to cluster the reduced first and second plurality of image data to generate first and second clustered datasets.
37. The system according to claim 36, wherein the clustering instruction is completed by k-means clustering.
38. The system according to claim 36, wherein the trained predictive model is trained on a clustered dataset representing the top 15% of the variance between the clustered dataset and the corresponding biomarker labels among the first and second clustered datasets.
39. The system according to claim 36, wherein the trained predictive model is trained on first and second clustered datasets of the biological samples and corresponding biomarker labels of the biological samples, the first and second clustered datasets comprising clustered datasets having silhouette coefficients within the top 50 percentile across all clusters of the first and second clustered datasets.
40. The system according to claim 39, wherein the corresponding biomarker label of the biological sample is determined by genome sequencing.
41. The system according to claim 30, wherein the output of the trained prediction model includes the mean prediction probability scores of the first and second prediction models.
42. The system according to claim 30, wherein the one or more regions include at least 100 regions or up to 1,000 regions, or at least 100 regions and up to 1,000 regions.
43. The system according to claim 30, wherein the one or more processors include one or more processors of a smartphone, tablet, laptop, desktop, server, cloud computing architecture, or any combination thereof.