Systems and methods for using sequencing data for pathogen detection - Patents.com
By using gene expression analysis to identify oncogenic pathogens in cancer, the method addresses the inefficiencies of separate assays, enabling accurate and personalized treatment decisions for HPV and EBV-associated cancers.
Patent Information
- Application Number
- JP2021550012
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-02-26
- Filing Date
- 2020-02-26
- Publication Date
- 2026-02-12
- Estimated Expiration
- 2040-02-26
AI Technical Summary
Conventional oncogenic pathogen diagnosis in cancer patients requires separate assays, increasing costs and delaying treatment plans, as it is typically performed after diagnosing the cancer type associated with the pathogen.
A method for identifying differentially expressed genes in cancerous tissue to distinguish between cancers with and without oncogenic pathogen infections using a classifier trained on gene expression data, allowing for tailored treatment decisions based on pathogen status without additional diagnostic assays.
Enables accurate detection of oncogenic pathogens in cancer with high specificity and sensitivity, facilitating personalized treatment strategies by identifying HPV and EBV infections in cervical and gastric cancers with 99% and 95% accuracy respectively, reducing the need for separate pathogen-specific assays.
Smart Images

Figure 0007813140000029 
Figure 0007813140000030 
Figure 0007813140000031
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 810,849, filed February 26, 2019, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0002] The present disclosure relates generally to detecting oncogenic pathogenic infections in cancer patients using expression profiles derived from cancerous tissue. [Background technology]
[0003] Precision oncology is the practice of tailoring cancer treatment to specific individuals, based on the unique pathological, genomic, epigenetic, and / or transcriptomic profiles of each individual tumor. In contrast, traditional cancer treatment is based solely on the type of cancer being treated. For example, traditionally, all breast cancers would be treated with one treatment regimen, while all lung cancers would be treated with a second treatment regimen. Precision oncology emerged from the frequent observation that different patients diagnosed with the same type of cancer, such as breast cancer, responded very differently to the same treatment regimen. Over time, researchers have identified genomic, epigenetic, and transcriptomic markers that facilitate some degree of prediction about how individual cancers will respond to specific therapies.
[0004] The use of targeted therapy has led to significant improvements in cancer patient outcomes, particularly in terms of progression-free survival (Radovich et al., Oncotarget, 7:56491-500 (2016)). Recent evidence reported from the IMPACT trial, which included genetic testing of advanced-stage tumors from 3,743 patients, with approximately 19% receiving targeted therapy matched based on tumor biology, showed a response rate of 16.2% in patients receiving matched therapy versus 5.2% in patients receiving mismatched therapy (Bankhead, “IMPACT Trial: Support for Targeted Cancer Tx Approaches,” MedPageToday, June 5, 2018). The IMPACT study also found that the 3-year overall survival rate for patients receiving molecularly matched therapy was more than double that of mismatched patients (15% vs. 7%). Id.; ASCO Post, "2018 ASCO: IMPACT Trial Matches Treatment to Genetic Changes in the Tumor to Improve Survival Across Multiple Cancer Conditions," The ASCO POST, June 6, 2018. Estimates of the percentage of patients whose care trajectory changes as a result of genetic testing vary widely, from approximately 10% to over 50%. Fernandes et al., Clinics, 72:588-94 (2017).
[0005] Therapies targeting specific genomic alterations are already standard of care in some tumor types, as suggested, for example, in the National Comprehensive Cancer Network (NCCN) guidelines for melanoma, colorectal cancer, and non-small cell lung cancer. Some well-known mutations in the NCCN guidelines can be identified in cancer patients using individual assays or small next-generation sequencing (NGS) panels. However, to maximize the number of cancer patients benefiting from personalized oncology, more comprehensive pathological, genomic, epigenetic, and / or transcriptomic analyses are needed to guide off-label drug use, combination therapy, or tissue-independent immunotherapy. (Schwaederle et al., JAMA Oncol., 2:1452-59 (2016); Schwaederle et al., J Clin Oncol., 32:3817-25 (2015); and Wheler et al., Cancer Res., 76:3690-701 (2016)).
[0006] The presence of oncogenic pathogen infection accounts for 10–12% of all cancers. For example, gastric cancer is the third most common cause of cancer death worldwide, with an estimated 700,000 deaths in 2012 (Ferlay, et al., “Cancer Incidence and Mortality Worldwide,” IARC CancerBase 11 [Internet], Lyon, France: International Agency for Research on Cancer (2013)). In addition to genetic factors, gastric carcinogenesis is thought to be related to multiple environmental factors, including Epstein-Barr virus (EBV) infection (Burke et al., Mod Pathol., 3:377–380 (1990)). In fact, a recent Cancer Genome Atlas study provided a molecular classification defining EBV-positive gastric cancer as a specific subtype (Cancer Genome Atlas Research Network, Nature, 513(7517):202–09 (2014)).
[0007] Therefore, the presence of such oncogenic pathogens affects the prognosis of related cancers.Therefore, if a subject has a type of cancer that is known to frequently occur in association with oncogenic pathogens, it is important to know the pathogen status of the subject, because it may change the treatment options of the subject.For example, many clinical trials investigating the benefits of reducing the dose of radiotherapy or chemotherapy for HPV-positive head and neck cancer have shown promising results.In addition, pathogen-associated tumors are likely to show higher levels of inflammation and immune infiltration, making them good candidates for immunotherapy.
[0008] A drawback of conventional oncogenic pathogen diagnosis is that to determine whether a subject is infected with a particular pathogen, a completely separate assay is performed separately from the assay used to initially diagnose the subject with cancer or to evaluate the stage of cancer. For example, in the case of EBV, separate laboratory methods such as in situ hybridization (ISH) or polymerase chain reaction (PCR) on excised tissue, biopsy, or blood, or enzyme-linked immunosorbent assay (ELISA) or immunofluorescence assay (IFA) on serum samples are performed to detect EBV infection. This increases the cost of diagnosis and is insufficient in some cases because pathogen testing is performed only after a type of cancer known to be associated with the oncogenic pathogen is diagnosed, delaying the creation of a subject's treatment plan until the results of the pathogen assay are obtained. Summary of the Invention
[0009] Given the above background, what is needed in the art are improved systems and methods for pathogen detection that directly determine the presence of a given pathogen without requiring a separate, independent assay for pathogen detection.
[0010] Thus, an improved method is provided for distinguishing between cancers associated with oncogenic pathogen infections that contribute to cancer pathology and cancers not associated with oncogenic pathogen infections. Improved methods for treating cancer patients based on whether their cancers are associated with oncogenic pathogen infections are also provided. The present disclosure addresses these needs, for example, by providing a method for identifying a set of genes that are differentially expressed in cancers associated with oncogenic pathogen infections compared to cancers not associated with oncogenic pathogen infections. The present disclosure also provides a method for training a classifier to distinguish between cancers associated with oncogenic pathogen infections and cancers not associated with oncogenic pathogen infections based on the identified genes that are differentially regulated in the two types of cancers. Thus, a method is also provided for using the trained classifier to classify cancers in patients as either associated with oncogenic pathogen infections or not associated with oncogenic pathogen infections. These methods then enable different treatments for patients based on whether their cancers are associated with oncogenic pathogen infections.
[0011] One aspect of the present disclosure provides a method for training a classifier to distinguish between a first cancer condition and a second cancer condition, where the first cancer condition is associated with infection by a first oncogenic pathogen and the second cancer condition is associated with a condition that does not include the oncogenic pathogen. The method includes acquiring, on a computer, a dataset including, for each respective subject in a plurality of subjects of a certain type, (i) a corresponding plurality of abundance values, where each respective abundance value in the corresponding plurality of abundance values quantifies an expression level of a corresponding gene in a plurality of genes in a tumor sample of the respective subject, and (ii) an indicator of a cancer condition for each subject, where the indicator identifies whether the respective subject has the first cancer condition or the second cancer condition, wherein the plurality of subjects includes a subset of the first subjects suffering from the first cancer condition and a subset of the second subjects suffering from the second condition.
[0012] The method then includes identifying a discriminatory gene set using the corresponding plurality of abundance values for each subject and each indicator of cancer status in the plurality of subjects, wherein the discriminatory gene set includes a subset of the plurality of genes.
[0013] In some embodiments, identifying the discriminatory gene set includes using a regression algorithm to regress the dataset based on all or a subset of the plurality of abundance values across the plurality of subjects for each indicator of cancer status across the plurality of subjects, thereby assigning a corresponding regression coefficient in the plurality of regression coefficients to each respective gene in the plurality of genes, and for the discriminatory gene set assigned coefficients by the regression algorithm that satisfy a coefficient threshold, selecting those genes in the plurality of genes.
[0014] In some alternative embodiments, identifying discriminatory gene sets includes dividing the dataset into a plurality of sets, each set in the plurality of sets including two or more subjects afflicted with a first cancer condition and two or more subjects afflicted with a second condition; independently regressing each respective gene in the plurality of sets using a regression algorithm based on all or a subset of a plurality of abundance values across the subjects in the respective set against respective indicators of cancer condition across the subjects in the respective set, thereby assigning a corresponding regression coefficient to each respective gene in the plurality of genes in the plurality of regression coefficients; and for discriminatory gene sets assigned coefficients by the regression algorithm that satisfy a coefficient threshold for at least a threshold percentage of the plurality of sets, selecting those genes in the plurality of genes. In some embodiments, the plurality of sets consists of 5 to 50 sets (e.g., 10 sets).
[0015] In some embodiments, the coefficient threshold is zero. The coefficient threshold is met if the absolute value of the corresponding regression coefficient is greater than zero.
[0016] In some embodiments, the regression algorithm disclosed above is logistic regression. In some such embodiments, the logistic regression assumes:
number
[0017] In some embodiments, the logistic regression is a logistic least absolute shrinkage and selection operator (LASSO) regression. In such embodiments, the logistic LASSO estimator
number
number
number
number
[0018] In some embodiments, the regression algorithm is logistic regression with L1 or L2 regularization.
[0019] The method further includes training a classifier (e.g., a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm) to use the abundance values of each of the discriminatory gene sets and the respective indicators of the cancer states across the plurality of subjects to discriminate between the first cancer state and the second cancer state as a function of the respective abundance values for the discriminatory gene sets.
[0020] Another aspect of the present disclosure provides a method for identifying a first cancer state and a second cancer state in a subject, wherein the first cancer state is associated with infection by a first oncogenic pathogen, and the second cancer state is associated with a condition that does not include oncogenic pathogens.The method includes: obtaining a dataset for the subject, the dataset includes a plurality of abundance values, and each abundance value in the plurality of abundance values quantifies the expression level of a corresponding gene in a plurality of genes in cancerous tissue from the subject.The method then includes inputting the dataset into a classifier that has been trained according to any one of the methodologies described herein.
[0021] Another aspect of the present disclosure provides a plurality of nucleic acid probes for identifying a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with an oncogenic pathogen infection and the second cancer condition is associated with a condition that does not contain an oncogenic pathogen. The nucleic acid probes have nucleic acid sequences that are complementary to or identical to the sequences of genes identified as being differentially expressed in cancers associated with oncogenic pathogen infection.
[0022] Another aspect of the present disclosure provides a method for identifying a first cancer state and a second cancer state in a subject with a first type of cancer, wherein the first cancer state is associated with infection by a first oncogenic pathogen, and the second cancer state is associated with a condition that does not include the oncogenic pathogen.The method includes obtaining a dataset for the subject, the dataset having a plurality of abundance values (e.g., relative mRNA expression values), and each abundance value in the plurality of abundance values quantifies the expression level of a corresponding gene in a discriminatory gene set in cancerous tissue from the subject.The method then includes inputting the dataset into a classifier trained to discriminate at least the first cancer state and the second cancer state based on the abundance values for the discriminatory gene set in the cancerous tissue of the subject, thereby determining the cancer state of the subject.
[0023] In some embodiments, the first type of cancer is breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophageal cancer, head and neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer.
[0024] In some embodiments, the dataset further comprises mutant allele counts for one or more mutant alleles at one or more loci in the genome of the cancerous tissue from the subject.
[0025] In some embodiments, the first cancerous condition is associated with infection by a first oncogenic pathogen selected from the group consisting of Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T-cell lymphotropic virus (HTLV-1), Kaposi-associated sarcoma virus (KSHV), and Merkel cell polyomavirus (MCV).
[0026] In some embodiments, the first cancer condition is selected from the group consisting of human papillomavirus (HPV)-associated cervical cancer, HPV-associated head and neck cancer, Epstein-Barr virus (EBV)-associated gastric cancer, EBV-associated nasopharyngeal carcinoma, EBV-associated Burkitt lymphoma, EBV-associated Hodgkin lymphoma, hepatitis B virus (HBV)-associated liver cancer, hepatitis C virus (HCV)-associated liver cancer, Kaposi's sarcoma virus (KSHV)-associated Kaposi's sarcoma, human T-cell lymphotropic virus (HTLV-1)-associated adult T-cell leukemia / lymphoma, and Merkel cell polyomavirus (MCV)-associated Merkel cell carcinoma.
[0027] In some embodiments, the first cancer condition is associated with infection by a human papillomavirus (HPV) oncogenic virus and the second cancer condition is associated with a non-HPV condition, and the discriminatory gene set comprises at least five genes selected from the genes set forth in Table 3. In some embodiments, the first cancer condition is cervical cancer associated with infection by a human papillomavirus (HPV). In some embodiments, the first cancer condition is head and neck cancer associated with infection by a human papillomavirus (HPV). In some embodiments, the discriminatory gene set comprises at least 10 genes selected from the genes set forth in Table 3. In some embodiments, the discriminatory gene set comprises at least 20 genes selected from the genes set forth in Table 3. In some embodiments, the discriminatory gene set comprises at least all 24 of the genes set forth in Table 3. In some embodiments, the dataset also comprises variant allele counts for TP53 (ENSG00000141510) and CDKN2A (ENSG00000147889) in the genome of cancerous tissue from the subject.
[0028] In some embodiments, the method also includes treating the subject for cervical cancer by administering a first therapy tailored to treat cervical cancer associated with HPV infection if the classifier results indicate that the human cancer patient is infected with an HPV oncogenic virus, and administering a second therapy tailored to treat cervical cancer not associated with HPV infection if the classifier results indicate that the human cancer patient is not infected with an HPV oncogenic virus. In some embodiments, the first therapy tailored to treat cervical cancer associated with HPV infection comprises a therapeutic vaccine or adoptive cell therapy. In some embodiments, the second therapy tailored to treat cervical cancer not associated with HPV infection is chemotherapy. In some embodiments, the chemotherapy comprises co-administration of cisplatin with a second therapeutic agent selected from the group consisting of 5-fluorouracil, paclitaxel, and bevacizumab.
[0029] In some embodiments, the method also includes treating the subject for head and neck cancer by administering a first therapy tailored to treat head and neck cancer associated with HPV infection if the classifier results indicate that the human cancer patient is infected with an HPV oncogenic virus, and administering a second therapy tailored to treat head and neck cancer not associated with HPV infection if the classifier results indicate that the human cancer patient is not infected with an HPV oncogenic virus. In some embodiments, the first therapy tailored to treat head and neck cancer associated with HPV infection comprises a therapeutic vaccine, an immune checkpoint inhibitor, or a PI3K inhibitor. In some embodiments, the second therapy tailored to treat head and neck cancer not associated with HPV infection comprises chemotherapy. In some embodiments, the chemotherapy comprises administration of cisplatin, and the second therapy also comprises concurrent radiation therapy or postoperative chemoradiotherapy.
[0030] In some embodiments, the first cancer condition is associated with infection by the Epstein-Barr Virus (EBV) oncogenic virus and the second cancer condition is associated with an EBV-free condition, and the discriminatory gene set comprises at least five genes selected from the genes set forth in Table 4. In some embodiments, the first cancer condition is gastric cancer associated with infection by Epstein-Barr Virus (EBV). In some embodiments, the discriminatory gene set comprises all nine genes set forth in Table 4. In some embodiments, the dataset also comprises mutant allele counts for TP53 (ENSG00000141510) and PIK3CA (ENSG00000121879) in the genome of cancerous tissue from the subject.
[0031] In some embodiments, the method also includes treating the subject for gastric cancer by administering a first therapy tailored to treat gastric cancer associated with EBV infection if the classifier results indicate the human cancer patient is infected with the EBV oncogenic virus, and administering a second therapy tailored to treat gastric cancer not associated with EBV infection if the classifier results indicate the human cancer patient is not infected with the EBV oncogenic virus. In some embodiments, the first therapy tailored to treat gastric cancer associated with EBV infection comprises an immune checkpoint inhibitor. In some embodiments, the second therapy tailored to treat gastric cancer not associated with EBV infection comprises chemotherapy. In some embodiments, the chemotherapy comprises administering a therapeutic agent selected from the group consisting of paclitaxel, carboplatin, cisplatin, 5-fluorouracil, and oxaliplatin.
[0032] In some embodiments, the method also includes treating the subject for cancer by administering a first therapy tailored for the treatment of a first type of cancer associated with infection by the first oncogenic pathogen if the results of the classifier indicate that the human cancer patient is infected with the first oncogenic pathogen, and administering a second therapy tailored for the treatment of a first type of cancer associated with an oncogenic pathogen-free condition if the results of the classifier indicate that the human cancer patient is not infected with the first oncogenic pathogen.
[0033] In some embodiments, the classifier is trained by a method comprising: (1) obtaining, for each respective subject in a plurality of subjects of a certain type, a dataset comprising: (i) a corresponding plurality of abundance values, each respective abundance value in the corresponding plurality of abundance values quantifying an expression level of a corresponding gene in a plurality of genes in a tumor sample of the respective subject; and (ii) an indicator of a cancer state for each subject, the indicator identifying whether the respective subject has a first cancer state or a second cancer state, wherein the plurality of subjects comprises a subset of the first subjects suffering from the first cancer state and a subset of the second subjects suffering from the second state; (2) identifying a discriminatory gene set using the corresponding plurality of abundance values and the respective indicators of cancer state for each subject in the plurality of subjects, the discriminatory gene set comprising a subset of the plurality of genes; and (3) training a classifier to discriminate between the first cancer state and the second cancer state as a function of the respective abundance values for the discriminatory gene set using the respective indicators of cancer state across the plurality of subjects.
[0034] Other embodiments are directed to systems, portable consumer devices, and computer-readable media associated with the methods described herein.
[0035] As disclosed herein, where applicable, any embodiment disclosed herein may be applied to any aspect.
[0036]
[0013] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details can be modified in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. [Brief explanation of the drawings]
[0037] [Figure 1A] 1 illustrates a block diagram of an exemplary computing device according to some embodiments of the present disclosure. [Figure 1B] 1 illustrates a block diagram of an exemplary computing device according to some embodiments of the present disclosure. [Figure 2A] 1 provides a flowchart of a process and features for training a classifier to distinguish between a first cancer condition associated with infection by a first oncogenic pathogen and a second cancer condition associated with a condition that does not contain the oncogenic pathogen, according to some embodiments of the present disclosure, where optional blocks are indicated by dashed boxes. [Figure 2B] 1 provides a flowchart of a process and features for training a classifier to distinguish between a first cancer condition associated with infection by a first oncogenic pathogen and a second cancer condition associated with a condition that does not contain the oncogenic pathogen, according to some embodiments of the present disclosure, where optional blocks are indicated by dashed boxes. [Figure 2C] 1 provides a flowchart of a process and features for training a classifier to distinguish between a first cancer condition associated with infection by a first oncogenic pathogen and a second cancer condition associated with a condition that does not contain the oncogenic pathogen, according to some embodiments of the present disclosure, where optional blocks are indicated by dashed boxes. [Figure 2D] 1 provides a flowchart of a process and features for training a classifier to distinguish between a first cancer condition associated with infection by a first oncogenic pathogen and a second cancer condition associated with a condition that does not contain the oncogenic pathogen, according to some embodiments of the present disclosure, where optional blocks are indicated by dashed boxes. [Figure 2E] 1 provides a flowchart of a process and features for training a classifier to distinguish between a first cancer condition associated with infection by a first oncogenic pathogen and a second cancer condition associated with a condition that does not contain the oncogenic pathogen, according to some embodiments of the present disclosure, where optional blocks are indicated by dashed boxes. [Figure 3]
[0013] Provided are flowcharts of processes and features for identifying a first cancerous condition associated with infection by a first oncogenic pathogen and a second cancerous condition associated with a condition that does not contain an oncogenic pathogen, and optionally treating the cancerous conditions based on the oncogenic pathogen status of the cancer, according to some embodiments of the present disclosure. [Figure 4A] 1 provides a breakdown of the composition of TGCA training and test datasets for training a classifier to distinguish between a first cancer condition associated with HPV oncogenic virus infection and a second cancer condition not associated with HPV oncogenic virus infection, according to some embodiments of the present disclosure. [Figure 4B] 1 illustrates cancerous tissue characteristics useful for distinguishing between a first cancerous condition associated with HPV oncogenic virus infection and a second cancerous condition not associated with HPV oncogenic virus infection, according to some embodiments of the present disclosure. [Figure 4C] 1 shows performance metrics for a support vector machine trained on a training dataset for identifying a first cancer condition associated with HPV oncogenic virus infection and a second cancer condition not associated with HPV oncogenic virus infection, according to some embodiments of the present disclosure. [Figure 4D] 1 shows performance metrics for a support vector machine trained on a validation dataset for identifying a first cancer condition associated with HPV oncogenic virus infection and a second cancer condition not associated with HPV oncogenic virus infection, according to some embodiments of the present disclosure. [Figure 5A] 1 provides a breakdown of the composition of TGCA training and testing datasets for training a classifier to distinguish between a first cancer condition associated with EBV oncogenic virus infection and a second cancer condition not associated with EBV oncogenic virus infection, according to some embodiments of the present disclosure. [Figure 5B] 1 illustrates cancerous tissue characteristics useful for distinguishing between a first cancerous condition associated with EBV oncogenic viral infection and a second cancerous condition not associated with HPV oncogenic viral infection, according to some embodiments of the present disclosure. [Figure 5C]1 shows performance metrics for a support vector machine trained on a training dataset for identifying a first cancer condition associated with EBV oncogenic virus infection and a second cancer condition not associated with EBV oncogenic virus infection, according to some embodiments of the present disclosure. [Figure 5D] 1 shows performance metrics for a support vector machine trained on a validation dataset for identifying a first cancer condition associated with EBV oncogenic virus infection and a second cancer condition not associated with EBV oncogenic virus infection, according to some embodiments of the present disclosure. [Figure 6A] Example 3, according to some embodiments of the present disclosure, shows a principal component analysis of expression signatures of genes identified as differentially expressed in head and neck cancer and cervical cancer associated with HPV viral infection in head and neck cancer and cervical cancer tissue samples. [Figure 6B] Example 4, according to some embodiments of the present disclosure, shows a principal component analysis of expression signatures of genes identified as differentially expressed in gastric cancer associated with EBV viral infection in head and neck cancer and cervical cancer tissue samples. [Figure 7A] 1 shows a reported example of HPV-positive head and neck squamous cell carcinoma, according to some embodiments of the present disclosure. [Figure 7B] 1 shows a reported example of HPV-positive cervical cancer, according to some embodiments of the present disclosure.
[0038] Like reference numerals refer to corresponding parts throughout the several views of the drawings. DETAILED DESCRIPTION OF THE INVENTION
[0039] The present disclosure provides systems and methods useful for distinguishing cancers associated with oncogenic pathogen infections that contribute to cancer pathology from cancers that are not associated with oncogenic pathogen infection. The present disclosure further provides systems and methods useful for treating cancer patients based on whether their cancer is associated with oncogenic pathogen infection.
[0040] Advantageously, the systems and methods described herein enable the detection of oncogenic pathogens in cancer without the need for additional diagnostic assays. Surprisingly, it has been found that oncogenic pathogen infections can be identified based on mRNA expression levels in tumor biopsies. Therefore, the present disclosure obviates the need for additional assays developed to identify nucleic acid or protein components of these pathogens. Rather, a single mRNA expression analysis can be performed to both characterize the transcriptional profile of a cancer and determine whether it is associated with oncogenic pathogen infection. For example, as reported in Example 3, a support vector machine classifier trained on only mRNA expression data and two allele states identified HPV infection in head and neck cancer and cervical cancer with 99% specificity and 99% sensitivity. Similarly, as reported in Example 4, a support vector machine classifier trained on only mRNA expression data and two allele states identified EBV infection in gastric cancer with 99% specificity and 95% sensitivity.
[0041] For example, in one embodiment, the present disclosure provides a method for training a classifier to distinguish between a first cancer condition and a second cancer condition, wherein the first cancer condition is associated with infection by a first oncogenic pathogen and the second cancer condition is associated with a condition that does not include the oncogenic pathogen. According to the method, referring to FIG. 4A, a dataset is obtained having a corresponding plurality of abundance values for each of a plurality of subjects of a certain type. Each respective abundance value quantifies the expression level of a corresponding gene in a plurality of genes in the tumor sample of each of the subjects. The dataset further includes an index of the cancer condition of each of the respective subjects tracked by the dataset. The index of the cancer condition identifies whether the subject has the first or second cancer condition (e.g., HPV-positive head and neck cancer or cervical cancer, or HPV-negative head and neck cancer or cervical cancer, as shown in FIG. 4A).
[0042] In some embodiments, each of the subjects has a particular cancer of the same origin (e.g., gastric cancer as shown in FIG. 5A ), and whether a subject is in a first or second cancer class is determined by whether the subject is also affected by an oncogenic pathogen known to be associated with that cancer (e.g., EBV virus in the case of FIG. 5A ), such that the prognosis of a subject with a cancer also affected by the oncogenic pathogen differs from the prognosis of a subject with a cancer not affected by the oncogenic pathogen. Some of the subjects tracked by the dataset (a subset of the first subjects) are affected by the first cancer condition, while some of the subjects tracked by the dataset (a subset of the second subjects) are affected by the second condition. A discriminatory gene set is then identified using the corresponding plurality of abundance values for each subject in the plurality of subjects and each indicator of the cancer condition. The discriminatory gene set includes a subset of the plurality of genes. Generally, the abundance levels (e.g., expression) of such genes distinguish between the first and second cancer conditions. Details regarding the discriminatory gene sets are disclosed below with reference to block 218 of Figure 2C. Figure 4B shows the discriminatory gene sets for HPV-associated cancers (head and neck cancer and cervical cancer), while Figure 5B shows the discriminatory gene sets for EBV-associated cancers (gastric cancer).
[0043] Using the respective abundance values for the discriminatory gene set and the respective indicators of cancer state across the plurality of subjects, a classifier is trained to distinguish between a first and a second cancer state as a function of the respective abundance values for the discriminatory gene set. In some optional embodiments, the trained classifier is used to classify the test subject into a first cancer or a second state (or determine the likelihood that the test subject has a first or second cancer state) by inputting the plurality of abundance values of the test into the trained classifier. In such embodiments, each respective abundance value in the plurality of abundance values of the test quantifies the expression level of a corresponding gene in the plurality of genes in the tumor sample of the test subject. In some optional embodiments, the results of the trained classifier are used to provide therapeutic intervention or imaging of the test subject based on a determination that the test subject has a first cancer state or a second cancer state (or the likelihood that the test subject has a first or second cancer state).
[0044] definition The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used in the description of the invention and in the claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. As used herein, the term "and / or" will also be understood to refer to and include any and all possible combinations of one or more of the associated listed items. It will also be understood that as used herein, the terms "comprises" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0045] As used herein, the term "if" may be interpreted to mean "if" or "when," or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "in response to determining" or "in response to detecting (a stated condition or event)" may be interpreted to mean "when determining" or "in response to determining," or "when detecting (a stated condition or event)" or "in response to detecting (a stated condition or event)," depending on the context.
[0046] It will also be understood that, although terms such as "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first subject can be referred to as a second subject, and similarly, a second subject can be referred to as a first subject, without departing from the scope of the present disclosure. A first subject and a second subject are both the same subject, but are not the same subject. Furthermore, the terms "subject," "user," and "patient" are used interchangeably herein.
[0047] As used herein, the term "subject" refers to any living or non-living organism, including, but not limited to, humans (e.g., male humans, female humans, fetuses, pregnant women, children, etc.), non-human mammals, or non-human animals. Any human or non-human animal can serve as a subject, including, but not limited to, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cows), equines (e.g., horses), caprines (e.g., sheep, goats), Suidae (e.g., pigs), Camelidae (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursinidae (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. In some embodiments, the subject is a male or female (e.g., a man, woman, or child) of any stage.
[0048] As used herein, the terms "control," "control sample," "reference," "reference sample," "normal," and "normal sample" refer to a sample from a subject who does not have a particular condition or is otherwise healthy. In one example, the method disclosed herein can be performed on a subject with a tumor, and the reference sample is a sample taken from the subject's healthy tissue. The reference sample can be obtained from the subject or a database. The reference can be, for example, a reference genome used to map sequence reads obtained from sequencing a sample from a subject. The reference genome can refer to a haploid or diploid genome to which sequence reads from a biological sample and a somatic sample can be aligned and compared. An example of a somatic sample can be DNA from white blood cells obtained from a subject. For a haploid genome, only one nucleotide can be present at each locus. For a diploid genome, heterozygous loci can be identified, and each heterozygous locus can have two alleles, with either allele allowing for alignment to the locus.
[0049] As used herein, the term "locus" refers to a position (e.g., site) within a genome, i.e., on a particular chromosome. In some embodiments, a locus refers to a single nucleotide position within a genome, i.e., on a particular chromosome. In some embodiments, a locus refers to a small group of nucleotide positions within a genome, such as those defined by mutations (e.g., substitutions, insertions, or deletions) of consecutive nucleotides within a cancer genome. Because normal mammalian cells have diploid genomes, a normal mammalian genome (e.g., a human genome) generally has two copies of every locus in the genome, or at least two copies of every locus on an autosome, i.e., one copy on a maternal autosome and one copy on a paternal autosome.
[0050] As used herein, the term "allele" refers to a particular sequence of one or more nucleotides at a chromosomal locus.
[0051] As used herein, the term "reference allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is either the predominant allele represented at that chromosomal locus within a population of a species (e.g., the "wild-type" sequence), or an allele that is predefined within a reference genome for the species.
[0052] As used herein, the term "variant allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is either not the predominant allele represented at that chromosomal locus within a population of a species (e.g., not the "wild-type" sequence) or is not a predefined allele within a reference genome for the species.
[0053] As used herein, the term "single nucleotide variation" or "SNV" refers to a substitution of one nucleotide with a different nucleotide at a position (e.g., site) in a nucleotide sequence, e.g., a sequence read from an individual. A substitution of a first nucleobase X with a second nucleobase Y can be indicated as "X>Y". For example, an SNV from cytosine to thymine can be indicated as "C>T".
[0054] As used herein, the term "mutation" or "mutant" refers to a detectable change in the genetic material of one or more cells. In certain instances, one or more mutations may be found in cancer cells and may identify the cancer cells (e.g., driver and passenger mutations). Mutations may be transmitted from a given cell to daughter cells. Those skilled in the art will understand that a genetic mutation in a parent cell (e.g., a driver mutation) may induce additional, different mutations (e.g., passenger mutations) in daughter cells. Mutations generally occur in nucleic acids. In certain instances, a mutation may be a detectable change in one or more deoxyribonucleic acids or fragments thereof. Mutations generally refer to nucleotides that have been added, deleted, substituted, inverted, or transposed to a new position in a nucleic acid. Mutations may be naturally occurring or experimentally induced. A mutation in a specific tissue sequence is an example of a "tissue-specific allele." For example, a tumor may have a mutation that results in an allele at a locus that does not occur in normal cells. Another example of a "tissue-specific allele" is a fetal-specific allele that occurs in fetal tissue but not in maternal tissue.
[0055] As used herein, the terms "cancer," "cancerous tissue," or "tumor" refer to an abnormal mass of tissue whose growth exceeds and is uncoordinated with that of normal tissue. Cancer or tumors can be defined as "benign" or "malignant" depending on the following characteristics: degree of cellular differentiation, including morphology and functionality, growth rate, local invasion, and metastasis. "Benign" tumors can be well differentiated, are characterized by slower growth than malignant tumors, and remain localized at the primary site. In addition, in some cases, benign tumors lack the ability to invade, infiltrate, or metastasize to distant sites. "Malignant" tumors can be poorly differentiated (anaplastic) and have characteristically rapid growth accompanied by progressive invasion, infiltration, and destruction of surrounding tissue. Furthermore, malignant tumors can have the ability to metastasize to distant sites. Thus, cancer cells are cells found within an abnormal mass of tissue whose growth is uncoordinated with that of normal tissue. Accordingly, a "tumor sample," as described herein, refers to a biological sample obtained from or derived from a tumor of a subject.
[0056] As used herein, "cancerous condition associated with oncogenic pathogen infection" refers, generally or with respect to a particular oncogenic pathogen, to a condition in which a cancer subject afflicted with a particular cancer is further afflicted with a pathogen (e.g., a virus) known to be associated with the particular cancer.
[0057] As used herein, a "cancer condition not associated with oncogenic pathogen infection" refers, generally or with respect to a particular oncogenic pathogen, to a condition in which a cancer subject afflicted with a particular cancer is not specifically afflicted with a pathogen (e.g., a virus) known to be associated with the particular cancer.
[0058] As used herein, "sequencing," "sequence determination," and like terms used herein generally refer to any and all biochemical processes that can be used to determine the order of biological macromolecules, such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as an mRNA transcript or a genomic locus.
[0059] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence produced by any sequencing process described herein or known in the art. A read can be generated from one end of a nucleic acid fragment (a "single-end read"), or in some cases, from both ends of a nucleic acid (e.g., a paired-end read, a double-end read). The length of a sequence read is often related to a particular sequencing technology. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, sequence reads are of an average, median, or average length of about 15 bp to 900 bp in length (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, sequence reads are of an average, median, or average length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. For example, nanopore sequencing can provide sequence reads that can vary in size from tens to hundreds or even thousands of base pairs. Illumina parallel sequencing can provide sequence reads that are less variable, e.g., most sequence reads can be less than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a series of nucleotides). For example, a sequence read can correspond to a series of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, a series of nucleotides at one or both ends of a nucleic acid fragment, or the nucleotides of an entire nucleic acid fragment. Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques, or using probes, e.g., hybridization arrays or capture probes, or amplification techniques such as polymerase chain reaction (PCR), linear amplification using a single primer, or isothermal amplification.
[0060] As used herein, the term "read segment" or "read" refers to any nucleotide sequence, including a sequence read obtained from an individual and / or a nucleotide sequence derived from an initial sequence read from a sample obtained from an individual. For example, a read segment can refer to an aligned sequence read, a folded sequence read, or a stitched read. Furthermore, a read segment can refer to an individual nucleotide base, such as a single base mutation.
[0061] As used herein, the terms "read depth," "sequencing depth," or "depth" refer to the total number of read segments from a sample obtained from an individual at a given location, region, or locus. A locus can be as small as a nucleotide, as large as a chromosome arm, or as large as the entire genome. Sequencing depth can be expressed as "Yx," e.g., 50x, 100x, etc., where "Y" refers to the number of times a locus is covered by sequence reads. In some embodiments, depth refers to the average sequencing depth across a genome, an exome, or a targeted sequencing panel. Sequencing depth can also apply to multiple loci, the whole genome, in which case Y refers to the average number of times a locus or haploid genome, whole genome, or whole exome is sequenced, respectively. When average depth is quoted, the actual depth for different loci included in a dataset can range beyond the range of values. Ultra-deep sequencing can refer to a sequencing depth at a locus of at least 100x.
[0062] As used herein, the term "sequencing breadth" refers to how many percentages of a particular reference exome (e.g., a human reference exome), a particular reference genome (e.g., a human reference genome), or a portion of an exome or genome have been analyzed. The denominator of the percentage can be the repeat-masked genome, so 100% can correspond to the entire reference genome excluding the masked portion. A repeat-masked exome or genome can refer to an exome or genome in which sequence repeats are masked (e.g., sequence reads align to unmasked portions of the exome or genome). Any portion of the exome or genome can be masked, and thus, any specific portion of the reference exome or genome can be focused on. Broad sequencing refers to sequencing and analyzing at least 0.1% of the exome or genome.
[0063] As used herein, the term "reference exome" refers to any particular known, sequenced, or characterized exome, whether partial or complete, of any tissue from any organism or pathogen that can be used to reference identified sequences from a subject. Exemplary reference exomes used for human subjects as well as many other organisms are provided in Examples 1 and 2.
[0064] As used herein, the term "reference genome" refers to any specific known, sequenced, or characterized genome, whether partial or complete, of any organism or pathogen that can be used to reference identified sequences from a subject. Exemplary reference genomes used for human subjects and many other organisms are provided in online genome browsers hosted by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz (UCSC). "Genome" refers to the complete genetic information of an organism or pathogen, expressed in nucleic acid sequences. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. A reference genome can be considered a representative set of genes for a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).
[0065] As used herein, the term "assay" refers to a technique for determining the characteristics of a substance, such as a nucleic acid, a protein, a cell, a tissue, or an organ. An assay (e.g., a first assay or a second assay) can include a technique for determining the copy number change of a nucleic acid in a sample, the methylation status of a nucleic acid in a sample, the fragment size distribution of a nucleic acid in a sample, the mutation status of a nucleic acid in a sample, or the fragmentation pattern of a nucleic acid in a sample. Any assay known to those skilled in the art can be used to detect any of the nucleic acid characteristics described herein. Nucleic acid characteristics can include sequence, genomic identity, copy number, methylation status at one or more nucleotide positions, nucleic acid size, the presence or absence of a mutation in a nucleic acid at one or more nucleotide positions, and the fragmentation pattern of a nucleic acid (e.g., the nucleotide positions at which the nucleic acid fragments). Assays or methods can have a particular sensitivity and / or specificity, and their relative usefulness as diagnostic tools can be measured using ROC-AUC statistics.
[0066] The term "classification" may refer to any number or other symbol associated with a particular characteristic of a sample. For example, a "+" symbol (or the word "positive") can indicate that the sample has been classified as having a deletion or amplification. In another example, the term "classification" may refer to the oncogenic pathogen infection status, the amount of tumor tissue in the subject and / or sample, the size of the tumor in the subject and / or sample, the stage of the tumor in the subject, the tumor burden in the subject and / or sample, and the presence of tumor metastasis in the subject. Classifications can be binary (e.g., positive or negative) or have more levels of classification (e.g., a scale of 1 to 10 or 0 to 1). The terms "cutoff" and "threshold" may refer to a predetermined number used in an operation. For example, a cutoff size may refer to the size above which a fragment is excluded. A threshold may be a value above or below which a particular classification is applied. Any of these terms may be used in any of these contexts.
[0067] As used herein, the term "relative abundance" can refer to the ratio of a first amount of nucleic acid fragments with a particular characteristic (e.g., aligning with a particular region of the exome) to a second amount of nucleic acid fragments with a particular characteristic (e.g., aligning with a particular region of the exome).In one example, relative abundance can refer to the ratio of the number of mRNA transcripts that encode a particular gene in a sample (e.g., aligning with a particular region of the exome) to the total number of mRNA transcripts in the sample.
[0068] As used herein, the term "untrained classifier" refers to a classifier that has not been trained on a training dataset.
[0069] As used herein, an "effective amount" or "therapeutically effective amount" is an amount sufficient to affect beneficial or desired clinical results during treatment. An effective amount can be administered to a subject in one or more doses. In terms of treatment, an effective amount is an amount sufficient to palliate, improve, stabilize, reverse, or slow the progression of a disease, or otherwise reduce the pathological consequences of a disease. An effective amount is generally determined by a physician on an individual basis and is within the skill of a person skilled in the art. When determining the appropriate dosage to achieve an effective amount, several factors are usually taken into consideration. These factors include the age, sex, and weight of the subject, the condition being treated, the severity of the condition, and the form and effective concentration of the therapeutic agent being administered.
[0070] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives.Sensitivity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly has a condition.For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects in a population that have cancer.In another example, sensitivity can characterize the ability of a method to correctly identify one or more markers that indicate cancer.
[0071] As used herein, the term " specificity " or " true negative rate " (TNR) refers to the number of true negatives divided by the total number of true negatives and false positives.Specificity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly does not have a condition.For example, specificity can characterize the ability of a method to correctly identify the number of subjects in a population that do not have cancer.In another example, specificity characterizes the ability of a method to correctly identify one or more markers that indicate cancer.
[0072] The terms used in this disclosure are for the purpose of describing particular cases only and are not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," or variations thereof are used in either the detailed description and / or claims, such terms are intended to be inclusive in the same manner as the term "comprising."
[0073] Some aspects are described below with reference to illustrative application examples. It should be understood that numerous specific details, relationships, and methods are set forth to provide a thorough understanding of the features described herein. However, those skilled in the art will readily recognize that the features described herein can be implemented without one or more of the specific details or in other ways. The features described herein are not limited by the illustrated order of acts or events, as some acts may occur in different orders and / or simultaneously with other acts or events. Furthermore, not all illustrated acts or events are required to implement a methodology in accordance with the features described herein.
[0074] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0075] Example System Embodiments Having provided an overview of some aspects of the present disclosure and some definitions used herein, details of an exemplary system will now be described in conjunction with FIG. 1. FIG. 1 is a block diagram illustrating a system 100 according to some implementations. The device 100 in some implementations includes one or more processing units CPU 102 (also referred to as a processor), one or more network interfaces 104, a user interface 106, non-persistent memory 111, persistent memory 112, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communication between the system components. The non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while the persistent memory 112 typically includes a CD-ROM, a digital versatile disc (DVD) or other optical storage, a magnetic cassette, a magnetic tape, a magnetic disk storage or other magnetic storage device, a magnetic disk storage device, an optical disk storage device, a flash memory device, or other non-volatile solid-state storage device. The persistent memory 112 optionally includes one or more storage devices located remotely from the CPU 102. The persistent memory 112 and the non-volatile memory devices within the non-persistent memory 112 comprise non-transitory computer-readable storage media. In some implementations, the non-persistent memory 111, or the non-transitory computer-readable storage media, possibly in combination with the persistent memory 112, stores the following programs, modules, data structures, or a subset thereof: · any operating system 116, including procedures for handling various basic system services and performing hardware-dependent tasks; · optional network communication modules (or instructions) 118 for connecting the system 100 with other devices and / or communication networks 105; an optional classifier training module 120 for training a classifier that distinguishes a first cancer condition associated with an oncogenic pathogen infection from a second cancer condition not associated with an oncogenic pathogen infection; an optional data store for a dataset for tumor samples from training subjects 122, including expression data from one or more training subjects 124, the expression data including a plurality of abundance data for each of a plurality of genes 126, supporting one or more genes 127, and a plurality of mutant alleles for each of cancer states 128; an optional classifier validation module 130 for validating a classifier that distinguishes a first cancer condition associated with an oncogenic pathogen infection from a second cancer condition not associated with an oncogenic pathogen infection; an optional data store for a dataset about tumor samples from a validation subject, the dataset including expression data from one or more training subjects, the expression data including a plurality of abundance data for each of a plurality of genes and cancer states; an optional patient classification module 134 for classifying cancer in a patient as either a first cancer condition associated with an oncogenic pathogen infection or a second cancer condition not associated with an oncogenic pathogen infection using a classifier, e.g., one trained using the classifier training module 120; an optional data store for a data construct about a cancer patient 136 that includes expression data from one or more cancer patients 140, the expression data including a plurality of abundance data for each of a plurality of genes 142; and Optional data store for a data construct about cancer patients 138, including variant allele data from one or more cancer patients 144, the variant allele data including multiple supports for variant alleles for each of one or more genes 146.
[0076] In various implementations, one or more of the above-identified elements are stored in one or more of the aforementioned memory devices and correspond to a set of instructions for performing the functions described above. The above-identified modules, data, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, data sets, or modules; thus, various subsets of these modules and data may be combined or otherwise reconfigured in various implementations. In some implementations, non-persistent memory 111 optionally stores a subset of the above-identified modules and data structures. Additionally, in some embodiments, memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system other than visualization system 100 that is addressable by visualization system 100 so that visualization system 100 may retrieve all or a portion of such data when needed.
[0077] While Figure 1 illustrates "system 100," the diagram is intended as a functional description of various features that may be present in a computer system, rather than as a structural schematic of the implementations described herein. In practice, and as will be recognized by those skilled in the art, items shown separately may be combined and some items may be separate. Additionally, while Figure 1 illustrates certain data and modules in non-persistent memory 111, some or all of these data and modules may be in persistent memory 112.
[0078] classifier training While a system according to the present disclosure is disclosed with reference to FIG. 1 , an overview of a method according to the present disclosure is provided in conjunction with FIG. 2A . In block 204 of FIG. 2A , a dataset is obtained. The dataset includes a corresponding plurality of abundance values for each respective subject in a plurality of subjects of a type. Each respective abundance value quantifies the expression level of a corresponding gene in a plurality of genes in the respective subject's tumor sample. The dataset further includes an indicator of a cancer state for each respective subject tracked by the dataset. The indicator of the cancer state identifies whether the subject has a first or second cancer state.
[0079] In some embodiments, each subject has a specific cancer (e.g., gastric cancer) of the same origin, and whether the subject is in the first or second cancer class is determined by whether the subject is also affected by an oncogenic pathogen known to be associated with this cancer, such that the prognosis of a subject with a cancer that also suffers from an oncogenic pathogen is different from the prognosis of a subject with a cancer that does not suffer from an oncogenic pathogen.For example, if the specific oncogenic pathogen is Epstein-Barr virus (EBV), each subject has a gastric cancer tumor, and whether the subject is in the first or second cancer class is determined by whether the subject is also affected by EBV.
[0080] In some embodiments, each subject has cancers related to a series of cancers, and whether the subject is in the first cancer class or the second cancer class is determined by whether the subject is also affected by an oncogenic pathogen known to be associated with any of the cancers in this series of cancers, such that the prognosis of the subjects with each cancer in the series of cancers affected by oncogenic pathogens is different from the prognosis of the subjects with each cancer not affected by oncogenic pathogens.For example, if a specific oncogenic pathogen is human papillomavirus (HPV), the series of cancers is head and neck squamous cell carcinoma and cervical cancer.That is, each subject has head and neck squamous cell carcinoma or cervical cancer, and whether the subject is in the first or second cancer class is determined by whether the subject is also infected with HPV.
[0081] In some embodiments, each of the subjects has a cancer listed in column 2 of the same row in Table 1 below, and whether a subject is in a first or second cancer class is delineated by whether the subject is also afflicted with an agent listed in column 1 of the same row in Table 1 below. See, e.g., Flora and Bonanni, Carcinogenesis 32(6), pp. 787-795, which is incorporated herein by reference. [Table 1]
[0082] As used herein, the term "human gut microbiota" refers to all microorganisms that inhabit the human digestive tract, a subset of which have been found to be carcinogenic. For example, pathogens that are hypothesized to cause or be associated with colon or colorectal cancer include sulfide-producing bacteria (e.g., Fusobacterium, Desulfovibrio, and Bilophila wadsworthia), Streptococcus bovis, and Fusobacterium nucleatum. For more information, see Dahmus et al., 2018, J Gastrointest Oncol., 9(4), pp. 769-77, the contents of which are incorporated herein in their entirety for all purposes.
[0083] Some of the subjects tracked by the dataset (a first subset of subjects) suffer from a first cancer condition, while some of the subjects tracked by the dataset (a second subset of subjects) suffer from a second condition. More details regarding such datasets are disclosed below with reference to block 202 of FIG. 2B.
[0084] Next, in block 218 of Figure 2A, a discriminatory gene set is identified using the corresponding plurality of abundance values for each subject in the plurality of subjects and each indicator of the cancer state. The discriminatory gene set includes a subset of the plurality of genes. Generally, the abundance levels (e.g., expression) of such genes distinguish between the first cancer state and the second cancer state. More details regarding discriminatory gene sets are disclosed below with reference to block 218 of Figure 2C.
[0085] Next, in block 242 of Figure 2A, the respective abundance values for the discriminatory gene sets and the respective indices of cancer states across the plurality of subjects are used to train a classifier to discriminate between the first and second cancer states as a function of the respective abundance values for the discriminatory gene sets. Details regarding training such a classifier based on discriminatory gene sets are disclosed below with reference to block 242 of Figure 2E.
[0086] Further, with reference to block 246 of FIG. 2A , in some optional embodiments, the trained classifier is used to classify a test subject into a first cancer or a second condition (or determine the likelihood that the test subject has a first or second cancer condition) by inputting the test plurality of abundance values into the classifier. In such embodiments, each respective abundance value in the test plurality of abundance values quantifies the expression level of a corresponding gene in the plurality of genes in the tumor sample of the test subject. The test subject is a subject whose test plurality of abundance values were not used to train the classifier. Furthermore, in typical examples, the test subject is a subject whose presence or absence of the first or second cancer condition has not been confirmed. Details regarding the diagnosis of a test subject using a trained classifier according to the present disclosure are disclosed below with reference to block 246 of FIG. 2E .
[0087] Further, with reference to block 248 of Figure 2A, in some optional embodiments, the results of the trained classifier are used to provide therapeutic intervention or imaging of the test subject based on a determination that the test subject has the first cancer condition or the second cancer condition (or the likelihood that the test subject has the first or second cancer condition). Details regarding such treatment options resulting from application of the trained classifier to abundance data of the plurality of test genes are disclosed below with reference to block 248 of Figure 2E.
[0088] Having provided an overview of the disclosed method in connection with FIG. 2A, attention is now directed to FIGS. 2B-2E, which provide further details regarding the disclosed method.
[0089] Block 202. Referring to block 202 of FIG. 2A, a method is provided for training a classifier to distinguish between a first cancer condition and a second cancer condition. As discussed above, the first cancer condition is associated with infection by a first oncogenic pathogen, and the second cancer condition is associated with a condition that does not include an oncogenic pathogen. Non-limiting examples of cancers known to be associated with oncogenic pathogen infection are described below with reference to FIG. 3. Thus, in some embodiments, the first cancer condition is a specific type of cancer associated with a specific oncogenic pathogen infection, e.g., as described below, and the second cancer condition is the same specific type of cancer that is not associated with a specific oncogenic pathogen infection. For example, in one embodiment, the first cancer condition is cervical cancer associated with HPV infection, and the second cancer condition is cervical cancer that is not associated with pathogen infection.
[0090] Block 204. Referring to block 204 of Figure 2A, a dataset is obtained that includes a corresponding plurality of abundance values for each respective subject in a plurality of subjects of a single species. Each respective abundance value in the corresponding plurality of abundance values quantifies the expression level of a corresponding gene in the plurality of genes in the tumor sample of each subject. The dataset further includes an indicator of a cancer state for each subject. The indicator of the cancer state identifies whether each subject has a first or second cancer state. The plurality of subjects includes a subset of first subjects suffering from the first cancer state and a subset of second subjects suffering from the second state.
[0091] Block 206. Referring to block 206, in some embodiments, the corresponding plurality of abundance values are obtained by RNA-seq. RNA-seq is a methodology for RNA profiling based on next-generation sequencing, which enables the measurement and comparison of gene expression patterns across multiple subjects. In some embodiments, millions of short series, called "sequence reads," are generated by sequencing random positions in cDNA prepared from input RNA obtained from the subject's tumor tissue. These reads can then be computer-mapped to a reference genome to reveal a "transcription map," and the number of sequence reads aligned to each gene provides a measure of its expression level (e.g., abundance). Next-generation sequencing is disclosed in Shendure, 2008, "Next-generation DNA sequencing," Nat. Biotechnology 26, pp. 1135-1145, which is incorporated herein by reference. RNA-seq is disclosed in Nagalakshmi et al., 2008, “The transcriptional landscape of the yeast genome defined by RNA sequencing,” Science 320, pp. 1344-1349, and Finotell and Camillo, 2014, “Measuring differential gene expression with RNA-seq: challenges and strategies for data analysis,” Briefings in Functional Genomics 14(2), pp. 130-142, each of which is incorporated herein by reference.
[0092] According to block 206, for each tumor sample from each subject in the plurality of subjects, the RNA in the sample of interest is first fragmented and reverse transcribed into complementary DNA (cDNA). The resulting cDNA is then amplified and subjected to next-generation DNA sequencing (NGS). In principle, any NGS technology can be used for RNA-seq. In some embodiments, an Illumina sequencer (see the internet at illumina.com) is used. See Wang, Z., et al., "RNA-Seq: a revolutionary tool for transcriptomics," Nat Rev Genet., 10(1):57-63 (2009), which is incorporated herein by reference. The millions of short reads generated for each such sample are then mapped to a reference genome, and the number of reads aligned to each gene, called "counts," provides a digital measurement of the gene expression level in the sample under investigation.
[0093] In some alternative embodiments, rather than using RNA-seq, microarrays are used to measure gene abundance values. Such microarrays have been widely used in studies such as Wang et al., 2009, "RNA-Seq: a revolutionary tool for transcriptomics," Nat Rev Genet 10, pp. 57-63; Roy et al., 2011, "A comparison of analog and next-generation transcriptomic tools for mammalian studies," Brief Funct Genomic 10:135-150; Shendure, 2008, "The beginning of the end for microarrays?" Nat Methods 5, pp. 585-587; Cloonan et al., 2008, "Stem cell transcriptome profiling via massive-scale mRNA sequencing," Nat. Methods 5, pp. 613-619; Mortazavi et al., 2008, "Mapping and quantifying mammalian transcriptomes by RNA-Seq," Nat. Methods 5, pp. 621-628; and Bullard et al. al., 2010, "Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments," BMC Bioinformatics 11, p. 94, each of which is incorporated herein by reference.
[0094] The first computational step in the RNA-seq data analysis pipeline is read mapping, in which reads are aligned to a reference genome or transcriptome by identifying genomic regions that match the read sequence. Any of a variety of alignment tools can be used for this task. See, e.g., Hatem et al., 2013, "Benchmarking short sequence mapping tools," BMC Bioinformatics 14, p. 184, and Engstrom et al., "Systematic evaluation of spliced alignment programs for RNA-seq data," Nat Methods 10, pp. 1185-1191, each of which is incorporated herein by reference. In some embodiments, the mapping process begins by building an index of either the reference genome or the reads, which is then used to search for a set of positions in the reference sequence to which the reads are likely to align. Once this subset of possible mapping positions is identified, alignments are performed at these candidate regions using slower, more sensitive algorithms. See, e.g., Hatem et al., 2013, "Benchmarking short sequence mapping tools," BMC Bioinformatics 14:184, and Flicek and Birney, 2009, "Sense from sequence reads: methods for alignment and assembly," Nat Methods 6 (Suppl. 11), S6-S12, each of which is incorporated herein by reference. In some embodiments, the mapping tool is a methodology that utilizes a hash table or a Burrows-Wheeler transform (BWT).See, e.g., Li and Homer, 2010, “A survey of sequence alignment algorithms for next-generation sequencing,” Brief Bioinformatics 11, pp. 473-483, which is incorporated herein by reference.
[0095] After mapping, counts are calculated using the reads aligned to each coding unit, such as exons, transcripts, or genes, to provide estimates of their abundance (e.g., expression) levels. In some embodiments, such counts consider the total number of reads that overlap with the exons of a gene. However, in some instances, because some sequence reads map outside the known exon boundaries, alternative embodiments consider the full length of the gene and also count reads from introns. Furthermore, in some embodiments, spliced reads are used to model the abundance of different splicing isoforms of a gene. See, e.g., Trapnell et al., 2010, "Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation," Nat Biotechnol 28, pp. 511-515, and Gatto et al., 2014, "Fine-Splice, enhanced splice junction detection and quantification: a novel pipeline based on the assessment of diverse RNA-Seq alignment solutions," Nucleic Acids Res 42, p. e71, each of which is incorporated herein by reference.
[0096] As explained above, quantifying gene abundance from RNA-seq data is typically implemented in an analysis pipeline through two computational steps: aligning reads to a reference genome or transcriptome, and then estimating gene and isoform abundance based on the aligned reads. Unfortunately, the reads generated by most commonly used RNA-seq techniques are generally much shorter than the transcripts from which they were sampled. As a result, it is not always possible to uniquely assign short sequence reads to specific genes in the presence of transcripts with similar sequences. Such sequence reads are called "multiple reads" because they are homologous to two or more regions of the reference genome. In some embodiments, such multiple reads are discarded, i.e., they do not contribute to the gene abundance count. In some embodiments, programs such as MMSEQ or RSEM are used to resolve ambiguities. See, e.g., Turro et al., 2011, "Haplotype and isoform specific expression estimation using multi-mapping RNAseq reads," Genome Biol 12, p. R13, and Nicolae et al., "Estimation of alternative splicing isoform frequencies from RNA-Seq data," Algorithms Mol Biol 6, p. 9, each of which is incorporated herein by reference.
[0097] Another aspect of RNA-seq is the normalization of sequence read counts, which in some embodiments involves normalization to take into account different sequencing depths. See, e.g., Lin et al., 2011, "Comparative studies of de novo assembly tools for next-generation sequencing technologies," Bioinformatics 27, pp. 2031-2037; Robinson Oshlack, 2010, "A scaling normalization method for differential expression analysis of RNA-seq data," Genome Biol 11, p. R25; and Li et al., 2012, "Normalization, testing, and false discovery rate estimation for RNA-sequencing data," Biostatistics 13, pp. 523-538, each of which is incorporated herein by reference. In some embodiments, sequence read counts are normalized to account for gene length bias. Finotell and Camillo, 2014, "Measuring differential gene expression with RNA-seq: challenges and strategies for data analysis," Briefings in Functional Genomics 14(2), pp. 130-142, which is incorporated herein by reference.
[0098] Block 208. Referring to block 208 of Figure 2B, in some embodiments, each subject in the plurality of subjects is afflicted with a first type of cancer. In other words, in some embodiments, each subject in database 122 is afflicted with the same type of cancer. In some such embodiments, each subject in the plurality of subjects has breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophageal cancer, head and neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer.
[0099] Block 210. Referring to block 208 of Figure 2B, in some embodiments, each subject in the plurality of subjects has a first type of cancer at a first stage. In other words, in some embodiments, each subject in database 122 has the same type of cancer, and the cancer is at the same stage. In some such embodiments, each subject in the plurality of subjects has breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophageal cancer, head and neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer. Further, in such embodiments, the stage of the cancer in each subject in the plurality of subjects is stage I, stage II, stage III, or stage IV cancer.
[0100] Blocks 212-214. Referring to block 212 of FIG. 2B and block 214 of FIG. 2C, the cohort used in the disclosed methods is of sufficient size to develop a classifier with performance suitable for screening subjects to identify whether they have a first or second cancer condition. Thus, in some embodiments, the plurality of subjects includes 100 subjects, the first subset of subjects (those with the first cancer condition) includes 20 subjects, and the second subset of subjects (those with the second cancer condition) includes 20 subjects. This is by way of example only. In other embodiments, the plurality of subjects includes 1000 subjects, the first subset of subjects includes 100 subjects, and the second subset of subjects includes 100 subjects. In still other embodiments, the plurality of subjects comprises 100, 500, 2000, 4000, or 10,000 subjects, wherein the first subset of subjects comprises 100, 500, or 1000 subjects, and the second subset of subjects consists of 100, 500, or 1000 subjects. In some embodiments, more of the subjects have the first cancer condition than the second cancer condition. For example, in some embodiments, more than 10 percent, more than 20 percent, more than 30 percent, more than 40 percent, more than 50 percent, more than 60 percent, more than 70 percent, more than 80 percent, or more than 90 percent of the subjects in dataset 122 have the first cancer condition, and the remainder have the second cancer condition.
[0101] Block 216. Referring to block 216 of Figure 2C, in some embodiments, the disclosed methods are used with training subjects that are humans. Each training subject in dataset 122 is from the same species, although the species need not be human. In some embodiments, the species is canine, bovine, porcine, or some other species.
[0102] Block 218. Referring to block 218 of FIG. 2C , once a dataset 122 is obtained that includes a corresponding plurality of abundance values for each respective subject in a plurality of subjects of a single species, the dataset 122 is used to identify a discriminatory gene set using the abundance values of each subject and each indicator of cancer status in the plurality of subjects of the dataset 122. The discriminatory gene set includes a subset of the plurality of genes. Particular methods for identifying a discriminatory gene set according to some embodiments of the present disclosure are described in more detail below with reference to blocks 226-240.
[0103] Blocks 220-224. Referring to block 220 of FIG. 2C , in some embodiments, the species under consideration is human, the plurality of genes (for which abundance data is considered) includes 10,000 or more genes, e.g., the xGen Exome Research Panel v1.0 (IDT) spans a 39 Mb target region containing 19,396 genes (see Nguyen, A., et al., “Multiplexed Hybrid Capture for Whole Exome Sequencing,” Technical Note, Integrated DNA Technologies, Inc., (2018), the contents of which are incorporated by reference in their entirety for all purposes), and the discriminatory gene set consists of 5-40 genes. Referring to block 222 of FIG. 2C , in some embodiments, the species under consideration is human, the plurality of genes includes 5,000 genes, and the discriminatory gene set consists of 5-25 genes. Other ranges are possible. For example, in some embodiments, the plurality of genes (abundance data taken into account) includes at least 200, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 15000, or 20000 genes, and the discriminating gene set consists of 5 to 500 genes, 5 to 100 genes, 5 to 50 genes, or 5 to 20 genes. Regardless of the range, the range of discriminating genes is smaller than the range of the original plurality of genes. In some embodiments, the discriminating set consists of at least one-quarter of the genes in the plurality of genes in dataset 122 (e.g., a reduction from 1000 genes to 250 genes or less). By selecting a smaller set of genes for the discriminatory gene set than is available in dataset 122, the algorithm for discriminating between the first and second states can be trained with smaller, more informative data (e.g., abundance data for fewer genes), which leads to more computationally efficient training of classifiers for discriminating between the first and second cancer states.Such improvements in computational efficiency due to reduced size of the discriminatory gene set can be advantageously used to speed up classifier training or improve the performance of such classifiers (e.g., through more extensive training of the classifier). In some embodiments, the discriminatory gene set consists of at least one-quarter, one-fifth, one-sixth, one-seventh, one-eighth, one-ninth, one-tenth, one-twentieth, one-thirtieth, one-fortieth, or one-fiftieth of the genes in dataset 122. Furthermore, reducing the number of genes used in the analysis improves the model by preventing overfitting of the data.
[0104] Block 226. Referring to block 226 of FIG. 2C , in some embodiments, identifying the discriminatory gene set includes using a regression algorithm to regress the dataset 122 based on all or a subset of the plurality of abundance values 126 across the plurality of training subjects 124 against each indicator of cancer state 128 across the plurality of training subjects 124, thereby assigning a corresponding regression coefficient in a plurality of regression coefficients to each respective gene in the plurality of genes. Thus, in such embodiments, cancer state is the dependent variable and the gene abundance values are the independent variables. In such embodiments, the genes from the plurality of genes selected for the discriminatory gene set are genes whose coefficients are assigned by the regression algorithm that satisfy a coefficient threshold. In such embodiments, genes whose coefficients satisfy the coefficient threshold are deemed sufficiently significant to have a significant impact on the dependent variable, cancer class, and are therefore retained for the discriminatory gene set. Details of such regression in certain embodiments of the present disclosure are presented below.
[0105] Blocks 228-232. Referring to block 228 of FIG. 2D , in some embodiments, identifying a discriminatory gene set includes dividing the dataset into a plurality of sets (e.g., 5-50 sets, exactly 10 sets, etc.). Each set in the plurality of sets includes two or more subjects afflicted with a first cancer condition and two or more subjects afflicted with a second cancer condition. Each respective set in the plurality of sets is then independently regressed using a regression algorithm based on all or a subset of a plurality of abundance values across the subjects of the respective set against respective indicators of cancer condition across the subjects of the respective set, thereby assigning a corresponding regression coefficient in a plurality of regression coefficients to each respective gene in the plurality of genes. Those genes assigned regression coefficients by the regression algorithm that satisfy a coefficient threshold for at least a threshold percentage of the plurality of sets are selected for the discriminatory gene set. Referring to block 230, in some embodiments, the coefficient threshold is zero. In some embodiments, the required threshold percentage is at least 40 percent of the plurality of sets. Thus, for illustrative purposes, consider the case of 10 sets. In such a case, for gene A to be included in the discriminatory gene set, when regressing each of the 10 sets against the cancer state, the regression coefficient for gene A must meet the regression threshold in 4 of the 10 sets. A regression threshold of zero means that a positive regression coefficient is required to meet the regression threshold, and the regression coefficient for gene A must be positive in at least 4 of the 10 sets. In some embodiments, the threshold is applied to the absolute value of the coefficient. However, in some embodiments described herein, the threshold is set to 0 because the LASSO regression is designed to return sparse coefficients. In some embodiments, the required threshold percentage is at least 50 percent, at least 60 percent, at least 70 percent, at least 80 percent, at least 90 percent, or all of the multiple sets.Referring to block 232, in some embodiments, the regression coefficient threshold is greater than zero (e.g., 0.1, 0.2, 0.3, or some other positive value). It will be appreciated that requiring a larger regression coefficient serves to increase the stringency of what is required for a gene to be included in the identification dataset. In various alternative embodiments, a regression coefficient meets the regression coefficient threshold if the absolute value of the regression coefficient during regression is non-zero, greater than 0.1, or greater than 0.2.
[0106] Blocks 234-240. Note that the dependent variable used in identifying the discriminatory gene set takes one of two labels: a first cancer state or a second cancer state. Thus, referring to block 234 of FIG. 2D , in some embodiments, the regression algorithm is a logistic regression that assumes:
number
[0107] Referring to block 238, in some embodiments, the logistic regression is a logistic least absolute shrinkage and selection operator (LASSO) regression. In such embodiments, a logistic LASSO estimator
number
number
number
number
[0108] In some embodiments, a regularization method other than LASSO is used to identify genes in the plurality of genes that distinguish between the first and second cancer states based on gene abundance values across the training subjects 124 of the dataset 122. For example, in some embodiments, an elastic net is used to identify genes in the plurality of genes that distinguish between the first and second cancer states based on gene abundance values across the training subjects 124 of the dataset 122. See Zou and Hastie, 2005, "Regularization and variable selection via the elastic net," JR Stat Soc Series B Stat Methodol 67, pp. 301-320, which is incorporated herein by reference. In some embodiments, a sparse Laplacian penalty is used to identify genes in a plurality of genes that distinguish between first and second cancer states based on gene abundance values across training subjects 124 of dataset 122. See Huang et al., 2011, "The sparse Laplacian shrinkage estimator for high-dimensional regression," Ann Stat 39, pp. 2021-2046, which is incorporated herein by reference.In some embodiments, elastic nets, group lasso (Yuan and Lin, 2006, "Model Selection and Estimation in Regression with Grouped Variables," Journal of the Royal Statistical Society. Series B Statistical Methodology 68(1), pp. 49-67), fused lasso (Tibshirani et al., 2005, "Sparsity and Smoothness via the Fused lasso," Journal of the Royal Statistical Society. Series B Statistical Methodology 67(1), pp. 91-108), quasi-norm and bridge regression (Fu, 1998, "The Bridge versus the Lasso," Journal of Computational and Graphical Statistics 7(3), pp. 397-416), or adaptive lasso are used to identify genes in a plurality of genes that distinguish between the first and second cancer states based on gene abundance values across training subjects 124 of dataset 122. Referring to block 240 of FIG. 2E, in some embodiments, the regression algorithm includes an L1 (LASSO) or L2 (Ridge) regularization term.
[0109] Blocks 242-244. The above disclosure details how gene abundance values 126 of subjects 124 in training set 122 are used to identify a discriminatory gene set whose abundance values collectively distinguish between first and second cancer states. Once this discriminatory gene set is identified, training set 122 is used to formally train a classifier that can distinguish between first and second cancer states for the test subject using abundance values of discriminatory genes measured from biological samples taken from the test subject. In typical embodiments, the cancer state of the test subject is unknown. That is, the test subject may be known to have a particular cancer, but it is not known whether the subject is suffering from a pathogen that adversely affects the prognosis of the subject's cancer. In typical embodiments, the biological sample used to measure gene abundance values for the test subject is a solid tumor within the test subject. Referring to block 242, in some embodiments, the respective abundance values for the discriminatory gene set and respective indicators of cancer state across multiple subjects are used to train a classifier to distinguish between the first and second cancer states as a function of the respective abundance values for the discriminatory gene set. In some embodiments, as disclosed in the Examples below, additional features are utilized in addition to the abundance values of the discriminatory gene set to train the classifier. For example, in some embodiments, the absence of specific mutations in selected genes is also used to train the classifier in conjunction with the abundance values for the discriminatory gene set.
[0110] Referring to block 244 of FIG. 2E, in some embodiments, by way of non-limiting example, the classifier used in block 242 is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine (SVM) algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision key algorithm, a clustering algorithm, or a combination thereof.
[0111] A logistic regression algorithm suitable for use as the classifier in block 242 is disclosed, for example, in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103-144, John Wiley & Son, New York, which is incorporated by reference.
[0112] Neural network algorithms, including convolutional neural network algorithms, suitable for use as the classifier in block 242 are disclosed, for example, in Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference. A neural network has a layered structure including a layer of input units (and bias) connected to a layer of output units by a layer of weights. In the case of regression, the layer of output units typically includes only one output unit. However, neural networks can process multiple quantitative responses in a seamless fashion. In a multi-layer neural network, there are input units (input layer), hidden units (hidden layer), and output units (output layer). In addition, there is a single bias unit connected to each unit other than the input units. Additional exemplary neural networks suitable for use as the classifier in block 242 are disclosed in Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, Inc., New York, and Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, each of which is incorporated herein by reference in its entirety.Additional exemplary neural networks suitable for use as the classifier in block 242 are described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC, and Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is incorporated herein by reference in its entirety.
[0113] An SVM algorithm suitable for use as the classifier in block 242 is described, for example, in Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge; and Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5th International Conference on Machine Learning and Machine Learning (ICHM) , Vol. 1, No. 1, pp. 111-115, 2002. thAnnual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York, Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, an SVM separates a given set of binary labeled data training set (here, the first and second cancer statuses of each subject in dataset 122) with a hyperplane that is maximally separated from the labeled data. When linear separation is not possible, an SVM works in conjunction with kernel techniques to automatically achieve a nonlinear mapping to feature space. The hyperplane found by the SVM in feature space corresponds to a nonlinear decision boundary in input space.
[0114] A naive Bayes classifier suitable for use as the classifier in block 242 is disclosed, for example, in Ng et al., 2002, “On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes,” Advances in Neural Information Processing Systems, 14, which is incorporated herein by reference.
[0115] Decision tree algorithms suitable for use as the classifier in block 242 are described, for example, in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is incorporated herein by reference. Tree-based methods partition the feature space into a set of rectangles and fit a model (such as a constant) in each. In some embodiments, the decision tree is a random forest regression. One particular algorithm that can be used as the classifier in block 244 is a classification and regression tree (CART). Other examples of particular decision tree algorithms that can be used as the classifier in block 244 include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and pp. 411-412, which are incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forests are described in Breiman, 1999, "Random Forests--Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated herein by reference in its entirety.
[0116] Clustering algorithms suitable for use as the classifier in block 242 are described, for example, in Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter "Duda 1973"), pages 211-256, which is incorporated herein by reference in its entirety. As described in section 6.7 of Duda 1973, the clustering problem is described as one of finding natural groupings in a dataset. To identify natural groupings, two problems are addressed. First, a method for measuring the similarity (or dissimilarity) between two samples is determined. This metric (similarity measure) is used to ensure that samples in one cluster are more similar to each other than to samples in the other cluster. Here, the similarity measure is at the abundance level of the discriminatory gene set across the training dataset 122. Next, a mechanism for dividing the data into clusters using the similarity measure is determined. Similarity measures are described in Section 6.7 of Duda 1973, and one way to begin a clustering study is to define a distance function and calculate a matrix of distances between all pairs of samples in a data set. If distance is an appropriate measure of similarity, the distance between samples in the same cluster will be significantly shorter than the distance between samples in different clusters. However, as described on page 215 of Duda 1973, clustering does not require the use of a distance metric. For example, one may compare two vectors x and x' using the nonmetric similarity function s(x, x'). Conventionally, s(x, x') is a symmetric function whose value is large if x and x' are "similar" in some way. An example of a nonmetric similarity function s(x, x') is provided on page 216 of Duda 1973.
[0117] Once a method for measuring "similarity" or "dissimilarity" between points in a data set has been chosen, clustering utilizes a criterion function that measures the clustering quality of any partition of the data. The partition of the data set that limits the criterion function is used to cluster the data. See page 217 of Duda 1973. Criterion functions are discussed in section 6.8 of Duda 1973. More recently, Duda et al., Pattern Classification, 2 nd edition, John Wiley & Sons, Inc., New York. Pages 537-563 provide a detailed discussion of clustering. Further details on clustering techniques suitable for use as the classifier in block 242 can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, NY; Everitt, 1993, Cluster analysis (3rd ed.), Wiley, New York, NY; and Backer, 1995, Computer-Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, NJ. Specific exemplary clustering techniques that can be used as the classifier in block 242 include, but are not limited to, hierarchical clustering (agglomerative clustering using nearest-neighbor, farthest-neighbor, average linkage, centroid, or sum-of-squares algorithms), k-means clustering, fuzzy k-means clustering, and Jarvis-Patrick clustering.
[0118] In some embodiments, the classifier used in block 242 is a nearest neighbor algorithm. For nearest neighbor, given a query point x0 (test object), the k training points x0 that are closest to x0 are found. (r), r, ..., k (here, training objects) are identified, and the point x is classified using its k nearest neighbors, where the distance between these neighbors is a function of the abundance values of the discriminative gene set. In some embodiments, the distance is calculated using Euclidean distance in feature space as d (i) =||x (i) -x (O) || is determined. Typically, when a nearest neighbor algorithm is used, the abundance data used to calculate the linear discriminant are standardized to have a mean of zero and a variance of one. Nearest neighbor rules can be refined to address issues of imbalanced class preferences, differential misclassification costs, and feature selection. Many of these refinements involve some form of weighted voting for neighbors. For more information on nearest neighbor analysis, see Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, each of which is incorporated herein by reference.
[0119] Blocks 246-248. The above disclosure describes training a classifier using abundance values of a discriminatory gene set.
[0120] Referring to block 246, in some embodiments, a trained classifier is used to classify a test subject to determine whether the test subject has a first cancer condition or a second cancer condition by inputting the test plurality of abundance values into the classifier. In such embodiments, each respective abundance value in the test plurality of abundance values quantifies the expression level of a corresponding gene in a plurality of genes, more specifically a discriminatory gene set, in the test subject's biological sample (e.g., a tumor sample). In response to this input, the classifier designates whether the test subject has the first cancer condition or the second cancer condition.
[0121] Referring to block 246, in some alternative embodiments, the trained classifier is used to determine the likelihood or probability that the subject has the first cancer condition or the second condition. This is done in such embodiments by inputting a plurality of test abundance values into the classifier. In such embodiments, each respective abundance value in the plurality of test abundance values quantifies the expression level of a corresponding gene in a plurality of genes (more specifically, a discriminatory gene set) in the test subject's biological sample (e.g., a tumor sample). In response to this input, the classifier assigns a likelihood or probability that the test subject has the first cancer condition or a likelihood or probability that the test subject has the second cancer condition.
[0122] Referring to block 248, in some embodiments, therapeutic intervention or imaging of the test subject is provided based on a determination that the test subject has the first cancer condition or the second cancer condition (or the likelihood that the test subject has the first or second cancer condition). Examples of such conditional treatments are provided below in conjunction with Figure 3. For example, non-limiting examples of ongoing clinical trials of therapies for specific cancer types associated with oncogenic pathogen infection are shown in Table 2 below.
[0123] RNA analysis pipeline In some embodiments, the methods and systems described herein are performed in conjunction with sequencing RNA molecules isolated from patient biological samples. In some embodiments, a FASTQ file or equivalent file format of sequencing data is the output of such a sequencing reaction.
[0124] In some embodiments, each FASTQ file contains reads, which may be paired-end or single-read, and may be short or long reads, each read representing one detected sequence of nucleotides in an mRNA molecule isolated from a patient sample and inferred by using a sequencing device to detect the sequence of nucleotides contained in a cDNA molecule generated from the mRNA molecule isolated during library preparation. Each read in a FASTQ file is also associated with a quality rating. The quality rating may reflect the likelihood that an error occurred during the sequencing procedure affecting the associated read.
[0125] Each FASTQ file can be processed by a bioinformatics pipeline. In various embodiments, the bioinformatics pipeline can filter the FASTQ data. Filtering the FASTQ data can include correcting sequencing errors and removing (trimming) low-quality sequences or bases, adapter sequences, contamination, chimeric reads, overrepresented sequences, biases caused by library preparation, amplification, or capture, and other errors. Potentially erroneous reads, individual nucleotides, or multiple nucleotides can be discarded based on a quality assessment associated with the read in the FASTQ file, the known error rate of the sequencing instrument, and / or a comparison of each nucleotide in the read with one or more nucleotides in other reads aligned to the same position in the reference genome. Filtering can be performed in part or in whole by various software tools. FASTQ files can be analyzed by sequencing data QC software, such as AfterQC, Kraken, RNA-SeQC, FastQC (Illumina, BaseSpace Labs, or see https: / / www.illumina.com / products / by-type / informatics-products / basespace-sequence-hub / apps / fastqc.html), or another similar software program, for quality control and rapid evaluation of reads. For paired-end reads, reads can be merged.
[0126] For each FASTQ file, each read in the file may be aligned to the position in the reference genome whose sequence most closely matches the sequence of nucleotides in the read. Many software programs designed to align reads are available, including those using the Bowtie, Burrows Wheeler Aligner (BWA), and Smith-Waterman algorithms. Alignment can be directed using a reference genome (e.g., GRCh38, hg38, GRCh37, or other reference genomes developed by the Genome Reference Consortium) by comparing the nucleotide sequence in each read with a portion of the nucleotide sequence in the reference genome to determine the portion of the reference genome sequence that most likely corresponds to the sequence in the read. The alignment may also take into account RNA splice sites. The alignment may generate a SAM file that stores the start and end positions of each read in the reference genome and the coverage (number of reads) of each nucleotide in the reference genome. SAM files can be converted to BAM files, BAM files can be sorted, and duplicate reads can be marked for deletion.
[0127] In one example, kallisto software can be used for alignment and quantification of RNA reads (see Nicolas L Bray, Harold Pimentel, Pall Melsted and Lior Pachter, Near-optimal probabilistic RNA-seq quantification, Nature Biotechnology 34, 525-527 (2016), doi:10.1038 / nbt.3519). In another embodiment, quantification of RNA reads can be performed using other software, such as Sailfish or Salmon (see Rob Patro, Stephen M. Mount, and Carl Kingsford (2014) Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms. Nature Biotechnology (doi:10.1038 / nbt.2862) or Patro, R., Duggal, G., Love, MI, Irizarry, RA, & Kingsford, C. (2017). Salmon provides fast and bias-aware quantification of transcript expression. Nature Methods.). These RNA-seq quantification methods may not require alignment. There are many software packages available for normalization, quantitative analysis, and differential expression analysis of RNA-seq data.
[0128] For each gene, the raw RNA read count for a given gene can be calculated. The raw read counts can be stored in a tabular file for each sample, with each column representing a gene and each entry representing the raw RNA read count for that gene. In one example, the Kallisto alignment software calculates the raw RNA read count for each read as the sum of the probabilities that the read aligns to the gene. Therefore, in this example, the raw count is not an integer.
[0129] The raw RNA read counts can then be normalized, for example, using full quantile normalization, to correct for GC content and gene length, and adjusted for sequencing depth, for example, using the size factor method. In one example, normalization of RNA read counts is performed according to the method disclosed in U.S. Patent Application No. 16 / 581,706 or PCT19 / 52801, entitled "Methods of Normalizing and Correcting RNA Expression Data," filed September 24, 2019, which are incorporated herein by reference in their entirety. The rationale for normalization is that the copy number of each cDNA molecule in the sequencing device may not reflect the distribution of mRNA molecules in patient samples. For example, during library preparation, amplification, and capture steps, certain portions of mRNA molecules may be over- or under-represented due to artifacts arising from various aspects of reverse transcription priming caused by random hexamers, amplification (PCR enrichment), rRNA depletion, and probe binding and errors generated during sequencing, which may be due to GC content, read length, gene length, and other characteristics of the sequence in each nucleic acid molecule. Each raw RNA read count for each gene can be adjusted to eliminate or reduce over- or under-representation caused by biases or artifacts in the NGS sequencing protocol. Normalized RNA read counts can be saved for each sample in a tabular file, with columns representing genes and each entry representing the normalized RNA read count for that gene.
[0130] The transcriptome value set may refer to either normalized RNA read counts or raw RNA read counts, as described above.
[0131] HPV classifier training In one aspect, the present disclosure provides a method for training a classifier to detect human papillomavirus (HPV) infection in cancer. The method includes obtaining abundance values 126, e.g., mRNA expression levels, for genes that are useful for assessing the HPV status of an HPV-associated cancer from a training set of subjects 124 with HPV-associated cancers and known HPV status. The method then includes, for each respective training subject, training a classifier on at least (i) the abundance values 126 and (ii) the HPV status of the patient's cancer, e.g., using a classifier training module 120. In some embodiments, the classifier is also trained on the status of one or more mutant alleles 127 in each training subject's cancer.
[0132] In some embodiments, each training subject has an HPV-related cancer selected from cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, and vulvar cancer. In some embodiments, the classifier is trained on data from patients who all have the same type of cancer, for example, cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, or vulvar cancer. However, since classifier training generally improves by increasing the size of the training data set, in some embodiments, the classifier is trained on data from patients who have two or more types of HPV-related cancer, for example, two, three, four, five, six, seven, or all eight of cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, and vulvar cancer. In a specific embodiment illustrated by Example 3, each training subject has either head and neck squamous cell carcinoma or cervical cancer.
[0133] In some embodiments, the classifier is trained on abundance values for a plurality of genes selected from those listed in Table 3, e.g., KRT86, CRISPLD1, DSG1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, ZFR2, RNF212, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1. As reported below, e.g., see Example 3, these 24 genes were found to be differentially expressed depending on the subject's HPV status in at least 8 of 10 training sets formed from expression data of cervical cancers with known HPV status or head and neck cancers in The Cancer Genome Atlas (TCGA). However, one skilled in the art will understand that in some cases, the use of different training datasets may lead to different results, e.g., one or more of these genes may not be informative in at least 80% of the training folds, and / or one or more genes found to be not informative in at least 80% of the training folds in the study reported in Example 3 may be informative. These differences may arise, for example, when different criteria are used to select the training population, e.g., different inclusion and / or exclusion criteria such as cancer type, personal characteristics (e.g., age, sex, ethnicity, family history, smoking status, etc.), or simply by using smaller or larger datasets.
[0134] Thus, in some embodiments, the classifier is trained on at least 5 of the genes listed in Table 3. In some embodiments, the classifier is trained on at least 10 of the genes listed in Table 3. In some embodiments, the classifier is trained on at least 15 of the genes listed in Table 3. In some embodiments, the classifier is trained on at least 20 of the genes listed in Table 3. In some embodiments, the classifier is trained on all 24 of the genes listed in Table 3. In some embodiments, the classifier is trained on 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or all 24 of the genes listed in Table 3. Additionally, in some embodiments, the classifier is also trained on abundance values for one or more genes not listed in Table 3. In some embodiments, the classifier is also trained on abundance values for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes not listed in Table 3. In some embodiments, the classifier is also trained on abundance values for 1 to 10 genes not listed in Table 3. In some embodiments, the classifier is also trained on abundance values for 1 to 5 genes not listed in Table 3. In other embodiments, the classifier is not also trained on abundance values for any genes not listed in Table 3.
[0135] Furthermore, those skilled in the art will also understand that some features, for example, the abundance values of specific genes, will be more informative than other features in a specific classifier.One measure of the predictive power of each feature in a classifier based on multiple features is the regression coefficient calculated for the feature during model training.Regression coefficients represent the relationship between each feature and the response of the model.Coefficient values represent the average change in response given a one-unit increase in feature value.Therefore, at least for variables of the same type, the magnitude of regression coefficients, for example, absolute value, correlates with the importance of the feature in the model.That is, the larger the magnitude of regression coefficients, the more important the variable is to the model. For example, as reported in Example 3, in a particular support vector machine (SVM) classifier trained on the abundance values of all 24 of the genes listed in Table 3, as well as the variant allele status for the TP53 and CDKN2A genes, only 6 of the 24 genes had regression coefficients of at least 0.5 magnitude - CDKN2A (1.13), SMC1B (1.02), EFNB3 (-0.97), KCNS1 (0.74), CCND1 (-0.65), and RNF212 (0.517).
[0136] Thus, one skilled in the art may select a Feature Set that includes fewer than all of the genes listed in Table 3 based, at least in part, on the importance of each feature in one or more classification models. For example, in some embodiments, one or more genes with lower predictive power in a classification model may be omitted during classifier training. For example, in some embodiments, the features used for training include at least gene expression features listed in Table 5, e.g., CDKN2A, SMC1B, EFNB3, KCNS1, CCND1, and RNF212, with a regression coefficient of at least 0.5. In some embodiments, the features used for training include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.4. In some embodiments, the features used for training include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.3. In some embodiments, the features used for training include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.2. In some embodiments, the features used for training include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.1.
[0137] Similarly, the size of the feature set can be affected by which features are included and / or excluded. For example, in some embodiments, if a particular feature with high predictive power is included in a classification model, fewer total features may be included in the model. For example, in some embodiments, if abundance values for SMC1B, CDKN2A, and EFNB3 are included in the model, abundance values for no more than two of the other genes whose abundance values are used as features in Table 5 need to be included in the model. Thus, in some embodiments, the features used to train the model include abundance values for SMC1B, CDKN2A, and EFNB3, and at least two other genes whose abundance values are used as features in Table 5. In some embodiments, the features used to train the model include abundance values for SMC1B, CDKN2A, and EFNB3, and at least five other genes whose abundance values are used as features in Table 5. In some embodiments, the features used to train the model include abundance values for SMC1B, CDKN2A, and EFNB3, and at least 10 other genes whose abundance values are used as features in Table 5. In some embodiments, the features used to train the model include abundance values for SMC1B, CDKN2A, and EFNB3, and at least 15 other genes whose abundance values are used as features in Table 5.
[0138] Similarly, in some embodiments, when features with high predictive power are excluded from the classification model, more of the other features may be included in the model. For example, in some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 15 of the other genes whose abundance values are used as features in Table 5 are included in the model. In some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 20 of the other genes whose abundance values are used as features in Table 5 are included in the model. In some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 15, 16, 17, 18, 19, 20, or all 21 of the other genes whose abundance values are used as features in Table 5 are included in the model.
[0139] Of course, other metrics are available to assess the importance of a feature in the model, such as the change in standardized regression coefficient and R-squared when the feature was last added to the model.
[0140] When selecting a feature set, those skilled in the art will also consider the degree to which the features are correlated with each other. Correlation is a statistical measure that indicates the degree to which two variables are linearly dependent on each other. Therefore, two correlated features provide redundant information to a predictive model, which can have a negative impact on the classifier. Therefore, there are several reasons to remove correlated features from a model. For example, the more features in a classifier, the more calculations that need to be performed, so removing correlated features can make the algorithm faster. Removing correlated features can also remove harmful biases that arise from correlation from the model. Finally, removing correlated features can make the model more interpretable.
[0141] Thus, one skilled in the art may select a Feature Set that includes fewer than all of the genes listed in Table 3 based, at least in part, on the correlation of each feature in one or more classification models. In some embodiments, the choice of removing one or other correlated feature from a Feature Set is informed by the predictive power of the two features, e.g., their respective regression coefficients. For example, the gene expression values for ENSG00000105278 (CXCL14) and ENSG00000077935 (SMC1B) are highly correlated in the Feature Set listed in Table 3 (correlation = 0.718983175). Thus, in some embodiments, the Feature Set does not include either CXCL14 or SMC1B. In some embodiments, CXCL14, but not SMC1B, is excluded from the Feature Set because, as reported in Table 5, SMC1B has a higher regression coefficient (1.02) than CXCL14 (-0.29) in the SVM model described in Example 3.
[0142] As reported in Table 6, 10 pairs of gene expression features have a correlation of at least 0.6. Thus, in some embodiments, features in at least one pair of features having a correlation of at least 0.6 are excluded from the model. In some embodiments, features in at least two pairs of features having a correlation of at least 0.6 are excluded from the model. In other embodiments, features in at least three, four, five, six, seven, eight, nine, or all ten pairs of features having a correlation of at least 0.6 are excluded from the model. In some embodiments, the excluded features are features in a pair of highly correlated features having a lower regression coefficient than reported in Table 5. For example, with reference to Table 6, the features with the lower regression coefficient in each highly correlated pair (e.g., corresponding to a correlation of at least 0.6) are as follows: Pair 1 = DSG1 Pair 2 = ZFR2 Pair 3 = RNF212 Pair 4 = SYCP2 Pair 5 = ZFR2 Pair 6 = MYO3A Pair 7 = SYCP2 Pair 8 = DSG1 Pair 9 = KCNS1 Pair 10 = ZFR2 Thus, in some embodiments, one or more of DSG1, ZFR2, RNF212, SYCP2, MYO3A, and KCNS1 are excluded from the Feature Set based on them being the least informative features in a pair of highly correlated features.
[0143] However, in some embodiments, this selection process does not allow both features of a highly correlated pair of features to be excluded from the Feature Set, for example, because both genes in at least one of the highly correlated pair of features are the least informative features. Thus, in some embodiments, one or more of SYCP2, MYO3A, and KCNS1 are not excluded from the Feature Set. Similarly, in some embodiments, this selection process does not allow highly informative features, for example, features with a regression coefficient of at least 0.5, to be excluded from the Feature Set. Thus, in some embodiments, one or both of RNF212 and KCNS1 are not excluded from the Feature Set.
[0144] Thus, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0145] Similarly, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, RNF212, MKRN3, MYL1, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0146] Similarly, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, RNF212, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0147] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 50 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 50 data constructs.
[0148] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs.
[0149] In some embodiments, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm, as described above with reference to Figure 2. In some embodiments, the classifier was trained according to the methodology described above with reference to Figure 2.
[0150] EBV classifier training In one aspect, the present disclosure provides a method for training a classifier to detect Epstein-Barr virus (EBV) infection in cancer. The method includes obtaining abundance values 126, e.g., mRNA expression levels, for genes that are useful for assessing the EBV status of an EBV-associated cancer from a training set of subjects 124 with EBV-associated cancer and known EBV status. The method then includes, for each respective training subject, training a classifier on at least (i) the abundance values 126 and (ii) the EBV status of the patient's cancer, e.g., using a classifier training module 120. In some embodiments, the classifier is also trained on the status of one or more mutant alleles 127 in each training subject's cancer.
[0151] In some embodiments, each training subject has an EBV-associated cancer selected from Burkitt lymphoma, sinonasal angiocentric T-cell lymphoma, non-Hodgkin lymphoma, Hodgkin lymphoma, nasopharyngeal carcinoma, and gastric cancer. In some embodiments, the classifier is trained on data from patients who all have the same type of cancer, such as Burkitt lymphoma, sinonasal angiocentric T-cell lymphoma, non-Hodgkin lymphoma, Hodgkin lymphoma, nasopharyngeal carcinoma, or gastric cancer. However, since classifier training generally improves by increasing the size of the training data set, in some embodiments, the classifier is trained on data from patients who have two or more types of EBV-associated cancer, such as two, three, four, five, or all six of Burkitt lymphoma, sinonasal angiocentric T-cell lymphoma, non-Hodgkin lymphoma, Hodgkin lymphoma, nasopharyngeal carcinoma, and gastric cancer. In a specific embodiment illustrated by Example 4, each training subject has gastric cancer.
[0152] In some embodiments, the classifier is trained on the abundance values of a plurality of genes selected from those listed in Table 4, for example, SCNN1A, CDX1, KCNK15, PRKCG, KRT7, NKD2, GPR158, CLDN3, and ZNF683. As reported below, for example, see Example 4, these nine genes were found to be differentially expressed according to the EBV status of subjects in at least 80% of the gastric cancer training sets in The Cancer Genome Atlas (TCGA). However, those skilled in the art will understand that in some cases, the use of different training data sets may lead to different results, for example, one or more of these genes may not be informative in at least 80% of the training folds, and / or one or more genes found to be not informative in at least 80% of the training folds in the study reported in Example 4 may be informative. These differences may arise, for example, when different criteria are used to select the training population, for example, by using different inclusion and / or exclusion criteria such as cancer type, personal characteristics (e.g., age, sex, ethnicity, family history, smoking status, etc.), or simply by using smaller or larger datasets.
[0153] Thus, in some embodiments, the classifier is trained on at least five of the genes listed in Table 4. In some embodiments, the classifier is trained on at least six of the genes listed in Table 4. In some embodiments, the classifier is trained on at least seven of the genes listed in Table 4. In some embodiments, the classifier is trained on at least eight of the genes listed in Table 4. In some embodiments, the classifier is trained on all nine of the genes listed in Table 4. Further, in some embodiments, the classifier is also trained on abundance values for one or more genes not listed in Table 4. In some embodiments, the classifier is also trained on abundance values for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes not listed in Table 4. In some embodiments, the classifier is also trained on abundance values for 1 to 10 genes not listed in Table 4. In some embodiments, the classifier is also trained on abundance values for 1 to 5 genes not listed in Table 4. In other embodiments, the classifier is not also trained on abundance values for any genes not listed in Table 4.
[0154] Furthermore, those skilled in the art will also understand that some features, for example, the abundance values of specific genes, will be more informative than other features in a specific classifier.One measure of the predictive power of each feature in a classifier based on multiple features is the regression coefficient calculated for the feature during model training.Regression coefficients represent the relationship between each feature and the response of the model.Coefficient values represent the average change in response given a one-unit increase in feature value.Therefore, at least for variables of the same type, the magnitude of regression coefficients, for example, absolute value, correlates with the importance of the feature in the model.That is, the larger the magnitude of regression coefficients, the more important the variable is to the model. For example, as reported in Example 4, in a particular support vector machine (SVM) classifier trained on abundance values for all nine of the genes listed in Table 4, as well as the variant allele status for the TP53 and PIK3CA genes, only four of the nine genes had regression coefficients with a magnitude of at least 0.75: SCNN1A (-1.26), KCNK15 (-1.04), KRT7 (-0.94), and CLDN3 (-1.68).
[0155] Thus, one skilled in the art may select a Feature Set that includes fewer than all of the genes listed in Table 4 based, at least in part, on the importance of each feature in one or more classification models. For example, in some embodiments, one or more genes with lower predictive power in a classification model may be omitted during classifier training. For example, in some embodiments, the features used for training include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.75, e.g., SCNN1A (-1.26), KCNK15 (-1.04), KRT7 (-0.94), and CLDN3 (-1.68). In some embodiments, the features used for training include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.6.
[0156] Similarly, the size of the feature set can be affected by which features are included and / or excluded. For example, in some embodiments, if a particular feature with high predictive power is included in a classification model, fewer total features may be included in the model. For example, in some embodiments, if abundance values for SCNN1A, KCNK15, KRT7, and CLDN3 are included in the model, the abundance values of no more than one of the other genes listed in Table 4 need to be included in the model. Thus, in some embodiments, the features used to train the model include abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and at least one other gene listed in Table 4. In some embodiments, the features used to train the model include abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and at least two other genes listed in Table 4. SCNN1A, KCNK15, KRT7, and CLDN3, and at least three other genes listed in Table 4. SCNN1A, KCNK15, KRT7, and CLDN3, and at least four other genes listed in Table 4.
[0157] Similarly, in some embodiments, when a feature with high predictive power is excluded from a classification model, more of other features can be included in the model.For example, in some embodiments, if the abundance value of one or more of SCNN1A, KCNK15, KRT7 and CLDN3 is not included in the model, the abundance value of at least four of the other genes listed in Table 4 is included in the model.In some embodiments, if the abundance value of one or more of SCNN1A, KCNK15, KRT7 and CLDN3 is not included in the model, the abundance value of all five of the other genes listed in Table 4 is included in the model.
[0158] Of course, other metrics are available to assess the importance of a feature in the model, such as the change in standardized regression coefficient and R-squared when the feature was last added to the model.
[0159] When selecting a feature set, one skilled in the art will also consider the degree to which the features are correlated with each other. Correlation is a statistical measure of how linearly dependent two variables are on each other. Therefore, two correlated features provide redundant information to a predictive model, which can adversely affect the classifier. Therefore, there are several reasons for excluding correlated features from a model. For example, the more features in a classifier, the more calculations that need to be performed, so removing correlated features can make the algorithm faster. Removing correlated features can also remove harmful biases resulting from correlation from the model. Finally, removing correlated features can make the model more interpretable. Therefore, one skilled in the art may select a feature set that includes fewer genes than all of those listed in Table 3, based at least in part on the correlation of each feature in one or more classification models. For example, statistical analysis of the SVM model trained in Example 4 revealed that the gene expression values for ENSG00000135480 (KRT7) and ENSG00000124249 (KCNK15) were highly correlated (0.650). Thus, in some embodiments, the abundance values for one of KRT7 and KCNK15 are excluded from the feature set.
[0160] For example, in one embodiment, the feature set includes abundance values for at least SCNN1A, CDX1, KCNK15, PRKCG, NKD2, GPR158, CLDN3, and ZNF683. In another embodiment, the feature set includes abundance values for at least SCNN1A, CDX1, PRKCG, KRT7, NKD2, GPR158, CLDN3, and ZNF683.
[0161] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 50 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 50 data constructs.
[0162] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs.
[0163] In some embodiments, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm, as described above with reference to Figure 2. In some embodiments, the classifier was trained according to the methodology described above with reference to Figure 2.
[0164] Classification method In some embodiments, the present disclosure provides a method for distinguishing a first cancer state and a second cancer state in a human subject, wherein the first cancer state is associated with infection by oncogenic pathogens, and the second cancer state is associated with a condition that does not include oncogenic pathogens.Generally, the method includes obtaining abundance data, such as relative expression levels, for a plurality of genes that are differentially expressed in cancerous tissues associated with oncogenic pathogen infection and the same type of cancerous tissue that is not associated with oncogenic pathogen infection.Then, the abundance data is input into a classifier that is trained to distinguish the first cancer state and the second cancer state based at least in part on the abundance of the genes that are differentially expressed in two types of cancerous tissues.An example of training such a classifier is provided above in conjunction with the description of Figure 2.
[0165] Many of the embodiments described below, in conjunction with Figure 3, relate to analyses performed using expression data derived from a cancer patient's exome, e.g., obtained from a sample of cancerous tissue in the patient. Generally, these embodiments are independent and therefore do not depend on a particular expression data generation method, e.g., sequencing, hybridization, and / or qPCR methodology. However, in some embodiments, the methods described below include one or more steps (301) of generating expression data.
[0166] In some embodiments, the methods include obtaining a sample of cancerous tissue (302). Methods for obtaining a sample of cancerous tissue are known in the art and depend on the type of cancer being sampled. For example, bone marrow biopsy and isolation of circulating tumor cells can be used to obtain samples of blood cancers, endoscopic biopsy can be used to obtain samples of gastrointestinal, bladder, and lung cancers, needle biopsy (e.g., fine needle aspiration, core needle aspiration, vacuum-assisted biopsy, and image-guided biopsy) can be used to obtain samples of subcutaneous tumors, skin biopsy (e.g., shave biopsy, punch biopsy, incisional biopsy, and excision biopsy) can be used to obtain samples of skin cancers, and surgical biopsy can be used to obtain samples of cancers affecting the patient's internal organs.
[0167] In some embodiments, mRNA is then isolated from the cancerous tissue sample (304). Many techniques for isolating RNA from tissue samples are known in the art, such as acid guanidine thiocyanate-phenol-chloroform extraction (see, e.g., Chomczynski and Sacchi, Nat Protoc, 1(2):581-85 (2006), the contents of which are incorporated herein by reference in their entirety for all purposes), and silica bead / glass fiber adsorption (see, e.g., Poeckh, T. et al., Anal Biochem., 373(2):253-62 (2008), the contents of which are incorporated herein by reference in their entirety for all purposes). The selection of any particular RNA isolation technique for use in conjunction with the embodiments described herein is within the skill of one of ordinary skill in the art, taking into account the tissue type, tissue condition (e.g., fresh, frozen, formalin-fixed, paraffin-embedded (FFPE)), and the type of nucleic acid analysis to be performed on the RNA sample.
[0168] In some embodiments, RNA is isolated from blood samples and / or tissue sections (e.g., tumor biopsies) using commercially available reagents, such as proteinase K, TURBO DNase-I, and / or RNA Clean XP beads. In some embodiments, the isolated RNA is subjected to a quality control protocol to determine the concentration and / or quantity of RNA molecules, including the use of fluorescent dyes and a fluorescence microplate reader, a standard spectrofluorometer, or a filter fluorometer.
[0169] In some embodiments, expression data is obtained directly from isolated mRNA, for example, by direct RNA sequencing (314). Methods for direct RNA sequencing are known in the art. See, for example, Ozsolak F., et al., Nature 461:814-18 (2009), and Garalde, DR, et al., Nat Methods, 15(3):201-206 (2018), the contents of which are incorporated herein by reference in their entirety for all purposes.
[0170] In other embodiments, expression data is obtained via a cDNA intermediate. Thus, in some embodiments, isolated RNA is used to create a cDNA library via cDNA synthesis (310). In some embodiments, a cDNA library is prepared from isolated RNA that is purified and selected for cDNA molecule size selection using commercially available reagents, e.g., Roche KAPA Hyper Beads. In another example, a New England Biolabs (NEB) kit can be used.
[0171] In some embodiments, cDNA library preparation involves ligating adapters to cDNA molecules. For example, UDI adapters, such as Roche SeqCap dual-end adapters, or UMI adapters (e.g., full-length or stubby Y adapters) can be ligated to cDNA molecules. Adapters are nucleic acid molecules that can function as barcodes to identify cDNA molecules according to the sample from which they are derived and / or to facilitate downstream bioinformatics processing and / or next-generation sequencing reactions. The sequence of nucleotides in the adapter can be sample-specific to distinguish the sample. Adapters can facilitate binding of cDNA molecules to anchor oligonucleotide molecules on a sequencing instrument flow cell and serve as seeds for the sequencing process by providing a starting point for the sequencing reaction.
[0172] The cDNA library can be amplified and purified using reagents such as Axygen MAG PCR cleanup beads. The concentration and / or amount of cDNA molecules can then be quantified using fluorescent dyes and a fluorescent microplate reader, a standard spectrofluorometer, or a filter fluorometer.
[0173] In some embodiments, for both direct RNA sequencing and prior to cDNA library construction, isolated RNA is first enriched for desired types of RNA (e.g., mRNA) or species (e.g., specific mRNA transcripts) prior to cDNA library construction (308). Methods for enriching for desired RNA molecules are also known in the art. For example, mRNA molecules can be enriched relative to other RNA molecules in a total RNA preparation using, for example, oligo-dT affinity techniques (see, e.g., Rio, DC, et al., Cold Spring Harb Protoc., 2010 Jul 1; 2010(7), the contents of which are incorporated herein by reference in their entirety for all purposes). Specific mRNA transcripts can also be isolated using, for example, hybridization probes that specifically bind to one or more mRNA sequences of interest.
[0174] In some embodiments, the cDNA libraries are pooled and treated with reagents to reduce off-target capture, such as human COT-1 and / or IDT xGen Universal Blockers, before being dried under vacuum. The pools are then resuspended in a hybridization mixture, such as IDT xGen Lockdown, and probes, such as IDT xGen Exome Research Panel v1.0 probes, IDT xGen Exome Research Panel v2.0 probes, other IDT probe panels, Roche probe panels, or other probes, can be added to each pool. The pools can be incubated in an incubator, PCR machine, water bath, or other temperature-controlled device to allow the probes to hybridize. The pools can then be mixed with streptavidin-coated beads or another means for capturing hybridized cDNA probe molecules, particularly cDNA molecules representing exons in the human genome. In another embodiment, polyA capture can be used. The pools can be amplified and purified once again using commercially available reagents, such as the KAPA HiFi Library Amplification Kit and Axygen MAG PCR cleanup beads, respectively.
[0175] The construction of a cDNA library from isolated mRNA is also known in the art.In some embodiments, the construction of a cDNA library is carried out by using reverse transcriptase to synthesize the first strand of DNA from isolated mRNA, followed by using DNA polymerase to synthesize the second strand.Examples of methods for cDNA synthesis are described in McConnell and Watson, 1986, FEBS Lett.195(1-2), pp.199-202; Lin and Ying, 2003, Methods Mol Biol.221, pp.129-143; and Oh et al., 2003, Exp Mol Med.35(6), pp.586-90, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0176] The cDNA library can also be analyzed to determine the fragment size of the cDNA molecules, which can be done via gel electrophoresis techniques and can include using a device such as the LabChip GX Touch. The pool can be cluster amplified using a kit (e.g., Illumina Paired-end Cluster Kits with PhiX spikes). In one example, the cDNA library preparation and / or whole exome capture steps can be performed in an automated system using a liquid-handling robot (e.g., SciClone NGSx).
[0177] Library amplification can be performed on a device such as an Illumina C-Bot2, and the resulting flow cells containing the amplified target capture cDNA libraries can be sequenced on a next-generation sequencer, such as an Illumina HiSeq4000 or Illumina NovaSeq 6000, to a user-selected specific on-target depth, e.g., 300x, 400x, 500x, 10,000x, etc. The next-generation sequencer can generate FASTQ, BCL, or other files for each patient sample or each flow cell.
[0178] When two or more patient samples are processed simultaneously on the same sequencer flow cell, the reads from multiple patient samples are initially contained in the same BCL file and then split into individual FASTQ files for each patient. The differences in the sequence of the adapters used for each patient sample can serve the purpose of barcoding, facilitating the association of each read with the correct patient sample and placement in the correct FASTQ file.
[0179] Methods for mRNA sequencing are known in the art. In some embodiments, mRNA sequencing is performed by whole exome sequencing (WES). Generally, WES is performed by isolating RNA from tissue samples, optionally selecting desired sequences and / or depleting unwanted RNA molecules, generating a cDNA library, and then sequencing the cDNA library (312), for example, using next-generation sequencing (NGS) technology. For a review of the use of whole exome sequencing technology in cancer diagnosis, see Serrati et al., 2016, Onco Targets Ther. 9, pp. 7355-7365, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0180] Next-generation sequencing methods are also known in the art and include sequencing by synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent Sequencing), single molecule real-time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing by synthesis with reversible dye terminators.
[0181] In some embodiments, sequence reads can be aligned to a reference exome or reference genome using methods known in the art to determine alignment position information. The alignment position information can indicate the start and end positions of a region in the reference genome corresponding to the starting and ending nucleotide bases of a given sequence read. The alignment position information can also include the length of the sequence read, which can be determined from the start and end positions. The region in the reference genome can be associated with a gene or a gene segment. Non-limiting examples of known software for assembling and managing transcriptome information from RNA-seq data include TopHat and Cufflinks; see Trapnell et al., 2012, Nat Protoc. 7(3), pp. 562-578, the contents of which are incorporated herein by reference in their entirety for all purposes. See also Hintzsche et al., 2016, Int J Genomics 7983236, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0182] In other embodiments, expression data is generated by hybridization of a cDNA library (313), for example, using a microarray. The use of microarray-based gene profiling to identify differential gene expression after pathogen infection is known in the art. See, for example, Adamas et al., 2008, Tree Physiol. 28(6), pp. 885-897, the contents of which are incorporated herein by reference in their entirety for all purposes. Similarly, in other embodiments, still other methods for quantifying expression based on a cDNA library, such as quantitative real-time PCR (RT-qPCR), are used. See, for example, Wagner, 2013, Methods Mol Biol. 1027, pp. 19-45, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0183] 3, in some embodiments, method 300 is performed, at least in part, on a computer system (e.g., computer system 100 of FIG. 1) having one or more processors and memory storing one or more programs for execution by the one or more processors to identify a first cancerous condition and a second cancerous condition in a subject, where the first cancerous condition is associated with infection by a first oncogenic pathogen and the second cancerous condition is associated with a condition that does not include the oncogenic pathogen. Some operations in method 300 are optionally combined and / or the order of some operations is optionally changed.
[0184] In some embodiments, the method includes obtaining a dataset for a subject, the dataset including a plurality of abundance values, each respective abundance value in the plurality of abundance values quantifying the expression level of a corresponding gene in a plurality of genes in cancerous tissue from the subject. In some embodiments, the obtained abundance values are determined according to any of the methodologies described with respect to submethod 301. In some embodiments, the abundance data is pre-generated and communicated to computer system 100 over a network, for example, using network interface 104. Method 300 then includes inputting the dataset into a classifier (316) trained to identify a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by an oncogenic pathogen and the second cancer condition is associated with a condition not involving an oncogenic pathogen. Examples of such classifiers are provided above in conjunction with FIG. 2. The method thereby determines (320) whether the subject has a first cancer condition associated with oncogenic pathogen infection or a second cancer condition not associated with oncogenic pathogen infection.
[0185] In some embodiments, method 300 also includes inputting mutant allele counts for one or more mutant alleles at one or more loci in the genome of cancerous tissue from the subject into a classifier. That is, in some embodiments, the classifier is also trained on data regarding the presence or absence of one or more mutant alleles in subjects with cancer associated with oncogenic pathogen infection or not associated with oncogenic pathogen infection. In some embodiments, the one or more mutant alleles are selected from mutant alleles in genes selected from the group consisting of TP53 (ENSG00000141510), CDKN2A (ENSG00000147889), and PIK3CA (ENSG00000121879).
[0186] In some embodiments, the subject has breast cancer, lung cancer, prostate cancer, colorectal cancer, renal cancer, uterine cancer, pancreatic cancer, esophageal cancer, head and neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer.
[0187] In some embodiments, the first cancerous condition is associated with infection by a first oncogenic pathogen selected from Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T-cell lymphotropic virus (HTLV-1), Kaposi-associated sarcoma virus (KSHV), and Merkel cell polyomavirus (MCV).
[0188] More specifically, in some embodiments, the first cancer condition is selected from human papillomavirus (HPV)-associated cervical cancer, HPV-associated head and neck cancer, Epstein-Barr virus (EBV)-associated gastric cancer, EBV-associated nasopharyngeal carcinoma, EBV-associated Burkitt lymphoma, EBV-associated Hodgkin lymphoma, hepatitis B virus (HBV)-associated liver cancer, hepatitis C virus (HCV)-associated liver cancer, Kaposi's sarcoma virus (KSHV)-associated Kaposi's sarcoma, human T-cell lymphotropic virus (HTLV-1)-associated adult T-cell leukemia / lymphoma, and Merkel cell polyomavirus (MCV)-associated Merkel cell carcinoma. For a summary of cancer conditions known to be associated with oncogenic viral infections, see de Flora, 2011, "The prevention of infection-associated cancers," Carcinogenesis 32, pp. 787-795.
[0189] Therefore, if the first cancer state is a specific type of cancer associated with a specific oncogenic pathogen, the second cancer state is the same specific type of cancer associated with the absence of infection with a specific oncogenic pathogen.For example, if the first cancer state is cervical cancer associated with human papillomavirus (HPV) infection, the second cancer state is cervical cancer not associated with human papillomavirus (HPV) infection.Furthermore, as described above, the classifier used to distinguish between the two cancer states is trained on a dataset that includes at least gene abundance values (e.g., mRNA expression profiles) from subjects known to have cervical cancer associated with human papillomavirus (HPV) infection and from subjects known to have cervical cancer not associated with human papillomavirus (HPV) infection.
[0190] In some embodiments, the method further comprises treating the subject with either a first therapy (322) tailored to treat a first cancerous condition associated with an oncogenic pathogenic infection, or a second therapy (324) tailored to treat a second cancerous condition not associated with an oncogenic pathogenic infection.
[0191] Thus, in one embodiment, a method for treating cancer in a human cancer patient is provided. The method includes obtaining a dataset for the patient, the dataset including a plurality of abundance values, to determine whether the patient is infected with an oncogenic pathogen linked to the pathology of the cancer, and inputting the dataset into a classifier trained to distinguish at least a first cancer state associated with infection with the oncogenic pathogen and a second cancer state not associated with infection with the oncogenic pathogen. Each abundance value in the dataset quantifies the expression level of a corresponding gene found to be differentially expressed in cancers associated with infection with the oncogenic pathogen and cancers not associated with infection with the oncogenic pathogen. In some embodiments, the genes whose abundance values are used to distinguish the cancer state for any particular type of cancer are selected according to any of the selection methodologies described above with reference to FIG. 2. Similarly, in some embodiments, the classifier used is trained according to any of the training methods described above with reference to FIG. 2.
[0192] In some embodiments, if the subject is determined to have a first cancerous condition associated with an oncogenic pathogen infection, the method includes assigning and / or administering immunotherapy to the subject. In some embodiments, if the subject is determined to have a second cancerous condition not associated with an oncogenic pathogen infection, the method includes assigning and / or administering chemotherapy to the subject.
[0193] Several clinical trials are underway for the treatment of virus-associated tumors, as summarized in Table 2. Accordingly, in some embodiments, the methods described herein involve assigning and / or administering a treatment for a particular cancer associated with a particular oncogenic viral infection, as described in Table 2. For example, in some embodiments, once a subject is determined to have phase 3 cervical cancer associated with HPV infection, the subject is assigned and / or administered a therapeutically effective dosing regimen of axalimogene filolisbac, a live attenuated Listeria monocytogenes transfected with a plasmid encoding the HPV-16 E7 protein fused to a truncated fragment of the Lm protein listeriolysin O. [Table 2-1] [Table 2-2]
[0194] HPV oncogenic virus infection In some embodiments, the methods described herein relate to the classification and / or treatment of cancers known to be associated with human papillomavirus (HPV) infection. As reported in Example 3 below, the 24 genes listed in Table 3 and shown in Figure 4B were found to be differentially expressed in at least 8 of 10 training sets formed from expression data of cervical cancers or head and neck cancers with known HPV status in The Cancer Genome Atlas (TCGA). Thus, in some embodiments, the expression levels of one or more of the genes listed in Table 3 are used to classify cervical cancers or head and neck cancers as either associated with HPV infection or not associated with HPV infection. In some embodiments, the expression levels of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or all 24 of the genes listed in Table 3 are used to classify cervical cancer or head and neck cancer as either associated with HPV infection or not associated with HPV infection. [Table 3]
[0195] In one embodiment, a method is provided for identifying a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by a human papillomavirus (HPV) oncogenic virus and the second cancer condition is associated with a condition that does not involve HPV. The method includes obtaining a dataset for the subject, e.g., as described above with reference to FIG. 3. The dataset includes a plurality of abundance values from the subject, wherein each respective abundance value in the plurality of abundance values quantifies the expression level of a corresponding gene in the plurality of genes in cancerous tissue from the subject. In some embodiments, the plurality of genes includes at least five genes selected from the genes listed in Table 3. The method then includes inputting the dataset into a classifier trained to identify at least the first cancer condition and the second cancer condition based on the abundance values of the plurality of genes. In some embodiments, the classifier is trained according to any of the methodologies described above with reference to FIG. 2.
[0196] In some embodiments, the first cancer condition is cervical cancer associated with HPV infection, and the second cancer condition is cervical cancer not associated with HPV infection. In some embodiments, the first cancer condition is head and neck cancer associated with HPV infection, and the second cancer condition is head and neck cancer not associated with HPV infection. In some embodiments, the head and neck cancer is a specific form or head and neck cancer, such as hypopharyngeal cancer, laryngeal cancer, lip and oral cavity cancer, metastatic squamous cell carcinoma with occult primary, nasopharyngeal cancer, oropharynx cancer, paranasal sinus and nasal cavity cancer, or salivary gland cancer.
[0197] In some embodiments, the plurality of genes includes at least 10 of the genes listed in Table 3. In some embodiments, the plurality of genes includes at least 15 of the genes listed in Table 3. In some embodiments, the plurality of genes includes at least 20 of the genes listed in Table 3. In some embodiments, the plurality of genes includes all of the genes listed in Table 3. In some embodiments, the plurality of genes includes one or more genes not listed in Table 3, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 3. In some embodiments, the plurality of genes includes 20 or fewer genes. In some embodiments, the plurality of genes includes 25 or fewer genes. In some embodiments, the plurality of genes includes 50 or fewer genes. In some embodiments, the plurality of genes includes no more than 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 genes.
[0198] In some embodiments, the dataset also includes a mutant allele count for one or more alleles at one or more loci in the genome of the cancerous tissue from the subject. In some embodiments, the mutant allele count is either 1, representing that the subject has the mutant allele, or 0, representing that the subject does not have the mutant allele. In some embodiments, the mutant allele is a somatic mutation derived from the subject's germline. In some embodiments, the mutant allele is a mutation derived from a cancer derived from the cancerous tissue. In some embodiments, the mutant allele is located in the TP53 (ENSG00000141510) or CDKN2A (ENSG00000147889) gene.
[0199] In some embodiments, the classifier is trained to determine the HPV status of subjects with HPV-related cancers selected from cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, and vulvar cancer. In some embodiments, the classifier is trained to determine the HPV status of test patients with specific HPV-related cancers, such as cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, or vulvar cancer. However, since classifier training generally improves by increasing the size of the training data set, in some embodiments, the classifier is trained on data from patients with two or more types of HPV-related cancers, such as two, three, four, five, six, seven, or all eight of cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, and vulvar cancer. In a specific embodiment illustrated by Example 3, the classifier is trained on subjects with either head and neck squamous cell carcinoma or cervical cancer. However, in some embodiments, a classifier trained on patients with one or more types of HPV-associated cancer is useful for determining the HPV status of patients with different types of HPV-associated cancer.
[0200] In some embodiments, the classifier features comprise abundance values for a plurality of genes selected from those set forth in Table 3, e.g., KRT86, CRISPLD1, DSG1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, ZFR2, RNF212, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1. As reported below, e.g., see Example 3, these 24 genes were found to be differentially expressed depending on the subject's HPV status in at least 8 of 10 training sets formed from expression data of cervical cancers with known HPV status or head and neck cancers in The Cancer Genome Atlas (TCGA). However, one skilled in the art will understand that in some cases, the use of different training datasets may yield different results, e.g., one or more of these genes may not be informative in at least 80% of the training folds, and / or one or more genes found to be not informative in at least 80% of the training folds in the study reported in Example 3 may be informative. These differences may arise, for example, when different criteria are used to select the training population, e.g., different inclusion and / or exclusion criteria such as cancer type, personal characteristics (e.g., age, sex, ethnicity, family history, smoking status, etc.), or simply by using smaller or larger datasets.
[0201] Thus, in some embodiments, the classifier feature comprises at least 5 of the genes listed in Table 3. In some embodiments, the classifier feature comprises at least 10 of the genes listed in Table 3. In some embodiments, the classifier feature comprises at least 15 of the genes listed in Table 3. In some embodiments, the classifier feature comprises at least 20 of the genes listed in Table 3. In some embodiments, the classifier feature comprises all 24 of the genes listed in Table 3. In some embodiments, the classifier feature comprises 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or all 24 of the genes listed in Table 3. Additionally, in some embodiments, the classifier feature comprises abundance values for one or more genes not listed in Table 3. In some embodiments, a classifier feature comprises abundance values for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes not listed in Table 3. In some embodiments, a classifier feature comprises abundance values for 1 to 10 genes not listed in Table 3. In some embodiments, a classifier feature comprises abundance values for 1 to 5 genes not listed in Table 3. In other embodiments, a classifier feature does not include abundance values for any genes not listed in Table 3.
[0202] Furthermore, those skilled in the art will also understand that some features, for example, the abundance values of specific genes, will be more informative than other features in a specific classifier.One measure of the predictive power of each feature in a classifier based on multiple features is the regression coefficient calculated for the feature during model training.Regression coefficients represent the relationship between each feature and the response of the model.Coefficient values represent the average change in response given a one-unit increase in feature value.Therefore, at least for variables of the same type, the magnitude of regression coefficients, for example, absolute value, correlates with the importance of the feature in the model.That is, the larger the magnitude of regression coefficients, the more important the variable is to the model. For example, as reported in Example 3, in a particular support vector machine (SVM) classifier trained on the abundance values of all 24 of the genes listed in Table 3, as well as the variant allele status for the TP53 and CDKN2A genes, only 6 of the 24 genes had regression coefficients of at least 0.5 magnitude - CDKN2A (1.13), SMC1B (1.02), EFNB3 (-0.97), KCNS1 (0.74), CCND1 (-0.65), and RNF212 (0.517).
[0203] Thus, one skilled in the art may select a Feature Set that includes fewer than all of the genes listed in Table 3 based, at least in part, on the importance of each feature in one or more classification models. For example, in some embodiments, one or more genes with lower predictive power in a classification model may be omitted during classifier training. For example, in some embodiments, the classifier features include at least gene expression features listed in Table 5, e.g., CDKN2A, SMC1B, EFNB3, KCNS1, CCND1, and RNF212, with a regression coefficient of at least 0.5. In some embodiments, the classifier features include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.4. In some embodiments, the classifier features include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.3. In some embodiments, the classifier features include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.2. In some embodiments, the classifier features include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.1.
[0204] Similarly, the size of the feature set can be affected by which features are included and / or excluded. For example, in some embodiments, if a particular feature with high predictive power is included in a classification model, fewer total features may be included in the model. For example, in some embodiments, if abundance values for SMC1B, CDKN2A, and EFNB3 are included in the model, abundance values for no more than two of the other genes whose abundance values are used as features in Table 5 need to be included in the model. Thus, in some embodiments, the classifier features include abundance values for SMC1B, CDKN2A, and EFNB3, and at least two other genes whose abundance values are used as features in Table 5. In some embodiments, the classifier features include abundance values for SMC1B, CDKN2A, and EFNB3, and at least five other genes whose abundance values are used as features in Table 5. In some embodiments, the classifier features include abundance values for SMC1B, CDKN2A, and EFNB3, and at least 10 other genes whose abundance values are used as features in Table 5. In some embodiments, the classifier features include abundance values for SMC1B, CDKN2A, and EFNB3, and at least 15 other genes whose abundance values are used as features in Table 5.
[0205] Similarly, in some embodiments, when features with high predictive power are excluded from the classification model, more of the other features may be included in the model. For example, in some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 15 of the other genes whose abundance values are used as features in Table 5 are included in the model. In some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 20 of the other genes whose abundance values are used as features in Table 5 are included in the model. In some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 15, 16, 17, 18, 19, 20, or all 21 of the other genes whose abundance values are used as features in Table 5 are included in the model.
[0206] Of course, other metrics are available to assess the importance of a feature in the model, such as the change in standardized regression coefficient and R-squared when the feature was last added to the model.
[0207] When selecting a feature set, those skilled in the art will also consider the degree to which the features are correlated with each other. Correlation is a statistical measure that indicates the degree to which two variables are linearly dependent on each other. Therefore, two correlated features provide redundant information to a predictive model, which can have a negative impact on the classifier. Therefore, there are several reasons to remove correlated features from a model. For example, the more features in a classifier, the more calculations that need to be performed, so removing correlated features can make the algorithm faster. Removing correlated features can also remove harmful biases that arise from correlation from the model. Finally, removing correlated features can make the model more interpretable.
[0208] Thus, one skilled in the art may select a Feature Set that includes fewer than all of the genes listed in Table 3 based, at least in part, on the correlation of each feature in one or more classification models. In some embodiments, the choice of removing one or other correlated feature from a Feature Set is informed by the predictive power of the two features, e.g., their respective regression coefficients. For example, the gene expression values for ENSG00000105278 (CXCL14) and ENSG00000077935 (SMC1B) are highly correlated in the Feature Set listed in Table 3 (correlation = 0.718983175). Thus, in some embodiments, the Feature Set does not include either CXCL14 or SMC1B. In some embodiments, CXCL14, but not SMC1B, is excluded from the Feature Set because, as reported in Table 5, SMC1B has a higher regression coefficient (1.02) than CXCL14 (-0.29) in the SVM model described in Example 3.
[0209] As reported in Table 6, 10 pairs of gene expression features have a correlation of at least 0.6. Thus, in some embodiments, features in at least one pair of features having a correlation of at least 0.6 are excluded from the model. In some embodiments, features in at least two pairs of features having a correlation of at least 0.6 are excluded from the model. In other embodiments, features in at least three, four, five, six, seven, eight, nine, or all ten pairs of features having a correlation of at least 0.6 are excluded from the model. In some embodiments, the excluded features are features in a pair of highly correlated features having a lower regression coefficient than reported in Table 5. For example, with reference to Table 6, the features with the lower regression coefficient in each highly correlated pair (e.g., corresponding to a correlation of at least 0.6) are as follows: Pair 1 = DSG1 Pair 2 = ZFR2 Pair 3 = RNF212 Pair 4 = SYCP2 Pair 5 = ZFR2 Pair 6 = MYO3A Pair 7 = SYCP2 Pair 8 = DSG1 Pair 9 = KCNS1 Pair 10 = ZFR2 Thus, in some embodiments, one or more of DSG1, ZFR2, RNF212, SYCP2, MYO3A, and KCNS1 are excluded from the Feature Set based on them being the least informative features in a pair of highly correlated features.
[0210] However, in some embodiments, this selection process does not allow both features of a highly correlated pair of features to be excluded from the Feature Set, for example, because both genes in at least one of the highly correlated pair of features are the least informative features. Thus, in some embodiments, one or more of SYCP2, MYO3A, and KCNS1 are not excluded from the Feature Set. Similarly, in some embodiments, this selection process does not allow highly informative features, for example, features with a regression coefficient of at least 0.5, to be excluded from the Feature Set. Thus, in some embodiments, one or both of RNF212 and KCNS1 are not excluded from the Feature Set.
[0211] Thus, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0212] Similarly, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, RNF212, MKRN3, MYL1, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0213] Similarly, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, RNF212, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0214] In some embodiments, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm, as described above with reference to Figure 2. In some embodiments, the classifier was trained according to the methodology described above with reference to Figure 2.
[0215] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 50 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 50 data constructs.
[0216] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs.
[0217] In some embodiments, the method further includes assigning and / or administering a therapy to the subject based on the classification of the cancer status, e.g., based on whether the subject's cancer is associated with HPV viral infection.
[0218] Thus, in one embodiment, a method for treating cervical cancer in a human cancer patient is provided. The method includes determining whether the human cancer patient is infected with a human papillomavirus (HPV) oncogenic virus by obtaining a dataset about the human cancer patient, the dataset including a plurality of abundance values, each of which quantifies the expression level of a corresponding gene in a plurality of genes, the plurality of genes including at least five genes selected from the genes listed in Table 3. The method then includes inputting the dataset into a classifier trained to identify, in the subject's cancerous tissue, at least a first cervical cancer state associated with HPV infection and a second cervical cancer state associated with an HPV-free state based on the abundance values of the plurality of genes. In some embodiments, the classifier is trained according to the methodology described above with reference to FIG. 2. The method then includes treating cervical cancer. If the classifier results indicate that the human cancer patient is infected with an HPV oncogenic virus, a first therapy tailored to treat cervical cancer associated with HPV infection is administered. If the results of the classifier indicate that the human cancer patient is not infected with the HPV oncogenic virus, a second therapy tailored for the treatment of cervical cancer not associated with HPV infection is administered.
[0219] In some embodiments, the plurality of genes includes at least 10 of the genes listed in Table 3. In some embodiments, the plurality of genes includes at least 15 of the genes listed in Table 3. In some embodiments, the plurality of genes includes at least 20 of the genes listed in Table 3. In some embodiments, the plurality of genes includes all of the genes listed in Table 3. In some embodiments, the plurality of genes includes one or more genes not listed in Table 3, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 3. In some embodiments, the plurality of genes includes 20 or fewer genes. In some embodiments, the plurality of genes includes 25 or fewer genes. In some embodiments, the plurality of genes includes 50 or fewer genes. In some embodiments, the plurality of genes includes no more than 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 genes.
[0220] In some embodiments, the dataset also includes a mutant allele count for one or more alleles at one or more loci in the genome of the cancerous tissue from the subject. In some embodiments, the mutant allele count is either 1, representing that the subject has the mutant allele, or 0, representing that the subject does not have the mutant allele. In some embodiments, the mutant allele is a somatic mutation derived from the subject's germline. In some embodiments, the mutant allele is a mutation derived from a cancer derived from the cancerous tissue. In some embodiments, the mutant allele is located in the TP53 (ENSG00000141510) or CDKN2A (ENSG00000147889) gene.
[0221] In some embodiments, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm, as described above with reference to Figure 2. In some embodiments, the classifier was trained according to the methodology described above with reference to Figure 2.
[0222] In some embodiments, the first therapy tailored to treat cervical cancer associated with HPV infection is a therapeutic vaccine. In some embodiments, the therapeutic vaccine is selected from axalimogene filolisbac (Advaxis), TG4001 (Transgene), GX-188E (Genexine), VGX-3100 (Inovio), MEDI-0457 (Inovio), INO-3106 (Inovio), TA-CIN (Cancer Research Technology), TA-HPV (Cancer Research Technology), ISA-101 (Isa), and PepCan (University of Arkansas).
[0223] In some embodiments, the first therapy tailored to treat cervical cancer associated with HPV infection is adoptive cell therapy. In some embodiments, the adoptive cell therapy comprises the administration of HPV-specific T cells, for example, as described in clinical trial ID NCT02379520 or NCT03197025 (Baylor College of Medicine).
[0224] In some embodiments, the first therapy tailored for the treatment of cervical cancer associated with HPV infection is an immune checkpoint inhibitor, hi some embodiments, the immune checkpoint inhibitor is nivolumab (Bristol-Myers Squibb).
[0225] In some embodiments, the first therapy tailored to treat cervical cancer associated with HPV infection is a PI3K inhibitor. In some embodiments, the PI3K inhibitor is AMG319 (Amgen) or BKM120 (Novartis).
[0226] Similarly, in one embodiment, a method for treating head and neck cancer in a human cancer patient is provided. The method includes determining whether the human cancer patient is infected with a human papillomavirus (HPV) oncogenic virus by obtaining a dataset about the human cancer patient, the dataset including a plurality of abundance values, each of which quantifies the expression level of a corresponding gene in a plurality of genes, the plurality of genes including at least five genes selected from the genes listed in Table 3. The method then includes inputting the dataset into a classifier trained to identify, in the subject's cancerous tissue, at least a first head and neck cancer condition associated with HPV infection and a second head and neck cancer condition associated with an HPV-free condition based on the abundance values of the plurality of genes. In some embodiments, the classifier is trained according to the methodology described above with reference to FIG. 2. The method then includes treating the head and neck cancer. If the classifier result indicates that the human cancer patient is infected with an HPV oncogenic virus, the method includes administering a first therapy tailored to treat head and neck cancer associated with HPV infection. If the results of the classifier indicate that the human cancer patient is not infected with the HPV oncogenic virus, the method includes administering a second therapy tailored for the treatment of head and neck cancer not associated with HPV infection.
[0227] In some embodiments, the plurality of genes includes at least 10 of the genes listed in Table 3. In some embodiments, the plurality of genes includes at least 15 of the genes listed in Table 3. In some embodiments, the plurality of genes includes at least 20 of the genes listed in Table 3. In some embodiments, the plurality of genes includes all of the genes listed in Table 3. In some embodiments, the plurality of genes includes one or more genes not listed in Table 3, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 3. In some embodiments, the plurality of genes includes 20 or fewer genes. In some embodiments, the plurality of genes includes 25 or fewer genes. In some embodiments, the plurality of genes includes 50 or fewer genes. In some embodiments, the plurality of genes includes no more than 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 genes.
[0228] In some embodiments, the dataset also includes a mutant allele count for one or more alleles at one or more loci in the genome of the cancerous tissue from the subject. In some embodiments, the mutant allele count is either 1, representing that the subject has the mutant allele, or 0, representing that the subject does not have the mutant allele. In some embodiments, the mutant allele is a somatic mutation derived from the subject's germline. In some embodiments, the mutant allele is a mutation derived from a cancer derived from the cancerous tissue. In some embodiments, the mutant allele is located in the TP53 (ENSG00000141510) or CDKN2A (ENSG00000147889) gene.
[0229] In some embodiments, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm, as described above with reference to Figure 2. In some embodiments, the classifier was trained according to the methodology described above with reference to Figure 2.
[0230] In some embodiments, the first therapy tailored to treat head and neck cancer associated with HPV infection is a therapeutic vaccine. In some embodiments, the therapeutic vaccine is selected from axalimogene filolisbac (Advaxis), TG4001 (Transgene), GX-188E (Genexine), VGX-3100 (Inovio), MEDI-0457 (Inovio), INO-3106 (Inovio), TA-CIN (Cancer Research Technology), TA-HPV (Cancer Research Technology), ISA-101 (Isa), and PepCan (University of Arkansas).
[0231] In some embodiments, the first therapy tailored to treat head and neck cancer associated with HPV infection is adoptive cell therapy. In some embodiments, the adoptive cell therapy comprises the administration of HPV-specific T cells, for example, as described in clinical trial ID NCT02379520 or NCT03197025 (Baylor College of Medicine).
[0232] In some embodiments, the first therapy tailored for the treatment of head and neck cancer associated with HPV infection is an immune checkpoint inhibitor, hi some embodiments, the immune checkpoint inhibitor is nivolumab (Bristol-Myers Squibb).
[0233] In some embodiments, the first therapy tailored for the treatment of head and neck cancer associated with HPV infection is a PI3K inhibitor. In some embodiments, the PI3K inhibitor is AMG319 (Amgen) or BKM120 (Novartis).
[0234] HPV Probe Set In some embodiments, the present disclosure provides probes for binding, enriching, and / or detecting nucleic acid molecules, such as mRNA transcripts isolated from cancerous tissue samples from subjects and / or cDNA molecules prepared from those mRNA transcripts, which are useful for determining whether the subject has a first cancerous condition associated with HPV oncogenic virus infection or a second cancerous condition not associated with HPV oncogenic virus infection. Generally, the probe comprises a DNA, RNA, or modified nucleic acid structure having a base sequence complementary to a nucleic acid molecule of interest. Thus, when a probe is designed to hybridize to mRNA molecules isolated from cancerous tissue, the probe will contain a nucleic acid sequence complementary to the coding strand of the gene from which the transcript is derived, i.e., the probe will contain an antisense sequence of the gene. However, when a probe is designed to hybridize to cDNA molecules, because the molecules in a cDNA library are double-stranded, the probe may contain either a sequence complementary to the coding sequence of the gene of interest (antisense sequence) or a sequence identical to the coding sequence of the gene of interest (sense sequence).
[0235] In some embodiments, the probe contains an additional nucleic acid sequence that does not share any homology with the gene sequence of interest. For example, in some embodiments, the probe also contains an identifier sequence, e.g., a nucleic acid sequence containing a unique molecular identifier (UMI), that is unique to a particular cancerous tissue sample or cancer patient. Examples of identifier sequences are described, for example, in Kivioja et al., 2011, Nat. Methods 9(1), pp. 72-74, and Islam et al., 2014, Nat. Methods 11(2), pp. 163-66, the contents of which are incorporated herein by reference in their entirety for all purposes. Similarly, in some embodiments, the probe also contains a primer nucleic acid sequence useful for amplifying the nucleic acid molecule of interest, e.g., using PCR. In some embodiments, the probe also contains a capture sequence designed to hybridize to an anti-capture sequence to recover the nucleic acid molecule of interest from the sample.
[0236] Similarly, in some embodiments, the probe comprises a non-nucleic acid affinity moiety covalently bound to a nucleic acid molecule complementary to a gene of interest for retrieving the nucleic acid molecule of interest. Non-limiting examples of non-nucleic acid affinity moieties include biotin, digoxigenin, and dinitrophenol. In some embodiments, the probe is attached to a solid surface or particle, such as a dipstick or magnetic bead, for retrieving the nucleic acid molecule of interest.
[0237] Thus, in one embodiment, the present disclosure provides a plurality of nucleic acid probes for distinguishing between a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by a human papillomavirus (HPV) oncogenic virus and the second cancer condition is associated with a condition that does not include HPV. The plurality of nucleic acid probes comprises at least five nucleic acid probes, each of the at least five nucleic acid probes comprising a respective nucleic acid sequence identical to or complementary to at least 10 contiguous bases of an RNA transcript of a different respective gene selected from the genes set forth in Table 3.
[0238] In some embodiments, the plurality of nucleic acid probes comprises at least 10 probes having sequences complementary to or identical to sequences from different genes listed in Table 3. In some embodiments, the plurality of nucleic acid probes comprises at least 15 probes having sequences complementary to or identical to sequences from different genes listed in Table 3. In some embodiments, the plurality of nucleic acid probes comprises at least 20 probes having sequences complementary to or identical to sequences from different genes listed in Table 3. In some embodiments, the plurality of nucleic acid probes comprises probes having sequences complementary to or identical to sequences from all of the genes listed in Table 3. In some embodiments, the plurality of nucleic acid probes comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 probes having sequences complementary to or identical to sequences from different genes listed in Table 3.
[0239] In some embodiments, the plurality of nucleic acid probes includes one or more probes that bind to sequences of genes not listed in Table 3. In some embodiments, the plurality of nucleic acid probes includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more probes that bind to sequences of genes not listed in Table 3. In some embodiments, the plurality of nucleic acid probes includes probes having sequences that bind to 20 or fewer genes. In some embodiments, the plurality of nucleic acid probes includes probes having sequences that bind to 25 or fewer genes. In some embodiments, the plurality of nucleic acid probes includes probes having sequences that bind to 50 or fewer genes. In some embodiments, the plurality of nucleic acid probes includes probes having sequences that bind to 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0240] In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical to or complementary to at least 15 contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 3. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical to or complementary to at least 30 contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 3. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical to or complementary to at least 50 contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 3. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical to or complementary to at least 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, or more contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 3.
[0241] EBV oncogenic virus infection In some embodiments, the methods described herein relate to the classification and / or treatment of cancers known to be associated with Epstein-Barr virus (EBV) infection. As reported in Example 4 below, the 24 genes listed in Table 4 and shown in Figure 5B were found to be differentially expressed in at least 8 of 10 training sets formed from expression data of gastric cancers with known EBV status in The Cancer Genome Atlas (TCGA). Thus, in some embodiments, the expression levels of one or more of the genes listed in Table 4 are used to classify gastric cancer as either associated with EBV infection or not associated with EBV infection. In some embodiments, the expression levels of at least two, three, four, five, six, seven, eight, or all nine of the genes listed in Table 4 are used to classify gastric cancer as either associated with EBV infection or not associated with EBV infection. [Table 4]
[0242] In one embodiment, a method is provided for identifying a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by the Epstein-Barr virus (EBV) oncogenic virus and the second cancer condition is associated with a condition that does not involve EBV. The method includes obtaining a dataset for the subject, for example, as described above with reference to FIG. 3. The dataset includes a plurality of abundance values from the subject, where each respective abundance value in the plurality of abundance values quantifies the expression level of a corresponding gene in the plurality of genes in cancerous tissue from the subject. In some embodiments, the plurality of genes includes at least five genes selected from the genes listed in Table 4. The method then includes inputting the dataset into a classifier trained to identify at least the first cancer condition and the second cancer condition based on the abundance values of the plurality of genes. In some embodiments, the classifier is trained according to any of the methodologies described above with reference to FIG. 2.
[0243] In some embodiments, the plurality of genes includes all of the genes listed in Table 4. In some embodiments, the plurality of genes includes one or more genes not listed in Table 4, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 4. In some embodiments, the plurality of genes includes 20 or fewer genes. In some embodiments, the plurality of genes includes 25 or fewer genes. In some embodiments, the plurality of genes includes 50 or fewer genes. In some embodiments, the plurality of genes includes 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0244] In some embodiments, the dataset also includes a mutant allele count for one or more alleles at one or more loci in the genome of the cancerous tissue from the subject. In some embodiments, the mutant allele count is either 1, representing that the subject has the mutant allele, or 0, representing that the subject does not have the mutant allele. In some embodiments, the mutant allele is a somatic mutation derived from the subject's germline. In some embodiments, the mutant allele is a mutation derived from a cancer derived from the cancerous tissue. In some embodiments, the mutant allele is located in the TP53 (ENSG00000141510) or PIK3CA (ENSG00000121879) gene.
[0245] In some embodiments, the classifier is trained to determine the EBV status of test subjects with an EBV-associated cancer selected from Burkitt lymphoma, sinonasal angiocentric T-cell lymphoma, non-Hodgkin's lymphoma, Hodgkin's lymphoma, nasopharyngeal carcinoma, and gastric cancer. In some embodiments, the classifier is trained to determine the EBV status of test patients with a specific EBV-associated cancer, such as Burkitt lymphoma, sinonasal angiocentric T-cell lymphoma, non-Hodgkin's lymphoma, Hodgkin's lymphoma, nasopharyngeal carcinoma, or gastric cancer. However, because classifier training generally improves with increasing the size of the training dataset, in some embodiments, the classifier is trained on data from patients with two or more types of EBV-associated cancer, such as two, three, four, five, or all six of Burkitt lymphoma, sinonasal angiocentric T-cell lymphoma, non-Hodgkin's lymphoma, Hodgkin's lymphoma, nasopharyngeal carcinoma, and gastric cancer. In a particular embodiment exemplified by Example 4, the classifier is trained on patients with gastric cancer. However, in some embodiments, a classifier trained on patients with one or more types of EBV-associated cancer is useful for determining the EBV status of patients with different types of EBV-associated cancer.
[0246] In some embodiments, the classifier features include abundance values for a plurality of genes selected from those listed in Table 4, such as SCNN1A, CDX1, KCNK15, PRKCG, KRT7, NKD2, GPR158, CLDN3, and ZNF683. As reported below, for example, see Example 4, these nine genes were found to be differentially expressed depending on the subject's EBV status in at least 80% of gastric cancer training sets in The Cancer Genome Atlas (TCGA). However, those skilled in the art will understand that in some cases, the use of different training datasets may lead to different results, for example, one or more of these genes may not be informative in at least 80% of training folds, and / or one or more genes found to be not informative in at least 80% of training folds in the study reported in Example 4 may be informative. These differences may arise, for example, when different criteria are used to select the training population, for example, by using different inclusion and / or exclusion criteria such as cancer type, personal characteristics (e.g., age, sex, ethnicity, family history, smoking status, etc.), or simply by using smaller or larger datasets.
[0247] Thus, in some embodiments, the classifier feature includes at least five of the genes listed in Table 4. In some embodiments, the classifier feature includes at least six of the genes listed in Table 4. In some embodiments, the classifier feature includes at least seven of the genes listed in Table 4. In some embodiments, the classifier feature includes at least eight of the genes listed in Table 4. In some embodiments, the classifier feature includes all nine of the genes listed in Table 4. Additionally, in some embodiments, the classifier feature also includes abundance values for one or more genes not listed in Table 4. In some embodiments, the classifier feature includes abundance values for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes not listed in Table 4. In some embodiments, the classifier feature includes abundance values for 1 to 10 genes not listed in Table 4. In some embodiments, the classifier feature includes abundance values for 1 to 5 genes not listed in Table 4. In other embodiments, the classifier feature does not include abundance values for any genes not listed in Table 4.
[0248] Furthermore, those skilled in the art will also understand that some features, for example, the abundance values of specific genes, will be more informative than other features in a specific classifier.One measure of the predictive power of each feature in a classifier based on multiple features is the regression coefficient calculated for the feature during model training.Regression coefficients represent the relationship between each feature and the response of the model.Coefficient values represent the average change in response given a one-unit increase in feature value.Therefore, at least for variables of the same type, the magnitude of regression coefficients, for example, absolute value, correlates with the importance of the feature in the model.That is, the larger the magnitude of regression coefficients, the more important the variable is to the model. For example, as reported in Example 4, in a particular support vector machine (SVM) classifier trained on abundance values for all nine of the genes listed in Table 4, as well as the variant allele status for the TP53 and PIK3CA genes, only four of the nine genes had regression coefficients with a magnitude of at least 0.75: SCNN1A (-1.26), KCNK15 (-1.04), KRT7 (-0.94), and CLDN3 (-1.68).
[0249] Thus, one skilled in the art may select a Feature Set that includes fewer than all of the genes listed in Table 4 based, at least in part, on the importance of each feature in one or more classification models. For example, in some embodiments, one or more genes with lower predictive power in a classification model may be omitted during classifier training. For example, in some embodiments, the classifier features include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.75, e.g., SCNN1A (-1.26), KCNK15 (-1.04), KRT7 (-0.94), and CLDN3 (-1.68). In some embodiments, the classifier features include at least gene expression features listed in Table 5 with a regression coefficient of at least 0.6.
[0250] Similarly, the size of the feature set can be affected by which features are included and / or excluded. For example, in some embodiments, if a particular feature with high predictive power is included in a classification model, fewer total features may be included in the model. For example, in some embodiments, if abundance values for SCNN1A, KCNK15, KRT7, and CLDN3 are included in the model, the abundance values of no more than one of the other genes listed in Table 4 need to be included in the model. Thus, in some embodiments, the classifier features include abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and at least one other gene listed in Table 4. In some embodiments, the classifier features include abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and at least two other genes listed in Table 4. In some embodiments, the classifier features include abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and at least three other genes listed in Table 4. In some embodiments, the classifier features include abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and at least four other genes listed in Table 4.
[0251] Similarly, in some embodiments, when a feature with high predictive power is excluded from a classification model, more of other features can be included in the model.For example, in some embodiments, if the abundance value of one or more of SCNN1A, KCNK15, KRT7 and CLDN3 is not included in the model, the abundance value of at least four of the other genes listed in Table 4 is included in the model.In some embodiments, if the abundance value of one or more of SCNN1A, KCNK15, KRT7 and CLDN3 is not included in the model, the abundance value of all five of the other genes listed in Table 4 is included in the model.
[0252] Of course, other metrics are available to assess the importance of a feature in the model, such as the change in standardized regression coefficient and R-squared when the feature was last added to the model.
[0253] When selecting a feature set, one skilled in the art will also consider the degree to which the features are correlated with each other. Correlation is a statistical measure of how linearly dependent two variables are on each other. Therefore, two correlated features provide redundant information to a predictive model, which can adversely affect the classifier. Therefore, there are several reasons for excluding correlated features from a model. For example, the more features in a classifier, the more calculations that need to be performed, so removing correlated features can make the algorithm faster. Removing correlated features can also remove harmful biases resulting from correlation from the model. Finally, removing correlated features can make the model more interpretable. Therefore, one skilled in the art may select a feature set that includes fewer genes than all of those listed in Table 3, based at least in part on the correlation of each feature in one or more classification models. For example, statistical analysis of the SVM model trained in Example 4 revealed that the gene expression values for ENSG00000135480 (KRT7) and ENSG00000124249 (KCNK15) were highly correlated (0.650). Thus, in some embodiments, the abundance values for one of KRT7 and KCNK15 are excluded from the feature set.
[0254] For example, in one embodiment, the feature set includes abundance values for at least SCNN1A, CDX1, KCNK15, PRKCG, NKD2, GPR158, CLDN3, and ZNF683. In another embodiment, the feature set includes abundance values for at least SCNN1A, CDX1, PRKCG, KRT7, NKD2, GPR158, CLDN3, and ZNF683.
[0255] In some embodiments, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm, as described above with reference to Figure 2. In some embodiments, the classifier was trained according to the methodology described above with reference to Figure 2.
[0256] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation dataset of at least 50 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 50 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 50 data constructs.
[0257] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation dataset of at least 100 data constructs, e.g., none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs.
[0258] In some embodiments, the method further includes assigning and / or administering a therapy to the subject based on the classification of the cancer status, for example, based on whether the subject's cancer is associated with EBV viral infection.
[0259] Thus, in one embodiment, a method for treating gastric cancer in a human cancer patient is provided. The method includes determining whether the human cancer patient is infected with the Epstein-Barr virus (EBV) oncogenic virus by obtaining a dataset about the human cancer patient, the dataset including a plurality of abundance values, each of which quantifies the expression level of a corresponding gene in a plurality of genes, the plurality of genes including at least five genes selected from the genes listed in Table 4. The method then includes inputting the dataset into a classifier trained to identify, in the subject's cancerous tissue, at least a first gastric cancer state associated with EBV infection and a second gastric cancer state associated with EBV-free status based on the abundance values of the plurality of genes. In some embodiments, the classifier is trained according to the methodology described above with reference to FIG. 2. The method then includes treating the gastric cancer. If the classifier results indicate that the human cancer patient is infected with the EBV oncogenic virus, a first therapy tailored to treat gastric cancer associated with EBV infection is administered. If the results of the classifier indicate that the human cancer patient is not infected with the EBV oncogenic virus, a second therapy tailored for the treatment of gastric cancer not associated with EBV infection is administered.
[0260] In some embodiments, the plurality of genes includes all of the genes listed in Table 4. In some embodiments, the plurality of genes includes one or more genes not listed in Table 4, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 4. In some embodiments, the plurality of genes includes 20 or fewer genes. In some embodiments, the plurality of genes includes 25 or fewer genes. In some embodiments, the plurality of genes includes 50 or fewer genes. In some embodiments, the plurality of genes includes 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0261] In some embodiments, the dataset also includes a mutant allele count for one or more alleles at one or more loci in the genome of the cancerous tissue from the subject. In some embodiments, the mutant allele count is either 1, representing that the subject has the mutant allele, or 0, representing that the subject does not have the mutant allele. In some embodiments, the mutant allele is a somatic mutation derived from the subject's germline. In some embodiments, the mutant allele is a mutation derived from a cancer derived from the cancerous tissue. In some embodiments, the mutant allele is located in the TP53 (ENSG00000141510) or PIK3CA (ENSG00000121879) gene.
[0262] In some embodiments, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm, as described above with reference to Figure 2. In some embodiments, the classifier was trained according to the methodology described above with reference to Figure 2.
[0263] In some embodiments, the first therapy tailored to treat gastric cancer associated with EBV infection is adoptive cell therapy, hi some embodiments, the adoptive cell therapy comprises ATA 129 (Atara), EBVST (Tessa), or CMD-003 (Cell Medica).
[0264] In some embodiments, the first therapy tailored for the treatment of gastric cancer associated with EBV infection is an immune checkpoint inhibitor, hi some embodiments, the immune checkpoint inhibitor is pembrolizumab (Merck) or nivolumab (Bristol-Myers Squibb).
[0265] In some embodiments, the first therapy tailored for the treatment of gastric cancer associated with EBV infection is a BTK inhibitor. In some embodiments, the BTK inhibitor is ibrutinib (Pharmacyclics).
[0266] report In some embodiments, the methods described herein include generating a patient report of the subject's cancer condition. The report may be presented to the patient, physician, medical professional, or researcher as a digital copy (e.g., a JSON object, a PDF file, or an image on a website or portal), a hard copy (e.g., printed on paper or another tangible medium), audio (e.g., recorded or streamed), or in another format.
[0267] The report includes information related to particular features of the patient's cancer, such as detected genetic mutations, epigenetic abnormalities, associated oncogenic pathogenic infections, and / or pathological abnormalities. In some embodiments, other features of the patient's sample and / or clinical records are also included in the report. In some embodiments, the report includes information about clinical trials for which the patient is eligible, therapies specific to the patient's cancer, and / or possible adverse treatment effects associated with particular features of the patient's cancer, such as the patient's genetic mutations, epigenetic abnormalities, associated oncogenic pathogenic infections, and / or pathological abnormalities, or other features of the patient's sample and / or clinical records.
[0268] In some embodiments, the results included in the report, and / or any additional results (e.g., from a bioinformatics pipeline), are used to query a database of clinical data, e.g., to determine whether there are trends indicating that a particular therapy has been effective in treating (e.g., slowing or halting the progression of cancer) other patients with the same or similar characteristics.
[0269] In some embodiments, results are used to design cell-based research of patient biology, for example, tumor organoid experiments.For example, organoid can be genetically engineered to have the same characteristics as specimen, and can be observed after exposure to therapy to determine whether the therapy can reduce the growth rate of organoid, and thus reduce the growth rate of the patient associated with specimen.Similarly, in some embodiments, results are used to direct research on tumor organoids that are directly derived from patients.An example of such experiment is described in US Provisional Patent Application No. 62 / 944,292, filed on December 5, 2019, the contents of which are incorporated herein by reference in their entirety for all purposes.
[0270] In some embodiments, the patient report includes a section regarding the subject's oncogenic pathogen infection status. For example, Figures 7A and 7B show examples of information provided at the time of diagnosis of HPV-positive head and neck cancer and HPV-positive cervical cancer, respectively.
[0271] EBV probe set In some embodiments, the present disclosure provides probes for binding, enriching, and / or detecting nucleic acid molecules, such as mRNA transcripts isolated from a cancerous tissue sample from a subject and / or cDNA molecules prepared from those mRNA transcripts, which are useful for determining whether the subject has a first cancerous condition associated with EBV oncogenic virus infection or a second cancerous condition not associated with EBV oncogenic virus infection. Generally, the probe comprises a DNA, RNA, or modified nucleic acid structure having a base sequence complementary to a nucleic acid molecule of interest. Thus, when a probe is designed to hybridize to mRNA molecules isolated from cancerous tissue, the probe will comprise a nucleic acid sequence complementary to the coding strand of the gene from which the transcript was derived; for example, the probe will comprise an antisense sequence of the gene. However, when a probe is designed to hybridize to cDNA molecules, because the molecules in a cDNA library are double-stranded, the probe may comprise either a sequence complementary to the coding sequence of the gene of interest (antisense sequence) or a sequence identical to the coding sequence of the gene of interest (sense sequence).
[0272] In some embodiments, the probe contains an additional nucleic acid sequence that does not share any homology with the gene sequence of interest. For example, in some embodiments, the probe also contains an identifier sequence, e.g., a nucleic acid sequence containing a unique molecular identifier (UMI), that is specific to a particular cancerous tissue sample or cancer patient. Examples of identifier sequences are described, for example, in Kivioja et al., 2011, Nat. Methods 9(1):72-74, and Islam et al., 2014, Nat. Methods 11(2), pp. 163-66, the contents of which are incorporated herein by reference in their entirety for all purposes. Similarly, in some embodiments, the probe also contains a primer nucleic acid sequence useful for amplifying the nucleic acid molecule of interest, e.g., using PCR. In some embodiments, the probe also contains a capture sequence designed to hybridize to an anti-capture sequence to recover the nucleic acid molecule of interest from the sample.
[0273] Similarly, in some embodiments, the probe comprises a non-nucleic acid affinity moiety covalently bound to a nucleic acid molecule complementary to a gene of interest for retrieving the nucleic acid molecule of interest. Non-limiting examples of non-nucleic acid affinity moieties include biotin, digoxigenin, and dinitrophenol. In some embodiments, the probe is attached to a solid surface or particle, such as a dipstick or magnetic bead, for retrieving the nucleic acid molecule of interest.
[0274] Thus, in one embodiment, the present disclosure provides a plurality of nucleic acid probes for identifying a first cancer condition and a second cancer condition in a human subject, wherein the first cancer condition is associated with infection by the Epstein-Barr virus (EBV) oncogenic virus and the second cancer condition is associated with a condition that does not involve EBV. The plurality of nucleic acid probes comprises at least five nucleic acid probes, each of the at least five nucleic acid probes comprising a respective nucleic acid sequence identical to or complementary to at least 10 contiguous bases of an RNA transcript of a different respective gene selected from the genes set forth in Table 4.
[0275] In some embodiments, the plurality of nucleic acid probes comprises at least 10 probes having sequences complementary to or identical to sequences from different genes listed in Table 4. In some embodiments, the plurality of nucleic acid probes comprises 2, 3, 4, 5, 6, 7, 8, or 9 probes having sequences complementary to or identical to sequences from different genes listed in Table 4.
[0276] In some embodiments, the plurality of nucleic acid probes includes one or more probes that bind to sequences of genes not listed in Table 4. In some embodiments, the plurality of nucleic acid probes includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more probes that bind to sequences of genes not listed in Table 4. In some embodiments, the plurality of nucleic acid probes includes probes having sequences that bind to 20 or fewer genes. In some embodiments, the plurality of nucleic acid probes includes probes having sequences that bind to 25 or fewer genes. In some embodiments, the plurality of nucleic acid probes includes probes having sequences that bind to 50 or fewer genes. In some embodiments, the plurality of nucleic acid probes includes probes having sequences that bind to 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0277] In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical to or complementary to at least 15 contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 4. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical to or complementary to at least 30 contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 4. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical to or complementary to at least 50 contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 4. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical to or complementary to at least 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, or more contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 4.
[0278] Digital and Laboratory Healthcare Platform In some embodiments, the above-described methods and systems are utilized in conjunction with or as part of a digital and laboratory healthcare platform generally directed to medical care and research. It should be understood that many uses of the above-described methods and systems in conjunction with such platforms are possible. One example of such a platform is described in U.S. Patent Application No. 16 / 657,804, entitled "Data Based Cancer Research and Treatment Systems and Methods," filed October 18, 2019, which is incorporated herein by reference in its entirety for all purposes.
[0279] For example, an implementation of one or more embodiments of the above-described methods and systems may include microservices constituting a digital and laboratory healthcare platform that supports diagnosis and treatment selection for cancers associated with oncogenic pathogen infection. The embodiment may include a single microservice for performing and providing diagnosis and treatment selection for cancers associated with oncogenic pathogen infection, or may include multiple microservices, each with a specific role that together implement one or more of the above-described embodiments. In one example, a first microservice may perform classification to provide a second microservice with a diagnosis for recommending an appropriate treatment for cancers associated with oncogenic pathogen infection. Similarly, the second microservice may perform treatment analysis to provide a recommended treatment, according to the above-described embodiments.
[0280] When the above embodiments are implemented in one or more microservices in conjunction with or as part of a digital and laboratory healthcare platform, one or more of such microservices may be part of an order management system that orchestrates the sequence of events as needed at the appropriate time and in the appropriate order necessary to instantiate the above embodiments. A microservices-based order management system is disclosed, for example, in U.S. Provisional Patent Application No. 62 / 873,693, entitled "Adaptive Order Fulfillment and Tracking Methods and Systems," filed July 12, 2019, which is incorporated herein by reference in its entirety for all purposes.
[0281] For example, continuing with the first and second microservices described above, the order management system may notify the first microservice that an order for classifying the oncogenic pathogen status of a cancer has been received and is ready for processing. When the classification is ready for delivery to the second microservice, the first microservice may execute and notify the order management system. Further, the order management system may identify that the execution parameters (preconditions) of the second microservice, including the completion of the first microservice, have been met and notify the second microservice that it can continue processing the order for recommending an appropriate treatment for the cancer associated with an oncogenic pathogen infection according to the above embodiment.
[0282] When the digital and laboratory healthcare platform further includes a genetic analysis system, the genetic analysis system may include a targeted panel and / or sequencing probes. Examples of targeted panels are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 902,950, entitled "System and Method for Expanding Clinical Options for Cancer Patients using Integrated Genomic Profiling," filed September 19, 2019, which is incorporated herein by reference in its entirety for all purposes. In one example, the targeted panel may enable delivery of next-generation sequencing results for detecting oncogenic pathogen infection according to one embodiment described above. Examples of next-generation sequencing probe designs are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 924,073, entitled "Systems and Methods for Next-Generation Sequencing Uniform Probe Design," filed October 21, 2019, which is incorporated herein by reference in its entirety for all purposes.
[0283] If the digital and laboratory healthcare platform further includes a bioinformatics pipeline, the above-described methods and systems can be utilized after the completion or substantial completion of the systems and methods utilized in the bioinformatics pipeline. In one example, the bioinformatics pipeline can receive the results of next-generation gene sequencing and return a set of binary files, such as one or more BAM files, that reflect DNA and / or RNA read counts aligned to a reference genome. The above-described methods and systems can be utilized, for example, to capture DNA and / or RNA read counts and, as a result, generate a classification of the subject's oncogenic pathogen status.
[0284] If the digital and laboratory healthcare platform further includes an RNA data normalizer, any RNA read counts may be normalized before processing the embodiment as described above. Examples of RNA data normalizers are disclosed, for example, in U.S. Patent Application No. 16 / 581,706, entitled "Methods of Normalizing and Correcting RNA Expression Data," filed September 24, 2019, which is incorporated herein by reference in its entirety for all purposes.
[0285] Where the digital and laboratory healthcare platform further includes a genetic data deconvolutor, any of the systems and methods for deconvolution may be utilized to analyze genetic data associated with a specimen having two or more biological components to determine the contribution of each component to the genetic data and / or to determine what genetic data would be associated with any component of the specimen if the specimen were purified. Examples of genetic data deconvoluters are disclosed, for example, in U.S. Patent Application Nos. 16 / 732,229 and PCT 19 / 69161, both entitled "Transcriptome Deconvolution of Metastatic Tissue Samples," filed December 31, 2019; U.S. Provisional Patent Application No. 62 / 924,054, entitled "Calculating Cell-type RNA Profiles for Diagnosis and Treatment," filed October 21, 2019; and U.S. Provisional Patent Application No. 62 / 944,995, entitled "Rapid Deconvolution of Bulk RNA Transcriptomes for Large Data Sets (Including Transcriptomes of Specimens Having Two or More Tissue Types)," filed December 6, 2019, which are incorporated by reference in their entireties for all purposes.
[0286] When the digital and laboratory healthcare platform further includes an automated RNA expression caller, the RNA expression level can be adjusted to be expressed as a value relative to a reference expression level, which is often done to prepare multiple RNA expression datasets for analysis and to avoid artifacts that occur when the datasets are different because they were not generated using the same method, device, and / or reagent. Examples of automated RNA expression callers are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 943,712, entitled "Systems and Methods for Automating RNA Expression Calls in a Cancer Prediction Pipeline," filed December 4, 2019, which is incorporated herein by reference in its entirety for all purposes.
[0287] The digital and laboratory healthcare platform may further include one or more insight engines for providing condition-related information, characteristics, or decisions that may be based on genetic and / or clinical data associated with the patient and / or specimen. Exemplary insight engines may include a tumor of unknown origin engine, a human leukocyte antigen (HLA) loss of homozygosity (LOH) engine, a tumor mutation burden engine, a PD-L1 status engine, a homologous recombination deficiency engine, a cellular pathway activation reporting engine, an immune infiltration engine, a microsatellite instability engine, a pathogen infection status engine, and the like. Examples of tumor of unknown origin engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 855,750, entitled "Systems and Methods for Multi-Label Cancer Classification," filed May 31, 2019, which is incorporated herein by reference in its entirety for all purposes. Examples of HLA LOH engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 889,510, entitled "Detection of Human Leukocyte Antigen Loss of Heterozygosity," filed August 20, 2019, which is incorporated by reference in its entirety for all purposes. Examples of tumor mutation burden (TMB) engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 804,458, entitled "Assessment of Tumor Burden Methodologies for Targeted Panel Sequencing," filed February 12, 2019, which is incorporated by reference in its entirety for all purposes.Examples of PD-L1 status engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 854,400, entitled "A Pan-Cancer Model to Predict The PD-L1 Status of a Cancer Cell Sample Using RNA Expression Data and Other Patient Data," filed May 30, 2019, which is incorporated by reference in its entirety for all purposes. Additional examples of PD-L1 status engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 824,039, entitled "PD-L1 Prediction Using H&E Slide Images," filed March 26, 2019, which is incorporated by reference in its entirety for all purposes. An example of a homologous recombination deficiency engine is disclosed, for example, in U.S. Provisional Patent Application No. 62 / 804,730, entitled "An Integrative Machine-Learning Framework to Predict Homologous Recombination Deficiency," filed February 12, 2019, which is incorporated by reference in its entirety for all purposes. An example of a cellular pathway activation reporting engine is disclosed, for example, in U.S. Provisional Patent Application No. 62 / 888,163, entitled "Cellular Pathway Report," filed August 16, 2019, which is incorporated by reference in its entirety for all purposes. An example of an immune infiltration engine is disclosed, for example, in U.S. Patent Application No. 16 / 533,676, entitled "A Multi-Modal Approach to Predicting Immune Infiltration Based on Integrated RNA Expression and Imaging Features," filed August 6, 2019, which is incorporated by reference in its entirety for all purposes.Additional examples of immune infiltration engines are disclosed, for example, in U.S. Patent Application No. 62 / 804,509, entitled "Comprehensive Evaluation of RNA Immune System for the Identification of Patients with an Immunologically Active Tumor Microenvironment," filed February 12, 2019, which is incorporated by reference in its entirety for all purposes. Examples of MSI engines are disclosed, for example, in U.S. Patent Application No. 16 / 653,868, entitled "Microsatellite Instability Determination System and Related Methods," filed October 15, 2019, which is incorporated by reference in its entirety for all purposes. Additional examples of MSI engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 931,600, entitled "Systems and Methods for Detecting Microsatellite Instability of a Cancer Using a Liquid Biopsy," filed November 6, 2019, which is incorporated by reference in its entirety for all purposes.
[0288] When the digital and laboratory healthcare platform further includes a report generation engine, the above-described methods and systems can be utilized to generate a summary report of a patient's genetic profile and the results of one or more insight engines for presentation to a physician. For example, the report can provide the physician with information about the extent to which the sequenced specimen contained tumor or normal tissue from a first organ, a second organ, a third organ, etc. For example, the report can provide a genetic profile for each tissue type, tumor, or organ in the specimen. The genetic profile represents the gene sequences present in the tissue type, tumor, or organ and can include information about mutations, expression levels, gene products, or other information that can be derived from genetic analysis of the tissue, tumor, or organ. The report can include tailored therapies and / or clinical trials based on some or all of the genetic profile or insight engine results and summaries. For example, therapy can be adapted according to the systems and methods disclosed in U.S. Provisional Patent Application No. 62 / 804,724, entitled "Therapeutic Suggestion Improvements Gained Through Genomic Biomarker Matching Plus Clinical History," filed February 12, 2019, which is incorporated herein by reference in its entirety for all purposes. For example, clinical trials can be adapted according to the systems and methods disclosed in U.S. Provisional Patent Application No. 62 / 855,913, entitled "Systems and Methods of Clinical Trial Evaluation," filed May 31, 2019, which is incorporated herein by reference in its entirety for all purposes.
[0289] The report may include a comparison of the results to a database of results from many specimens. Examples of methods and systems for comparing results to a database of results are disclosed in U.S. Provisional Patent Application No. 62 / 786,739, entitled "A Method and Process for Predicting and Analyzing Patient Cohort Response, Progression and Survival," filed December 31, 2018, which is incorporated by reference in its entirety for all purposes. The information may optionally be used in combination with similar information from additional specimens and / or clinical response information to discover biomarkers or design clinical trials.
[0290] When the digital and laboratory healthcare platform further includes the application of one or more of the embodiments herein to the organoids developed in connection with the platform, the method and system can be used to further evaluate the genetic sequencing data derived from the organoids to provide information about the extent to which the sequenced organoids contain a first cell type, a second cell type, a third cell type, etc. For example, the report can provide a genetic profile for each cell type in the specimen. The genetic profile represents the gene sequences present in a given cell type and can include information about mutations, expression levels, gene products, or other information that can be derived from genetic analysis of the cells. The report can include therapies that are matched based on some or all of the deconvoluted information. These therapies can be tested on the organoids, their derivatives, and / or similar organoids to determine the sensitivity of the organoids to these therapies. For example, organoids can be cultured and tested according to the systems and methods disclosed in U.S. Patent Application No. 16 / 693,117, entitled "Tumor Organoid Culture Compositions, Systems, and Methods," filed November 22, 2019; U.S. Provisional Patent Application No. 62 / 924,621, entitled "Systems and Methods for Predicting Therapeutic Sensitivity," filed October 22, 2019; and U.S. Provisional Patent Application No. 62 / 944,292, entitled "Large Scale Phenotypic Organoid Analysis," filed December 5, 2019, which are incorporated herein by reference in their entireties for all purposes.
[0291] When the digital and laboratory healthcare platform further includes one or more of the above applications in combination with or as part of a medical device or laboratory-developed test directed to healthcare and research generally, the results of such laboratory-developed test or medical device can be enhanced and personalized through the use of artificial intelligence. Examples of laboratory-developed tests, particularly tests that can be enhanced by artificial intelligence, are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 924,515, entitled "Artificial Intelligence Assisted Precision Medicine Enhancements to Standardized Laboratory Diagnostic Testing," filed October 22, 2019, which is incorporated herein by reference in its entirety for all purposes.
[0292] It should be understood that the above examples are illustrative and not limiting of the use of the systems and methods described herein in combination with digital and laboratory healthcare platforms. [Example]
[0293] Example 1 - The Cancer Genome Atlas (TCGA) The data used to train the classifiers shown in Examples 2 and 3 below was obtained from The Cancer Genome Atlas (TCGA). Briefly, the TCGA dataset is a publicly available dataset containing over 2 petabytes of genomic data for over 11,000 cancer patients, including clinical information about the cancer patients, metadata about the samples collected from such patients (e.g., sample portion weight, etc.), histopathology slide images from the sample portions, and molecular information obtained from the samples (e.g., mRNA / miRNA expression, protein expression, copy number, etc.). The TCGA dataset contains data on 33 different cancers: breast (ductal carcinoma, panlobular carcinoma), central nervous system (glioblastoma multiforme, low-grade glioma), endocrine (adrenocortical carcinoma, papillary thyroid carcinoma, paraganglioma, and pheochromocytoma), gastrointestinal (bile duct carcinoma, colorectal adenocarcinoma, esophageal carcinoma, hepatocellular carcinoma, pancreatic ductal adenocarcinoma, and gastric carcinoma), gynecological (cervical carcinoma, ovarian serous cystadenocarcinoma, uterine carcinosarcoma, and endometrial carcinoma), head and neck (head and neck squamous cell carcinoma, uveal melanoma), hematological (acute myeloid leukemia, thymoma), skin (cutaneous melanoma), soft tissue (sarcoma), breast (lung adenocarcinoma, lung squamous cell carcinoma, and mesothelioma), and urological (chromophobe renal carcinoma, clear cell renal carcinoma, papillary renal carcinoma, prostate adenocarcinoma, testicular germ cell carcinoma, and urothelial bladder carcinoma).
[0294] Example 2 - RNA Expression Profiling Referring to Figure 3, expression profiles of genes useful for determining HPV viral status were determined from head and neck cancer tumor samples.
[0295] Head and neck tumor biopsies were obtained from cancer patients using the biopsy techniques described herein according to block 302 of Figure 3. The biopsies were flash frozen in liquid nitrogen immediately after removal from the patient.
[0296] mRNA was isolated from tumor samples according to block 304 in Figure 3. Briefly, the sample tissue block was removed from liquid nitrogen, and a 5mm x 5mm x 5mm block of the sample was extracted and dissected using a cold knife. The dissected sample was mixed with TRIzol reagent (Chomczynski and Sacchi, 1987, Anal Biochem. 162(1), pp. 156-59, the contents of which are incorporated herein by reference in their entirety for all purposes) and homogenized using a tissue homogenizer for three short cycles, e.g., 60 seconds, 30 seconds, and 30 seconds. Chloroform was added to the homogenized tumor sample, and the reaction mixture was mixed. After phase separation, the aqueous phase of the reaction mixture was removed and mixed with an equal volume of isopropanol to precipitate the RNA. The reaction mixture was centrifuged to pellet the RNA, and the supernatant was removed. The pellet was washed twice with cold ethanol and then air-dried. The extracted RNA was then resuspended in RNase-free water.
[0297] Next, referring to block 306 of Figure 3, mRNA in the isolated RNA was quantified by whole-exome sequencing. According to block 308 of Figure 3, mRNA was isolated from the extracted RNA by heating the extracted RNA to disrupt secondary structures, and then annealing the RNA to magnetic oligo(dT)-conjugated beads by incubating the RNA with the oligo(dT)-conjugated beads in hybridization buffer at room temperature. The beads were recovered and washed twice with hybridization buffer. The hybridized mRNA was then eluted by heating and recovered from the reaction mixture.
[0298] A cDNA library was constructed from the isolated mRNA according to block 310 in Figure 3. Briefly, divalent cations were added to the isolated mRNA to fragment the molecules at high temperature. The fragmented mRNA was precipitated by incubation in ethanol at pH 5.2 at -80°C using glycogen as a carrier molecule. The mRNA was pelleted by centrifugation, washed with 70% ethanol, air-dried, and resuspended in RNase-free water. First-strand DNA synthesis was performed using random primers and reverse transcriptase. Second-strand DNA synthesis was then performed using DNA polymerase in the presence of RNase H to form double-stranded cDNA. The 5'-overhangs created by second-strand synthesis were repaired using T4 and Klenow DNA polymerase to form blunt ends. The 3' ends of the blunt-ended cDNA were adenylated using Klenow DNA polymerase. Adapters were ligated to the ends of the adenylated cDNA using T4 DNA ligase, and the cDNA templates were purified and sized by agarose gel electrophoresis. If necessary, the purified cDNA template is enriched by PCR amplification, thereby forming the final cDNA library.
[0299] According to block 312 of Figure 3, whole-exome sequencing of the cDNA library was performed using Integrated DNA Technologies (IDT) XGEN® LOCKDOWN® technology with the xGen Exome Research Panel. Briefly, the xGen Exome Research Panel covers a 51 Mb end-to-end tiled probe space of the human genome, enabling deep and uniform coverage of target capture across the exome. The cDNA library was hybridized to biotinylated DNA capture probes covering a reference human exome. The hybridized probes were recovered by binding to streptavidin beads. Post-capture PCR was performed to enrich the captured sequences. The amplified products were then sequenced using sequencing-by-synthesis (SBS) technology (Bently et al., 2008, Nature 456(7218), pp. 53-59, the contents of which are incorporated herein by reference in their entirety for all purposes).
[0300] The RNA sequencing data is then normalized using gene length data, guanine-cytosine (GC) content data, and sepsis determination depth data to normalize the gene length data of at least one gene to reduce systematic bias, normalize the GC content data of at least one gene to reduce systematic bias, and normalize the sequencing depth data for each sample, as described in U.S. Provisional Patent Application No. 62 / 735,349 and U.S. Patent Application No. 16 / 581,706, the contents of which are incorporated herein by reference in their entirety for all purposes. As described in U.S. Provisional Patent Application No. 62 / 735,349 and U.S. Patent Application No. 16 / 581,706, the RNA sequencing data is also corrected against a standard gene expression dataset by comparing sequence data for at least one gene in the gene expression dataset with sequence data in a standard gene expression dataset. The normalized and corrected RNA expression data for the 24 genes identified in Table 3, as well as the patient's CDKN2A and TP53 allele status, were then input into the HPV detection classifier trained in Example 3 to determine the patient's HPV viral status.
[0301] Example 3 - Detection of human papillomavirus Referring to Figures 4A-4D, a classifier for determining HPV viral status was trained using gene expression derived from tumor RNA-seq data of a training population in which each subject in the training population was diagnosed with head and neck squamous cell carcinoma or cervical cancer.
[0302] A training dataset was obtained according to block 204 of FIG. 2A. Here, the dataset included corresponding abundance values for each subject in TCGA with cervical cancer or head and neck cancer whose HPV status was known, as described in Example 1. As shown in FIG. 4A, 427 subjects in TCGA met these selection criteria and served as the subjects for the training dataset. Of the 427 subjects, 263 had head and neck cancer, and 164 had cervical cancer. Of the 263 subjects with head and neck cancer, 32 were HPV-positive and 231 were HPV-negative. Of the 164 subjects with cervical cancer, 156 were HPV-positive and 8 were HPV-negative. Thus, of the 427 subjects, 188 subjects were considered to have a primary cancer condition (having HPV and having head and neck cancer or cervical cancer), and the remaining 239 subjects were considered to have a secondary cancer condition (not having HPV but having head and neck cancer or cervical cancer).
[0303] Next, according to block 218 of FIG. 2C and block 228 of FIG. 2D, gene expression values derived from the whole-exome RNA data in the TCGA dataset for the 427 subjects were used to identify discriminatory gene sets by regression, with the gene expression values derived from the whole-exome mRNA expression data for the 427 subjects in the TCGA dataset serving as independent variables, and an indicator of whether each subject had a first cancer condition (having HPV and having head and neck cancer or cervical cancer) or a second cancer condition (having no HPV but having head and neck cancer or cervical cancer) serving as the dependent variable. More specifically, according to block 228 of FIG. 2D, the dataset consisting of 427 subjects was divided into 10 sets (10-fold). Each set included two or more subjects with the first cancer condition and two or more subjects with the second cancer condition. Each of the 10 sets (folds) was independently subjected to a regression in which the whole-exome mRNA expression data for subjects in each set served as the independent variable and an indicator of whether each subject in each set had a first or second cancer condition served as the dependent variable. Each regression (fold) was performed using L1 (lasso) regularization, as per block 238 of Figure 2E. Because L1 regularization leads to sparse coefficients, only a small subset of genes had non-zero coefficients for each set. Only genes with non-zero coefficients in 80% or more of the sets were included in the final model. In other words, only genes with non-zero regression coefficients for at least eight of the 10 sets (folds) were accepted into the discriminatory set of genes based on their expression data. The list of genes that met this requirement is shown in Figure 4B, with the feature type being "gene expression." Additionally, Figure 6A shows a principal component analysis of the abundance values of the genes listed in Figure 4B across the training set. FIG. 6A shows that a plot of the first and second PCA values for each of the subjects in the training set separates into two distinct groups corresponding to the first cancer state (group 602) and the second cancer state (604), demonstrating the power of the gene abundance values described in FIG. 4B to distinguish between the first and second cancer states.
[0304] In some embodiments, additional genes were included in the set of discriminatory genes based on the presence or absence of mutations (e.g., number of mutations) in the additional genes. In this example, as detailed in Figure 4B, the genes CDKN2A and TP53 were included in the discriminatory gene set, and the features for these genes were the number of times mutations in these genes were observed in each of the respective 427 subjects in the training set.
[0305] Next, according to block 242 of FIG. 2E, a classifier was trained to distinguish between first and second cancer states as a function of the abundance values of the discriminatory gene sets using the respective abundance values for the discriminatory gene sets and the respective indicators of cancer state across the 427 subjects. In the first model, the classifier used was a logistic regression classifier with L1 regularization, trained on 427 subjects but using only the TCGA gene abundance levels for the genes listed in FIG. 4B whose feature is "gene expression." In the second model, the classifier used was a logistic regression classifier with L1 regularization, trained on 427 subjects using the TCGA gene abundance levels for the genes listed in FIG. 4B whose feature is "gene expression," and the TCGA mutation counts for the two genes in FIG. 4B whose feature is "number of mutations." In the third model, the classifier used was a support vector machine (SVM) classifier from Scikit-learn, as disclosed in Pedregosa et al. 2011, "Machine Learning in Python," JMLR 12, pp. 2825-2830, incorporated herein by reference, and was trained on 427 subjects, but using only TCGA gene abundance levels for the genes listed in Figure 4B, where the feature is "gene expression." When validated against data from a cohort of 133 subjects with cervical or head and neck cancer and known HPV status, the classifier performed with a specificity of 92.5% and a sensitivity of 89.7%.
[0306] In the fourth model, the classifier used was the same SVM classifier, trained on 427 subjects using TCGA gene abundance levels for the genes listed in Figure 4B, where the feature is "gene expression," and TCGA mutation counts for the two genes listed in Figure 4B, where the feature is "number of mutations." The performance of this trained classifier is reported in Figure 4C. The regression coefficients and correlation statistics for each of the features used in the model are shown in Tables 5 and 6, respectively. The SVM parameters used were: class weight: none, decision function shape: ovo, gamma: scale, kernel: linear, probability: true, shrinkage: false, and tol: 1. As shown in Figure 4C, the trained SVM predicts cancer type for the 427 subjects. The classifier correctly identified the HPV infection status of 122 of the 133 subjects with cervical or head and neck cancer and known HPV status, with a specificity of 95% and a sensitivity of 87.5% for a training set of 427 subjects, i.e., whether the subject had a first cancer type (HPV and head and neck cancer or cervical cancer) or a second cancer type (no HPV but head and neck cancer or cervical cancer). [Table 5] [Table 6-1] [Table 6-2] [Table 6-3] [Table 6-4] [Table 6-5] [Table 6-6] [Table 6-7] [Table 6-8] [Table 6-9]
[0307] To validate the model, the trained SVM classifier reported in Figure 4C was tested on a validation population that was not used to train the classifier.As detailed in Figure 4A, the validation dataset included a corresponding plurality of abundance values for each subject in the dataset referred to as the "test" dataset described in Example 2, who had cervical cancer or head and neck cancer with known HPV status.As shown in Figure 4A, 133 subjects that met these selection criteria were selected from the validation dataset to serve as a plurality of subjects for the validation dataset.Of the 133 validation subjects, 93 had head and neck cancer and 40 had cervical cancer.Of the 93 subjects with head and neck cancer, 28 were HPV positive and 65 were HPV negative.Of the 40 subjects with cervical cancer, 28 were HPV positive and 12 were HPV negative. Therefore, of the 133 validated subjects, 56 were considered to have the primary cancer condition (having HPV and having head and neck cancer or cervical cancer), and the remaining 77 were considered to have the secondary cancer condition (not having HPV but having head and neck cancer or cervical cancer).
[0308] Each of the 133 validation subjects was run against the trained SVM, and its performance is reported in Figure 4C. The SVM assigned the cancer to either the first or second cancer class. Specifically, gene abundance values for the genes listed in Figure 4B, where the feature type is "gene expression," and mutation counts for the two genes listed in Figure 4B, where the feature type is "number of mutations," were measured from the tumor samples of each of the 133 validation subjects. This data for each validation subject was input individually into the trained SVM model in Figure 5C. As shown in Figure 4D, the trained SVM had a specificity of 95% and a sensitivity of 88% for cancer class across the 133 validation subjects. Adding covariates for the number of mutations in the genes TP53 and CDKN2A to the SVM was found to improve accuracy but improve the AUC from 0.97 to 0.98. This example demonstrates that the trained SVM model accurately predicts viral infection in tumors using RNA expression data.
[0309] This example confirms that viral infection is generally associated with an upregulated immune response. It further demonstrates that viral detection based on whole transcriptome data is a useful clinical tool in its own right and can be combined with existing diagnostic methods to provide insight into the viral status and tumor microenvironment in a single test.
[0310] Example 4 - Detection of Epstein-Barr Virus Referring to Figures 5A-5D, a classifier for determining EBV viral status was trained using gene expression derived from tumor RNA-seq data of a training population in which each subject in the training population was diagnosed with gastric cancer.
[0311] According to block 204 of FIG. 2A, a training dataset was obtained. Here, the dataset included a corresponding plurality of abundance values for each subject in TCGA with gastric cancer and known EBV status, as described in Example 1. As shown in FIG. 5A, 212 subjects in TCGA met these selection criteria and served as a plurality of subjects in the training dataset. Of the 212 subjects, 21 were EBV-positive and 191 were EBV-negative. Therefore, of the 212 subjects, 21 subjects were considered to be in the first cancer state (having EBV and having gastric cancer), and the remaining 191 subjects were considered to be in the second cancer state (not having EBV but having gastric cancer).
[0312] Next, according to block 218 of FIG. 2C and block 228 of FIG. 2D, gene expression values derived from the whole-exome RNA data in the TCGA dataset for the 212 subjects were used to identify discriminatory gene sets by regression, with the gene expression values derived from the whole-exome mRNA expression data for the 212 subjects in the TCGA dataset serving as independent variables, and an indicator of whether each subject had a first cancer condition (having EBV and having gastric cancer) or a second cancer condition (not having EBV but having gastric cancer) serving as a dependent variable. More specifically, according to block 228 of FIG. 2D, the dataset consisting of 212 subjects was divided into 10 sets (10-fold). Each set included two or more subjects with the first cancer condition and two or more subjects with the second cancer condition. Each of the 10 sets (folds) was independently subjected to a regression in which the whole-exome mRNA expression data for subjects in each set served as the independent variable and an indicator of whether each subject in each set had a first or second cancer condition served as the dependent variable. Each regression (fold) was performed using L1 (lasso) regularization, as per block 238 of Figure 2E. Because L1 regularization leads to sparse coefficients, only a small subset of genes had non-zero coefficients for each set. Only genes with non-zero coefficients in 80% or more of the sets were included in the final model. In other words, only genes with non-zero regression coefficients for at least eight of the 10 sets (folds) were accepted into the discriminatory set of genes based on their expression data. The list of genes that met this requirement is shown in Figure 5B, with the feature type being "gene expression." Additionally, Figure 6B shows a principal component analysis of the abundance values of the genes listed in Figure 5B across the training set. Figure 6B shows that a plot of the first and second PCA values for each of the subjects in the training set separates into two distinct groups corresponding to the first cancer state (group 606) and the second cancer state (606), demonstrating the power of the gene abundance values described in Figure 5B to distinguish between the first and second cancer states.
[0313] In some embodiments, additional genes were included in the set of discriminatory genes based on the presence or absence of mutations (e.g., number of mutations) in the additional genes. In this example, as detailed in Figure 5B, the genes PIK3CA and TP53 were included in the discriminatory gene set, and the features for these genes were the number of times mutations in these genes were observed in each of the 212 subjects in the training set.
[0314] Next, according to block 242 of FIG. 2E, a classifier was trained to distinguish between first and second cancer states as a function of the abundance values of the discriminatory gene sets using the respective abundance values for the discriminatory gene sets across the 212 subjects and the respective indicators of cancer state. In the first model, the classifier used was a logistic regression classifier with L1 regularization, trained on 212 subjects but using only the TCGA gene abundance levels for the genes listed in FIG. 5B whose feature is "gene expression." In the second model, the classifier used was a logistic regression classifier with L1 regularization, trained on 212 subjects using the TCGA gene abundance levels for the genes listed in FIG. 5B whose feature is "gene expression," and the TCGA mutation counts for the two genes in FIG. 5B whose feature is "number of mutations." In the third model, the classifier used was a support vector machine (SVM) classifier from Scikit-learn, as disclosed in Pedregosa et al. 2011, "Machine Learning in Python," JMLR 12, pp. 2825-2830, incorporated herein by reference, and was trained on 212 subjects, but using only the TCGA gene abundance levels for the genes listed in Figure 5B, where the feature is "gene expression." When validated against data from a cohort of 55 subjects with gastric cancer and known EBV status, the classifier correctly identified the EBV status of 54 or 55 validation subjects with 100% specificity and 75% sensitivity.
[0315] In the fourth model, the same SVM classifier was used, trained on 212 subjects using TCGA gene abundance levels for the genes listed in Figure 4B, where the feature is "gene expression," and TCGA mutation counts for the two genes listed in Figure 4B, where the feature is "number of mutations." The performance of this trained classifier is reported in Figure 5C. The regression coefficients and correlation statistics for each of the features used in the model are shown in Tables 7 and 8, respectively. The SVM parameters used were: class weight: none, decision function shape: ovo, gamma: scale, kernel: linear, probability: true, shrinkage: false, and tol: 1. As shown in Figure 5C, the trained SVM predicted the cancer type of the 212 subjects—that is, whether the subject had the first cancer type (EBV and gastric cancer) or the second cancer type (not EBV but gastric cancer)—with 99% specificity and 95% sensitivity for the training set of 212 subjects. The classifier was then validated against data from a cohort of 55 subjects with gastric cancer and known EBV status. The classifier correctly identified the EBV infection status of 54 of the 55 validation subjects, with a specificity of 100% and a sensitivity of 75%. [Table 7] [Table 8-1] [Table 8-2]
[0316] To validate the model, the trained SVM classifier reported in Figure 5C was tested on a validation population that was not used to train the classifier. As detailed in Figure 5A, the validation dataset included a corresponding plurality of abundance values for each subject in the dataset referred to as the "test" dataset described in Example 2, who had gastric cancer with a known EBV status. As shown in Figure 5A, 55 subjects who met these selection criteria and served as a plurality of subjects in the validation dataset were selected from the validation dataset. Of the 55 validation subjects, 4 were EBV-positive and 51 were EBV-negative. Therefore, of the 55 validation subjects, 4 were considered to have the first cancer status (having EBV and having gastric cancer), and the remaining 51 were considered to have the second cancer status (not having EBV but having gastric cancer).
[0317] Each of the 55 validation subjects was run against the trained SVM, the performance of which is reported in Figure 5C. The SVM assigned the cancer to either the first or second cancer class. Specifically, gene abundance values for the genes listed in Figure 5B, where the feature type is "gene expression," and mutation counts for the two genes listed in Figure 5B, where the feature type is "number of mutations," were measured from the tumor sample for each of the 55 validation subjects. This data for each validation subject was individually input into the trained SVM model in Figure 5C. As shown in Figure 5D, the trained SVM had 75% specificity and 100% sensitivity for cancer class using such data across the 55 validation subjects. This example demonstrates that the trained SVM model accurately predicts viral infection in tumors using RNA expression data. This example confirms that viral infection is generally associated with an upregulated immune response. This example further demonstrates that viral detection based on whole transcriptome data is a useful clinical tool in its own right and can be combined with existing diagnostic methods to provide insights into the viral status and tumor microenvironment in a single test.
[0318] Example 5 - Obtaining normalized RNA count data In this example, patient samples were subjected to RNA whole-exome short-read next-generation sequencing (NGS) to generate RNA sequencing data, which were then processed through a bioinformatics pipeline to generate RNA-seq expression profiles for each patient sample. Specifically, total nucleic acids (DNA and RNA) from solid tumors were extracted from macro-dissected FFPE tissue sections and digested with proteinase K to remove proteins. RNA was purified from total nucleic acids using TURBO DNase-I to remove DNA, and the reaction mixture was then washed using RNA clean XP beads to remove enzyme proteins. The isolated RNA was subjected to a quality control protocol using RiboGreen fluorescent dye to determine the concentration of RNA molecules.
[0319] Library preparation was performed using the KAPA Hyper Prep Kit, where 100 ng of RNA was heat-fragmented to an average size of 200 bp in the presence of magnesium. The library was then reverse transcribed into cDNA, and Roche SeqCap dual-end adapters were ligated to the cDNA. The cDNA library was then purified and size-selected using KAPA Hyper Beads. The library was then PCR-amplified for 10 cycles and purified using Axygen MAG PCR cleanup beads. Quality control was performed using a PicoGreen fluorescent kit to determine the cDNA library concentration. The cDNA libraries were then pooled into 6-plex hybridization reactions. Each pool was treated with human COT-1 and IDTxGen Universal Blockers and vacuum-dried. The RNA pools were then resuspended in IDT xGen Lockdown hybridization mix, and IDT xGen Exome Research Panel v1.0 probes were added to each pool. The pools were incubated to allow the probes to hybridize. The pool was then mixed with streptavidin-coated beads to capture hybridized cDNA molecules. The pool was amplified and purified again using the KAPA HiFi Library Amplification Kit and Axygen MAG PCR cleanup beads, respectively. A final quality control step, involving PicoGreen pool quantification and LabChip GX Touch, was performed to assess pool fragment size. The pool was cluster-amplified using Illumina Paired-end Cluster Kits with PhiX spikes on an Illumina C-Bot2, and the resulting flow cells containing the amplified target capture cDNA libraries were sequenced on an Illumina HiSeq 4000 to an average unique on-target depth of 500x, generating FASTQ files.
[0320] In this example, cDNA library preparation was performed in an automated system using a liquid handling robot (SciClone NGSx).
[0321] Each FASTQ file contained a list of paired-end reads generated by an Illumina sequencer, each associated with a quality assessment. The reads in each FASTQ file were processed through a bioinformatics pipeline. FASTQ files were analyzed using FASTQC for quality control and rapid read assessment. For each FASTQ file, each read in the file was aligned to the reference genome (GRch37) using Kallisto alignment software. This alignment generated a SAM file, which was then converted to a BAM file. The BAM files were sorted, and duplicates were marked for removal.
[0322] For each gene, the raw RNA read count for a given gene was calculated by the Kallisto alignment software as the sum of the probabilities that the read aligns to the gene for each read. Therefore, in this example, the raw count is not an integer. The raw read counts were saved for each patient in a tabular file, with columns representing genes and each entry representing the raw RNA read count for that gene.
[0323] The raw RNA read counts were then normalized to correct for GC content and gene length using full quantile normalization, and to adjust for sequencing depth via the size factor method. The normalized RNA read counts were saved for each patient in a tabular file, with columns representing genes and each entry representing the raw RNA read count for that gene.
[0324] Referenced and Alternative Embodiments All references cited herein are incorporated by reference herein in their entirety for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety for all purposes.
[0325] The present invention can be implemented as a computer program product including a computer program mechanism embedded in a non-transitory computer-readable storage medium. For example, the computer program product can include program modules as shown in any combination of Figures 1A and 1B and / or as described in Figures 2A, 2B, 2C, 2D, 2E, and 3. These program modules can be stored on a CD-ROM, DVD, magnetic disk storage product, USB key, or any other non-transitory computer-readable data or program storage product.
[0326] As will be apparent to those skilled in the art, many modifications and variations of this application can be made without departing from the spirit and scope of the present application. The specific embodiments described herein are provided by way of example only. The embodiments have been chosen and described to best explain the principles of the invention and its practical uses, so as to enable those skilled in the art to best utilize the invention and various embodiments, with various modifications suited to the particular uses contemplated. The present invention is to be limited only by the terms of the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
1. 1. A method for providing information about a cancerous condition in a human subject, wherein a first cancerous condition is cervical cancer associated with infection by a human papillomavirus (HPV) oncogenic virus or head and neck cancer associated with infection by HPV, and a second cancerous condition is cervical cancer associated with a non-HPV condition or head and neck cancer associated with a HPV-free condition, the method comprising: (A) obtaining a dataset for a subject, the dataset comprising a plurality of abundance values from the subject; each respective abundance value in the plurality of abundance values quantifies the level of expression of each gene in the plurality of genes in a sample of cancerous tissue from the subject; The plurality of genes includes at least all 24 of the genes listed in Table 3; and (B) inputting the dataset into a classifier trained to distinguish between at least the first cancer state and the second cancer state based on the abundance values of the plurality of genes.
2. 2. The method of claim 1, wherein the plurality of genes comprises 300 or fewer genes.
3. 3. The method of claim 1 or 2, further comprising determining the plurality of abundance values by RNA sequencing of the sample of the cancerous tissue from the subject.
4. 4. The method of any one of claims 1 to 3, wherein the dataset further comprises variant allele counts for one or more alleles at one or more loci in the genome of the cancerous tissue from the subject.
5. 5. The method of claim 4, wherein the one or more alleles are selected from alleles in the TP53 (ENSG00000141510) or CDKN2A (ENSG00000147889) gene.
6. 6. The method of any one of claims 1 to 5, wherein the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm.
7. providing information on cervical cancer or head and neck cancer associated with HPV infection if the result of the classifier indicates that the human subject is infected with HPV oncogenic virus; and providing information on cervical cancer or head and neck cancer not associated with HPV infection if the result of the classifier indicates that the human subject is not infected with HPV oncogenic virus. The method of any one of claims 1 to 6, further comprising providing information on cervical cancer or head and neck cancer by:
8. The method of claim 7, wherein the information on cervical cancer associated with HPV infection is information on a therapeutic vaccine or adoptive cell therapy.
9. The method according to claim 7 or 8, wherein the information on cervical cancer not associated with HPV infection is information on chemotherapy.
10. The method of claim 7, wherein the information on head and neck cancer associated with HPV infection is information on a therapeutic vaccine, an immune checkpoint inhibitor, or a PI3K inhibitor.
11. The method according to claim 7 or 10, wherein the information on head and neck cancer not associated with HPV infection is information on chemotherapy.
12. The method of claim 9 or 11, wherein the chemotherapy information comprises cisplatin information.
13. 13. The method of claim 12, wherein the information on cervical cancer further comprises information on a therapeutic agent selected from the group consisting of 5-fluorouracil, paclitaxel, and bevacizumab.
14. The method of claim 12 , wherein the information on head and neck cancer further includes information on concurrent radiotherapy or postoperative chemoradiotherapy.
Citation Information
Patent Citations
Prognosis and Treatment of Squamous Cell Carcinoma
JP2018532745A
Biomarkers for the detection of head and neck tumors
US20110229876A1
Device and Methods for the Detection of Cervical Disease
US20120156698A1