Systems and methods for using sequencing data for pathogen detection
Patent Information
- Application Number
- JP2025032870
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-02-26
- Filing Date
- 2025-03-03
- Publication Date
- 2025-12-25
AI Technical Summary
Conventional cancer treatment methods do not account for the presence of oncogenic pathogens, requiring separate assays that increase diagnosis costs and delay treatment plans, necessitating an integrated system for pathogen detection within cancer diagnosis.
A method using mRNA expression analysis to identify discriminant gene sets that differentiate between cancers associated with and not associated with oncogenic pathogens, enabling a classifier to determine cancer states and tailor treatments accordingly.
Enables direct pathogen detection without additional assays, achieving high specificity and sensitivity in identifying HPV and EBV infections, allowing personalized cancer treatment based on pathogen status.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 810,849, filed on February 26, 2019, the content of which is hereby incorporated by reference in its entirety for all purposes.
[0002] The present disclosure generally relates to using expression profiles derived from cancerous tissue to detect oncogenic pathogenic infections in cancer patients.
Background Art
[0003] Precision oncology is the practice of tailoring cancer treatment to an individual, based on, for example, the unique pathology, genome, epigenetics, and / or transcriptome profile of an individual tumor. In contrast, conventional cancer treatment is based simply on the type of cancer being treated. For example, conventionally, all breast cancers would be treated with a first treatment regimen, while all lung cancers would be treated with a second treatment regimen. Precision oncology arose from many observations that different patients diagnosed with the same type of cancer, e.g., breast cancer, responded very differently to the same treatment regimen. Over time, researchers have identified genomic, epigenetic, and transcriptomic markers that facilitate, to some extent, prediction of how individual cancers will respond to specific therapies.
[0004] The use of targeted therapies has led to significant improvements in the outcomes of cancer patients, particularly in terms of progression-free survival. Radovich et al., Oncotarget, 7:56491-500 (2016). Recent evidence reported from the IMPACT trial, which included genetic testing of advanced tumors from 3,743 patients and in which approximately 19% of patients received a matched targeted therapy based on tumor biology, showed a response rate of 16.2% in patients who received a matched treatment versus 5.2% in patients who received an unmatched treatment. Bankhead, “IMPACT Trial: Support for Targeted Cancer Tx Approaches,” MedPageToday, June 5, 2018. It was also found in the IMPACT study that the 3-year overall survival rate for patients who received a molecularly matched treatment was more than twice that of patients who did not match (15% vs. 7%). Id.; ASCO Post, “2018 ASCO: IMPACT Trial Matches Treatment to Genetic Changes in the Tumor to Improve Survival Across Multiple Cancer conditions,” The ASCO POST, June 6, 2018. Estimates of the proportion of patients whose care trajectories are changed by genetic testing vary widely, from approximately 10% to over 50%. Fernandes et al., Clinics, 72:588-94 (2017).
[0005] Therapies targeting specific genomic alterations are already standard of care in some tumor types, as suggested in the National Comprehensive Cancer Network (NCCN) guidelines for, for example, melanoma, colorectal cancer, and non-small cell lung cancer. Some well-known mutations in the NCCN guidelines can be identified in cancer patients using individual assays or small next-generation sequencing (NGS) panels. However, more comprehensive pathologic, genomic, epigenetic, and / or transcriptome analysis is needed to facilitate the use of off-label drugs, combination therapies, or tissue-independent immunotherapies in order for the largest number of cancer patients to benefit from personalized oncology. Schwaederle et al., JAMA Oncol., 2:1452-59 (2016), Schwaederle et al., J Clin Oncol., 32:3817-25 (2015), and Wheler et al., Cancer Res., 76:3690-701 (2016).
[0006] The presence of oncogenic pathogen infections accounts for 10-12% of all cancers. For example, gastric cancer is a common cause of cancer-related death, the third most common cancer worldwide, and an estimated >700,000 people died from gastric cancer in 2012. Ferlay, et al., “Cancer Incidence and Mortality Worldwide,” IARC CancerBase 11 [Internet], Lyon, France: International Agency for Research on Cancer (2013). In addition to genetic factors, gastric carcinogenesis is thought to be associated with multiple environmental factors, including Epstein-Barr virus (EBV) infection. Burke et al., Mod Pathol., 3:377-380 (1990). Indeed, recent Cancer Genome Atlas research has provided a molecular classification that defines EBV-positive gastric cancer as a distinct subtype. Cancer Genome Atlas Research Network, Nature, 513(7517):202-09 (2014).
[0007] Thus, the presence of such carcinogenic pathogens affects the prognosis of the associated cancer. Therefore, if a subject has a type of cancer known to frequently occur in relation to a carcinogenic pathogen, knowledge of the subject's pathogen status is important because it can change the subject's treatment options. For example, many clinical trials investigating the benefits of dose reduction of radiotherapy or chemotherapy for HPV-positive head and neck cancers have shown promising results. In addition, pathogen-associated tumors are likely to exhibit higher levels of inflammation and immune infiltration and are excellent candidates for immunotherapy.
[0008] A drawback of conventional carcinogenic pathogen diagnosis is that, to determine whether a subject is infected with a particular pathogen, a completely separate and independent assay is performed apart from the assay used initially to diagnose the subject with cancer or to assess the stage of the cancer. For example, in the case of EBV, separate laboratory methods such as in situ hybridization (ISH) or polymerase chain reaction (PCR) for resected tissue, biopsy, or blood, or enzyme-linked immunosorbent assay (ELISA) or immunofluorescence assay (IFA) for serum samples are performed to detect EBV infection. This is inadequate because it increases the cost of diagnosis and, in some cases, delays the creation of the subject's treatment plan until after a type of cancer known to be associated with a carcinogenic pathogen has been diagnosed and the results of the pathogen assay are obtained since the pathogen test is only performed after the cancer diagnosis. Summary of the Invention
[0009] In view of the above background, what is needed in the art is an improved system and method for pathogen detection that directly determines the presence of a given pathogen detection without the need for a separate independent assay for pathogen detection.
[0010] Accordingly, provided are improved methods for discriminating between cancers associated with oncogenic pathogen infections that contribute to cancer pathology and cancers not associated with oncogenic pathogen infections. Also provided are improved methods for treating cancer patients based on whether their cancers are associated with oncogenic pathogen infections. The present disclosure addresses these needs, for example, by providing methods for identifying a set of genes that are differentially expressed in cancers associated with oncogenic pathogen infections as compared to cancers not associated with oncogenic pathogen infections. The present disclosure also provides methods for training a classifier to discriminate between cancers associated with oncogenic pathogen infections and cancers not associated with oncogenic pathogen infections based on identified genes that are differentially regulated in the two types of cancers. Accordingly, provided is a method for classifying cancers in a patient as either associated with or not associated with oncogenic pathogen infections using the trained classifier. These methods then enable different treatments for patients based on whether their cancers are associated with oncogenic pathogen infections.
[0011] One aspect of the present disclosure provides a method for training a classifier to discriminate between a first cancer state and a second cancer state, wherein the first cancer state is associated with infection by a first oncogenic pathogen and the second cancer state is associated with a state free of oncogenic pathogens. The method includes, on a computer, obtaining a dataset that includes, for each of a plurality of subjects of a certain type, (i) a corresponding plurality of abundance values, wherein each of the corresponding plurality of abundance values quantifies the expression level of a corresponding gene among a plurality of genes in a tumor sample of the respective subject, and (ii) an indicator of the cancer state of the respective subject, which identifies whether the respective subject has the first cancer state or the second cancer state, wherein the plurality of subjects includes a first subset of subjects suffering from the first cancer state and a second subset of subjects suffering from the second state.
[0012] Next, the method includes identifying a discriminant gene set using corresponding abundance values of each of a plurality of targets and respective indicators of cancer status in the plurality of targets, the discriminant gene set including a subset of a plurality of genes.
[0013] In some embodiments, identifying the discriminant gene set includes using a regression algorithm to regress a data set based on all or a subset of the plurality of abundance values across the plurality of targets against the respective indicators of cancer status across the plurality of targets, thereby assigning, for each respective gene in the plurality of genes, the corresponding regression coefficient in the plurality of regression coefficients, and selecting, among the plurality of genes, those genes for the discriminant gene set for which coefficients are assigned by a regression algorithm that satisfies a coefficient threshold.
[0014] In some alternative embodiments, identifying the discriminant gene set includes dividing the data set into a plurality of sets, each set in the plurality of sets including two or more targets suffering from a first cancer status and two or more targets suffering from a second status, independently regressing each respective set in the plurality of sets using a regression algorithm based on all or a subset of the plurality of abundance values across the targets of each respective set against the respective indicator of cancer status across the targets of each respective set, thereby assigning, for each respective gene in the plurality of genes, the corresponding regression coefficient in the plurality of regression coefficients, and selecting, among the plurality of genes, those genes for the discriminant gene set for which coefficients are assigned by a regression algorithm that satisfies a coefficient threshold for at least a threshold percentage of the plurality of sets. In some embodiments, the plurality of sets consists of 5 to 50 sets (e.g., 10 sets).
[0015] In some embodiments, the coefficient threshold is zero. The coefficient threshold is satisfied if the absolute value of the corresponding regression coefficient is greater than zero.
[0016] In some embodiments, the regression algorithm disclosed above is logistic regression. In some such embodiments, logistic regression assumes the following:
Number
[0017] In some embodiments, logistic regression is logistic least absolute shrinkage and selection operator (LASSO) regression. In such embodiments, the logistic LASSO estimator
Number
Number
Number
Number
[0018] In some embodiments, the regression algorithm is logistic regression with L1 or L2 regularization.
[0019] The method further uses the respective abundance values of each of the discriminative gene sets and the respective indicators of the cancer states across a plurality of subjects to classify the first cancer state and the second cancer state as functions of the respective abundance values for the discriminative gene sets (e.g., logistic regression algorithm, neural network algorithm, convolutional neural network algorithm, support vector machine algorithm, naive bayes algorithm, nearest neighbor algorithm, boosting tree algorithm, random forest algorithm, decision tree algorithm, or clustering algorithm).
[0020] Another aspect of the present disclosure provides a method for discriminating between a first cancer state and a second cancer state in a subject, where the first cancer state is associated with infection by a first carcinogenic pathogen and the second cancer state is associated with a state free of carcinogenic pathogens. The method includes obtaining a dataset for the subject, the dataset including a plurality of abundance values, where each respective abundance value among the plurality of abundance values quantifies the expression level of a corresponding gene among a plurality of genes in a cancerous tissue derived from the subject. Next, the method includes inputting the dataset into a classifier trained according to any one of the methodologies described herein.
[0021] Another aspect of the present disclosure provides a plurality of nucleic acid probes for discriminating a first cancer state and a second cancer state in a human subject, wherein the first cancer state is associated with an oncogenic pathogen infection and the second cancer state is associated with a state free of oncogenic pathogens. The nucleic acid probes have nucleic acid sequences that are complementary or identical to the sequences of genes identified as being differentially expressed in cancers associated with oncogenic pathogen infections.
[0022] Another aspect of the present disclosure provides a method for discriminating a first cancer state and a second cancer state in a subject having a first type of cancer, wherein the first cancer state is associated with an infection by a first oncogenic pathogen and the second cancer state is associated with a state free of oncogenic pathogens. The method includes obtaining a dataset for the subject, the dataset having a plurality of abundance values (e.g., relative mRNA expression values), each respective abundance value of the plurality of abundance values quantifying the expression level of a corresponding gene in an identified gene set in the cancerous tissue derived from the subject. Next, the method includes inputting the dataset into a classifier trained to discriminate at least the first cancer state and the second cancer state based on the abundance values for the identified gene set in the cancerous tissue of the subject, thereby determining the cancer state of the subject.
[0023] In some embodiments, the first type of cancer is breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, head and neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer.
[0024] In some embodiments, the dataset further includes mutant allele counts for one or more mutant alleles at one or more loci in the genome of the cancerous tissue derived from the subject.
[0025] In some embodiments, the first cancer state is associated with an infection by a first oncogenic pathogen selected from the group consisting of Epstein-Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T-cell lymphotropic virus (HTLV-1), Kaposi's sarcoma-associated herpesvirus (KSHV), and Merkel cell polyomavirus (MCV).
[0026] In some embodiments, the first cancer state is selected from the group consisting of cervical cancer associated with human papillomavirus (HPV), head and neck cancer associated with HPV, gastric cancer associated with Epstein-Barr virus (EBV), nasopharyngeal cancer associated with EBV, Burkitt lymphoma associated with EBV, Hodgkin lymphoma associated with EBV, liver cancer associated with hepatitis B virus (HBV), liver cancer associated with hepatitis C virus (HCV), Kaposi sarcoma associated with Kaposi's sarcoma-associated herpesvirus (KSHV), adult T-cell leukemia / lymphoma associated with human T-cell lymphotropic virus (HTLV-1), and Merkel cell carcinoma associated with Merkel cell polyomavirus (MCV).
[0027] In some embodiments, the first cancer state is associated with an infection by a human papillomavirus (HPV) oncogenic virus, and the second cancer state is associated with a state that does not include HPV, identified The distinct gene set comprises at least 5 genes selected from the genes listed in Table 3. In some embodiments, the first cancer state is cervical cancer associated with infection by human papillomavirus (HPV). In some embodiments, the first cancer state is head and neck cancer associated with infection by human papillomavirus (HPV). In some embodiments, the discriminatory gene set comprises at least 10 genes selected from the genes listed in Table 3. In some embodiments, the discriminatory gene set comprises at least 20 genes selected from the genes listed in Table 3. In some embodiments, the discriminatory gene set comprises all 24 of the genes listed in Table 3 at least. In some embodiments, the dataset also comprises the mutant allele counts for TP53 (ENSG00000141510) and CDKN2A (ENSG00000147889) in the genome of the cancerous tissue from the subject.
[0028] In some embodiments, the method also comprises, when the results of the classifier indicate that the human cancer patient is infected with an HPV oncogenic virus, administering a first therapy adjusted for the treatment of cervical cancer associated with HPV infection, and, when the results of the classifier indicate that the human cancer patient is not infected with an HPV oncogenic virus, administering a second therapy adjusted for the treatment of cervical cancer not associated with HPV infection, thereby treating the subject for cervical cancer. In some embodiments, the first therapy adjusted for the treatment of cervical cancer associated with HPV infection comprises a therapeutic vaccine or adoptive cell therapy. In some embodiments, the second therapy adjusted for the treatment of cervical cancer not associated with HPV infection is chemotherapy. In some embodiments, the chemotherapy comprises co-administration of cisplatin and a second therapeutic agent selected from the group consisting of 5-fluorouracil, paclitaxel, and bevacizumab.
[0029] In some embodiments, the method also includes, when the classifier result indicates that a human cancer patient is infected with an oncogenic HPV virus, performing a first therapy adjusted for the treatment of head and neck cancer associated with HPV infection, and, when the classifier result indicates that a human cancer patient is not infected with an oncogenic HPV virus, performing a second therapy adjusted for the treatment of head and neck cancer not associated with HPV infection, to treat a subject for head and neck cancer. In some embodiments, the first therapy adjusted for the treatment of head and neck cancer associated with HPV infection includes a therapeutic vaccine, an immune checkpoint inhibitor, or a PI3K inhibitor. In some embodiments, the second therapy adjusted for the treatment of head and neck cancer not associated with HPV infection includes chemotherapy. In some embodiments, the chemotherapy includes administration of cisplatin, and the second therapy also includes concurrent radiotherapy or postoperative chemoradiotherapy.
[0030] In some embodiments, the first cancer state is associated with infection by Epstein-Barr virus (EBV) oncogenic virus, the second cancer state is associated with an EBV-free state, and the discriminatory gene set includes at least five genes selected from the genes listed in Table 4. In some embodiments, the first cancer state is gastric cancer associated with infection by Epstein-Barr virus (EBV). In some embodiments, the discriminatory gene set includes all nine genes listed in Table 4. In some embodiments, the dataset also includes variant allele counts for TP53 (ENSG00000141510) and PIK3CA (ENSG00000121879) in the genome of the cancerous tissue from the subject.
[0031] In some embodiments, the method also includes, when the classifier result indicates that a human cancer patient is infected with an oncogenic EBV virus, performing a first therapy adjusted for the treatment of gastric cancer associated with EBV infection, and, when the classifier result indicates that a human cancer patient is not infected with an oncogenic EBV virus, for the treatment of gastric cancer not associated with EBV infection Implementing a second therapy adjusted thereto, including treating a subject for gastric cancer. In some embodiments, the first therapy adjusted for the treatment of gastric cancer associated with EBV infection includes an immune checkpoint inhibitor. In some embodiments, the second therapy adjusted for the treatment of gastric cancer not associated with EBV infection includes chemotherapy. In some embodiments, chemotherapy includes administering a therapeutic agent selected from the group consisting of paclitaxel, carboplatin, cisplatin, 5-fluorouracil, and oxaliplatin.
[0032] In some embodiments, the method also includes, when the result of the classifier indicates that the human cancer patient is infected with a first oncogenic pathogen, implementing a first therapy adjusted for the treatment of a first type of cancer associated with infection by the first oncogenic pathogen, and, when the result of the classifier indicates that the human cancer patient is not infected with the first oncogenic pathogen, implementing a second therapy adjusted for the treatment of a first type of cancer associated with a state free of oncogenic pathogens, including treating a subject for cancer.
[0033] In some embodiments, a classifier obtains a dataset including (1) for each of a plurality of objects of a certain kind, (i) a corresponding plurality of abundance values, where each abundance value in the corresponding plurality of abundance values quantifies the expression level of a corresponding gene in a plurality of genes in a tumor sample of the respective object, and (ii) an indicator of the cancer state of each object, which identifies whether each object has a first cancer state or a second cancer state, the plurality of objects including a subset of first objects suffering from the first cancer state and a subset of second objects suffering from the second state, and (2) identifies a discriminant gene set using the corresponding plurality of abundance values and the respective indicators of the cancer state for each object among the plurality of objects, the discriminant gene set including a subset of the plurality of genes, and (3) trains the classifier to identify the first cancer state and the second cancer state as a function of the respective abundance values for the discriminant gene set using the respective abundance values and the respective indicators of the cancer state for the discriminant gene set across the plurality of objects, and is trained by a method including these steps.
[0034] Other embodiments are directed to systems, portable consumer devices, and computer-readable media related to the methods described herein.
[0035] As disclosed herein, where applicable, any of the embodiments disclosed herein may be applied to any aspect.
[0036] Additional aspects and advantages of the present disclosure will be readily apparent to those skilled in the art from the following detailed description, where only exemplary embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other and different embodiments and some of the details thereof may be modified in various obvious respects without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive. BRIEF DESCRIPTION OF THE DRAWINGS
[0037]
Fig. 1A
Fig. 1B
Fig. 2A
Fig. 2B
Fig. 2C
Fig. 2D
Fig. 2E
Fig. 3
Fig. 4A
Fig. 4B
Fig. 4C
Fig. 4D
Fig. 5A
Fig. 5B
Fig. 5C
Fig. 5D
Fig. 6A
Fig. 6B
Fig. 7A
Fig. 7B
[0038] Throughout several views of the drawings, like reference numerals refer to corresponding parts. **DETAILED DESCRIPTION OF THE INVENTION**
[0039] The present disclosure provides systems and methods useful for distinguishing cancers associated with oncogenic pathogen infections that contribute to cancer pathology from cancers not associated with oncogenic pathogen infections. The present disclosure further provides systems and methods useful for treating cancer patients based on whether the cancer is associated with an oncogenic pathogen infection.
[0040] Advantageously, the systems and methods described herein enable the detection of oncogenic pathogens in cancer without the need for additional diagnostic assays. Surprisingly, it has been found that oncogenic pathogen infections can be identified based on mRNA expression levels in tumor biopsies. Thus, additional assays developed to identify nucleic acid or protein components of these pathogens are rendered unnecessary by the present disclosure. Rather, a single mRNA expression analysis can be performed to both characterize the transcriptional profile of the cancer and determine whether it is associated with oncogenic pathogen infection. For example, as reported in Example 3, a support vector machine classifier trained on only mRNA expression data and two allelic states identified HPV infection in head and neck cancer and cervical cancer with 99% specificity and 99% sensitivity. Similarly, as reported in Example 4, a support vector machine classifier trained on only mRNA expression data and two allelic states identified EBV infection in gastric cancer with 99% specificity and 95% sensitivity.
[0041] For example, in one aspect, the present disclosure provides a method for training a classifier to discriminate between first and second cancer states, wherein the first cancer state is associated with infection by a first oncogenic pathogen and the second cancer state is associated with a state free of oncogenic pathogens. According to the method, with reference to FIG. 4A, a data set is obtained having corresponding plurality of abundance values for each respective subject in a plurality of subjects of a certain type. Each respective abundance value quantifies the expression level of a corresponding gene in a plurality of genes in a tumor sample of each respective subject. The data set further includes an indicator of the cancer state of each respective subject tracked by the data set. The indicator of the cancer state identifies whether the subject has a first or second cancer state (e.g., HPV-positive head and neck cancer or cervical cancer, or HPV-negative head and neck cancer or cervical cancer, as shown in FIG. 4A).
[0042] In some embodiments, each of the subjects has a particular cancer of the same origin (e.g., gastric cancer as shown in FIG. 5A), and whether the subject is in a first cancer class or a second cancer class is described by whether the subject also has a cancer-causing pathogen associated with this cancer, such that the prognosis of a subject with a cancer who also has a cancer-causing pathogen is different from the prognosis of a subject with a cancer who does not have a cancer-causing pathogen. Some of the subjects (a first subset of subjects) tracked by the dataset have a first cancer state, while some of the subjects (a second subset of subjects) tracked by the dataset have a second state. Next, the discriminant gene set is identified using the corresponding plurality of abundance values for each subject and an indicator for each cancer state in a plurality of subjects. The discriminant gene set includes a subset of a plurality of genes. Generally, the abundance level (e.g., expression) of such genes discriminates between the first and second cancer states. Details regarding the discriminant gene set are disclosed below with reference to block 218 of FIG. 2C. FIG. 4B shows a discriminant gene set for HPV-related cancers (head and neck cancers and cervical cancers), while FIG. 5B shows a discriminant gene set for EBV-related cancers (gastric cancers). A subset) have a first cancer state, while some of the subjects (a second subset of subjects) tracked by the dataset have a second state. Next, the discriminant gene set is identified using the corresponding plurality of abundance values for each subject and an indicator for each cancer state in a plurality of subjects. The discriminant gene set includes a subset of a plurality of genes. Generally, the abundance level (e.g., expression) of such genes discriminates between the first and second cancer states. Details regarding the discriminant gene set are disclosed below with reference to block 218 of FIG. 2C. FIG. 4B shows a discriminant gene set for HPV-related cancers (head and neck cancers and cervical cancers), while FIG. 5B shows a discriminant gene set for EBV-related cancers (gastric cancers).
[0043] Using the respective abundance values for each of a plurality of identification gene sets across a plurality of subjects and the respective indicators of cancer states, a classifier is trained to identify first and second cancer states as a function of the respective abundance values for the identification gene sets. In some optional embodiments, the trained classifier is used to classify a test subject into a first cancer or a second state (or to determine the likelihood that the test subject has a first or second cancer state) by inputting a plurality of abundance values of the test into the trained classifier. In such embodiments, each respective abundance value among the plurality of abundance values of the test quantifies the expression level of a corresponding gene among a plurality of genes in a tumor sample of the test subject. In some optional embodiments, based on the determination that the test subject has a first cancer state or a second cancer state (or the likelihood that the test subject has a first or second cancer state) using the results of the trained classifier, a therapeutic intervention or imaging of the test subject is provided.
[0044] Definitions The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the invention. When used in the description and claims of the present invention, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, the term "and / or" refers to any and all possible combinations of one or more of the recited related items and is to be construed as inclusive. Further, as used herein, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0045] As used herein, the term "if" may be construed to mean "in case" or "when" or "depending on a determination" or "depending on a detection", as the context may require. Similarly, the phrases "if determined" or "if (a stated condition or event) is detected" may be construed to mean "when determining" or "depending on a determination" or "when (a stated condition or event) is detected" or "depending on a detection (of a stated condition or event)", as the context may require.
[0046] Also, terms such as first, second, etc. may be used herein to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the present disclosure, a first object may be referred to as a second object, and similarly, a second object may be referred to as a first object. The first object and the second object are both the same object, but not the same object. Further, the terms "object", "user", and "patient" are used interchangeably herein.
[0047] As used herein, the term "object" refers to any living or non-living organism including, but not limited to, humans (e.g., male humans, female humans, fetuses, pregnant women, children, etc.), non-human mammals, or non-human animals. Any human or non-human animal including, but not limited to, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, genus Bos (e.g., cows), Equidae (e.g., horses), goats and sheep (e.g., sheep, goats), Suidae (e.g., pigs), Camelidae (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), Ursidae plantigrade carnivores (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, sharks can serve as an object. In some embodiments, the object is a male or female (e.g., male, female, or child) at any stage.
[0048] As used herein, the terms "control", "control sample", "reference", "reference sample", "normal", and "normal sample" refer to a sample derived from a subject that does not have a particular condition or is healthy if not. In one example, the methods disclosed herein can be performed on a subject having a tumor, and the reference sample is a sample taken from healthy tissue of the subject. The reference sample can be obtained from the subject or a database. The reference can be, for example, a reference genome used to map sequence reads obtained from sequencing a sample derived from a subject. The reference genome can refer to a haploid or diploid genome that can align and compare sequence reads from biological samples and constitutional samples. An example of a constitutional sample can be DNA of white blood cells obtained from a subject. For a haploid genome, only one nucleotide can be present at each locus. For a diploid genome, heterozygous loci can be identified, each heterozygous locus can have two alleles, and either allele can allow for a match for alignment to the locus.
[0049] As used herein, the term "locus" refers to a position (e.g., site) within the genome, i.e., on a particular chromosome. In some embodiments, the locus refers to a single nucleotide position within the genome, i.e., on a particular chromosome. In some embodiments, the locus refers to a small group of nucleotide positions within the genome, such as defined by a mutation (e.g., substitution, insertion, or deletion) of contiguous nucleotides within the cancer genome. Since normal mammalian cells have a diploid genome, a normal mammalian genome (e.g., the human genome) generally has two copies of all loci in the genome, or at least two copies of all loci on the autosomes, i.e., one copy on the maternal autosome and one copy on the paternal autosome.
[0050] As used herein, the term "allele" refers to a particular sequence of one or more nucleotides at a chromosomal locus.
[0051] As used herein, the term "reference allele" refers to the sequence of one or more nucleotides at a chromosomal locus that is either the dominant allele (e.g., the "wild-type" sequence) represented at that chromosomal locus within a population of the species, or any allele predefined within the reference genome for the species.
[0052] As used herein, the term "mutant allele" refers to the sequence of one or more nucleotides at a chromosomal locus that is not the dominant allele represented at that chromosomal locus within a population of the species (e.g., not the "wild-type" sequence), or is not any allele predefined within the reference genome for the species.
[0053] As used herein, the term "single nucleotide variant" or "SNV" refers to a substitution of one nucleotide for a different nucleotide at a position (e.g., site) in a nucleotide sequence, e.g., a sequence read from an individual. The substitution from the first nucleobase X to the second nucleobase Y can be denoted as "X>Y". For example, an SNV from cytosine to thymine can be denoted as "C>T".
[0054] As used herein, the terms "mutation" or "variant" refer to a detectable change in the genetic material of one or more cells. In certain instances, one or more mutations may be found in cancer cells and can identify cancer cells (e.g., driver and passenger mutations). Mutations can be transmitted from a parent cell to daughter cells. One of ordinary skill in the art will understand that a genetic mutation (e.g., a driver mutation) in a parent cell can induce additional different mutations (e.g., passenger mutations) in daughter cells. Mutations generally occur in nucleic acids. In certain instances, a mutation can be a detectable change in one or more deoxyribonucleic acids or fragments thereof. Mutations generally refer to nucleotides that have been added, deleted, substituted, inverted, or transposed to a new position in a nucleic acid. Mutations can be spontaneous mutations or experimentally induced mutations. A mutation in the sequence of a particular tissue is an example of a "tissue-specific allele". For example, a tumor can have mutations that result in alleles at loci that do not occur in normal cells. Another example of a "tissue-specific allele" is a fetal-specific allele that occurs in fetal tissue but not in maternal tissue.
[0055] As used herein, the terms "cancer", "cancerous tissue", or "tumor" refer to an abnormal mass of tissue whose growth exceeds that of normal tissue and is unregulated. A cancer or tumor can be defined as "benign" or "malignant" according to the following characteristics: the degree of cell differentiation, including morphology and functionality, growth rate, local invasion, and metastasis. "Benign" tumors can be well-differentiated, are characterized by slower growth than malignant tumors, and remain localized to the site of origin. In addition, in some cases, benign tumors do not have the ability to invade, infiltrate, or metastasize to distant sites. "Malignant" tumors can be poorly differentiated (anaplastic), have a characteristically rapid growth that is accompanied by progressive invasion, infiltration, and destruction of surrounding tissues. Furthermore, malignant tumors can have the ability to metastasize to distant sites. Thus, cancer cells are cells found within an abnormal mass of tissue whose growth is not coordinated with that of normal tissue. Thus, a "tumor sample" refers to a biological sample obtained from or derived from a subject's tumor as described herein.
[0056] As used herein, "cancer state associated with oncogenic pathogen infection" generally, or with respect to a particular oncogenic pathogen, refers to a state in which a cancer subject suffering from a particular cancer is further suffering from a pathogen (e.g., virus) known to be associated with the particular cancer.
[0057] As used herein, "cancer state not associated with oncogenic pathogen infection" generally, or with respect to a particular oncogenic pathogen, refers to a state in which a cancer subject suffering from a particular cancer is not particularly suffering from a pathogen (e.g., virus) known to be associated with the particular cancer.
[0058] As used herein, the terms "sequencing", "sequence determination" and similar terms used herein generally refer to any biochemical process that can be used to determine the order of a biopolymer such as a nucleic acid or a protein. For example, sequencing data can include all or part of the nucleotide bases in a nucleic acid molecule such as an mRNA transcript or a genomic locus.
[0059] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence produced by any sequencing process described herein or known in the art. A read can be generated from one end of a nucleic acid fragment ("single-end read") and, in some cases, can also be generated from both ends of the nucleic acid (e.g., paired-end read, double-end read). The length of a sequence read is often a function of a particular sequencing technology is related to. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, the sequence reads are of an average, median, or average length of about 15 bp to 900 bp (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, the sequence reads are of an average, median, or average length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. For example, nanopore sequencing can provide sequence reads that can vary in size from tens to hundreds, thousands of base pairs. Illumina parallel sequencing can provide sequence reads that do not vary as much, for example, most sequence reads can be less than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a series of nucleotides). For example, a sequence read can correspond to a series of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, or to a series of nucleotides at one or both ends of a nucleic acid fragment, or to the nucleotides of an entire nucleic acid fragment. Sequence reads can be obtained in a variety of ways, for example, using sequencing techniques or amplification techniques such as probes, e.g., hybridization arrays or capture probes, or polymerase chain reaction (PCR), or linear amplification using a single primer, or isothermal amplification.
[0060] As used herein, the term "read segment" or "read" refers to any nucleotide sequence that includes the nucleotide sequence derived from an array read obtained from an individual and / or the first array read derived from a sample obtained from an individual. For example, a read segment can refer to an aligned array read, a folded array read, or a stitched read. Further, a read segment can refer to an individual nucleotide base, such as a single nucleotide polymorphism.
[0061] As used herein, the term "read depth", "sequencing depth", or "depth" refers to the total number of read segments derived from a sample obtained from an individual at a given position, region, or locus. The locus can be as small as a nucleotide or as large as a chromosomal arm or the entire genome. The sequencing depth can be expressed as "Yx", for example 50x, 100x, etc., where "Y" refers to the number of times the locus is covered by array reads. In some embodiments, the depth refers to the average sequencing depth across the genome, across the exome, or across a targeted sequencing panel. The sequencing depth can also be applied to multiple loci, the entire genome, in which case "Y" refers to the average number of times the locus or haploid genome, the entire genome, or the entire exome is sequenced respectively. When the average depth is cited, the actual depth for the different loci included in the dataset can span beyond the range of values. Ultra-deep sequencing can refer to at least 100x in the sequencing depth at a locus.
[0062] As used herein, the term "sequencing breadth" refers to what percentage of a particular reference exome (e.g., the human reference exome), a particular reference genome (e.g., the human reference genome), or a portion of an exome or genome was analyzed. The denominator of the percentage can be the repetitive-masked genome, and thus 100% can correspond to all of the reference genome excluding the masked portions. A repetitive-masked exome or genome can refer to an exome or genome in which sequence repeats are masked (e.g., sequence reads align to unmasked portions of the exome or genome). Any portion of the exome or genome can be masked, and thus, any particular portion of the reference exome or genome can be focused on. Broad sequencing refers to sequencing and analyzing at least 0.1% of an exome or genome. refers to sequencing and analyzing at least 0.1% of an exome or genome.
[0063] As used herein, the term "reference exome" refers to any particular known, sequenced, or characterized exome of any organism or pathogen-derived tissue, whether partial or complete, that can be used to reference a specified sequence from a subject. Exemplary reference exomes used for not only human subjects but also many other organisms are provided in Examples 1 and 2.
[0064] As used herein, the term "reference genome" refers to any particular known, sequenced, or characterized genome of any organism or pathogen, whether partial or complete, that can be used to reference a specified sequence from a subject. Exemplary reference genomes used for human subjects and many other organisms are the National It is provided in online genome browsers hosted by the Center for Biotechnology Information (「NCBI」) or the University of California, Santa Cruz (UCSC). A 「genome」 refers to the complete genetic information of an organism or pathogen expressed as a nucleic acid sequence. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genomic sequence from an individual or multiple individuals. In some embodiments, the reference genome is an assembled or partially assembled genomic sequence from one or more human individuals. The reference genome can be considered a representative example of a set of genes of a species. In some embodiments, the reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hg16), NCBI build 35 (UCSC equivalent: hg17), NCBI build 36.1 (UCSC equivalent: hg18), GRCh37 (UCSC equivalent: hg19), and GRCh38 (UCSC equivalent: hg38).
[0065] As used herein, the term 「assay」 refers to a technique for determining the properties of a substance, such as a nucleic acid, protein, cell, tissue, or organ. An assay (e.g., the first assay or the second assay) can include a technique for determining a change in the copy number of a nucleic acid in a sample, the methylation state of a nucleic acid in a sample, the fragment size distribution of a nucleic acid in a sample, the mutation state of a nucleic acid in a sample, or the fragmentation pattern of a nucleic acid in a sample. Any assay known to those of skill in the art can be used to detect any of the nucleic acid properties described herein. The properties of a nucleic acid can include sequence, genomic identity, copy number, methylation state at one or more nucleotide positions, size of the nucleic acid, the presence or absence of a mutation in the nucleic acid at one or more nucleotide positions, and the fragmentation pattern of the nucleic acid (e.g., the nucleotide positions at which the nucleic acid fragments). An assay or method can have a particular sensitivity and / or specificity, and their relative utility as diagnostic tools can be measured using ROC-AUC statistics.
[0066] The term "classification" may refer to any number or other symbol associated with a particular characteristic of a sample. For example, the "+" symbol (or the word "positive") can indicate that the sample is classified as having a deletion or amplification. In another example, the term "classification" can refer to the carcinogenic pathogen infection status, the amount of tumor tissue in a subject and / or sample, the size of the tumor in a subject and / or sample, the stage of the tumor in a subject, the tumor burden in a subject and / or sample, and the presence of tumor metastasis in a subject. The classification can be binary (e.g., positive or negative) or can have more levels of classification (e.g., a scale of 1-10 or 0-1). The terms "cutoff" and "threshold" can refer to a predetermined number used in an operation. For example, the cutoff size can refer to the size above which a fragment is excluded. The threshold can be a value above or below which a particular classification applies. Any of these terms can be used in any of these contexts.
[0067] As used herein, the term "relative abundance" can refer to the ratio of a first amount of nucleic acid fragments having a particular characteristic (e.g., aligning to a particular region of an exome) to a second amount of nucleic acid fragments having the particular characteristic (e.g., aligning to a particular region of an exome). In one example, the relative abundance can refer to the ratio of the number of mRNA transcripts encoding a particular gene (e.g., aligning to a particular region of an exome) in a sample to the total number of mRNA transcripts in the sample.
[0068] As used herein, the term "untrained classifier" refers to a classifier that has not been trained on a training dataset.
[0069] As used herein, the term "effective amount" or "therapeutically effective amount" is an amount sufficient to affect a beneficial or desired clinical result during treatment. An effective amount can be administered to a subject in one or more dosages. For treatment, an effective amount is an amount sufficient to alleviate, ameliorate, stabilize, reverse, or delay the progression of a disease or otherwise reduce the pathological consequences of the disease. An effective amount is generally determined individually by a physician and is within the skill of the art. When determining an appropriate dosage to achieve an effective amount, several factors are usually considered. These factors include the age, sex, and weight of the subject, the condition being treated, the severity of the condition, and the form and effective concentration of the therapeutic agent being administered.
[0070] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly has a condition. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects within a population that have cancer. In another example, sensitivity can characterize the ability of a method to correctly identify one or more markers indicative of cancer.
[0071] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to correctly identify the proportion of a population that truly does not have a condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects within a population that do not have cancer. In another example, specificity can characterize the ability of a method to correctly identify one or more markers indicative of cancer.
[0072] The terms used in this disclosure are for the purpose of describing particular cases only and are not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Further, the terms "including," "includes," "having," "has," "with," or variants thereof, as used in any detailed description and / or claims, are intended to be inclusive in the same manner as the term "comprising."
[0073] Some aspects will be described below with reference to application examples for illustration. It should be understood that numerous specific details, relationships, and methods are shown in order to provide a complete understanding of the features described herein. However, one of ordinary skill in the art will readily recognize that the features described herein may be practiced without one or more of the specific details, or in other ways. Since some acts can occur in a different order and / or concurrently with other acts or events, the features described herein are not limited by the illustrated order of acts or events. Furthermore, not all acts or events illustrated are required to implement the methodology in accordance with the features described herein.
[0074] Referring now to embodiments in detail, examples thereof are shown in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks are not described in detail so as not to obscure aspects of the embodiments.
[0075] Examples of System Embodiments Since an overview of some aspects of the present disclosure and some definitions used in the present disclosure have been provided, next, the details of an exemplary system will be described in conjunction with FIG. 1. FIG. 1 is a block diagram showing a system 100 according to some implementations. The device 100 in some implementations includes one or more processing units CPU 102 (also referred to as processors), one or more network interfaces 104, a user interface 106, a non-persistent memory 111, a persistent memory 112, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include a circuit (sometimes referred to as a chipset) for interconnecting and controlling communications between system components. The non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while the persistent memory 112 typically includes CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. The persistent memory 112 optionally includes one or more storage devices located remotely from the CPU 102. The persistent memory 112 and the non-volatile memory devices within the non-persistent memory 112 comprise non-transitory computer-readable storage media. In some implementations, the non-persistent memory 111, or the non-transitory computer-readable storage media, stores the following programs, modules, data structures, or subsets thereof, optionally in combination with the persistent memory 112. · Any operating system 116 including procedures for processing various basic system services and executing tasks that are hardware-dependent; · Any network communication module (or instructions) 118 for connecting the system 100 to other devices and / or communication network 105; ·Any classifier training module 120 for training a classifier to distinguish a first cancer state associated with an oncogenic pathogen infection from a second cancer state not associated with an oncogenic pathogen infection; ·Any data storage for a dataset for a tumor sample from a training subject 122, including expression data from one or more training subjects 124, the expression data including abundance data for each of a plurality of genes 126, supporting for each of one or more genes 127 and a plurality of mutant alleles for each cancer state 128; ·Any classifier verification module 130 for verifying a classifier that distinguishes a first cancer state associated with an oncogenic pathogen infection from a second cancer state not associated with an oncogenic pathogen infection; ·Any data storage for a dataset for a tumor sample from a verification subject, including expression data from one or more training subjects, the expression data including abundance data for each of a plurality of genes and each cancer state; ·Any patient classification module 134 for classifying cancer in a patient as either a first cancer state associated with an oncogenic pathogen infection or a second cancer state not associated with an oncogenic pathogen infection, using a classifier, e.g., one trained using classifier training module 120; ·Any data storage for a data construct for a cancer patient 136, including expression data from one or more cancer patients 140 where the expression data includes abundance data for each of a plurality of genes 142; and ·Any data storage for a data construct for a cancer patient 138, including mutant allele data from one or more cancer patients 144, the mutant allele data including a plurality of supports for mutant alleles for each of one or more genes 146.
[0076] In various implementations, one or more of the elements identified above are stored in one or more of the aforementioned memory devices and correspond to a set of instructions for performing the functions described above. The modules, data, or programs (e.g., sets of instructions) identified above need not be implemented as separate software programs, procedures, data sets, or modules, and thus, various subsets of these modules and data can be combined or otherwise reconfigured in various implementations. In some implementations, the non-persistent memory 111 optionally stores a subset of the modules and data structures identified above. Further, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the elements identified above are stored in a computer system other than the visualization system 100 and are addressable by the visualization system 100 such that the visualization system 100 can retrieve all or a portion of such data when needed.
[0077] FIG. 1 shows the "System 100", but the figure is intended as a functional illustration of various features that may exist in a computer system rather than a structural schematic of the implementations described herein. In fact, and as will be recognized by those of skill in the art, the items shown separately can be combined and some items can be separated. Further, FIG. 1 shows certain data and modules within the non-persistent memory 111, but some or all of these data and modules can be in the persistent memory 112.
[0078] Classifier Training The system according to the present disclosure is disclosed with reference to FIG. 1, while an overview of the method according to the present disclosure is provided in conjunction with FIG. 2A. In block 204 of FIG. 2A, a data set is obtained. The data set includes corresponding abundance values for each of a plurality of objects in a certain type of plurality of objects. Each of the abundance values quantifies the expression level of a corresponding gene in a plurality of genes in the tumor sample of each object. The data set further includes an indicator of the cancer state of each object tracked by the data set. The indicator of the cancer state identifies whether the object has a first or second cancer state.
[0079] In some embodiments, each of the objects has a specific cancer (e.g., gastric cancer) with the same origin, and depicting whether the object is in a first cancer class or a second cancer class is based on whether the object also suffers from a carcinogenic pathogen known to be associated with this cancer, such that the prognosis of an object with cancer who also suffers from a carcinogenic pathogen is different from the prognosis of an object with cancer who does not suffer from a carcinogenic pathogen. For example, if the specific carcinogenic pathogen is Epstein - Barr virus (EBV), each of the objects has a gastric tumor, and determining whether each object is in a first or second cancer class is based on whether the object also suffers from EBV.
[0080] In some embodiments, each of the objects has a cancer associated with a series of cancers, and depicting whether the object is in a first cancer class or a second cancer class is based on whether the prognosis of those objects having each cancer in a series of cancers suffering from a carcinogenic pathogen is different from the prognosis of those objects having each cancer not suffering from a carcinogenic pathogen in this series of cancers Whether the subject is also affected by the oncogenic pathogen known to be related to any of the cancers in each of the cancers in.For example, if the specific oncogenic pathogen is human papillomavirus (HPV), the series of cancers is head and neck squamous cell carcinoma and cervical cancer.That is, each subject has head and neck squamous cell carcinoma or cervical cancer, and whether each subject is the first or second cancer class is determined by whether the subject is also infected with HPV.
[0081] In some embodiments, each of the subjects has a cancer listed in column 2 of the same row in Table 1 below, and what delineates whether a subject is in a first or second cancer class is whether the subject is also afflicted with an agent listed in column 1 of the same row in Table 1 below. See, e.g., Flora and Bonanni, Carcinogenesis 32(6), pp. 787-795, which is incorporated herein by reference. [Table 1]
[0082] As used herein, the term "human gut microbiota" refers to all of the microorganisms that inhabit the human digestive tract, a subset of which have been found to be carcinogenic, e.g., diseases that are hypothesized to cause or correlate with colon or colorectal cancer. Prodrugs include sulfide-producing bacteria (e.g., Fusobacterium, Desulfovibrio, and Bilophila wadsworthia), Streptococcus bovis, and Fusobacterium nucleatum. For details, see Dahmus et al., 2018, J Gastrointest Oncol., 9(4), pp. 769-77, the contents of which are incorporated herein in their entirety for all purposes.
[0083] Some of the subjects tracked by the dataset (a first subset of subjects) have the first cancer state, while some of the subjects tracked by the dataset (a second subset of subjects) have the second state. Details regarding such a dataset are disclosed below with reference to block 202 of FIG. 2B.
[0084] Next, in block 218 of FIG. 2A, the discriminant gene set is identified using the corresponding plurality of abundance values of each subject and the respective indicators of the cancer state in a plurality of subjects. The discriminant gene set includes a subset of a plurality of genes. Generally, the abundance levels (e.g., expression) of such genes distinguish between the first cancer state and the second cancer state. Details regarding the discriminant gene set are disclosed below with reference to block 218 of FIG. 2C.
[0085] Next, in block 242 of FIG. 2A, a classifier is trained to distinguish between the first and second cancer states as a function of the respective abundance values of the discriminant gene set using the respective abundance values of the discriminant gene set across a plurality of subjects and the respective indicators of the cancer state. Details regarding the training of such a classifier based on the discriminant gene set are disclosed below with reference to block 242 of FIG. 2E.
[0086] Further, referring to block 246 of FIG. 2A, in some optional embodiments, a trained classifier is used to classify a test subject into a first cancer or a second condition (or determine the likelihood that the test subject has a first or second cancerous condition) by inputting a plurality of abundance values of a test into the classifier. In such embodiments, each respective abundance value among the plurality of abundance values of the test quantifies the expression level of a corresponding gene among a plurality of genes in a tumor sample of the test subject. The test subject is a subject whose plurality of abundance values of the test were not used to train the classifier. Further, in a typical example, the test subject is a subject whose status as having a first or second cancerous condition has not been confirmed. Details regarding the diagnosis of a test subject using a trained classifier according to the present disclosure are disclosed below with reference to block 246 of FIG. 2E.
[0087] Further, referring to block 248 of FIG. 2A, in some optional embodiments, the results of a trained classifier are used to provide a therapeutic intervention or imaging for a test subject based on the determination that the test subject has a first cancerous condition or a second cancerous condition (or the likelihood that the test subject has a first or second cancerous condition). Details regarding such treatment options resulting from the application of a trained classifier to abundance data of a plurality of test genes are disclosed below with reference to block 248 of FIG. 2E.
[0088] Since an overview of the disclosed method has been provided in relation to FIG. 2A, attention is now turned to FIGS. 2B - 2E, which provide further details regarding the disclosed method.
[0089] Block 202. Referring to block 202 of FIG. 2A, a method is provided for training a classifier to distinguish between a first cancerous condition and a second cancerous condition. As discussed above, the first cancerous condition is associated with an infection by a first carcinogenic pathogen, and the second cancerous condition is associated with a condition that does not include a carcinogenic pathogen. Cancers not known to be associated with carcinogenic pathogen infection A limited example will be described below with reference to FIG. 3. Thus, in some embodiments, the first cancer state is a particular type of cancer associated with a particular carcinogenic pathogen infection, as described below, and the second cancer state is the same particular type of cancer not associated with a particular carcinogenic pathogen infection. For example, in one embodiment, the first cancer state is cervical cancer associated with HPV infection, and the second cancer state is cervical cancer not associated with pathogen infection.
[0090] Block 204. Referring to block 204 of FIG. 2A, a data set is obtained that includes a corresponding plurality of abundance values for each respective subject in a plurality of subjects of a single species. Each respective abundance value in the corresponding plurality of abundance values quantifies the expression level of a corresponding gene in a plurality of genes in the tumor sample of each respective subject. The data set further includes an indicator of the cancer state of each respective subject. The indicator of the cancer state identifies whether each respective subject has a first or second cancer state. The plurality of subjects includes a first subset of subjects suffering from the first cancer state and a second subset of subjects suffering from the second state.
[0091] Block 206. Referring to block 206, in some embodiments, the corresponding plurality of abundance values are obtained by RNA-seq. RNA-seq is a methodology for RNA profiling based on next-generation sequencing that enables the measurement and comparison of gene expression patterns across multiple subjects. In some embodiments, millions of short sequences called "sequence reads" are generated by sequencing random positions of cDNA prepared from input RNA obtained from a subject's tumor tissue. These reads are then computationally mapped to a reference genome to reveal a "transcription map," and the number of sequence reads aligned to each gene provides a measure of its expression level (e.g., abundance). Next-generation sequencing is disclosed in Shendure, 2008, "Next-generation DNA sequencing," Nat. Biotechnology 26, pp. 1135-1145, which is incorporated herein by reference. RNA-seq is disclosed in Nagalakshmi et al., 2008, "The transcriptional landscape of the yeast genome defined by RNA sequencing," Science 320, pp. 1344-1349, and Finotell and Camillo, 2014, "Measuring differential gene expression with RNA-seq: challenges and strategies for data analysis," Briefings in Functional Genomics 14(2), pp. 130-142, each of which is incorporated herein by reference.
[0092] According to block 206, for each tumor sample of each subject among a plurality of subjects, the RNA in the sample of interest is first fragmented and reverse transcribed into complementary DNA (cDNA). The resulting cDNA is then amplified and subjected to next-generation DNA sequencing (NGS). In principle, for RNA-seq, any NGS technology can be used. In some embodiments, an Illumina sequencing device (see the internet at illumina.com) is used. Wang, Z., et al., “RNA-Seq: a revolutionary tool for transcriptomics,” Nat Rev Genet., 10(1):57-63(2009) is hereby incorporated by reference, which is incorporated herein by reference. Then, millions of short reads generated for each such sample are mapped to a reference genome, and the number of reads aligned to each gene, called “count,” provides a digital measurement of the gene expression level in the sample under investigation.
[0093] In some alternative embodiments, instead of using RNA-seq, micro arrays are used to measure gene abundance values. Such microarrays are described in Wang et al., 2009, “RNA-Seq: a revolutionary tool for transcriptomics,” Nat Rev Genet 10, pp.57-63, Roy et al., 2011, “A comparison of analog and next-generation transcriptomic tools for mammalian studies,” Brief Funct Genomic 10:135-150, Shendure, 2008, "The beginning of the end for microarrays?," Nat Methods 5, pp. 585-587, Cloonan et al., 2008, "Stem cell transcriptome profiling via massive-scale mRNA sequencing," Nat.Methods 5, pp. 613-619, Mortazavi et al., 2008, "Mapping and quantifying mammalian transcriptomes by RNA-Seq," Nat Methods 5, pp. 621-628, and Bullard et al., 2010, "Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments" BMC Bioinformatics 11, p. 94 are disclosed, and each of them is incorporated herein by reference.
[0094] The first computational step in an RNA-seq data analysis pipeline is read mapping, where reads are aligned to a reference genome or transcriptome by identifying the gene regions that match the read sequences. For this task, any of a variety of alignment tools can be used. For example, Hatem et al., 2013, "Benchmarking short sequence mapping tools," BMC Bioinformatics 14, p. 184, and Engstrom et al., "Systematic evaluation of See spliced alignment programs for RNA-seq data, Nat Methods 10, pp. 1185-1191, each of which is incorporated herein by reference. In some embodiments, the mapping process begins by constructing an index of either the reference genome or the reads, and then uses that to search for a set of positions in the reference sequences where the reads are likely to align. Once a subset of these possible mapping positions is identified, alignment is performed on these candidate regions using a slower and more sensitive algorithm. See, for example, Hatem et al., 2013, “Benchmarking short sequence mapping tools,” BMC Bioinformatics 14: p. 184, and Flicek and Birney, 2009, “Sense from sequence reads: methods for alignment and assembly,” Nat Methods 6 (Suppl. 11), S6-S12, each of which is incorporated herein by reference. In some embodiments, the mapping tool is a methodology that utilizes a hash table or the Burrows-Wheeler transform (BWT). See, for example, Li and Homer, 2010, “A survey of sequence alignment algorithms for next-generation sequencing,” Brief Bioinformatics 11, pp. 473-483, which is incorporated herein by reference.
[0095] After mapping, the reads aligned to each coding unit, such as an exon, transcript, or gene, are used to provide an estimate of its abundance (e.g., expression) level Use to calculate the count. In some embodiments, such a count takes into account the total number of reads that overlap with exons of a gene. However, in some examples, since a portion of the sequence reads are mapped outside the boundaries of known exons, alternative embodiments take into account the full length of the gene and also count reads derived from introns. Further, in some embodiments, spliced reads are used to model the abundance of different splicing isoforms of a gene. For example, see Trapnell et al., 2010, “Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation,” Nat Biotechnol 28, pp. 511-515, and Gatto et al, 2014, “Fine-Splice, enhanced splice junction detection and quantification: a novel pipeline based on the assessment of diverse RNA-Seq alignment solutions,” Nucleic Acids Res 42, p.e71, each of which is incorporated herein by reference.
[0096] As described above, quantification of gene abundances from RNA-seq data is typically implemented in an analysis pipeline through two computational steps: alignment of reads to a reference genome or transcriptome, and subsequent estimation of gene and isoform abundances based on the aligned reads. Unfortunately, reads generated by the most commonly used RNA-Seq technologies are generally much shorter than the transcripts from which they were sampled. As a result, in the presence of transcripts with similar sequences, it is not always possible to uniquely assign short sequence reads to a particular gene. Such sequence reads are referred to as “multi-reads” because they are homologous to two or more regions of the reference genome. In some embodiments, such multi-reads are discarded, i.e., they do not contribute to gene abundance counts. In some embodiments, programs such as MMSEQ or RSEM are used to resolve the ambiguity. See, e.g., Turro et al., 2011, “Haplotype and isoform specific expression estimation using multi-mapping RNAseq reads,” Genome Biol 12, p. R13, and Nicolae et al., “Estimation of alternative splicing isoform frequencies from RNA-Seq data,” Algorithms Mol Biol 6, p. 9, each of which is incorporated herein by reference.
[0097] Another aspect of RNA-seq is the normalization of sequence read counts. In some embodiments, this includes normalization to account for different sequencing depths. For example, see Lin et al., 2011, “Comparative studies of de novo assembly tools for next-generation sequencing technologies,” Bioinformatics 27, pp. 2031-2037, Robinson Oshlack, 2010, “A scaling normalization method for differential expression analysis of RNA-seq data,” Genome Biol 11, p. R25, and Li et al., 2012, “Normalization, testing, and false discovery rate estimation for RNA-sequencing data, Biostatistics 13, pp. 523-538 which are each incorporated herein by reference. In some embodiments, the sequence read counts are normalized to account for gene length bias. See Finotell and Camillo, 2014, “Measuring differential gene expression with RNA-seq: challenges and strategies for data analysis,” Briefings in Functional Genomics 14(2), pp. 130-142 which is incorporated herein by reference.
[0098] Block 208. Referring to block 208 of FIG. 2B, in some embodiments, each of the plurality of subjects has cancer of a first type. In other words, in some embodiments, each subject in database 122 has cancer of the same type. In some such embodiments, each of the plurality of subjects has breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, head and neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer.
[0099] Block 210. Referring to block 208 of FIG. 2B, in some embodiments, each of the plurality of subjects has cancer of a first type at a first stage. In other words, in some embodiments, each subject in database 122 has cancer of the same type, and the cancer is at the same stage. In some such embodiments, each of the plurality of subjects has breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, head and neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer. Further, in such embodiments, the stage of the cancer in each of the plurality of subjects is stage I, II, III, or IV cancer.
[0100] Blocks 212-214. Referring to block 212 of FIG. 2B and block 214 of FIG. 2C, the cohorts used in the disclosed method are of a size sufficient to develop a classifier having performance suitable for screening subjects to determine whether they have a first or second cancerous state. Thus, in some embodiments, the plurality of subjects includes 100 subjects, a first subset of subjects (those having a first cancerous state) includes 20 subjects, and a second subset of subjects (those having a second cancerous state) includes 20 subjects. This is just one example. In other embodiments, the plurality of subjects includes 1000 subjects, the first subset of subjects includes 100 subjects, and the second subset of subjects includes 100 subjects. In still other embodiments, the plurality of subjects includes 100, 500, 2000, 4000, or 10000 subjects, the first subset of subjects includes 100, 500, or 1000 subjects, and the second subset of subjects is composed of 100, 500, or 1000 subjects. In some embodiments, more of the subjects have a first cancerous state than a second cancerous state. For example, in some embodiments, more than 10%, more than 20%, more than 30%, more than 40%, more than 50%, more than 60%, more than 70%, more than 80%, or more than 90% of the subjects in dataset 122 have a first cancerous state and the remainder have a second cancerous state.
[0101] Block 216. Referring to block 216 of FIG. 2C, in some embodiments, the disclosed method is used with training subjects that are human. Each training subject in dataset 122 is of the same species, although the species need not be human. In some embodiments, the species is dog, cow, pig, or some other species.
[0102] Block 218. Referring to block 218 of FIG. 2C, a dataset 122 is obtained that includes corresponding plurality of abundance values for each respective subject in a plurality of subjects of a single species. When this occurs, the dataset 122 is used to identify a discriminant gene set using the abundance value of each target and each indicator of the cancer state in a plurality of targets of the dataset 122. The discriminant gene set includes a subset of a plurality of genes. A specific method for identifying a discriminant gene set according to some embodiments of the present disclosure is described in detail below with reference to blocks 226-240.
[0103] Blocks 220-224. Referring to block 220 of FIG. 2C, in some embodiments, the species under consideration is human, and the plurality of genes (for which abundance data is considered) includes 10,000 or more genes. For example, the xGen Exome Research Panel v1.0 (IDT) spans a 39 Mb target region that includes 19,396 genes (Nguyen, A., et al., “Multiplexed See “Hybrid Capture for Whole Exome Sequencing,” Technical Note, Integrated DNA Technologies, Inc., (2018), the contents of which are hereby incorporated by reference in their entirety for all purposes. The set of identified genes consists of 5 to 40 genes. Referring to block 222 in FIG. 2C, in some embodiments, the species is human, the plurality of genes includes 5000 genes, and the set of identified genes consists of 5 to 25 genes. Other ranges are possible. For example, in some embodiments, the plurality of genes (where abundance data is considered) includes at least 200, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 15000, or 20000 genes, and the set of identified genes consists of 5 genes to 500 genes, 5 genes to 100 genes, 5 genes to 50 genes, or 5 genes to 20 genes. Regardless of the range, the range of identified genes is smaller than the range of the original plurality of genes. In some embodiments, the identified set consists of at least one quarter of the genes of the plurality of genes in dataset 122 (e.g., a reduction from 1000 genes to 250 or fewer genes). By selecting a gene set for the identified gene set that is smaller than what is available in dataset 122, an algorithm for distinguishing between the first and second states can be trained with smaller, more useful data (e.g., abundance data for fewer genes), which leads to a more computationally efficient training of a classifier that distinguishes between the first and second cancer states. Such an improvement in computational efficiency due to a decrease in the size of the identified gene set can advantageously be used to speed up classifier training or to improve the performance of such a classifier (e.g., through more extensive training of the classifier).In some embodiments, the set of identifying genes consists of at least one quarter, one fifth, one sixth, one seventh, one eighth, one ninth, one tenth, one twentieth, one thirtieth, one fortieth, or one fiftieth of the plurality of genes within the dataset 122. Further, reducing the number of genes used in the analysis improves the model by preventing overfitting of the data.
[0104] Block 226. Referring to block 226 of FIG. 2C, in some embodiments, identification of the set of identifying genes involves using a regression algorithm to regress the dataset 122 based on all or a subset of the plurality of abundance values 126 across the plurality of training subjects 124 against the respective indicators of the cancer state 128 across the plurality of training subjects 124, thereby assigning the corresponding regression coefficients in the plurality of regression coefficients to each respective gene in the plurality of genes. Thus, in such embodiments, the cancer state is the dependent variable and the abundance values of the genes are the independent variables. In such embodiments, the genes from the plurality of genes selected for the set of identifying genes are those to which coefficients have been assigned by a regression algorithm that meets a coefficient threshold. In such embodiments, genes whose coefficients meet the coefficient threshold are considered to be sufficiently significant to have a substantial effect on the dependent variable, the cancer class, and are thus retained for the set of identifying genes Details of such regression in certain embodiments of the present disclosure are presented below.
[0105] Blocks 228-232. Referring to block 228 in FIG. 2D, in some embodiments, identifying a set of discriminant genes includes dividing a dataset into a plurality of sets (e.g., 5-50 sets, exactly 10 sets, etc.). Each set in the plurality of sets includes two or more subjects suffering from a first cancer state and two or more subjects suffering from a second cancer state. Next, each respective set in the plurality of sets is independently regressed using a regression algorithm based on all or a subset of the plurality of abundance values across the subjects in that set for each respective indicator of the cancer state across the subjects in that set, thereby assigning the corresponding regression coefficient in the plurality of regression coefficients to each respective gene in the plurality of genes. Those genes to which regression coefficients are assigned by a regression algorithm that satisfies a coefficient threshold for at least a threshold percentage of the plurality of sets are selected for the discriminant gene set. Referring to block 230, in some embodiments, the coefficient threshold is zero. In some embodiments, the required threshold percentage is at least 40 percent of the plurality of sets. Thus, for purposes of illustration, consider the case where there are 10 sets. In such a case, for gene A to be included in the discriminant gene set, the regression coefficient for gene A must satisfy the regression threshold in 4 of the 10 sets upon regression of each of the 10 sets against the cancer state. If the regression threshold is zero, this means that a positive regression coefficient is required to satisfy the regression threshold, and the regression coefficient for gene A must be positive in at least 4 of the 10 sets. In some embodiments, the threshold is applied to the absolute value of the coefficient. However, in some embodiments described herein, since LASSO regression is designed to return sparse coefficients, the threshold is set to 0. In some embodiments, the required threshold percentage is at least 50 percent, at least 60 percent, at least 70 percent, at least 80 percent, at least 90 percent, or all of the plurality of sets.Referring to block 232, in some embodiments, the regression coefficient threshold is greater than zero (e.g., 0.1, 0.2, 0.3, or some other positive value). It will be appreciated that requiring a larger regression coefficient helps to increase the rigor of what is required for a gene to be included in the discriminant data set. In various alternative embodiments, if the absolute value of the regression coefficient at regression is other than zero, greater than 0.1, or greater than 0.2, the regression coefficient meets the regression coefficient threshold.
[0106] Blocks 234 - 240. Note that the dependent variable used in identifying the set of discriminant genes employs one of two labels, the first cancer state or the second cancer state. Thus, referring to block 234 in FIG. 2D, in some embodiments, the regression algorithm is a logistic regression that assumes the following.
Number
[0107] Referring to block 238, in some embodiments, the logistic regression is a logistic least absolute shrinkage and selection operator (LASSO) regression. In such embodiments, the logistic LASSO estimator
Number
Number
Number
Number
[0108] In some embodiments, a regularization method other than LASSO is used to identify genes in a plurality of genes that discriminate between the first and second cancer states based on gene abundance values across the training subjects 124 of the dataset 122. For example, in some embodiments, the elastic net is used to identify genes in a plurality of genes that discriminate between the first and second cancer states based on gene abundance values across the training subjects 124 of the dataset 122. Zou and Hastie, 2005, “Regularization and variable selection via the elastic net,” See R Stat Soc Series B Stat Methodol 67, pp. 301 - 320, which is incorporated herein by reference. In some embodiments, a sparse Laplacian penalty is used to identify genes in a plurality of genes that distinguish between a first and a second cancer state based on gene abundance values across training subjects 124 of dataset 122. Huang et al., 2011, “The sparse Laplacian shrinkage estimator for high - dimensional regression, Ann Stat 39, pp. 2021 - 2046 is referred to, which is incorporated herein by reference. In some embodiments, elastic net, group LASSO (Yuan and Lin, 2006, “Model Selection and Estimation in Regression with Grouped Variables,” Journal of the Royal Statistical Society. Series B Statistical Methodology 68(1), pp. 49 - 67), fused LASSO (Tibshirani et al., 2005, “Sparsity and Smoothness via the Fused lasso,” Journal of the Royal Statistical Society. Series B Statistical Methodology 67(1), pp. 91 - 108), quasi - norm and bridge regression (Fu, 1998, “The Bridge versus the Lasso,” Journal of Computational and Graphical Statistics 7(3), pp. 397 - 416), or adaptive LASSO is used to identify genes in a plurality of genes that distinguish between a first and a second cancer state based on gene abundance values across training subjects 124 of dataset 122. Referring to block 240 of FIG. 2E, in some embodiments, the regression algorithm includes an L1 (LASSO) or L2 (Ridge) regularization term.
[0109] Blocks 242 to 244. The above disclosure details how the gene abundance values 126 of the subject 124 in the training set 122 are used to identify a discriminant gene set whose abundance values collectively distinguish between the first and second cancer states. Once this discriminant gene set is identified, the training set 122 is used to formally train a classifier that can distinguish between the first and second cancer states of a test subject using the abundance values of the discriminant genes measured from a biological sample taken from the test subject. In a typical embodiment, the cancer state of this test subject is unknown. That is, it may be known that the test subject has a specific cancer, but it is not known whether the subject is suffering from a pathogen that has an adverse effect on the prognosis of the subject's cancer. In a typical embodiment, the biological sample used to measure the gene abundance values of the test subject is a solid tumor within the test subject. Referring to block 242, in some embodiments, for each abundance value of the discriminant gene set and each indicator of the cancer state across a plurality of subjects, a classifier is trained to distinguish between the first cancer state and the second cancer state as a function of each abundance value of the discriminant gene set. In some embodiments, as disclosed in the following examples, additional features are utilized in addition to the abundance values of the discriminant gene set to train the classifier. For example, in some embodiments, the absence of a specific mutation in a selected gene is also used in combination with the abundance values of the discriminant gene set to train the classifier.
[0110] Referring to block 244 of FIG. 2E, in some embodiments, by way of non-limiting example, the classifier used in block 242 is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine (SVM) algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a decision tree algorithm, a clustering algorithm, or a combination thereof.
[0111] A logistic regression algorithm suitable for use as a classifier in block 242 is disclosed, for example, in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103-144, John Wiley & Son, New York, which is incorporated by reference.
[0112] A neural network algorithm including a convolutional neural network algorithm suitable for use as a classifier in block 242 is disclosed, for example, in Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408, Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Disclosed in Mach Learn Res 10, pp. 1 - 40 and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference. A neural network has a layered structure including a layer of input units (and biases) connected by layers of weights to a layer of output units. In the case of regression, the layer of output units typically includes only one output unit. However, a neural network can process multiple quantitative responses in a seamless manner. In a multi - layer neural network, there are input units (input layer), hidden units (hidden layer), and output units (output layer). Further, there is a single bias unit connected to each unit other than the input units. Additional exemplary neural networks suitable for use as the classifier of block 242 are disclosed in Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, Inc., New York, and Hastie et al., 2001, The Elements of Statistical Learning, Springer - Verlag, New York, each of which is incorporated herein by reference in its entirety. Additional exemplary neural networks suitable for use as the classifier of block 242 are described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC, and Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is incorporated herein by reference in its entirety.
[0113] The SVM algorithm suitable for use as a classifier for block 242 is described, for example, in Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge, Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5 th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152, Vapnik, 1998, Statistical Learning Theory, Wiley, New York, Mount , 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265, and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is hereby incorporated by reference in its entirety. When used for classification, the SVM separates a given set of binary-labeled data training sets (here, the first and second cancer states of each subject in dataset 122) with a hyperplane that is maximally distant from the labeled data. If linear separation is not possible, the SVM functions in combination with a `kernel' technique that automatically implements a non-linear mapping into the feature space. The hyperplane found by the SVM in the feature space corresponds to a non-linear decision boundary in the input space.
[0114] A Naive Bayes classifier suitable for use as the classifier of block 242 is, for example, Ng et al., 2002, “On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes,” Advances in Neural Information Processing Systems, 14, which is hereby incorporated by reference.
[0115] A decision tree algorithm suitable for use as the classifier of block 242 is, for example, Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395 - 396, which is hereby incorporated by reference. The tree - based method partitions the feature space into a set of rectangles and fits a model (such as a constant) in each. In some embodiments, the decision tree is a random forest regression. One particular algorithm that can be used as the classifier in block 244 is the Classification and Regression Tree (CART). Other examples of specific decision - tree algorithms that can be used as the classifier in block 244 include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396 - 408 and pp. 411 - 412, which is hereby incorporated by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer - Verlag, New York, Chapter 9, which is hereby incorporated by reference in its entirety. Random forest is described in Breiman, 1999, “Random Forests - Random Features,” Technical Report 567, Statistics Department, U.C. Berkeley, September 1999, which is hereby incorporated by reference in its entirety.
[0116] Clustering algorithms suitable for use as the classifier in block 242 are described, for example, on pages 211 - 256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter “Duda 1973”), which is hereby incorporated by reference in its entirety. As described in section 6.7 of Duda 1973, the clustering problem is the data set One way of finding natural groupings in data is described. To identify natural groupings, two problems are addressed. First, a method for measuring the similarity (or dissimilarity) between two samples is determined. This measure (scale of similarity) is used to confirm that samples in one cluster are more similar to each other than samples in the other cluster. Here, the scale of similarity is at the abundance level of the discriminant gene set across the training dataset 122. Next, a mechanism for dividing the data into clusters using the scale of similarity is determined. The scale of similarity is described in Section 6.7 of Duda 1973, and one way to start a clustering investigation is to define a distance function and calculate a matrix of distances between all pairs of samples in the dataset. If distance is an appropriate scale of similarity, the distances between samples in the same cluster will be significantly shorter than the distances between samples in different clusters. However, as described on page 215 of Duda 1973, it is not necessary to use a distance measure in clustering. For example, a non-metric similarity function s(x, x') can be used to compare two vectors x and x'. Conventionally, s(x, x') is a symmetric function that increases in value when x and x' are "similar" in some way. Examples of non-metric similarity functions s(x, x') are provided on page 216 of Duda 1973.
[0117] Once a method for measuring "similarity" or "dissimilarity" between points in the dataset is selected, clustering utilizes a criterion function that measures the clustering quality of any partition of the data. The partition of the dataset that maximizes the criterion function is used to cluster the data. See page 217 of Duda 1973. Criterion functions are discussed in Section 6.8 of Duda 1973. More recently, Duda et al., Pattern Classification, 2 ndThis was published by John Wiley & Sons, Inc., New York. Pages 537 - 563 provide a detailed explanation of clustering. Details of clustering techniques suitable for use as the classifier in block 242 can be found in Kaufman and Rousseeuw, 1990, Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, N.Y., Everitt, 1993, Cluster analysis (3d ed.), Wiley, New York, N.Y., and Backer, 1995, Computer - Assisted Reasoning in Cluster Analysis, Prentice Hall, Upper Saddle River, N.J. Specific exemplary clustering techniques that can be used as the classifier in block 242 include, but are not limited to, hierarchical clustering (agglomerative clustering using the nearest - neighbor algorithm, farthest - neighbor algorithm, average - linkage algorithm, centroid algorithm, or sum - of - squares algorithm), k - means clustering, fuzzy k - means clustering algorithm, and Jarvis - Patrick clustering.
[0118] In some embodiments, the classifier used for block 242 is the nearest - neighbor algorithm. Regarding the nearest - neighbor, when a query point x0 (test subject) is given, the k training points x (r) , r,..., k (here, training subjects) that are at the closest distance to x0 are identified, and the point x0 is classified using the k - nearest neighbors. Here, the distances of these neighbors are a function of the abundance values of the discriminant gene set. In some embodiments, the Euclidean distance in the feature space is used to calculate the distance d (i) = ||x (i) - x (O)||is determined as. Typically, when a nearest neighbor algorithm is used, the abundance data used in the calculation of the linear discriminant is standardized so that the mean is zero and the variance is 1. The nearest neighbor rule can be improved to address issues of unbalanced class preference, cost of sequential misclassification, and feature selection. Many of these improvements involve some form of weighted voting for the neighbors. For details on nearest neighbor analysis, see Duda, Pattern Classification, S econd Edition, 2001, John Wiley & Sons, Inc, and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, each of which is incorporated herein by reference.
[0119] Blocks 246-248. The above disclosure describes the training of a classifier using the abundance values of an identified gene set.
[0120] Referring to Block 246, in some embodiments, a trained classifier is used to classify a test subject by inputting a plurality of abundance values of the test into the classifier to determine whether the test subject has a first cancer state or a second cancer state. In such embodiments, each respective abundance value among the plurality of abundance values of the test quantifies the expression level of the corresponding gene in a biological sample (e.g., a tumor sample) of the test subject, more specifically in an identified gene set. In response to this input, the classifier designates whether the test subject has either the first cancer state or the second cancer state.
[0121] Referring to block 246, in some alternative embodiments, a trained classifier is used to determine the likelihood or probability that a subject has a first cancer state or a second state. This is done in such embodiments by inputting a plurality of abundance values of the test into the classifier. In such embodiments, each respective abundance value among the plurality of abundance values of the test quantifies the expression level of a corresponding gene among a plurality of genes (more specifically, a set of discriminant genes) in a biological sample (e.g., a tumor sample) of the test subject. In response to this input, the classifier specifies the likelihood or probability that the test subject has a first cancer state, or the likelihood or probability that the test subject has a second cancer state.
[0122] Referring to block 248, in some embodiments, a therapeutic intervention or imaging of a test subject is provided based on the determination that the test subject has a first cancer state or a second cancer state (or the likelihood that the test subject has a first or second cancer state). Examples of such conditional treatment are provided below in conjunction with FIG. 3. For example, non-limiting examples of ongoing clinical trials of therapies for specific cancer types associated with oncogenic pathogen infections are shown in Table 2 below.
[0123] RNA analysis pipeline In some embodiments, the methods and systems described herein are performed in conjunction with sequencing RNA molecules isolated from a biological sample of a patient. In some embodiments, a FASTQ file or equivalent file format of the sequencing data is the output of such a sequencing reaction.
[0124] In some embodiments, each FASTQ file contains reads, which can be paired-end or single-read, and can be short-read or long-read, and each read is isolated from a patient sample and represents one detected sequence of nucleotides in an mRNA molecule inferred by using a sequencing device to detect the sequence of nucleotides contained in a cDNA molecule generated from the mRNA molecules isolated during library preparation. Each read in the FASTQ file is also associated with a quality assessment. The quality assessment may reflect the likelihood that an error occurred during the sequencing procedure that affected the associated read.
[0125] Each FASTQ file can be processed by a bioinformatics pipeline. In various embodiments, the bioinformatics pipeline can filter the FASTQ data. Filtering of the FASTQ data can include correction of sequencing device errors, as well as low-quality sequences or bases, adapter sequences, contamination, chimeric reads It may include overrepresented sequences, biases caused by library preparation, amplification, or capture, and removal (trimming) of other errors. Entire reads, individual nucleotides, or multiple nucleotides where errors may occur can be discarded based on quality assessments associated with the reads in the FASTQ file, known error rates of the sequencing instrument, and / or comparison of each nucleotide in the read with one or more nucleotides in other reads aligned to the same position in the reference genome. Filtering can be performed in part or in whole by various software tools. The FASTQ file can be analyzed by sequencing data QC software such as, for example, AfterQC, Kraken, RNA-SeQC, FastQC (see Illumina, BaseSpace Labs, or https: / / www.illumina.com / products / by-type / informatics-products / basespace-sequence-hub / apps / fastqc.html), or another similar software program for quality control and rapid assessment of the reads. For paired-end reads, the reads can be merged.
[0126] For each FASTQ file, each read in the file can be aligned to a position in a reference genome that has a sequence that best matches the nucleotide sequence in the read. There are many software programs designed to align reads, such as programs that use the Bowtie, Burrows Wheeler Aligner (BWA), or Smith-Waterman algorithm. The alignment can be directed using a reference genome (e.g., GRCh38, hg38, GRCh37, or other reference genomes developed by the Genome Reference Consortium) by comparing the nucleotide sequence in each read to a portion of the nucleotide sequence in the reference genome to determine the portion of the reference genome sequence that is most likely to correspond to the sequence in the read. The alignment may take into account RNA splice sites. The alignment can generate a SAM file that stores the start and end positions of each read in the reference genome and the coverage (number of reads) of each nucleotide in the reference genome. The SAM file can be converted to a BAM file, sorted, or marked for deletion of duplicate reads.
[0127] In one example, kallisto software can be used for alignment and quantification of RNA reads (see Nicolas L Bray, Harold Pimentel, Pall Melsted and Lior Pachter, Near-optimal probabilistic RNA-seq quantification, Nature Biotechnology 34, 525-527 (2016), doi:10.1038 / nbt.3519). In another embodiment, quantification of RNA reads can be performed using another software, such as Sailfish or Salmon (see Rob Patro, Stephen M. Mount, and Carl Kingsford (2014) Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms. Nature Biotechnology (doi:10.1038 / nbt.2862) or Patro, R., Duggal, G., Love, M.I., Irizarry, R.A., & Kingsford, C. (2017). Salmon provides fast and bias-aware quantification of transcript expression. Nature Methods. These RNA-seq quantification methods may not require alignment. There are a number of software packages that can be used for normalization, quantitative analysis, and differential expression analysis of RNA-seq data.
[0128] For each gene, raw RNA read counts for a given gene can be calculated. The raw read counts can be stored in a tabular file for each sample, with columns representing genes and each entry representing the raw RNA read count for that gene. In one example, kallisto alignment software calculates the raw RNA read count as the sum of the probabilities that a read aligns to a gene. Thus, in this example, the raw counts are not integers.
[0129] Next, the raw RNA read counts can be normalized, for example using full quantile normalization, to correct for GC content and gene length, and adjusted for sequencing depth, for example using the size factor method. In one example, normalization of the RNA read counts is performed according to the methods disclosed in U.S. Patent Application No. 16 / 581,706, filed September 24, 2019, or PCT 19 / 52801, titled Methods of Normalizing and Correcting RNA Expression Data, which are hereby incorporated by reference in their entirety. The rationale for normalization is that the copy number of each cDNA molecule in the sequencing device may not reflect the distribution of mRNA molecules in the patient sample. For example, during library preparation, amplification, and capture steps, probe binding and errors generated during sequencing, which can be caused by random hexamers, amplification (PCR enrichment), rRNA depletion, and other characteristics of GC content, read length, gene length, and sequence in each nucleic acid molecule, may result in artifacts that cause specific portions of mRNA molecules to be over- or under-represented. Each raw RNA read count for each gene can be adjusted to eliminate or reduce over- or under-representation caused by biases or artifacts in the NGS sequencing protocol. The normalized RNA read counts can be stored in a tabular file for each sample, with columns representing genes and each entry representing the normalized RNA read count for that gene. of Normalizing and Correcting RNA Expression Data, which are hereby incorporated by reference in their entirety. The rationale for normalization is that the copy number of each cDNA molecule in the sequencing device may not reflect the distribution of mRNA molecules in the patient sample. For example, during library preparation, amplification, and capture steps, probe binding and errors generated during sequencing, which can be caused by random hexamers, amplification (PCR enrichment), rRNA depletion, and other characteristics of GC content, read length, gene length, and sequence in each nucleic acid molecule, may result in artifacts that cause specific portions of mRNA molecules to be over- or under-represented. Each raw RNA read count for each gene can be adjusted to eliminate or reduce over- or under-representation caused by biases or artifacts in the NGS sequencing protocol. The normalized RNA read counts can be stored in a tabular file for each sample, with columns representing genes and each entry representing the normalized RNA read count for that gene.
[0130] As described above, the transcriptome value set can refer to either normalized RNA read counts or raw RNA read counts.
[0131] HPV classifier training In one aspect, the present disclosure provides a method for training a classifier to detect human papillomavirus (HPV) infection in cancer. The method includes obtaining abundance values 126, such as mRNA expression levels, for genes that are useful for assessing the HPV status of HPV-related cancers from a training set of 124 subjects having HPV-related cancers and known HPV status. The method then includes training the classifier for each respective training subject with at least (i) the abundance values 126 and (ii) the HPV status of the patient's cancer, for example, using a classifier training module 120. In some embodiments, the classifier is also trained with respect to the status of one or more mutant alleles 127 in the cancer of each training subject.
[0132] In some embodiments, each training subject has an HPV-related cancer selected from cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, and vulvar cancer. In some embodiments, the classifier is trained on data from patients all having the same type of cancer, for example, cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, or vulvar cancer. However, since classifier training is generally improved by increasing the size of the training dataset, in some embodiments, the classifier is trained on data from patients having two or more types of HPV-related cancers, for example, two, three, four, five, six, seven, or all eight of cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, and vulvar cancer. In a particular embodiment illustrated by Example 3, each training subject has either head and neck squamous cell carcinoma or cervical cancer. either.
[0133] In some embodiments, the classifier is trained on abundance values for a plurality of genes selected from, for example, those listed in Table 3, such as KRT86, CRISPLD1, DSG1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, ZFR2, RNF212, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1. As reported below, for example, referring to Example 3, these 24 genes were found to be differentially expressed according to the HPV status of the subject in at least 8 of 10 training sets formed from expression data of cervical cancer or head and neck cancer for which the HPV status was known in The Cancer Genome Atlas (TCGA). However, one of ordinary skill in the art will understand that, in some cases, the use of different training data sets can result in different outcomes, for example, one or more of these genes may not be beneficial in at least 80% of the training folds, and / or one or more genes that were found not to be beneficial in at least 80% of the training folds in the study reported in Example 3 may be beneficial. These differences can occur, for example, when different criteria are used to select the training population, such as various inclusion and / or exclusion criteria, such as cancer type, individual characteristics (e.g., age, gender, ethnicity, family history, smoking status, etc.), or simply by using a smaller or larger data set.
[0134] Thus, in some embodiments, the classifier is trained on at least 5 of the genes listed in Table 3. In some embodiments, the classifier is trained on at least 10 of the genes listed in Table 3. In some embodiments, the classifier is trained on at least 15 of the genes listed in Table 3. In some embodiments, the classifier is trained on at least 20 of the genes listed in Table 3. In some embodiments, the classifier is trained on all 24 of the genes listed in Table 3. In some embodiments, the classifier is trained on 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or all 24 of the genes listed in Table 3. Further, in some embodiments, the classifier is also trained on abundance values for one or more genes not listed in Table 3. In some embodiments, the classifier is also trained on abundance values for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 3. In some embodiments, the classifier is also trained on abundance values for 1 - 10 genes not listed in Table 3. In some embodiments, the classifier is also trained on abundance values for 1 - 5 genes not listed in Table 3. In other embodiments, the classifier is not also trained on abundance values for any genes not listed in Table 3.
[0135] Furthermore, one of ordinary skill in the art will also understand that for some features, such as abundance values for certain genes, may be more informative than other features in a particular classifier. One measure of the predictive power of each feature in a classifier based on multiple features is the regression coefficient calculated for the feature during training of the model. The regression coefficient represents the relationship between each feature and the response of the model. The coefficient value represents the average change in the response given a one-unit increase in the feature value. Thus, for at least variables of the same type, the magnitude of the regression coefficient, e.g., the absolute value, correlates with the importance of the feature in the model. That is, the larger the magnitude of the regression coefficient, the more important the variable is to the model. For example, as reported in Example 3, for all abundance values of all 24 genes listed in Table 3, as well as for the mutant allele states for the TP53 and CDKN2A genes in a particular support vector machine (SVM) classifier trained on these, only 6 out of the 24 genes had regression coefficients with magnitudes of at least 0.5 - CDKN2A (1.13), SMC1B (1.02), EFNB3 (-0.97), KCNS1 (0.74), CCND1 (-0.65), and RNF212 (0.517).
[0136] Accordingly, one of ordinary skill in the art can select a feature set that includes fewer genes than all those listed in Table 3, based at least in part on the importance of each feature in one or more classification models. For example, in some embodiments, one or more genes having lower predictive power in a classification model can be omitted during classifier training. For example, in some embodiments, the features used for training include at least gene expression features having a regression coefficient of at least 0.5, such as those listed in Table 5, e.g., CDKN2A, SMC1B, EFNB3, KCNS1, CCND1, and RNF212. In some embodiments, the features used for training include at least gene expression features having a regression coefficient of at least 0.4, such as those listed in Table 5. In some embodiments, the features used for training include at least gene expression features having a regression coefficient of at least 0.3, such as those listed in Table 5. In some embodiments, the features used for training include at least gene expression features having a regression coefficient of at least 0.2, such as those listed in Table 5. In some embodiments, the features used for training include at least gene expression features having a regression coefficient of at least 0.1, such as those listed in Table 5.
[0137] Similarly, the size of the feature set can be affected by which features are included and / or excluded. For example, in some embodiments, if certain features with high predictive power are included in the classification model, fewer total features may be included in the model. For example, in some embodiments, if the abundance values for SMC1B, CDKN2A, and EFNB3 are included in the model, it may be necessary to include the abundance values for no more than two of the other genes whose abundance values are used as features in Table 5 in the model. Thus, in some embodiments, the features used to train the model include the abundance values for SMC1B, CDKN2A, and EFNB3, and the abundance values for at least two other genes whose abundance values are used as features in Table 5. In some embodiments, the features used to train the model include the abundance values for SMC1B, CDKN2A, and EFNB3, and the abundance values for at least five other genes whose abundance values are used as features in Table 5. In some embodiments, the features used to train the model include the abundance values for SMC1B, CDKN2A, and EFNB3, and the abundance values for at least ten other genes whose abundance values are used as features in Table 5. In some embodiments, the features used to train the model include the abundance values for SMC1B, CDKN2A, and EFNB3, and the abundance values for at least fifteen other genes whose abundance values are used as features in Table 5.
[0138] Similarly, in some embodiments, if features with high predictive power are excluded from the classification model, more of the other features may be included in the model. For example, in some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 15 of the other features used as features in Table 5 are included in the model. In some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 20 of the other genes used as features in Table 5 are included in the model. In some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 15, 16, 17, 18, 19, 20, or all 21 of the other genes used as features in Table 5 are included in the model.
[0139] Of course, other metrics can also be utilized to evaluate the importance of features in the model, such as the change in the standardized regression coefficient and R-squared when the feature was last added to the model.
[0140] When selecting a feature set, one of ordinary skill in the art will also consider the degree to which the features are correlated with each other. Correlation is a statistical measure that indicates how linearly dependent two variables are on each other. As such, two correlated features provide redundant information to the prediction model, which can potentially have an adverse effect on the classifier. For this reason, there are several reasons to exclude correlated features from the model. For example, the more features there are in the classifier, the more computations that need to be performed, so removing correlated features speeds up the algorithm. Removing correlated features can also remove the harmful bias resulting from the correlation from the model. Finally, removing correlated features can make the model more interpretable.
[0141] Therefore, one of ordinary skill in the art may select a feature set containing fewer genes than all those listed in Table 3, at least in part based on the correlation of each feature in one or more classification models. In some embodiments, the selection to remove one or the other feature from the correlated feature sets is informed by the predictive power of the two features, e.g., as given by their respective regression coefficients. For example, the gene expression values for ENSG00000105278 (CXCL14) and ENSG00000077935 (SMC1B) are highly correlated (correlation = 0.718983175) in the feature set described in Table 3. Thus, in some embodiments, the feature set does not include either CXCL14 or SMC1B. In some embodiments, as reported in Table 5, SMC1B has a higher regression coefficient (1.02) than CXCL14 (-0.29) in the SVM model described in Example 3, so CXCL14 rather than SMC1B is excluded from the feature set.
[0142] As reported in Table 6, ten pairs of gene expression features have a correlation of at least 0.6. Thus, in some embodiments, the features in at least one pair of features having a correlation of at least 0.6 are excluded from the model. In some embodiments, the features in at least two pairs of features having a correlation of at least 0.6 are excluded from the model. In other embodiments, the features in at least three pairs, four pairs, five pairs, six pairs, seven pairs, eight pairs, nine pairs, or all ten pairs of features having a correlation of at least 0.6 are excluded from the model. In some embodiments, the excluded features are the features in a pair of highly correlated features having a lower regression coefficient than reported in Table 5. For example, referring to Table 6, the features having a lower regression coefficient in each pair with high correlation (e.g., corresponding to a correlation of at least 0.6) are as follows. · Pair 1 = DSG1 · Pair 2 = ZFR2 · Pair 3 = RNF212 · Pair 4 = SYCP2 · Pair 5 = ZFR2 · Pair 6 = MYO3A · Pair 7 = SYCP2 · Pair 8 = DSG1 · Pair 9 = KCNS1 · Pair 10 = ZFR2 Thus, in some embodiments, one or more of DSG1, ZFR2, RNF212, SYCP2, MYO3A, and KCNS1 are excluded from the feature set based on them being the least beneficial features in pairs of features that are highly correlated.
[0143] However, in some embodiments, this selection process does not allow for excluding both features of a pair of highly correlated features based on, for example, both genes being the least beneficial features in at least one of the pairs of highly correlated features. Thus, in some embodiments, one or more of SYCP2, MYO3A, and KCNS1 are not excluded from the feature set. Similarly, in some embodiments, this selection process does not allow for excluding features that are very beneficial, for example, features having a regression coefficient of at least 0.5, from the feature set. Thus, in some embodiments, one or both of RNF212 and KCNS1 are not excluded from the feature set.
[0144] Thus, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0145] Similarly, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, RNF212, MKRN3, MYL1, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0146] Similarly, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, RNF212, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0147] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation data set of at least 50 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation data set of at least 50 data constructs.
[0148] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation data set of at least 100 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation data set of at least 100 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation data set of at least 100 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation data set of at least 100 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation data set of at least 100 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation data set of at least 100 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation data set of at least 100 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation data set of at least 100 data constructs.
[0149] In some embodiments, as described above with reference to FIG. 2, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. In some embodiments, the classifier is trained according to the above methodology with reference to FIG. 2.
[0150] EBV Classifier Training In one aspect, the present disclosure provides a method for training a classifier to detect Epstein-Barr virus (EBV) infection in cancer. The method includes obtaining abundance values 126, such as mRNA expression levels, for genes that are useful for evaluating the EBV status of EBV-related cancers from a training set of 124 subjects having EBV-related cancers and known EBV status. The method then includes training the classifier for each respective training subject with respect to at least (i) the abundance values 126 and (ii) the EBV status of the patient's cancer, for example, using a classifier training module 120. In some embodiments, the classifier is also trained with respect to the status of one or more mutant alleles 127 in the cancer of each training subject.
[0151] In some embodiments, each training subject is selected from Burkitt lymphoma, nasal-type extranodal NK / T-cell lymphoma, non-Hodgkin lymphoma, Hodgkin lymphoma, nasopharyngeal carcinoma, and gastric carcinoma has an EBV-related cancer. In some embodiments, the classifier is trained on data from patients all having the same type of cancer, such as Burkitt lymphoma, nasal-type extranodal NK / T-cell lymphoma, non-Hodgkin lymphoma, Hodgkin lymphoma, nasopharyngeal carcinoma, or gastric cancer. However, since classifier training is generally improved by increasing the size of the training dataset, in some embodiments, the classifier is trained on data from patients having two or more types of EBV-related cancers, such as two, three, four, five, or all six of Burkitt lymphoma, nasal-type extranodal NK / T-cell lymphoma, non-Hodgkin lymphoma, Hodgkin lymphoma, nasopharyngeal carcinoma, and gastric cancer. In a particular embodiment illustrated by Example 4, each training subject has gastric cancer.
[0152] In some embodiments, the classifier is trained on abundance values for a plurality of genes selected from those listed in Table 4, such as SCNN1A, CDX1, KCNK15, PRKCG, KRT7, NKD2, GPR158, CLDN3, and ZNF683. As reported below, for example, referring to Example 4, these nine genes were found to be differentially expressed according to the EBV status of the subject in at least 80% of the gastric cancer training set in The Cancer Genome Atlas (TCGA). However, one of ordinary skill in the art will understand that, in some cases, the use of different training datasets may result in different outcomes, such as one or more of these genes not being beneficial in at least 80% of the training folds, and / or one or more genes that were found not to be beneficial in at least 80% of the training folds in the study reported in Example 4 may be beneficial. These differences may occur, for example, when different criteria are used to select the training population, such as various inclusion and / or exclusion criteria such as cancer type, individual characteristics (e.g., age, gender, ethnicity, family history, smoking status, etc.), or simply by using a smaller or larger dataset.
[0153] Thus, in some embodiments, the classifier is trained on at least 5 of the genes listed in Table 4. In some embodiments, the classifier is trained on at least 6 of the genes listed in Table 4. In some embodiments, the classifier is trained on at least 7 of the genes listed in Table 4. In some embodiments, the classifier is trained on at least 8 of the genes listed in Table 4. In some embodiments, the classifier is trained on all 9 of the genes listed in Table 4. Further, in some embodiments, the classifier is also trained on abundance values for one or more genes not listed in Table 4. In some embodiments, the classifier is also trained on abundance values for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes not listed in Table 4. In some embodiments, the classifier is also trained on abundance values for 1 - 10 genes not listed in Table 4. In some embodiments, the classifier is also trained on abundance values for 1 - 5 genes not listed in Table 4. In other embodiments, the classifier is not also trained on abundance values for any genes not listed in Table 4.
[0154] Furthermore, one of ordinary skill in the art will also understand that some features, for example, abundance values for certain genes, may be more informative than other features in a particular classifier. One measure of the predictive power of each feature in a classifier based on multiple features is the regression coefficient calculated for the feature during training of the model. The regression coefficient represents the relationship between each feature and the response of the model. The coefficient value represents the average change in the response given a 1 - unit increase in the feature value. Thus, for at least the same type of variable, the magnitude of the regression coefficient, e.g., the absolute value, correlates with the importance of the feature in the model. That is, the larger the magnitude of the regression coefficient, the more important the variable is to the model. For example, as reported in Example 4, as listed in Table 4 For all nine abundance values of the gene, as well as for the mutant allele states for the TP53 and PIK3CA genes, in a particular support vector machine (SVM) classifier trained thereon, only four of the nine genes had regression coefficients of at least magnitude 0.75 - SCNN1A (-1.26), KCNK15 (-1.04), KRT7 (-0.94), and CLDN3 (-1.68).
[0155] Accordingly, one of ordinary skill in the art can select a feature set that includes fewer genes than all those listed in Table 4, based at least in part on the importance of each feature in one or more classification models. For example, in some embodiments, one or more genes having lower predictive power in the classification model can be omitted during classifier training. For example, in some embodiments, the features used for training include at least the gene expression features listed in Table 5 having a regression coefficient of at least 0.75, such as SCNN1A (-1.26), KCNK15 (-1.04), KRT7 (-0.94), and CLDN3 (-1.68). In some embodiments, the features used for training include at least the gene expression features listed in Table 5 having a regression coefficient of at least 0.6.
[0156] Similarly, the size of the feature set can be affected by which features are included and / or excluded. For example, in some embodiments, if certain features with high predictive power are included in the classification model, fewer total features may be included in the model. For example, in some embodiments, if the abundance values for SCNN1A, KCNK15, KRT7, and CLDN3 are included in the model, the abundance values for one or fewer of the other genes listed in Table 4 need to be included in the model. Thus, in some embodiments, the features used to train the model include the abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and at least one other gene listed in Table 4. In some embodiments, the features used to train the model include the abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and at least two other genes listed in Table 4. SCNN1A, KCNK15, KRT7, and CLDN3, and at least three other genes listed in Table 4. SCNN1A, KCNK15, KRT7, and CLDN3, and at least four other genes listed in Table 4.
[0157] Similarly, in some embodiments, if features with high predictive power are excluded from the classification model, more of the other features may be included in the model. For example, in some embodiments, if the abundance values for one or more of SCNN1A, KCNK15, KRT7, and CLDN3 are not included in the model, the abundance values for at least four of the other genes listed in Table 4 are included in the model. In some embodiments, if the abundance values for one or more of SCNN1A, KCNK15, KRT7, and CLDN3 are not included in the model, the abundance values for all five of the other genes listed in Table 4 are included in the model.
[0158] Of course, other metrics can also be utilized to evaluate the importance of features in the model, such as the change in the standardized regression coefficient and R-squared when the feature was last added to the model.
[0159] When selecting a feature set, one of ordinary skill in the art will also consider the degree to which the features are correlated with each other. Correlation is a statistical measure that indicates the degree to which two variables are linearly dependent on each other. As such, two correlated features provide redundant information to the prediction model, which can have an adverse effect on the classifier. For this reason, there are several reasons to exclude correlated features from the model. For example, the more features there are in the classifier, the more calculations need to be performed, so removing correlated features speeds up the algorithm. Removing correlated features can also remove the harmful bias resulting from the correlation from the model. Finally, removing correlated features can make the model more interpretable. Thus, one of ordinary skill in the art may select a feature set containing fewer genes than all those listed in Table 3, at least in part based on the correlation of each feature in one or more classification models. For example, statistical analysis of the SVM model trained in Example 4 revealed that the gene expression values for ENSG00000135480 (KRT7) and ENSG00000124249 (KCNK15) are highly correlated (0.650). Thus, in some embodiments, the abundance value for one of KRT7 and KCNK15 is excluded from the feature set.
[0160] For example, in one embodiment, the feature set includes abundance values for at least SCNN1A, CDX1, KCNK15, PRKCG, NKD2, GPR158, CLDN3, and ZNF683. In another embodiment, the feature set includes abundance values for at least SCNN1A, CDX1, PRKCG, KRT7, NKD2, GPR158, CLDN3, and ZNF683.
[0161] In some embodiments, the classifier has at least 70% specificity and at least 70% sensitivity for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has at least 75% specificity and at least 75% sensitivity for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has at least 80% specificity and at least 80% sensitivity for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has at least 85% specificity and at least 85% sensitivity for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has at least 90% specificity and at least 90% sensitivity for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has at least 95% specificity and at least 95% sensitivity for a validation data set of at least 50 data constructs. For example, none of the data constructs in the validation data set were used in training the classifier. In some embodiments, the classifier has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more specificity for a validation data set of at least 50 data constructs. In some embodiments, the classifier has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more sensitivity for a validation data set of at least 50 data constructs.
[0162] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs.
[0163] In some embodiments, as described above with reference to FIG. 2, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. In some embodiments, the classifier is trained according to the above methodology with reference to FIG. 2.
[0164] Classification method In some embodiments, the present disclosure provides a method for discriminating a first cancer state and a second cancer state in a human subject, wherein the first cancer state is associated with an infection by a carcinogenic pathogen and the second cancer state is associated with a state not containing a carcinogenic pathogen. Generally, the method includes obtaining abundance data, such as relative expression levels, for a plurality of genes that are differentially expressed in cancerous tissue associated with carcinogenic pathogen infection and the same type of cancerous tissue not associated with carcinogenic pathogen infection. The abundance data is then input, at least in part, to a classifier trained to discriminate the first cancer state and the second cancer state based on the abundance of genes differentially expressed in the two types of cancerous tissue. An example of the training of such a classifier is provided above in conjunction with the description of FIG. 2.
[0165] Many of the embodiments described below are related to analyses performed, in conjunction with FIG. 3, for example, using expression data derived from the exome of a cancer patient obtained from a sample of cancerous tissue in a patient. Generally, these embodiments are independent and thus do not depend on a particular expression data generation method, such as sequencing, hybridization, and / or qPCR methodology. However, in some embodiments, the methods described below include one or more steps (301) for generating expression data.
[0166] In some embodiments, these methods include obtaining a sample of cancerous tissue (302). Methods for obtaining a sample of cancerous tissue are known in the art and depend on the type of cancer being sampled. For example, a sample of blood cancer can be obtained using bone marrow biopsy and isolation of circulating tumor cells, a sample of cancers of the gastrointestinal tract, bladder, and lung can be obtained using endoscopic biopsy, a sample of subcutaneous tumors can be obtained using needle biopsy (e.g., fine needle aspiration, core needle aspiration, vacuum-assisted biopsy, and image-guided biopsy), a sample of skin cancer can be obtained using skin biopsy, e.g., shave biopsy, punch biopsy, incisional biopsy, and excisional biopsy, and a sample of cancer affecting a patient's internal organs can be obtained using surgical biopsy.
[0167] Next, in some embodiments, mRNA is isolated from the sample of cancerous tissue (304). Many techniques for isolating RNA from tissue samples are known in the art. For example, acid guanidinium thiocyanate-phenol-chloroform extraction (see, e.g., Chomczynski and Sacchi, Nat Protoc, 1(2):581-85(2006), the content of which is hereby incorporated by reference in its entirety for all purposes), and silica bead / glass fiber adsorption (see, e.g., Poeckh, T. et al., Anal Biochem., 373(2):253-62(2008), the content of which is hereby incorporated by reference in its entirety for all purposes). The selection of any particular RNA isolation technique for use in conjunction with the embodiments described herein is within the skill of one of ordinary skill in the art considering the type of tissue, the state of the tissue, e.g., fresh, frozen, formalin-fixed, paraffin-embedded (FFPE), and the type of nucleic acid analysis to be performed on the RNA sample.
[0168] In some embodiments, RNA is isolated from a blood sample and / or a tissue section (e.g., a tumor biopsy) using commercially available reagents such as proteinase K, TURBO DNase-I, and / or RNA Clean XP beads. In some embodiments, the isolated RNA is subjected to a quality control protocol for determining the concentration and / or amount of RNA molecules, including the use of a fluorescent dye and a fluorescence microplate reader, a standard spectrofluorometer, or a filter fluorometer.
[0169] In some embodiments, expression data is obtained directly from the isolated mRNA, e.g., by direct RNA sequencing (314). Methods for direct RNA sequencing are known in the art. See, for example, Ozsolak F., et al., Nature 461:814-18 (2009), and Garalde, D.R., et al., Nat Methods, 15(3):201-206 (2018), the contents of which are hereby incorporated by reference in their entirety for all purposes.
[0170] In other embodiments, expression data is obtained via a cDNA intermediate. Thus, in some embodiments, a cDNA library is created via cDNA synthesis using the isolated RNA (310). In some embodiments, the cDNA library is prepared from the isolated RNA that is purified and selected for cDNA molecule size selection using commercially available reagents such as Roche KAPA Hyper Beads. In another example, a New England Biolabs (NEB) kit can be used.
[0171] In some embodiments, cDNA library preparation includes ligation of adapters to cDNA molecules. For example, UDI adapters, such as Roche SeqCap dual end adapters, or UMI adapters (e.g., full-length or Staby Y adapters) can be ligated to cDNA molecules. The adapters can be nucleic acid molecules that function as barcodes to identify cDNA molecules according to the samples from which they are derived, and / or as barcodes to facilitate downstream bioinformatics processing and / or next-generation sequencing reactions. The nucleotide sequence in the adapter can be unique to the sample for differentiating samples. The adapter can facilitate the binding of cDNA molecules to anchor oligonucleotide molecules on the sequencing device flow cell and can function as a seed for the sequencing process by providing a starting point for the sequencing reaction.
[0172] The cDNA library can be amplified and purified using reagents, such as Axygen MAG PCR clean-up beads. Subsequently, the concentration and / or amount of cDNA molecules can be quantified using a fluorescent dye and a fluorescence microplate reader, a standard spectrofluorometer, or a filter fluorometer.
[0173] In some embodiments, for both direct RNA sequencing and prior to cDNA library construction, the isolated RNA is first enriched for the desired type of RNA (e.g., mRNA) or species (e.g., a particular mRNA transcript) prior to cDNA library construction (308). Methods for enriching for the desired RNA molecules are also known in the art. For example, mRNA molecules can be enriched relative to other RNA molecules in a total RNA preparation using, for example, oligo dT affinity techniques (see, e.g., Rio, D.C., et al., Cold Spring Harb Protoc., 2010 Jul 1;2010(7), the contents of which are hereby incorporated by reference in their entirety for all purposes). Particular mRNA transcripts can also be isolated using, for example, hybridization probes that specifically bind to one or more mRNA sequences of interest.
[0174] In some embodiments, the cDNA library is pooled and treated with reagents to reduce off-target capture, such as human COT-1 and / or IDT xGen Universal Blockers, before being dried in vacuo. The pool is then resuspended in a hybridization mixture, such as IDT xGen Lockdown, and a probe, such as an IDT xGen Exome Research Panel v1.0 probe, an IDT xGen Exome Research Panel v2.0 probe, other IDT probe panels, Roche probe panels, or other probes, can be added to each pool. The pool can be incubated in an incubator, a PCR machine, a water bath, or other temperature-regulated device to hybridize the probe. The pool can then be mixed with beads coated with streptavidin or another means for capturing hybridized cDNA probe molecules, particularly cDNA molecules representing exons of the human genome. In another embodiment, polyA capture can be used. The pool can be amplified and purified again, using commercially available reagents, such as the KAPA HiFi Library Amplification kit and Axygen MAG PCR clean-up beads, respectively.
[0175] The construction of cDNA libraries from isolated mRNA is also known in the art. In some embodiments, cDNA library construction is performed by first-strand DNA synthesis from isolated mRNA using reverse transcriptase, followed by second-strand synthesis using DNA polymerase. Examples of methods for cDNA synthesis are described in McConnell and Watson, 1986, FEBS Lett. 195(1-2), pp. 199-202, Lin and Ying, 2003, Methods Mol Biol. 221, pp. 129-143, and Oh et al., 2003, Exp Mol Med. 35(6), pp. 586-90, the contents of which are hereby incorporated by reference in their entirety for all purposes.
[0176] The cDNA library can also be analyzed to determine the fragment size of the cDNA molecules, which can be done via gel electrophoresis techniques and can include the use of devices such as the LabChip GX Touch. The pool can be cluster amplified using kits (e.g., Illumina Paired-end Cluster Kits with PhiX spike). In one example, the cDNA library preparation and / or whole exome capture step can be performed in an automated system using a liquid handling robot (e.g., SciClone NGSx).
[0177] Library amplification can be performed on a device, such as an Illumina C-Bot2 The resulting flow cell containing the amplified target-captured cDNA library can be sequenced on a next-generation sequencing device, such as an Illumina HiSeq4000 or Illumina NovaSeq 6000, to a user-selected unique on-target depth, e.g., 300x, 400x, 500x, 10,000x, etc. The next-generation sequencing device can generate FASTQ, BCL, or other files for each patient sample or each flow cell.
[0178] If two or more patient samples are processed simultaneously on the same sequencing device flow cell, the reads from the multiple patient samples are initially included in the same BCL file and then split into individual FASTQ files for each patient. Differences in the sequences of the adapters used for each patient sample can serve the purpose of barcoding and facilitate associating each read with the correct patient sample and placing it in the correct FASTQ file.
[0179] Methods for mRNA sequencing are known in the art. In some embodiments, mRNA sequencing is performed by whole exome sequencing (WES). Generally, WES involves isolating RNA from a tissue sample, optionally selecting a desired sequence and / or depleting unwanted RNA molecules, generating a cDNA library, and then sequencing the cDNA library (312) using, for example, next generation sequencing (NGS) technology. For a review of the use of whole exome sequencing technology in cancer diagnosis, see Serrati et al., 2016, Onco Targets Ther. 9, pp. 7355-7365, the contents of which are hereby incorporated by reference in their entirety for all purposes.
[0180] Next generation sequencing methods are also known in the art and include synthetic technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single molecule real time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing by synthesis with reversible dye terminators.
[0181] In some embodiments, the sequence reads can be aligned to a reference exome or reference genome using methods known in the art to determine alignment position information. The alignment position information can indicate the start and end positions of the region in the reference genome corresponding to the start and end nucleotide bases of a given sequence read. The alignment position information can also include the sequence read length that can be determined from the start and end positions. The region in the reference genome can be associated with a gene or segment of a gene. Non-limiting examples of known software for assembling and managing transcriptome information from RNA-seq data include TopHat and Cufflinks. See Trapnell et al., 2012, Nat Protoc. 7(3), pp. 562-578, the content of which is hereby incorporated by reference in its entirety for all purposes. Also see Hintzsche et al., 2016, Int J Genomics 7983236, the content of which is hereby incorporated by reference in its entirety for all purposes.
[0182] In other embodiments, the expression data is generated, for example, by hybridization (313) of a cDNA library using a microarray. The use of microarray-based gene profiling to identify differential gene expression after pathogen infection is known in the art. See, for example, Adomas et al., 2008, Tree Physiol. 28(6), pp. 885-897, the content of which is hereby incorporated by reference in its entirety for all purposes. Similarly, in other embodiments, still other methods for quantifying expression based on a cDNA library, such as quantitative real-time PCR (RT-qPCR), are used. See, for example, Wagner, 2013, Methods Mol Biol. 1027, pp. 19-45, the content of which is hereby incorporated by reference in its entirety for all purposes.
[0183] As shown with respect to FIG. 3, in some embodiments, method 300 is executed, at least in part, by a computer system (e.g., computer system 100 of FIG. 1) having one or more processors and a memory storing one or more programs for execution by the one or more processors to identify a first cancer state and a second cancer state in a subject, the first cancer state being associated with an infection by a first carcinogenic pathogen and the second cancer state being associated with a state free of carcinogenic pathogens. Some operations in method 300 are optionally combined and / or the order of some operations is optionally changed.
[0184] In some embodiments, the method includes obtaining a dataset for the subject, the dataset including a plurality of abundance values, each respective abundance value of the plurality of abundance values quantifying the expression level of a corresponding gene among a plurality of genes in cancerous tissue from the subject. In some embodiments, the obtained abundance values are determined according to any of the methodologies described with respect to sub-method 301. In some embodiments, the abundance data is pre-generated and communicated to computer system 100 via a network, e.g., using network interface 104. The method 300 then includes inputting (316) the dataset into a classifier trained to identify a first cancer state and a second cancer state in a human subject, the first cancer state being associated with an infection by a carcinogenic pathogen and the second cancer state being associated with a state free of carcinogenic pathogens. An example of such a classifier is provided above in conjunction with FIG. 2. Thereby, the method determines (320) whether the subject has a first cancer state associated with carcinogenic pathogen infection or a second cancer state not associated with carcinogenic pathogen infection.
[0185] In some embodiments, method 300 also includes inputting into the classifier a variant allele count for one or more variant alleles at one or more loci in the genome of a cancerous tissue from a subject. That is, in some embodiments, the classifier is also trained on data regarding the presence or absence of one or more variant alleles in subjects with cancer associated with or not associated with an oncogenic pathogen infection. In some embodiments, the one or more variant alleles are selected from variant alleles in genes selected from the group consisting of TP53 (ENSG00000141510), CDKN2A (ENSG00000147889), and PIK3CA (ENSG00000121879).
[0186] In some embodiments, the subject has breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, head and neck cancer, ovarian cancer, hepatobiliary cancer, cervical cancer, thyroid cancer, or bladder cancer.
[0187] In some embodiments, the first cancer state is associated with an infection by a first oncogenic pathogen selected from Epstein - Barr virus (EBV), hepatitis B virus (HBV), hepatitis C virus (HCV), human papillomavirus (HPV), human T - cell lymphotropic virus type 1 (HTLV - 1), Kaposi's sarcoma - associated herpesvirus (KSHV), and Merkel cell polyomavirus (MCV).
[0188] More specifically, in some embodiments, the first cancer state is human papillomavirus Selected from cervical cancer associated with human papillomavirus (HPV), head and neck cancer associated with HPV, gastric cancer associated with Epstein-Barr virus (EBV), nasopharyngeal cancer associated with EBV, Burkitt lymphoma associated with EBV, Hodgkin lymphoma associated with EBV, liver cancer associated with hepatitis B virus (HBV), liver cancer associated with hepatitis C virus (HCV), Kaposi sarcoma associated with Kaposi sarcoma-associated herpesvirus (KSHV), adult T-cell leukemia / lymphoma associated with human T-lymphotropic virus type 1 (HTLV-1), and Merkel cell carcinoma associated with Merkel cell polyomavirus (MCV). For a summary of cancer states known to be associated with oncogenic viral infections, see de Flora, 2011, “The prevention of infection-associated cancers,” Carcinogenesis 32, pp. 787-795.
[0189] Thus, if the first cancer state is a particular type of cancer associated with a particular oncogenic pathogen, the second cancer state is the same particular type of cancer associated with the absence of infection by the particular oncogenic pathogen. For example, if the first cancer state is cervical cancer associated with human papillomavirus (HPV) infection, the second cancer state is cervical cancer not associated with human papillomavirus (HPV) infection. Further, as described above, the classifier used to distinguish the two cancer states is trained on a dataset including at least gene abundance values (e.g., mRNA expression profiles) from subjects known to have cervical cancer associated with human papillomavirus (HPV) infection and from subjects known to have cervical cancer not associated with human papillomavirus (HPV) infection.
[0190] In some embodiments, the method further comprises treating the subject with either a first therapy (322) tailored for the treatment of a first cancer state associated with an oncogenic pathogenic infection or a second therapy (324) tailored for the treatment of a second cancer state not associated with an oncogenic pathogenic infection.
[0191] Accordingly, in one embodiment, a method for treating cancer in a human cancer patient is provided. The method includes obtaining a dataset for a patient, where the dataset includes a plurality of abundance values, to determine whether the patient is infected with an oncogenic pathogen linked to the pathology of the cancer, and inputting the dataset into a classifier trained to identify at least a first cancer state associated with infection by the oncogenic pathogen and a second cancer state not associated with infection by the oncogenic pathogen. The value of each abundance in the dataset quantifies the expression level of the corresponding gene found to be differentially expressed in cancers associated with infection by the oncogenic pathogen and cancers not associated with infection by the oncogenic pathogen. In some embodiments, the genes for which the abundance values are used to identify the cancer state for any particular type of cancer are selected according to any of the selection methodologies described above with reference to FIG. 2. Similarly, in some embodiments, the classifier used is trained according to any of the training methods described above with reference to FIG. 2.
[0192] In some embodiments, if the subject is determined to have a first cancer state associated with oncogenic pathogen infection, the method includes assigning and / or administering immunotherapy to the subject. In some embodiments, if the subject is determined to have a second cancer state not associated with oncogenic pathogen infection, the method includes assigning and / or administering chemotherapy to the subject.
[0193] As summarized in Table 2, several clinical trials are underway for the treatment of virus-related tumors. Accordingly, in some embodiments, the methods described herein include assigning and / or administering treatment for a particular cancer associated with a particular oncogenic virus infection, as described in Table 2. For example, in some embodiments, if the subject is determined to have stage 3 cervical cancer associated with HPV infection, the subject An effective dosing regimen for axalimogene filolisbac, a live attenuated Listeria monocytogenes transfected with a plasmid encoding an HPV-16E7 protein fused to a cleavage fragment of the protein listeriolysin O, is assigned and / or implemented. [Table 2-1] [Table 2-2]
[0194] HPV oncogenic virus infection In some embodiments, the methods described herein relate to the classification and / or treatment of cancers known to be associated with human papillomavirus (HPV) infection. As reported in Example 3 below, 24 genes listed in Table 3 and shown in FIG. 4B were found to be differentially expressed in at least 8 of 10 training sets formed from expression data of cervical or head and neck cancers having known HPV status in The Cancer Genome Atlas (TCGA). Thus, in some embodiments, the expression level of one or more of the genes listed in Table 3 is used for the classification of cervical or head and neck cancers as either associated with or not associated with HPV infection. In some embodiments, the expression levels of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or all 24 of the genes listed in Table 3 are used for the classification of cervical or head and neck cancers as either associated with or not associated with HPV infection. [Table 3]
[0195] In one embodiment, a method for discriminating between a first cancer state and a second cancer state in a human subject is provided, where the first cancer state is associated with infection by a human papillomavirus (HPV) oncogenic virus and the second cancer state is associated with a state that does not include HPV. The method includes obtaining a dataset for the subject, as described above with reference to FIG. 3, for example. The dataset includes a plurality of abundance values from the subject, and each respective abundance value among the plurality of abundance values quantifies the expression level of a corresponding gene among a plurality of genes in cancerous tissue from the subject. In some embodiments, the plurality of genes includes at least five genes selected from the genes listed in Table 3. The method then includes inputting the dataset into a classifier trained to discriminate at least the first cancer state and the second cancer state based on the abundance values of the plurality of genes. In some embodiments, the classifier is trained according to any of the methodologies described above with respect to FIG. 2.
[0196] In some embodiments, the first cancer state is cervical cancer associated with HPV infection and the second cancer state is cervical cancer not associated with HPV infection. In some embodiments, the first cancer state is head and neck cancer associated with HPV infection and the second cancer state is head and neck cancer not associated with HPV infection. In some embodiments, the head and neck cancer is a specific form or type of head and neck cancer, such as hypopharyngeal cancer, laryngeal cancer, lip and oral cavity cancer, metastatic squamous cell carcinoma with occult primary, nasopharyngeal cancer, oropharyngeal cancer, paranasal sinus and nasal cavity cancer, or salivary gland cancer.
[0197] In some embodiments, the plurality of genes comprise at least 10 of the genes listed in Table 3. In some embodiments, the plurality of genes comprise at least 15 of the genes listed in Table 3. In some embodiments, the plurality of genes comprise at least 20 of the genes listed in Table 3. In some embodiments, the plurality of genes comprise all of the genes listed in Table 3. In some embodiments, the plurality of genes comprise one or more genes not listed in Table 3, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 3. In some embodiments, the plurality of genes comprise 20 or fewer genes. In some embodiments, the plurality of genes comprise 25 or fewer genes. In some embodiments, the plurality of genes comprise 50 or fewer genes. In some embodiments, the plurality of genes comprise 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0198] In some embodiments, the dataset also includes mutant allele counts for one or more alleles at one or more genetic loci in the genome of the cancerous tissue from the subject. In some embodiments, the mutant allele count is either 1, which represents the state in which the subject carries the mutant allele, or 0, which represents the state in which the subject does not carry the mutant allele. In some embodiments, the mutant allele is a somatic mutation derived from the germline of the subject. In some embodiments, the mutant allele is a cancer-derived mutation from the cancerous tissue. In some embodiments, the mutant allele is located in the TP53 (ENSG00000141510) or CDKN2A (ENSG00000147889) gene.
[0199] In some embodiments, the classifier is trained to determine the HPV status of a subject having an HPV - related cancer selected from cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, and vulvar cancer. In some embodiments, the classifier is trained to determine the HPV status of a test patient having a specific HPV - related cancer, e.g., cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, or vulvar cancer. However, since classifier training is generally improved by increasing the size of the training dataset, in some embodiments, the classifier is trained on data from patients having two or more types of HPV - related cancers, e.g., two, three, four, five, six, seven, or all eight of cervical cancer, head and neck squamous cell carcinoma, ovarian cancer, penile cancer, pharyngeal cancer, anal cancer, vaginal cancer, and vulvar cancer. In a specific embodiment illustrated by Example 3, the classifier is trained on a subject having either head and neck squamous cell carcinoma or cervical cancer. However, in some embodiments, a classifier trained on patients having one or more types of HPV - related cancers is useful for determining the HPV status of patients having different types of HPV - related cancers.
[0200] In some embodiments, the features of the classifier are selected from those described in Table 3, e.g., KRT86, CRISPLD1, DSG1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, ZFR2, RNF212, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1 It includes abundance values for a plurality of genes. As reported below, for example, referring to Example 3, these 24 genes were found to be differentially expressed according to the HPV status of the subject in at least 8 out of 10 training sets formed from the expression data of cervical cancer or head and neck cancer with known HPV status in The Cancer Genome Atlas (TCGA). However, those skilled in the art will understand that, in some cases, the use of different training data sets may result in different outcomes, for example, one or more of these genes may not be beneficial in at least 80% of the training folds, and / or one or more genes that were found not to be beneficial in at least 80% of the training folds in the study reported in Example 3 may be beneficial. These differences may occur, for example, when different criteria are used to select the training population, such as various inclusion and / or exclusion criteria such as cancer type, individual characteristics (e.g., age, gender, ethnicity, family history, smoking status, etc.), or simply by using a smaller or larger data set.
[0201] Thus, in some embodiments, the features of the classifier include at least 5 of the genes listed in Table 3. In some embodiments, the features of the classifier include at least 10 of the genes listed in Table 3. In some embodiments, the features of the classifier include at least 15 of the genes listed in Table 3. In some embodiments, the features of the classifier include at least 20 of the genes listed in Table 3. In some embodiments, the features of the classifier include all 24 of the genes listed in Table 3. In some embodiments, the features of the classifier include 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or all 24 of the genes listed in Table 3. Further, in some embodiments, the features of the classifier include abundance values for one or more genes not listed in Table 3. In some embodiments, the features of the classifier include abundance values for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 3. In some embodiments, the features of the classifier include abundance values for 1 to 10 genes not listed in Table 3. In some embodiments, the features of the classifier include abundance values for 1 to 5 genes not listed in Table 3. In other embodiments, the features of the classifier do not include abundance values for any genes not listed in Table 3.
[0202] Furthermore, one skilled in the art will also understand that for some features, such as the abundance values for certain genes, they may be more beneficial than other features in a particular classifier. One measure of the predictive power of each feature in a classifier based on multiple features is the regression coefficient calculated for the feature during the training of the model. The regression coefficient represents the relationship between each feature and the response of the model. The coefficient value represents the average change in the response given a one-unit increase in the feature value. Therefore, for at least variables of the same type, the magnitude of the regression coefficient, for example, the absolute value, correlates with the importance of the feature in the model. That is, the larger the magnitude of the regression coefficient, the more important the variable is to the model. For example, as reported in Example 3, in a particular support vector machine (SVM) classifier trained on the abundance values of all 24 genes listed in Table 3, as well as the mutant allele states for the TP53 and CDKN2A genes, only 6 out of the 24 genes had regression coefficients with a magnitude of at least 0.5 - CDKN2A (1.13), SMC1B (1.02), EFNB3 (-0.97), KCNS1 (0.74), CCND1 (-0.65), and RNF212 (0.517).
[0203] Therefore, one skilled in the art can, at least in part, based on the importance of each feature in one or more classification models, obtain a feature set that includes fewer genes than all those listed in Table 3 It can be selected. For example, in some embodiments, one or more genes having lower predictive power in the classification model can be omitted during classifier training. For example, in some embodiments, the features of the classifier include at least the gene expression features described in Table 5 having a regression coefficient of at least 0.5, such as CDKN2A, SMC1B, EFNB3, KCNS1, CCND1, and RNF212. In some embodiments, the features of the classifier include at least the gene expression features described in Table 5 having a regression coefficient of at least 0.4. In some embodiments, the features of the classifier include at least the gene expression features described in Table 5 having a regression coefficient of at least 0.3. In some embodiments, the features of the classifier include at least the gene expression features described in Table 5 having a regression coefficient of at least 0.2. In some embodiments, the features of the classifier include at least the gene expression features described in Table 5 having a regression coefficient of at least 0.1.
[0204] Similarly, the size of the feature set can be affected by which features are included and / or excluded. For example, in some embodiments, if certain features with high predictive power are included in the classification model, fewer total features may be included in the model. For example, in some embodiments, if the abundance values for SMC1B, CDKN2A, and EFNB3 are included in the model, it may be necessary to include the abundance values for no more than two of the other genes whose abundance values are used as features in Table 5 in the model. Thus, in some embodiments, the features of the classifier include the abundance values for SMC1B, CDKN2A, and EFNB3, and the abundance values for at least two other genes whose abundance values are used as features in Table 5. In some embodiments, the features of the classifier include the abundance values for SMC1B, CDKN2A, and EFNB3, and the abundance values for at least five other genes whose abundance values are used as features in Table 5. In some embodiments, the features of the classifier include the abundance values for SMC1B, CDKN2A, and EFNB3, and the abundance values for at least ten other genes whose abundance values are used as features in Table 5. In some embodiments, the features of the classifier include the abundance values for SMC1B, CDKN2A, and EFNB3, and the abundance values for at least fifteen other genes whose abundance values are used as features in Table 5.
[0205] Similarly, in some embodiments, if features with high predictive power are excluded from the classification model, more of the other features can be included in the model. For example, in some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 15 of the other features used as features in Table 5 are included in the model. In some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 20 of the other genes used as features in Table 5 are included in the model. In some embodiments, if the abundance values for one or more of SMC1B, CDKN2A, and EFNB3 are not included in the model, the abundance values for at least 15, 16, 17, 18, 19, 20, or all 21 of the other genes used as features in Table 5 are included in the model.
[0206] Of course, other metrics can also be utilized to evaluate the importance of features in the model, such as the change in the standardized regression coefficient and R-squared when the feature was last added to the model.
[0207] When selecting a feature set, one of ordinary skill in the art will also consider the degree to which the features are correlated with each other. Correlation is a statistical measure that indicates how linearly dependent two variables are on each other. Thus, two correlated features provide redundant information to the prediction model, which can potentially have an adverse effect on the classifier. For this reason, there are several reasons to exclude correlated features from the model. For example, the more features there are in the classifier, the more computations that need to be performed, so removing correlated features speeds up the algorithm. Removing correlated features can also remove the harmful bias resulting from the correlation from the model. Finally, removing correlated features can make the model more interpretable.
[0208] Therefore, one of ordinary skill in the art can select a feature set that includes fewer genes than all those listed in Table 3, based at least in part on the correlation of each feature in one or more classification models. In some embodiments, the selection to remove one or the other feature of the correlated feature sets is informed by the predictive power of the two features, e.g., as given by their respective regression coefficients. For example, the gene expression values for ENSG00000105278 (CXCL14) and ENSG00000077935 (SMC1B) are highly correlated (correlation = 0.718983175) in the feature sets described in Table 3. Thus, in some embodiments, the feature set does not include either CXCL14 or SMC1B. In some embodiments, as reported in Table 5, SMC1B has a higher regression coefficient (1.02) than CXCL14 (-0.29) in the SVM model described in Example 3, so CXCL14 rather than SMC1B is excluded from the feature set.
[0209] As reported in Table 6, ten pairs of gene expression features have a correlation of at least 0.6. Thus, in some embodiments, the features in at least one pair of features having a correlation of at least 0.6 are excluded from the model. In some embodiments, the features in at least two pairs of features having a correlation of at least 0.6 are excluded from the model. In other embodiments, the features in at least three, four, five, six, seven, eight, nine, or all ten pairs of features having a correlation of at least 0.6 are excluded from the model. In some embodiments, the excluded features are the features in a pair of highly correlated features having a lower regression coefficient than that reported in Table 5. For example, referring to Table 6, the features having a lower regression coefficient in each pair with high correlation (e.g., corresponding to a correlation of at least 0.6) are as follows. · Pair 1 = DSG1 · Pair 2 = ZFR2 · Pair 3 = RNF212 · Pair 4 = SYCP2 · Pair 5 = ZFR2 · Pair 6 = MYO3A · Pair 7 = SYCP2 · Pair 8 = DSG1 · Pair 9 = KCNS1 · Pair 10 = ZFR2 Thus, in some embodiments, one or more of DSG1, ZFR2, RNF212, SYCP2, MYO3A, and KCNS1 are excluded from the feature set based on them being the least beneficial features in pairs of features that are highly correlated.
[0210] However, in some embodiments, this selection process does not allow for the exclusion of both features of a pair of highly correlated features based on, for example, both genes being the least beneficial features in at least one of the pair of highly correlated features. Thus, in some embodiments, one or more of SYCP2, MYO3A, and KCNS1 are not excluded from the feature set. Similarly, in some embodiments, this selection process does not allow for the exclusion of features that are very beneficial, such as features having a regression coefficient of at least 0.5, from the feature set. Thus, in some embodiments, one or both of RNF212 and KCNS1 are not excluded from the feature set.
[0211] Thus, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0212] Similarly, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, RNF212, MKRN3, MYL1, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0213] Similarly, in one embodiment, the feature set includes abundance values for at least KRT86, CRISPLD1, SESN3, DAMTS20, IRX1, SMC1B, CDKN2A, EFNB3, CXCL14, RNF212, MKRN3, SYCP2, MYL1, MYO3A, RNASE10, GALNT13, C19orf26, MUC4, PCDHGB1, CCND1, LCE1F, and KCNS1.
[0214] In some embodiments, as described above with reference to FIG. 2, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. In some embodiments, the classifier is trained according to the above methodology with reference to FIG. 2.
[0215] In some embodiments, the classifier has at least 70% specificity and at least 70% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 75% specificity and at least 75% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 80% specificity and at least 80% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 85% specificity and at least 85% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 90% specificity and at least 90% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 95% specificity and at least 95% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more specificity for a validation dataset of at least 50 data constructs. In some embodiments, the classifier has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more sensitivity for a validation dataset of at least 50 data constructs.
[0216] In some embodiments, the classifier has at least 70% specificity and at least 70% sensitivity for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 75% specificity and at least 75% sensitivity for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 80% specificity and at least 80% sensitivity for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 85% specificity and at least 85% sensitivity for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 90% specificity and at least 90% sensitivity for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 95% specificity and at least 95% sensitivity for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in training the classifier. In some embodiments, the classifier has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more specificity for a validation dataset of at least 100 data constructs. In some embodiments, the classifier has at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more sensitivity for a validation dataset of at least 100 data constructs.
[0217] In some embodiments, the method further comprises assigning and / or implementing a therapy to a subject based on a classification of the cancerous state, e.g., based on whether the subject's cancer is associated with a human papillomavirus (HPV) viral infection.
[0218] Accordingly, in one embodiment, a method for treating cervical cancer in a human cancer patient is provided. The method includes determining whether a human cancer patient is infected with a human papillomavirus (HPV) oncogenic virus by obtaining a dataset for the human cancer patient, the dataset including a plurality of abundance values, each respective abundance value of the plurality of abundance values quantifying an expression level of a corresponding gene among a plurality of genes, the plurality of genes including at least 5 genes selected from the genes listed in Table 3. The method then includes inputting the dataset into a classifier trained to identify at least a first cervical cancer state associated with HPV infection and a second cervical cancer state associated with an HPV-free state based on the abundance values of the plurality of genes in the cancerous tissue of the subject. In some embodiments, the classifier is trained according to the methodology described above with reference to FIG. 2. The method then includes treating the cervical cancer. If the result of the classifier indicates that the human cancer patient is infected with an HPV oncogenic virus, a first therapy adjusted for the treatment of cervical cancer associated with HPV infection is implemented. If the result of the classifier indicates that the human cancer patient is not infected with an HPV oncogenic virus, a second therapy adjusted for the treatment of cervical cancer not associated with HPV infection is implemented.
[0219] In some embodiments, the plurality of genes includes at least 10 of the genes listed in Table 3. In some embodiments, the plurality of genes includes at least 15 of the genes listed in Table 3. In some embodiments, the plurality of genes is in Table 3 It includes at least 20 of the genes described. In some embodiments, the plurality of genes includes all of the genes described in Table 3. In some embodiments, the plurality of genes includes one or more genes not described in Table 3, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not described in Table 3. In some embodiments, the plurality of genes includes 20 or fewer genes. In some embodiments, the plurality of genes includes 25 or fewer genes. In some embodiments, the plurality of genes includes 50 or fewer genes. In some embodiments, the plurality of genes includes 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0220] In some embodiments, the dataset also includes mutant allele counts for one or more alleles at one or more genetic loci in the genome of the cancerous tissue from the subject. In some embodiments, the mutant allele count is either 1, which represents the state where the subject carries the mutant allele, or 0, which represents the state where the subject does not carry the mutant allele. In some embodiments, the mutant allele is a somatic mutation derived from the germline of the subject. In some embodiments, the mutant allele is a cancer-derived mutation from the cancerous tissue. In some embodiments, the mutant allele is located in the TP53 (ENSG00000141510) or CDKN2A (ENSG00000147889) gene.
[0221] In some embodiments, as described above with reference to FIG. 2, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. In some embodiments, the classifier is trained according to the above methodology with reference to FIG. 2.
[0222] In some embodiments, the first therapy adapted for the treatment of cervical cancer associated with HPV infection is a therapeutic vaccine. In some embodiments, the therapeutic vaccines are selected from axalimogene filolisbac (Advaxis), TG4001 (Transgene), GX-188E (Genexine), VGX-3100 (Inovio), MEDI-0457 (Inovio), INO-3106 (Inovio), TA-CIN (Cancer Research Technology), TA-HPV (Cancer Research Technology), ISA-101 (Isa), and PepCan (University of Arkansas).
[0223] In some embodiments, the first therapy adapted for the treatment of cervical cancer associated with HPV infection is adoptive cell therapy. In some embodiments, adoptive cell therapy includes the administration of HPV-specific T cells, as described, for example, in clinical trial IDs NCT02379520 or NCT03197025 (Baylor College of Medicine).
[0224] In some embodiments, the first therapy adapted for the treatment of cervical cancer associated with HPV infection is an immune checkpoint inhibitor. In some embodiments, the immune checkpoint inhibitor is nivolumab (Bristol-Myers Squibb).
[0225] In some embodiments, adjusted for the treatment of cervical cancer associated with HPV infection The first therapy is a PI3K inhibitor. In some embodiments, the PI3K inhibitor is AMG319 (Amgen) or BKM120 (Novartis).
[0226] Similarly, in one embodiment, a method for treating head and neck cancer in a human cancer patient is provided. The method includes determining whether a human cancer patient is infected with a human papillomavirus (HPV) oncogenic virus by obtaining a dataset for the human cancer patient, the dataset including a plurality of abundance values, each respective abundance value of the plurality of abundance values quantifying the expression level of a corresponding gene in a plurality of genes, the plurality of genes including at least five genes selected from the genes listed in Table 3. The method then includes inputting the dataset into a classifier trained to identify at least a first head and neck cancer state associated with HPV infection and a second head and neck cancer state associated with a non-HPV-containing state based on the abundance values of the plurality of genes in the cancerous tissue of the subject. In some embodiments, the classifier is trained according to the methodology described above with reference to FIG. 2. The method then includes treating the head and neck cancer. If the results of the classifier indicate that the human cancer patient is infected with an HPV oncogenic virus, the method includes administering a first therapy adjusted for the treatment of head and neck cancer associated with HPV infection. If the results of the classifier indicate that the human cancer patient is not infected with an HPV oncogenic virus, the method includes administering a second therapy adjusted for the treatment of head and neck cancer not associated with HPV infection.
[0227] In some embodiments, the plurality of genes includes at least 10 of the genes listed in Table 3. In some embodiments, the plurality of genes includes at least 15 of the genes listed in Table 3. In some embodiments, the plurality of genes includes at least 20 of the genes listed in Table 3. In some embodiments, the plurality of genes includes all of the genes listed in Table 3. In some embodiments, the plurality of genes includes one or more genes not listed in Table 3, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 3. In some embodiments, the plurality of genes includes 20 or fewer genes. In some embodiments, the plurality of genes includes 25 or fewer genes. In some embodiments, the plurality of genes includes 50 or fewer genes. In some embodiments, the plurality of genes includes 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0228] In some embodiments, the dataset also includes a mutant allele count for one or more alleles at one or more genetic loci in the genome of the cancerous tissue from the subject. In some embodiments, the mutant allele count is either 1, representing that the subject carries the mutant allele, or 0, representing that the subject does not carry the mutant allele. In some embodiments, the mutant allele is a somatic mutation derived from the germline of the subject. In some embodiments, the mutant allele is a cancer-derived mutation from the cancerous tissue. In some embodiments, the mutant allele is located in the TP53 (ENSG00000141510) or CDKN2A (ENSG00000147889) gene.
[0229] In some embodiments, as described above with reference to FIG. 2, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. In some embodiments, the classifier is trained according to the above methodology with reference to FIG. 2.
[0230] In some embodiments, the first therapy adapted for the treatment of head and neck cancer associated with HPV infection is a therapeutic vaccine. In some embodiments, the therapeutic vaccine is selected from axalimogene filolisbac (Advaxis), TG4001 (Transgene), GX-188E (Genexine), VGX-3100 (Inovio), MEDI-0457 (Inovio), INO-3106 (Inovio), TA-CIN (Cancer Research Technology), TA-HPV (Cancer Research Technology), ISA-101 (Isa), and PepCan (University of Arkansas).
[0231] In some embodiments, the first therapy adapted for the treatment of head and neck cancer associated with HPV infection is adoptive cell therapy. In some embodiments, the adoptive cell therapy includes the administration of HPV-specific T cells, as described, for example, for clinical trials ID NCT02379520 or NCT03197025 (Baylor College of Medicine).
[0232] In some embodiments, the first therapy adapted for the treatment of head and neck cancer associated with HPV infection is an immune checkpoint inhibitor. In some embodiments, the immune checkpoint inhibitor is nivolumab (Bristol-Myers Squibb).
[0233] In some embodiments, the first therapy adapted for the treatment of head and neck cancer associated with HPV infection is a PI3K inhibitor. In some embodiments, the PI3K inhibitor is AMG319 (Amgen) or BKM120 (Novartis).
[0234] HPV probe set In some embodiments, the present disclosure provides probes for binding, enriching, and / or detecting nucleic acid molecules, e.g., mRNA transcripts isolated from a cancerous tissue sample from a subject and / or cDNA molecules prepared from those mRNA transcripts, which are useful for determining whether the subject has a first cancer state associated with HPV oncogenic virus infection or a second cancer state not associated with HPV oncogenic virus infection. Generally, the probes include DNA, RNA, or modified nucleic acid structures having a base sequence complementary to the nucleic acid molecule of interest. Thus, if the probe is designed to hybridize to an mRNA molecule isolated from a cancerous tissue, the probe will include a nucleic acid sequence complementary to the coding strand of the gene from which the transcript is derived, i.e., the probe will include the antisense sequence of the gene. However, if the probe is designed to hybridize to a cDNA molecule, since the molecules of the cDNA library are double-stranded, the probe may include either a sequence complementary to the coding sequence (antisense sequence) of the gene of interest or a sequence identical to the coding sequence of the gene of interest (sense sequence).
[0235] In some embodiments, the probe includes additional nucleic acid sequences that do not share any homology to the target gene sequence. For example, in some embodiments, the probe also includes a specifier sequence, e.g., a nucleic acid sequence that includes a unique molecular identifier (UMI) that is unique to a particular cancerous tissue sample or cancer patient. Examples of specifier sequences are described, e.g., in Kivioja et al., 2011, Nat. Methods 9(1), pp. 72-74, and Islam et al., 2014, Nat. Methods 11(2), pp. 163-66, the contents of which are hereby incorporated by reference in their entirety for all purposes. Similarly, in some embodiments, the probe also includes primer nucleic acid sequences useful for amplifying the target nucleic acid molecule, e.g., using PCR. In some embodiments, the probe also includes a capture sequence designed to hybridize to an anti-capture sequence for recovering the target nucleic acid molecule from the sample.
[0236] Similarly, in some embodiments, the probe includes a non-nucleic acid affinity moiety covalently attached to a nucleic acid molecule complementary to the target gene for recovering the target nucleic acid molecule. Non-limiting examples of non-nucleic acid affinity moieties include biotin, digoxigenin, and dinitrophenol. In some embodiments, the probe is attached to a solid surface or particle, e.g., a dipstick or magnetic bead, for recovering the target nucleic acid.
[0237] Accordingly, in one embodiment, the present disclosure provides a plurality of nucleic acid probes for discriminating between a first cancer state and a second cancer state in a human subject, wherein the first cancer state is associated with infection by a human papillomavirus (HPV) oncogenic virus and the second cancer state is associated with a state that does not include HPV. The plurality of nucleic acid probes includes at least five nucleic acid probes, and each of the at least five nucleic acid probes includes a respective nucleic acid sequence that is identical or complementary to at least 10 consecutive bases of an RNA transcript of a different respective gene selected from the genes listed in Table 3.
[0238] In some embodiments, the plurality of nucleic acid probes comprises at least 10 probes having sequences that are complementary or identical to sequences from different genes listed in Table 3. In some embodiments, the plurality of nucleic acid probes comprises at least 15 probes having sequences that are complementary or identical to sequences from different genes listed in Table 3. In some embodiments, the plurality of nucleic acid probes comprises at least 20 probes having sequences that are complementary or identical to sequences from different genes listed in Table 3. In some embodiments, the plurality of nucleic acid probes comprises probes having sequences that are complementary or identical to sequences from all of the genes listed in Table 3. In some embodiments, the plurality of nucleic acid probes comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 probes having sequences that are complementary or identical to sequences from different genes listed in Table 3.
[0239] In some embodiments, the plurality of nucleic acid probes comprises one or more probes that bind to sequences of genes not listed in Table 3. In some embodiments, the plurality of nucleic acid probes comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more probes that bind to sequences of genes not listed in Table 3. In some embodiments, the plurality of nucleic acid probes comprises probes having sequences that bind to 20 or fewer genes. In some embodiments, the plurality of nucleic acid probes comprises probes having sequences that bind to 25 or fewer genes. In some embodiments, the plurality of nucleic acid probes comprises probes having sequences that bind to 50 or fewer genes. In some embodiments, the plurality of nucleic acid probes comprises probes having sequences that bind to 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0240] In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical or complementary to at least 15 contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 3. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical or complementary to at least 30 contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 3. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical or complementary to at least 15 contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 3. In some embodiments, each probe in the plurality of probes comprises a nucleic acid sequence that is identical to or complementary to at least 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, or more contiguous bases of an RNA transcript of interest, e.g., a transcript from a gene listed in Table 3.
[0241] EBV oncogenic virus infection In some embodiments, the methods described herein relate to the classification and / or treatment of cancers known to be associated with Epstein-Barr Virus (EBV) infection. As reported in Example 4 below, the 24 genes listed in Table 4 and shown in FIG. 5B were found to be differentially expressed in at least 8 of 10 training sets formed from expression data of gastric cancers with known EBV status in The Cancer Genome Atlas (TCGA). Thus, in some embodiments, the expression levels of one or more of the genes listed in Table 4 are used to classify gastric cancer as either associated with EBV infection or not associated with EBV infection. In some embodiments, the expression levels of at least 2, 3, 4, 5, 6, 7, 8, or all 9 of the genes listed in Table 4 are used to classify gastric cancer as either associated with EBV infection or not associated with EBV infection.
Table 4
[0242] In one embodiment, a method for discriminating a first cancer state and a second cancer state in a human subject is provided, wherein the first cancer state is associated with infection by Epstein-Barr virus (EBV), an oncogenic virus, and the second cancer state is associated with a state not including EBV. The method includes obtaining a data set for the subject, as described above with reference to FIG. 3. The data set includes a plurality of abundance values from the subject, and each respective abundance value among the plurality of abundance values quantifies the expression level of a corresponding gene among a plurality of genes in a cancerous tissue from the subject. In some embodiments, the plurality of genes includes at least five genes selected from the genes listed in Table 4. The method then includes inputting the data set into a classifier trained to discriminate at least the first cancer state and the second cancer state based on the abundance values of the plurality of genes. In some embodiments, the classifier is trained according to any of the methodologies described above with respect to FIG. 2.
[0243] In some embodiments, the plurality of genes includes all of the genes listed in Table 4. In some embodiments, the plurality of genes includes one or more genes not listed in Table 4, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 4. In some embodiments, the plurality of genes includes 20 or fewer genes. In some embodiments, the plurality of genes includes 25 or fewer genes. In some embodiments, the plurality of genes includes 50 or fewer genes. In some embodiments, the plurality of genes includes 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0244] In some embodiments, the dataset also includes mutant allele counts for one or more alleles at one or more loci in the genome of the cancerous tissue from the subject. In some embodiments, the mutant allele count is either 1, representing the state where the subject carries the mutant allele, or 0, representing the state where the subject does not carry the mutant allele. In some embodiments, the mutant allele is a somatic mutation derived from the germline of the subject. In some embodiments, the mutant allele is a cancer-derived mutation from the cancerous tissue. In some embodiments, the mutant allele is located in the TP53 (ENSG00000141510) or PIK3CA (ENSG00000121879) gene.
[0245] In some embodiments, the classifier is trained to determine the EBV status of a test subject having an EBV-associated cancer selected from Burkitt lymphoma, nasal-type extranodal NK / T-cell lymphoma, non-Hodgkin lymphoma, Hodgkin lymphoma, nasopharyngeal carcinoma, and gastric cancer. In some embodiments, the classifier is trained to determine the EBV status of a test patient having a specific EBV-associated cancer, e.g., Burkitt lymphoma, nasal-type extranodal NK / T-cell lymphoma, non-Hodgkin lymphoma, Hodgkin lymphoma, nasopharyngeal carcinoma, or gastric cancer. However, since classifier training is generally improved by increasing the size of the training dataset, in some embodiments, the classifier is trained on data from patients having two or more types of EBV-associated cancers, e.g., two, three, four, five, or all six of Burkitt lymphoma, nasal-type extranodal NK / T-cell lymphoma, non-Hodgkin lymphoma, Hodgkin lymphoma, nasopharyngeal carcinoma, and gastric cancer. In a specific embodiment illustrated by Example 4, the classifier is trained on patients having gastric cancer. However, in some embodiments, a classifier trained on patients having one or more types of EBV-associated cancers is useful for determining the EBV status of patients having a different type of EBV-associated cancer.
[0246] In some embodiments, the features of the classifier include abundance values for a plurality of genes selected from those listed in Table 4, for example, SCNN1A, CDX1, KCNK15, PRKCG, KRT7, NKD2, GPR158, CLDN3, and ZNF683. As reported below, for example, referring to Example 4, these nine genes were found to be differentially expressed depending on the EBV status of the subject in at least 80% of the gastric cancer training set in the Cancer Genome Atlas (TCGA). However, one of ordinary skill in the art will appreciate that, in some cases, the use of different training data sets may result in different results, for example, one or more of these genes may not be beneficial in at least 80% of the training folds, and / or one or more genes that were found not to be beneficial in at least 80% of the training folds in the study reported in Example 4 may be beneficial. These differences may occur, for example, when different criteria are used to select the training population, such as various inclusion and / or exclusion criteria such as cancer type, individual characteristics (e.g., age, gender, ethnicity, family history, smoking status, etc.), or simply using a smaller or larger data set. These differences may occur, for example, when different criteria are used to select the training population, such as various inclusion and / or exclusion criteria such as cancer type, individual characteristics (e.g., age, gender, ethnicity, family history, smoking status, etc.), or simply using a smaller or larger data set. For example, in the Cancer Genome Atlas (TCGA) gastric cancer training set, at least 80% of these nine genes were found to be differentially expressed depending on the EBV status of the subject. However, one of ordinary skill in the art will understand that, in some cases, using different training data sets may lead to different results. For example, one or more of these genes may not be beneficial in at least 80% of the training folds, and / or one or more genes that were found not to be beneficial in at least 80% of the training folds in the study reported in Example 4 may be beneficial. These differences may occur, for example, when different criteria are used to select the training population, such as various inclusion and / or exclusion criteria such as cancer type, individual characteristics (e.g., age, gender, ethnicity, family history, smoking status, etc.), or simply using a smaller or larger data set.
[0247] Thus, in some embodiments, the classifier features include at least 5 of the genes listed in Table 4. In some embodiments, the classifier features include at least 6 of the genes listed in Table 4. In some embodiments, the classifier features include at least 7 of the genes listed in Table 4. In some embodiments, the classifier features include at least 8 of the genes listed in Table 4. In some embodiments, the classifier features include all 9 of the genes listed in Table 4. Further, in some embodiments, the classifier features also include abundance values for one or more genes not listed in Table 4. In some embodiments, the classifier features include abundance values for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes not listed in Table 4. In some embodiments, the classifier features include abundance values for 1 to 10 genes not listed in Table 4. In some embodiments, the classifier features include abundance values for 1 to 5 genes not listed in Table 4. In other embodiments, the classifier features do not include abundance values for any genes not listed in Table 4.
[0248] Furthermore, one of ordinary skill in the art will also understand that for some features, such as abundance values for certain genes, they may be more beneficial than other features in a particular classifier. One measure of the predictive power of each feature in a classifier based on multiple features is the regression coefficient calculated for the feature during model training. The regression coefficient represents the relationship between each feature and the response of the model. The coefficient value represents the average change in the response that gives a one-unit increase in the feature value. Therefore, for at least variables of the same type, the magnitude of the regression coefficient, for example the absolute value, correlates with the importance of the feature in the model. That is, the larger the magnitude of the regression coefficient, the more important the variable is to the model. For example, as reported in Example 4, in a particular support vector machine (SVM) classifier trained on the abundance values of all nine of the genes listed in Table 4, as well as the mutant allele status for the TP53 and PIK3CA genes, only four of the nine genes had regression coefficients with magnitudes of at least 0.75 - SCNN1A (-1.26), KCNK15 (-1.04), KRT7 (-0.94), and CLDN3 (-1.68).
[0249] Accordingly, one of ordinary skill in the art can select a feature set that includes fewer genes than all those listed in Table 4, at least in part based on the importance of each feature in one or more classification models. For example, in some embodiments, one or more genes having lower predictive power in the classification model can be omitted during classifier training. For example, in some embodiments, the features of the classifier include at least the gene expression features listed in Table 5 that have regression coefficients of at least 0.75, such as SCNN1A (-1.26), KCNK15 (-1.04), KRT7 (-0.94), and CLDN3 (-1.68). In some embodiments, the features of the classifier include at least the gene expression features listed in Table 5 that have regression coefficients of at least 0.6.
[0250] Similarly, the size of the feature set can be affected by which features are included and / or excluded. For example, in some embodiments, if certain features with high predictive power are included in the classification model, fewer total features may be included in the model. For example, in some embodiments, if the abundance values for SCNN1A, KCNK15, KRT7, and CLDN3 are included in the model, it may be necessary to include the abundance values for one or fewer of the other genes listed in Table 4 in the model. Thus, in some embodiments, the features of the classifier include the abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and at least one other gene listed in Table 4. In some embodiments, the features of the classifier include the abundance values for SCNN1A, KCNK15, KRT7, and CLD N3, and the abundance values for at least two other genes listed in Table 4. In some embodiments, the features of the classifier include the abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and the abundance values for at least three other genes listed in Table 4. In some embodiments, the features of the classifier include the abundance values for SCNN1A, KCNK15, KRT7, and CLDN3, and the abundance values for at least four other genes listed in Table 4.
[0251] Similarly, in some embodiments, if features with high predictive power are excluded from the classification model, more of the other features may be included in the model. For example, in some embodiments, if the abundance values for one or more of SCNN1A, KCNK15, KRT7, and CLDN3 are not included in the model, the abundance values for at least four of the other genes listed in Table 4 are included in the model. In some embodiments, if the abundance values for one or more of SCNN1A, KCNK15, KRT7, and CLDN3 are not included in the model, the abundance values for all five of the other genes listed in Table 4 are included in the model.
[0252] Of course, other metrics can also be used to evaluate the importance of features in a model, such as the standardized regression coefficients and changes in R-squared when the feature was last added to the model.
[0253] When selecting a feature set, one of ordinary skill in the art will also consider the degree to which the features are correlated with each other. Correlation is a statistical measure that indicates the degree to which two variables are linearly dependent on each other. Thus, two correlated features provide redundant information to a predictive model, which can have an adverse effect on the classifier. For this reason, there are several reasons to exclude correlated features from the model. For example, the more features there are in a classifier, the more calculations that need to be performed, so removing correlated features speeds up the algorithm. Removing correlated features can also remove the harmful bias resulting from the correlation from the model. Finally, removing correlated features can make the model more interpretable. Thus, one of ordinary skill in the art may select a feature set containing fewer genes than all those listed in Table 3, at least in part based on the correlation of each feature in one or more classification models. For example, statistical analysis of the SVM model trained in Example 4 revealed that the gene expression values for ENSG00000135480 (KRT7) and ENSG00000124249 (KCNK15) are highly correlated (0.650). Thus, in some embodiments, the abundance value for one of KRT7 and KCNK15 is excluded from the feature set.
[0254] For example, in one embodiment, the feature set includes abundance values for at least SCNN1A, CDX1, KCNK15, PRKCG, NKD2, GPR158, CLDN3, and ZNF683. In another embodiment, the feature set includes abundance values for at least SCNN1A, CDX1, PRKCG, KRT7, NKD2, GPR158, CLDN3, and ZNF683.
[0255] In some embodiments, as described above with reference to FIG. 2, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. In some embodiments, the classifier is trained according to the above methodology with reference to FIG. 2.
[0256] In some embodiments, the classifier has at least 70% specificity and at least 70% sensitivity for a validation data set of at least 50 data constructs, for example Then, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has at least 75% specificity and at least 75% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has at least 80% specificity and at least 80% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has at least 85% specificity and at least 85% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has at least 90% specificity and at least 90% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has at least 95% specificity and at least 95% sensitivity for a validation dataset of at least 50 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 50 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 50 data constructs.
[0257] In some embodiments, the classifier has a specificity of at least 70% and a sensitivity of at least 70% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 75% and a sensitivity of at least 75% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 80% and a sensitivity of at least 80% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 85% and a sensitivity of at least 85% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 90% and a sensitivity of at least 90% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 95% and a sensitivity of at least 95% for a validation dataset of at least 100 data constructs. For example, none of the data constructs in the validation dataset were used in the training of the classifier. In some embodiments, the classifier has a specificity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs. In some embodiments, the classifier has a sensitivity of at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or more for a validation dataset of at least 100 data constructs.
[0258] In some embodiments, the method further comprises assigning and / or performing a therapy to a subject based on a classification of the cancerous condition, e.g., based on whether the subject's cancer is associated with Epstein-Barr virus (EBV) infection.
[0259] Accordingly, in one embodiment, a method for treating gastric cancer in a human cancer patient is provided. The method includes determining whether a human cancer patient is infected with an Epstein-Barr virus (EBV) oncogenic virus by obtaining a dataset for the human cancer patient, the dataset including a plurality of abundance values, each respective abundance value of the plurality of abundance values quantifying an expression level of a corresponding gene in a plurality of genes, the plurality of genes including at least five genes selected from the genes listed in Table 4. The method then includes inputting the dataset into a classifier trained to identify at least a first gastric cancer state associated with EBV infection and a second gastric cancer state associated with a state not including EBV based on the abundance values of the plurality of genes in the cancerous tissue of the subject. In some embodiments, the classifier is trained according to the methodology described above with reference to FIG. 2. The method then includes treating the gastric cancer. If the result of the classifier indicates that the human cancer patient is infected with an EBV oncogenic virus, a first therapy adjusted for the treatment of gastric cancer associated with EBV infection is performed. If the result of the classifier indicates that the human cancer patient is not infected with an EBV oncogenic virus, a second therapy adjusted for the treatment of gastric cancer not associated with EBV infection is performed.
[0260] In some embodiments, the plurality of genes includes all of the genes listed in Table 4. In some embodiments, the plurality of genes includes one or more genes not listed in Table 4, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more genes not listed in Table 4. In some embodiments, the plurality of genes includes 20 or fewer genes. In some embodiments, the plurality of genes includes 25 or fewer genes. In some embodiments, the plurality of genes includes 50 or fewer genes. In some embodiments, the plurality of genes includes 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0261] In some embodiments, the dataset also includes variant allele counts for one or more alleles at one or more genetic loci in the genome of the cancerous tissue from the subject. In some embodiments, the variant allele count is either 1, representing that the subject carries the variant allele, or 0, representing that the subject does not carry the variant allele. In some embodiments, the variant allele is a somatic mutation derived from the germline of the subject. In some embodiments, the variant allele is a cancer-derived mutation from the cancerous tissue. In some embodiments, the variant allele is located in the TP53 (ENSG00000141510) or PIK3CA (ENSG00000121879) gene.
[0262] In some embodiments, as described above with reference to FIG. 2, the classifier is a logistic regression algorithm, a neural network algorithm, a convolutional neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a decision tree algorithm, or a clustering algorithm. In some embodiments, the classifier is trained according to the above methodology with reference to FIG. 2.
[0263] In some embodiments, the first therapy adapted for the treatment of gastric cancer associated with EBV infection is adoptive cell therapy. In some embodiments, the adoptive cell therapy includes ATA 129 (Atara), EBVST (Tessa), or CMD-003 (Cel l Medica).
[0264] In some embodiments, the first therapy adapted for the treatment of gastric cancer associated with EBV infection is an immune checkpoint inhibitor. In some embodiments, the immune checkpoint inhibitor is pembrolizumab (Merck) or nivolumab (Bristol-Myers Squibb).
[0265] In some embodiments, the first therapy adapted for the treatment of gastric cancer associated with EBV infection is a BTK inhibitor. In some embodiments, the BTK inhibitor is ibrutinib (Pharmacyclics).
[0266] Report In some embodiments, the methods described herein include generating a patient report regarding the cancer status of a subject. The report may be presented to the patient, physician, healthcare provider, or researcher as a digital copy (e.g., a JSON object, a PDF file, or an image on a website or portal), a hard copy (e.g., printed on paper or another tangible medium), as audio (e.g., a recording or a stream), or in another format.
[0267] The report includes information related to specific characteristics of the patient's cancer, such as detected genetic mutations, epigenetic abnormalities, associated oncogenic pathogenic infections, and / or pathological abnormalities. In some embodiments, other characteristics of the patient's sample and / or clinical records are also included in the report. In some embodiments, the report includes information about clinical trials for which the patient is eligible, therapies specific to the patient's cancer, and / or potentially adverse therapeutic effects associated with specific characteristics of the patient's cancer, such as the patient's genetic mutations, epigenetic abnormalities, associated oncogenic pathogenic infections, and / or pathological abnormalities, or other characteristics of the patient's sample and / or clinical records.
[0268] In some embodiments, the results included in the report, and / or any additional results (e.g., from a bioinformatics pipeline), are used to query a database of clinical data, for example, to determine whether a particular therapy tends to be effective in treating other patients with the same or similar characteristics (e.g., slowing or stopping cancer progression).
[0269] In some embodiments, the results are used to design cell-based studies of the patient's biology, such as tumor organoid experiments. For example, the organoids can be genetically engineered to have the same characteristics as the specimen, and after exposure to the therapy, it can be observed whether the therapy may reduce the growth rate of the organoids and thus the growth rate of the patient associated with the specimen. Similarly, in some embodiments, the results are used to direct studies on tumor organoids directly derived from the patient. Examples of such experiments are described in U.S. Provisional Patent Application No. 62 / 944,292, filed December 5, 2019, the content of which is hereby incorporated by reference in its entirety for all purposes.
[0270] In some embodiments, the patient report includes a section regarding the subject's oncogenic pathogen infection status. For example, Figures 7A and 7B show examples of information provided upon diagnosis of HPV-positive head and neck cancer and HPV-positive cervical cancer, respectively.
[0271] EBV Probe Set In some embodiments, the present disclosure provides a method for the detection of a nucleic acid molecule, e.g., a cancerous tissue sample from a subject. The present invention provides probes for binding, enriching, and / or detecting mRNA transcripts isolated from and / or cDNA molecules prepared from those mRNA transcripts, which are informative as to whether a subject has a first cancer condition associated with EBV oncogenic virus infection or a second cancer condition not associated with EBV oncogenic virus infection. Generally, the probe comprises a DNA, RNA, or modified nucleic acid structure having a base sequence complementary to a nucleic acid molecule of interest. Thus, when a probe is designed to hybridize to an mRNA molecule isolated from a cancerous tissue, the probe comprises a nucleic acid sequence complementary to the coding strand of the gene from which the transcript is derived, for example, the probe will comprise an antisense sequence of the gene. However, when a probe is designed to hybridize to a cDNA molecule, since the molecules of a cDNA library are double-stranded, the probe may comprise either a sequence complementary to the coding sequence of the gene of interest (antisense sequence) or a sequence identical to the coding sequence of the gene of interest (sense sequence).
[0272] In some embodiments, the probe includes additional nucleic acid sequences that do not share any homology to the target gene sequence. For example, in some embodiments, the probe also includes nucleic acid sequences that include identifier sequences, such as unique molecular identifiers (UMIs) that are unique to a particular cancerous tissue sample or cancer patient. Examples of identifier sequences are described, for example, in Kivioja et al., 2011, Nat. Methods 9(1):72-74, and Islam et al., 2014, Nat. Methods 11(2), pp. 163-66, the contents of which are hereby incorporated by reference in their entirety for all purposes. Similarly, in some embodiments, the probe also includes primer nucleic acid sequences that are useful for amplifying the target nucleic acid molecule, for example, using PCR. In some embodiments, the probe also includes a capture sequence designed to hybridize to an anti-capture sequence for recovering the target nucleic acid molecule from the sample.
[0273] Similarly, in some embodiments, the probe includes a non-nucleic acid affinity moiety covalently attached to a nucleic acid molecule that is complementary to the target gene for recovering the target nucleic acid molecule. Non-limiting examples of non-nucleic acid affinity moieties include biotin, digoxigenin, and dinitrophenol. In some embodiments, the probe is attached to a solid surface or particle, such as a dipstick or magnetic bead, for recovering the target nucleic acid.
[0274] Accordingly, in one embodiment, the present disclosure provides a plurality of nucleic acid probes for discriminating between a first cancer state and a second cancer state in a human subject, wherein the first cancer state is associated with infection by the Epstein-Barr virus (EBV) oncogenic virus and the second cancer state is associated with a state that does not include EBV. The plurality of nucleic acid probes includes at least five nucleic acid probes, and each of the at least five nucleic acid probes includes a respective nucleic acid sequence that is identical or complementary to at least 10 consecutive bases of an RNA transcript of a different respective gene selected from the genes listed in Table 4.
[0275] In some embodiments, the plurality of nucleic acid probes comprises at least 10 probes having sequences that are complementary or identical to sequences from different genes listed in Table 4. In some embodiments, the plurality of nucleic acid probes comprises 2, 3, 4, 5, 6, 7, 8, or 9 probes having sequences that are complementary or identical to sequences from different genes listed in Table 4.
[0276] In some embodiments, the plurality of nucleic acid probes comprises one or more probes that bind to sequences of genes not listed in Table 4. In some embodiments, the plurality of nucleic acid probes comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more probes that bind to sequences of genes not listed in Table 4. In some embodiments, the plurality of nucleic acid probes comprises probes having sequences that bind to 20 or fewer genes. In some embodiments, the plurality of nucleic acid probes comprises probes having sequences that bind to 25 or fewer genes. In some embodiments, the plurality of nucleic acid probes comprises probes having sequences that bind to 50 or fewer genes. In some embodiments, the plurality of nucleic acid probes comprises probes having sequences that bind to 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, or 300 or fewer genes.
[0277] In some embodiments, each probe among the plurality of probes comprises a nucleic acid sequence that is identical or complementary to at least 15 consecutive bases of the RNA transcript of interest, e.g., the transcript derived from the genes listed in Table 4. In some embodiments, each probe among the plurality of probes comprises a nucleic acid sequence that is identical or complementary to at least 30 consecutive bases of the RNA transcript of interest, e.g., the transcript derived from the genes listed in Table 4. In some embodiments, each probe among the plurality of probes comprises a nucleic acid sequence that is identical or complementary to at least 50 consecutive bases of the RNA transcript of interest, e.g., the transcript derived from the genes listed in Table 4. In some embodiments, each probe among the plurality of probes comprises a nucleic acid sequence that is identical or complementary to at least 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, or more consecutive bases of the RNA transcript of interest, e.g., the transcript derived from the genes listed in Table 4.
[0278] Digital and Laboratory Healthcare Platforms In some embodiments, the methods and systems described above are generally utilized in combination with, or as part of, digital and laboratory healthcare platforms directed to medical and research. It should be understood that many uses of the methods and systems described above in combination with such platforms are possible. An example of such a platform is described in U.S. Patent Application No. 16 / 657,804, filed on October 18, 2019, entitled "Data Based Cancer Research and Treatment Systems and Methods", which is hereby incorporated by reference in its entirety for all purposes.
[0279] For example, the implementation of one or more embodiments of the above methods and systems may include microservices that constitute a digital and laboratory healthcare platform that supports diagnosis and treatment options for cancers associated with oncogenic pathogen infections. An embodiment may include a single microservice for performing and providing diagnosis and treatment options for cancers associated with oncogenic pathogen infections, or may include a plurality of microservices, each having a specific role in implementing one or more of the above embodiments together. In one example, a first microservice may perform a classification to provide a diagnosis to a second microservice for recommending an appropriate treatment for a cancer associated with an oncogenic pathogen infection. Similarly, the second microservice may perform a treatment analysis to provide a recommended treatment according to the above embodiments.
[0280] When the above embodiments are implemented in one or more microservices, either with or as part of a digital and laboratory healthcare platform, one or more of such microservices may be part of an order management system that adjusts the order of events as needed in an appropriate time and appropriate order required to instantiate the above embodiments. The microservice-based order management system is described, for example, in U.S. Provisional Patent Application No. 62 / 873,693, filed on July 12, 2019, titled "Adaptive Order Fulfillment and Tracking Methods and Systems", which is hereby incorporated by reference in its entirety for all purposes. disclosed in Application No. 62 / 873,693, which is hereby incorporated by reference in its entirety for all purposes.
[0281] For example, continuing with the above first and second microservices, the order management system may notify the first microservice that an order for classifying the carcinogenic pathogen state of cancer has been received and is ready for processing. When the delivery of the classification for the second microservice is ready, the first microservice may be executed and notified to the order management system. Further, the order management system includes that the first microservice has been completed, identifies that the execution parameters (prerequisites) of the second microservice are satisfied, and notifies the second microservice that the processing of an order for recommending an appropriate treatment method for cancer related to carcinogenic pathogen infection according to the above embodiment can be continued.
[0282] When the digital and laboratory healthcare platforms further include a gene analysis system, the gene analysis system may include a targeted panel and / or sequencing probes. An example of a target panel is, for example, "System and Method for Expanding Clinical Options for Cancer Patients using Integrated Genomic Profiling", which is disclosed in U.S. Provisional Patent Application No. 62 / 902,950 filed on September 19, 2019, which is hereby incorporated by reference in its entirety for all purposes. In one example, the targeted panel may enable the delivery of the results of next-generation sequencing for detecting carcinogenic pathogen infection according to one of the above embodiments. An example of the design of next-generation sequencing probes is, for example, "Systems and Methods for Next Generation Sequencing Uniform Probe Design", which is disclosed in U.S. Provisional Patent Application No. 62 / 924,073 filed on October 21, 2019, which is hereby incorporated by reference in its entirety for all purposes.
[0283] If the digital and laboratory healthcare platforms further include a bioinformatics pipeline, the methods and systems described above can be utilized after the completion or substantial completion of the systems and methods utilized in the bioinformatics pipeline. In one example, the bioinformatics pipeline can receive next-generation gene sequencing results and return a set of binary files, such as one or more BAM files, reflecting DNA and / or RNA read counts aligned to a reference genome. The methods and systems described above can be utilized, for example, to incorporate DNA and / or RNA read counts and generate a classification of the subject's carcinogenic pathogen state as a result.
[0284] If the digital and laboratory healthcare platforms further include an RNA data normalizer, any RNA read counts can be normalized prior to processing the embodiments as described above. Examples of RNA data normalizers are disclosed, for example, in U.S. Patent Application No. 16 / 581,706, filed September 24, 2019, entitled "Methods of Normalizing and Correcting RNA Expression Data", which is hereby incorporated by reference in its entirety for all purposes.
[0285] If the digital and laboratory healthcare platforms further include a gene data deconvoluter, any systems and methods for deconvolution can be utilized to analyze gene data related to a specimen having two or more biological components to determine the contribution of each component to the gene data and / or to determine which genetic data is associated with any component of the specimen if the specimen is purified. Gene Examples of data deconvolutors are, for example, both U.S. Patent Application No. 16 / 732,229, filed December 31, 2019, and PCT 19 / 69161, both entitled "Transcriptome Deconvolution of Metastatic Tissue Samples", U.S. Provisional Patent Application No. 62 / 924,054, filed October 21, 2019, entitled "Calculating Cell-type RNA Profiles for Diagnosis and Treatment", and U.S. Provisional Patent Application No. 62 / 944,995, filed December 6, 2019, entitled "Rapid Deconvolution of Bulk RNA Transcriptomes for Large Data Sets (Including Transcriptomes of Specimens Having Two or More Tissue Types)", which are disclosed in each and are hereby incorporated by reference in their entirety for all purposes.
[0286] If the digital and laboratory healthcare platforms further include automated RNA expression callers, the RNA expression levels can be adjusted to be expressed as values relative to a reference expression level, which is often done to prepare multiple RNA expression data sets for analysis and avoid artifacts that occur when there are differences in the data sets, since they are not generated using the same methods, devices, and / or reagents. Examples of automated RNA expression callers are, for example, U.S. Provisional Patent Application No. 62 / 943,712, filed December 4, 2019, entitled "Systems and Methods for Automating RNA Expression Calls in a Cancer Prediction Pipeline", which is disclosed therein and is hereby incorporated by reference in its entirety for all purposes.
[0287] Digital and laboratory healthcare platforms may further include one or more insight engines for providing information, characteristics, or determinations related to medical conditions that can be obtained based on genetic and / or clinical data related to patients and / or specimens. Exemplary insight engines may include, for example, an unknown primary tumor engine, a human leukocyte antigen (HLA) loss of heterozygosity (LOH) engine, a tumor mutation burden engine, a PD-L1 status engine, a homologous recombination deficiency engine, a cell pathway activation reporting engine, an immune infiltration engine, a microsatellite instability engine, a pathogen infection status engine, and the like. An example of an unknown primary tumor engine is disclosed, for example, in U.S. Provisional Patent Application No. 62 / 855,750, filed May 31, 2019, entitled "Systems and Methods for Multi-Label Cancer Classification," which is hereby incorporated by reference in its entirety for all purposes. An example of an HLA LOH engine is disclosed, for example, in U.S. Provisional Patent Application No. 62 / 889,510, filed Aug. 20, 2019, entitled "Detection of Human Leukocyte Antigen Loss of Heterozygosity," which is hereby incorporated by reference in its entirety for all purposes. An example of a tumor mutation burden (TMB) engine is disclosed, for example, in U.S. Provisional Patent Application No. 62 / 804,458, filed Feb. 12, 2019, entitled "Assessment of Tumor Burden Methodologies for Targeted Panel Sequencing," which is hereby incorporated by reference in its entirety for all purposes.Examples of PD-L1 status engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 854,400, filed May 30, 2019, entitled "A Pan-Cancer Model to Predict The PD-L1 Status of a Cancer Cell Sample Using RNA Expression Data and Other Patient Data", which is hereby incorporated by reference in its entirety for all purposes. Additional examples of PD-L1 status engines are, for example, "PD-L1 Predict." ion Using H&E Slide Images", disclosed in U.S. Provisional Patent Application No. 62 / 824,039, filed Mar. 26, 2019, which is hereby incorporated by reference in its entirety for all purposes. Examples of homologous recombination deficiency engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 804,730, filed Feb. 12, 2019, entitled "An Integrative Machine-Learning Framework to Predict Homologous Recombination Deficiency", which is hereby incorporated by reference in its entirety for all purposes. Examples of cell pathway activation reporter engines are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 888,163, filed Aug. 16, 2019, entitled "Cellular Pathway Report", which is hereby incorporated by reference in its entirety for all purposes. Examples of immune infiltration engines are disclosed, for example, in U.S. Patent Application No. 16 / 533,676, filed Aug. 6, 2019, entitled "A Multi-Modal Approach to Predicting Immune Infiltration Based on Integrated RNA Expression and Imaging Features", which is hereby incorporated by reference in its entirety for all purposes. Additional examples of immune infiltration engines are, for example, "Comprehensive Evaluation of RNA Immune entitled "System for the Identification of Patients with an Immunologically Active Tumor Microenvironment", disclosed in U.S. Patent Application No. 62 / 804,509, filed on February 12, 2019, which is hereby incorporated by reference in its entirety for all purposes. Examples of MSI engines include, for example, "Microsatellite Instability Determination" entitled "System and Related Methods", disclosed in U.S. Patent Application No. 16 / 653,868, filed on October 15, 2019, which is hereby incorporated by reference in its entirety for all purposes. Additional examples of MSI engines include, for example, "Systems and Methods for Detecting Microsatellite Instability of a Cancer Using a Liquid Biopsy", disclosed in U.S. Provisional Patent Application No. 62 / 931,600, filed on November 6, 2019, which is hereby incorporated by reference in its entirety for all purposes.
[0288] If the digital and laboratory healthcare platforms further include a report generation engine, the above methods and systems can be utilized to create a summary report of the patient's genetic profile and the results of one or more insight engines for presentation to a physician. For example, the report can provide the physician with information about the extent to which the sequenced specimen included tumors or normal tissue from a first organ, a second organ, a third organ, etc. For example, the report can provide a genetic profile for each of the tissue types, tumors, or organs in the specimen. The gene profile represents the gene sequences present in the tissue type, tumor, or organ and can include mutations, expression levels, information about gene products, or other information that can be derived from the genetic analysis of the tissue, tumor, or organ. The report can include therapies and / or clinical trials tailored based on some or all of the gene profile or the results and summaries of the insight engines. For example, the therapy can be titled "Therapeutic Suggestion Improvements Gained Through Genomic Biomarker Matching Plus Clinical History" and can be tailored according to the systems and methods disclosed in U.S. Provisional Patent Application No. 62 / 804,724, filed on February 12, 2019, which is hereby incorporated by reference in its entirety for all purposes. For example, the clinical trial can be titled "Systems and Methods of Clinical Trial Evaluati on" and can be tailored according to the systems and methods disclosed in U.S. Provisional Patent Application No. 62 / 855,913, filed on May 31, 2019, which is hereby incorporated by reference in its entirety for all purposes.
[0289] The report can include a comparison of the results with a database of results from many specimens. Examples of methods and systems for comparing the results with a database of results are "A Method entitled "and Process for Predicting and Analyzing Patient Cohort Response, Progression and Survival", and disclosed in U.S. Provisional Patent Application No. 62 / 786,739, filed on December 31, 2018, which is hereby incorporated by reference in its entirety for all purposes. The information can be used in combination with additional specimens and / or similar information derived from clinical response information, optionally to discover biomarkers or design clinical trials.
[0290] If digital and laboratory healthcare platforms further include the application of one or more of the embodiments herein to organoids developed in relation to the platform, the methods and systems can be used to further evaluate gene sequencing data derived from the organoids to provide information about the extent to which the sequenced organoids contained a first cell type, a second cell type, a third cell type, etc. For example, the report can provide a genetic profile for each of the cell types in the specimen. The genetic profile represents the gene sequences present in a given cell type and can include mutations, expression levels, information about gene products, or other information derived from the genetic analysis of the cell. The report can include therapies matched based on some or all of the deconvolved information. These therapies can be tested in the organoids, derivatives of the organoids, and / or similar organoids to determine the sensitivity of the organoids to those therapies. For example, the organoids are described in "Tumor Organoid Culture Compositions, Systems, and Methods", U.S. Patent Application No. 16 / 693,117, filed on November 22, 2019, "Systems and Methods for Predicting Therapeutic Sensitivity", U.S. Provisional Patent Application No. 62 / 924,621, filed on October 22, 2019, and "Large Scale Phenotypic Organoid entitled "Analysis", and may be cultured and tested according to the systems and methods disclosed in U.S. Provisional Patent Application No. 62 / 944,292, filed December 5, 2019, which is hereby incorporated by reference in its entirety for all purposes.
[0291] When digital and laboratory healthcare platforms further include one or more of the above applications, generally in combination with or as part of a medical device or laboratory development test targeting medical and research, the results of such laboratory development tests or medical devices can be enhanced and personalized through the use of artificial intelligence. Examples of laboratory development tests, particularly tests that can be enhanced by artificial intelligence, are disclosed, for example, in U.S. Provisional Patent Application No. 62 / 924,515, entitled "Artificial Intelligence Assisted Precision Medicine Enhancements to Standardized Laboratory Diagnostic Testing", filed October 22, 2019, which is hereby incorporated by reference in its entirety for all purposes.
[0292] It should be understood that the above examples are illustrative and do not limit the use of the systems and methods described herein in combination with digital and laboratory healthcare platforms.
Example
[0293] Example 1 - The Cancer Genome Atlas (TCGA) The data used for training the classifiers shown in Examples 2 and 3 below was obtained from The Cancer Genome Atlas (TCGA). Briefly, the TCGA dataset is a publicly available dataset containing over 2 petabytes of genomic data for over 11,000 cancer patients, including clinical information about cancer patients, metadata about samples collected from such patients (e.g., weight of the sample portion, etc.), histopathological slide images derived from the sample portions, and molecular information obtained from the samples (e.g., mRNA / miRNA expression, protein expression, copy number, etc.). The TCGA dataset includes data on 33 different cancers, breast (ductal carcinoma, lobular carcinoma in situ), central nervous system (glioblastoma multiforme, low-grade glioma), endocrine (adrenocortical carcinoma, papillary thyroid carcinoma, paraganglioma, and pheochromocytoma), gastrointestinal (cholangiocarcinoma, colorectal adenocarcinoma, esophageal cancer, hepatocellular carcinoma, pancreatic ductal adenocarcinoma, and gastric cancer), gynecological (cervical cancer, ovarian serous cystadenocarcinoma, uterine carcinosarcoma, and endometrial carcinoma of the uterine corpus), head and neck (head and neck squamous cell carcinoma, uveal melanoma), blood (acute myeloid leukemia, thymoma), skin (cutaneous melanoma), soft tissue (sarcoma), thoracic (lung adenocarcinoma, lung squamous cell carcinoma, and mesothelioma), and urinary (chromophobe renal cell carcinoma, clear cell renal cell carcinoma, papillary renal carcinoma, prostate adenocarcinoma, testicular germ cell carcinoma, and urothelial bladder carcinoma).
[0294] Example 2 - RNA Expression Profiling Referring to FIG. 3, the expression profiles of genes useful for determining the status of the HPV virus were determined from tumor samples of head and neck cancers.
[0295] According to block 302 of FIG. 3, using the biopsy technique described herein, a tumor biopsy of head and neck cancer was obtained from a cancer patient. The biopsy was immediately snap-frozen in liquid nitrogen after removal from the patient.
[0296] According to block 304 in FIG. 3, mRNA was isolated from the tumor sample. Briefly, the sample tissue block was taken out of liquid nitrogen, a 5 mm x 5 mm x 5 mm block of the sample was taken out and dissected using a cold knife. The dissected sample was mixed with TRIzol reagent (Chomczynski and Sacchi, 1987, Anal Biochem. 162(1), pp. 156 - 59, the content of which is hereby incorporated by reference in its entirety for all purposes) and homogenized using a tissue homogenizer for three short cycles, for example, 60 seconds, 30 seconds, and 30 seconds. Chloroform was added to the homogenized tumor sample and the reaction solution was mixed. After phase separation, the aqueous phase of the reaction solution was removed and mixed with an equal volume of isopropanol to precipitate the RNA. The reaction solution was centrifuged to pellet the RNA and the supernatant was removed. The pellet was washed twice with cold ethanol and then air-dried. Subsequently, the extracted RNA was resuspended in RNase-free water.
[0297] Next, referring to block 306 in FIG. 3, the mRNA in the isolated RNA was quantified by total exome sequencing. According to block 308 in FIG. 3, the extracted RNA was heated to disrupt the secondary structure, and then the RNA was annealed to magnetic oligo(dT) beads by incubating the RNA in hybridization buffer with oligo(dT) beads having denatured RNA at room temperature to isolate mRNA from the extracted RNA. The beads were recovered and washed twice with hybridization buffer. Subsequently, the hybridized mRNA was eluted by heating and recovered from the reaction solution.
[0298] According to block 310 in FIG. 3, a cDNA library was constructed from the isolated mRNA. Briefly, divalent cations were added to the isolated mRNA to fragment the molecules at a high temperature. The fragmented mRNA was fragmented at pH 5 using glycogen as a carrier molecule. It was precipitated by incubating at -80 °C in ethanol of 2. The mRNA was pelleted by centrifugation, washed with 70% ethanol, air-dried, and then resuspended in RNase-free water. First-strand DNA synthesis was performed using random primers and reverse transcriptase. Next, second-strand DNA synthesis was performed using DNA polymerase in the presence of RNaseH to form double-stranded cDNA. The 5'-overhangs created by second-strand synthesis were repaired using T4 and Klenow DNA polymerases to form blunt ends. The 3' ends of the blunt-ended cDNA were adenylated using Klenow DNA polymerase. An adapter was ligated to the ends of the adenylated cDNA using T4 DNA ligase, and the cDNA template was purified and sized by agarose electrophoresis. If necessary, the purified cDNA template was concentrated by PCR amplification, thereby forming the final cDNA library.
[0299] According to block 312 of FIG. 3, whole exome sequencing of the cDNA library was performed using the Integrated DNA Technologies (IDT) XGEN® LOCKDOWN® technology with the xGen Exome Research Panel. Briefly, the xGen Exome Research Panel covers a 51 Mb end-to-end tiled probe space of the human genome, providing deep and uniform coverage of target capture across the entire exome. The cDNA library was hybridized to biotinylated DNA capture probes that cover the reference human exome. The hybridized probes were recovered by binding to streptavidin beads. Post-capture PCR was performed to enrich the captured sequences. The amplified products were then sequenced using synthesis-by-synthesis (SBS) technology (Bently et al., 2008, Nature 456(7218), pp. 53-59, the content of which is hereby incorporated by reference in its entirety for all purposes).
[0300] Next, the RNA sequencing data is normalized by using gene length data, guanine-cytosine (GC) content data, and sequencing depth data for sepsis determination to reduce systematic bias by normalizing the gene length data of at least one gene, reducing systematic bias by normalizing the GC content data of at least one gene, and normalizing the sequencing depth data for each sample, as described in U.S. Provisional Patent Application No. 62 / 735,349 and U.S. Patent Application No. 16 / 581,706, the contents of which are hereby incorporated by reference in their entirety for all purposes. As described in U.S. Provisional Patent Application No. 62 / 735,349 and U.S. Patent Application No. 16 / 581,706, the RNA sequencing data is also corrected against a standard gene expression dataset by comparing the sequence data for at least one gene in the gene expression dataset to the sequence data in the standard gene expression dataset. Next, the normalized and corrected RNA expression data for the 24 genes identified in Table 3, as well as the status of the patient's CDKN2A and TP53 alleles, were input into the HPV detection classifier trained in Example 3 to determine the patient's HPV virus status.
[0301] Example 3 - Detection of Human Papillomavirus Referring to FIGS. 4A-4D, a classifier for determining the status of the HPV virus was trained using gene expression from tumor RNA-seq data of a training population in which each subject in the training population was diagnosed with head and neck squamous cell carcinoma or cervical cancer.
[0302] According to block 204 of FIG. 2A, a training dataset was obtained. Here, the dataset included the corresponding plurality of abundance values for each subject in TCGA having cervical cancer or head and neck cancer with known HPV status as described in Example 1. As shown in FIG. 4A, in TCGA, 427 subjects met these selection criteria and served as the plurality of subjects in the training dataset. Of the 427 subjects, 263 had head and neck cancer Among them, 164 had cervical cancer. Of the 263 subjects with head and neck cancer, 32 were HPV-positive and 231 were HPV-negative. Of the 164 subjects with cervical cancer, 156 were HPV-positive and 8 were HPV-negative. Therefore, among the 427 subjects, 188 subjects were considered to be in the first cancer state (suffering from HPV and having head and neck cancer or cervical cancer), and the remaining 239 subjects were considered to be in the second cancer state (not suffering from HPV but having head and neck cancer or cervical cancer).
[0303] Next, according to block 218 of FIG. 2C and block 228 of FIG. 2D, using gene expression values derived from whole-exome RNA data in the TCGA dataset for 427 subjects, a discriminant gene set was identified by regression, and the gene expression values obtained from the whole-exome mRNA expression data for 427 subjects in the TCGA dataset functioned as independent variables, and the indicator of whether each subject was in the first cancer state (suffering from HPV and having head and neck cancer or cervical cancer) or the second cancer state (not suffering from HPV but having head and neck cancer or cervical cancer) functioned as the dependent variable. More specifically, according to block 228 of FIG. 2D, a dataset consisting of 427 subjects was divided into 10 sets (10-fold cross-validation). Each set included two or more subjects suffering from the first cancer state and two or more subjects suffering from the second cancer state. For each set of subjects, the whole-exome mRNA expression data functioned as independent variables, and the indicator of whether each subject in each set had the first or second cancer state functioned as the dependent variable in a regression. Each of the 10 sets (folds) was independently subjected to the regression. Each regression (fold) was performed using L1 (LASSO) regularization according to block 238 of FIG. 2E. Since L1 regularization leads to sparse coefficients, only a very small subset of genes with non-zero coefficients for each set was obtained. Only genes with non-zero coefficients in more than 80% of the sets were included in the final model. In other words, only genes with non-zero regression coefficients for at least 8 out of the 10 sets (folds) were recognized in the discriminant gene set based on their expression data. The list of genes that met this requirement is shown in FIG. 4B, and the feature type is "gene expression". Furthermore, FIG. 6A shows a principal component analysis of the abundance values of the genes described in FIG. 4B across the training set. FIG. 6A shows that the plot of the first and second PCA values for each subject in the training set is divided into two distinguishable groups corresponding to the first cancer state (group 602) and the second cancer state (604), indicating the power of the abundance values of the genes described in FIG. 4B to distinguish between the first and second cancer states.
[0304] In some embodiments, additional genes were included in the set of discriminant genes based on the presence or absence of mutations in the additional genes (e.g., the number of mutations). In this example, as detailed in FIG. 4B, genes CDKN2A and TP53 were included in the discriminant gene set, and the characteristics for these genes were the number of times mutations were observed in each of the 427 subjects in the training set for each of these genes.
[0305] Next, according to block 242 of FIG. 2E, using each abundance value for the set of discriminant genes and each indicator of the cancer state across 427 subjects, a classifier was trained to discriminate the first and second cancer states as a function of the abundance value of each discriminant gene set. In the first model, the classifier used was a logistic regression classifier with L1 regularization, and the training was for 427 subjects, but only the TCGA gene abundance levels for the genes described in FIG. 4B where the feature is "gene expression" were used. In the second model, the classifier used was a logistic regression classifier with L1 regularization, and the training was for 427 subjects, and the TCGA gene abundance levels for the genes described in FIG. 4B where the feature is "gene expression", and the TCGA mutation counts for the two genes in FIG. 4B where the feature is "number of mutations" were used. In the third model, the classifier used was a support vector machine (SVM) classifier from Scikit - learn, as disclosed in Pedregos a et al. 2011, “Machine Learning in Python,” JMLR 12, pp. 2825 - 2830, and the training was for 427 subjects, but only the TCGA gene abundance levels for the genes described in FIG. 4B where the feature is "gene expression" were used. When validated against data from a cohort of 133 subjects with cervical or head and neck cancer and known HPV status, the classifier performed with 92.5% specificity and 89.7% sensitivity.
[0306] In the fourth model, the classifier used was the same SVM classifier, and the training was performed on 427 subjects. For the genes described in Figure 4B where the feature is "gene expression", the TCGA gene abundance levels, and for the two genes in Figure 4B where the feature is "number of mutations", the TCGA mutation counts were used. The performance of this trained classifier is reported in Figure 4C. The regression coefficients and correlation statistics for each of the features used in the model are shown in Table 5 and Table 6 respectively. The SVM parameters used were: class weight: none, decision function shape: ovo, gamma: scale, kernel: linear, probability: true, shrinking: false, and tol: 1. As shown in Figure 4C, the trained SVM predicts the cancer type of 427 subjects. That is, whether the subject has the first cancer type (suffering from HPV and having head and neck cancer or cervical cancer), or the second cancer type (not suffering from HPV but having head and neck cancer or cervical cancer), and it has 99% specificity and 99% sensitivity for the training set of 427 subjects. Next, the classifier was validated on data from a cohort of 133 subjects with cervical cancer or head and neck cancer and known HPV status. The classifier correctly identified the HPV infection status of 122 out of 133 validation subjects, with a specificity of 95% and a sensitivity of 87.5%.
Table 5
Table 6-1
Table 6-2
Table 6-3
Table 6-4
Table 6-5
Table 6-6
Table 6-7
Table 6-8
Table 6-9
[0307] To validate the model, the trained SVM classifier reported in Figure 4C was tested against a validation cohort not used in the training of the classifier. As detailed in Figure 4A, the validation dataset included the corresponding abundance values for each subject in a dataset called the "test" dataset described in Example 2, which had cervical or head and neck cancer with known HPV status. As shown in Figure 4A, 133 subjects were selected from the validation dataset that met these selection criteria and served as a plurality of subjects in the validation dataset. Of the 133 validation subjects, 93 had head and neck cancer and 40 had cervical cancer. Of the 93 subjects with head and neck cancer, 28 were HPV positive and 65 were HPV negative. Of the 40 subjects with cervical cancer, 28 were HPV positive and 12 were HPV negative. Thus, of the 133 validation subjects, 56 validation subjects were considered to have a first cancer state (suffering from HPV and having head and neck cancer or cervical cancer), and the remaining 77 validation subjects were considered to have a second cancer state (not suffering from HPV but having head and neck cancer or cervical cancer).
[0308] Each of the 133 validation subjects was run against the trained SVM, and its performance was reported in Figure 4C, and was assigned by the SVM to either the first or second cancer class That is, the gene abundance value for the genes described in Fig. 4B where the feature type is "gene expression", and the mutation count for the two genes described in Fig. 4B where the feature type is "number of mutations", were measured from tumor samples for each of the 133 validation subjects, and this data for each validation subject was individually input into the trained SVM model of Fig. 5C. As shown in Fig. 4D, the trained SVM had a specificity of 95% and a sensitivity of 88% for the cancer class across the 133 validation subjects. It was found that adding the covariance of the number of mutations of the genes TP53 and CDKN2A to the SVM did not change the accuracy, but the AUC improved from 0.97 to 0.98. This example shows that the trained SVM model can accurately predict viral infection in tumors using RNA expression data.
[0309] This example confirmed that viral infection is generally associated with upregulation of the immune response. This example further shows that virus detection based on whole transcriptome data is a useful clinical tool in itself and, in combination with existing diagnostic methods, can provide insights into the viral state and the tumor microenvironment in a single test.
[0310] Example 4 - Detection of Epstein - Barr virus Referring to Figs. 5A - 5D, a classifier for determining the EBV virus state was trained using gene expression derived from tumor RNA - seq data of a training population in which each subject in the training population was diagnosed with gastric cancer.
[0311] According to block 204 of FIG. 2A, a training data set was obtained. Here, the data set included corresponding plurality of abundance values for each subject in TCGA having gastric cancer with a known EBV status as described in Example 1. As shown in FIG. 5A, in TCGA, 212 subjects satisfied these selection criteria and served as the plurality of subjects in the training data set. Among the 212 subjects, 21 were EBV-positive and 191 were EBV-negative. Thus, among the 212 subjects, 21 subjects were regarded as being in the first cancer state (suffering from EBV and having gastric cancer), and the remaining 191 subjects were regarded as being in the second cancer state (not suffering from EBV but having gastric cancer).
[0312] Next, according to block 218 in FIG. 2C and block 228 in FIG. 2D, using the gene expression values from the whole-exome RNA data in the TCGA dataset for 212 subjects, a discriminant gene set was identified by regression. The gene expression values obtained from the whole-exome mRNA expression data for 212 subjects in the TCGA dataset functioned as independent variables, and the indicator of whether each subject was in the first cancer state (suffering from EBV and having gastric cancer) or the second cancer state (not suffering from EBV but having gastric cancer) functioned as the dependent variable. More specifically, according to block 228 in FIG. 2D, a dataset consisting of 212 subjects was divided into 10 sets (10-fold). Each set included two or more subjects suffering from the first cancer state and two or more subjects suffering from the second cancer state. The whole-exome mRNA expression data for the subjects in each set functioned as independent variables, and each set of the 10 sets (divisions) was independently subjected to a regression in which the indicator of whether each subject in each set had the first or second cancer state functioned as the dependent variable. Each regression (division) was performed using L1 (LASSO) regularization according to block 238 in FIG. 2E. Since L1 regularization leads to sparse coefficients, only a very small subset of genes with non-zero coefficients for each set was obtained. Only genes with non-zero coefficients in more than 80% of the sets were included in the final model. In other words, only genes with non-zero regression coefficients for at least 8 out of 10 sets (divisions) were recognized as the discriminant gene set based on their expression data. The list of genes that met this requirement is shown in FIG. 5B, and the feature type is "gene expression". Further, FIG. 6B shows a principal component analysis of the abundance values of the genes described in FIG. 5B across the training set. FIG. 6B shows that the plots of the first and second PCA values for each of the subjects in the training set are divided into two distinguishable groups corresponding to the first cancer state (group 606) and the second cancer state (606), indicating the power of the abundance values of the genes described in FIG. 5B to distinguish the first and second cancer states. For each of the subjects in the training set, the first and second PCA values are plotted and divided into two distinguishable groups corresponding to the first cancer state (group 606) and the second cancer state (606), indicating the power of the abundance values of the genes described in FIG. 5B to distinguish the first and second cancer states.
[0313] In some embodiments, additional genes were included in the set of discriminant genes based on the presence or absence (e.g., the number of mutations) of mutations in the additional genes. In this example, as shown in detail in FIG. 5B, the genes PIK3CA and TP53 were included in the discriminant gene set, and the characteristics for these genes were the number of times mutations were observed in each of the 212 subjects in the training set for each of these genes.
[0314] Next, according to block 242 of FIG. 2E, using each abundance value for the set of discriminant genes and each indicator of the cancer state across 212 subjects, a classifier was trained to discriminate between the first and second cancer states as a function of the abundance value of each of the set of discriminant genes. In the first model, the classifier used was a logistic regression classifier with L1 regularization, and the training was for 212 subjects, but only the TCGA gene abundance levels for the genes described in FIG. 5B where the characteristic was “gene expression” were used. In the second model, the classifier used was a logistic regression classifier with L1 normalization, and the training was for 212 subjects, and the TCGA gene abundance levels for the genes described in FIG. 5B where the characteristic was “gene expression” and the TCGA mutation counts for the two genes in FIG. 5B where the characteristic was “number of mutations” were used. In the third model, the classifier used was a support vector machine (SVM) classifier from Scikit - learn, as disclosed in Pedregosa et al. 2011, “Machine Learning in Python,” JMLR 12, pp. 2825 - 2830, and the training was for 212 subjects, but only the TCGA gene abundance levels for the genes described in FIG. 5B where the characteristic was “gene expression” were used. When validated against data from a cohort of 55 subjects with gastric cancer and known EBV status, the classifier accurately identified the EBV status of 54 or 55 validation subjects with 100% specificity and 75% sensitivity.
[0315] In the fourth model, the classifier used was the same SVM classifier, and the training was performed on 212 subjects. For the genes described in Figure 4B where the feature is "gene expression", the TCGA gene abundance levels, and for the two genes in Figure 4B where the feature is "number of mutations", the TCGA mutation counts were used. The performance of this trained classifier is reported in Figure 5C. The regression coefficients and correlation statistics for each of the features used in the model are shown in Table 7 and Table 8, respectively. The SVM parameters used were: class weight: none, decision function shape: ovo, gamma: scale, kernel: linear, probability: true, shrinking: false, and tol: 1. As shown in Figure 5C, the trained SVM predicts the cancer type of 212 subjects. That is, whether the subject has the first cancer type (affected by EBV and having gastric cancer) or the second cancer type (not affected by EBV but having gastric cancer), and it has 99% specificity and 95% sensitivity for the training set of 212 subjects. Subsequently, the classifier was validated against data from a cohort of 55 subjects with gastric cancer and known EBV status. The classifier correctly identified the EBV infection status of 54 out of 55 validation subjects, with a specificity of 100% and a sensitivity of 75%.
Table 7
Table 8-1
Table 8-2
[0316] To validate the model, the trained SVM classifier reported in Figure 5C was tested against a validation cohort that was not used in the training of the classifier. As detailed in Figure 5A, the validation dataset included the corresponding abundance values for each subject in a dataset called the “test” dataset described in Example 2, which had gastric cancer with a known EBV status. As shown in Figure 5A, 55 subjects who met these selection criteria and served as the multiple subjects of the validation dataset were selected from the validation dataset. Of the 55 validation subjects, 4 were EBV positive and 51 were EBV negative. Thus, of the 55 validation subjects, 4 validation subjects were considered to have the first cancer state (suffering from EBV and having gastric cancer), and the remaining 51 validation subjects were considered to have the second cancer state (not suffering from EBV but having gastric cancer).
[0317] Each of the 55 validation subjects was run against the trained SVM, and its performance was reported in Figure 5C, and was assigned to either the first or second cancer class by the SVM. That is, the gene abundance values for the genes described in Figure 5B where the feature type is “gene expression”, and the mutation counts for the two genes described in Figure 5B where the feature type is “number of mutations” were measured from the tumor samples for each of the 55 validation subjects, and this data for the validation subjects was individually input into the trained SVM model of Figure 5C. As shown in Figure 5D, the trained SVM had a specificity of 75% and a sensitivity of 100% for the cancer class using such data across the 55 validation subjects. This example shows that the trained SVM model can accurately predict viral infection in tumors using RNA expression data. This example confirms that viral infection is generally associated with upregulation of the immune response. This example further shows that virus detection based on whole transcriptome data is a useful clinical tool in itself and, in combination with existing diagnostic methods, can provide insights into the viral state and the tumor microenvironment in a single test.
[0318] Example 5 - Acquisition of Normalized RNA Count Data In this example, patient samples were processed by RNA whole exome short read next-generation sequencing (NGS) to generate RNA sequencing data, and the RNA sequencing data was processed with a bioinformatics pipeline to generate the RNA-seq expression profile of each patient sample. Specifically, total nucleic acids (DNA and RNA) of solid tumors were extracted from macro-dissected FFPE tissue sections, digested with proteinase K to remove proteins, and RNA was purified from the total nucleic acids with TURBO DNase-I to remove DNA. After that, the reaction solution was washed with RNA clean XP beads to remove the enzyme protein. The isolated RNA was subjected to a quality control protocol using RiboGreen fluorescent dye to determine the concentration of RNA molecules.
[0319] Library preparation was performed using the KAPA Hyper Prep Kit, and 100 ng of RNA was thermally fragmented to an average size of 200 bp in the presence of magnesium. The library was then reverse transcribed into cDNA, and the Roche SeqCap dual - end adapter was ligated to the cDNA. The cDNA library was then purified, and size selection was performed using KAPA Hyper Beads. The library was then PCR amplified for 10 cycles and purified using Axygen MAG PCR clean - up beads. Quality control was performed using the PicoGreen fluorescence kit to determine the cDNA library concentration. The cDNA libraries were then pooled for 6 - plex hybridization reactions. Each pool was treated with human COT - 1 and IDTxGen Universal Blockers and vacuum dried. The RNA pool was then resuspended in the IDT xGen Lockdown hybridization mixture, and the IDT xGen Exome Research Panel v1.0 probes were added to each pool. The pools were incubated to hybridize the probes. The pools were then mixed with streptavidin - coated beads to capture the hybridized molecules of cDNA. The pools were amplified and purified again using the KAPA HiFi Library Amplification kit and Axygen MAG PCR clean - up beads respectively. Final quality control steps including quantification of the PicoGreen pool and evaluation of the size of the pool fragments using LabChip GX Touch were performed. The pools were cluster amplified using Illumina Paired - end Cluster Kits with PhiX spike on an Illumina C - Bot2, and the resulting flow cell containing the amplified target - captured cDNA library was sequenced to an average native on - target depth of 500x on an Illumina HiSeq 4000, generating FASTQ files.
[0320] In this example, cDNA library preparation was performed on an automated system using a liquid handling robot (SciClone NGSx).
[0321] Each FASTQ file contained a list of paired-end reads generated by an Illumina sequencer, each associated with a quality assessment. Reads in each FASTQ file were processed by a bioinformatics pipeline. The FASTQ files were analyzed using FASTQ C. For each FASTQ file, each read in the file was aligned to a reference genome (GRch37) using kallisto alignment software. This alignment generated a SAM file, each SAM file was converted to BAM, the BAM files were sorted, and duplicates were marked for removal.
[0322] For each gene, the raw RNA read count for a given gene was calculated by kallisto alignment software as the sum of the probabilities that each read aligned to the gene. Thus, in this example, the raw counts were not integers. The raw read counts were saved in a tabular file for each patient, with columns representing genes and each entry representing the raw RNA read count for that gene.
[0323] The raw RNA read counts were then normalized to correct for GC content and gene length using full quantile normalization and to adjust for sequencing depth via the size factor method. The normalized RNA read counts were saved in a tabular file for each patient, with columns representing genes and each entry representing the raw RNA read count for that gene.
[0324] Cited and alternative embodiments All references cited in this specification are hereby incorporated by reference in their entirety for all purposes as if each individual publication or patent or patent application were specifically and individually indicated to be incorporated by reference.
[0325] The present invention can be implemented as a computer program product including a computer program mechanism embedded in a non-transitory computer-readable storage medium. For example, the computer program product can include program modules as shown in any combination of FIGS. 1A, 1B and / or as described in FIGS. 2A, 2B, 2C, 2D, 2E, and 3. These program modules can be stored on a CD-ROM, DVD, magnetic disk storage product, USB key, or any other non-transitory computer-readable data or program storage product.
[0326] As will be apparent to those skilled in the art, many modifications and variations of this application can be made without departing from the spirit and scope of this application. The specific embodiments described herein are provided by way of example only. The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling those skilled in the art to best utilize the invention and various embodiments with various changes suitable for the particular applications contemplated. The invention should be limited only by the terms of the appended claims along with the full scope of equivalents to which such claims are entitled.
Claims
1. 1. A method for identifying a first cancerous condition and a second cancerous condition in a human subject, wherein the first cancerous condition is associated with infection by a human papillomavirus (HPV) oncogenic virus and the second cancerous condition is associated with a non-HPV condition, the method comprising: obtaining abundance data for at least five genes listed in Table 3 from a tumor sample of said human subject; and inputting the abundance data into a classifier trained to distinguish between the first cancer state and the second cancer state based, at least in part, on the abundance of the at least five genes listed in Table 3.
2. 1. A method for identifying a first cancerous condition and a second cancerous condition in a human subject, wherein the first cancerous condition is associated with infection by Epstein-Barr virus (EBV) oncogenic virus and the second cancerous condition is associated with a condition that does not involve EBV, the method comprising: obtaining abundance data for at least five genes listed in Table 4 from a tumor sample of said human subject; and inputting the abundance data into a classifier trained to distinguish between the first cancer state and the second cancer state based, at least in part, on the abundance of the at least five genes listed in Table 4.