Systems and methods for multi-label cancer classification

A multi-label classification method using RNA and genome sequencing combined with pathology data improves tumor origin determination, addressing data inconsistencies and enhancing diagnostic accuracy and treatment recommendations for tumors of unknown origin.

JP2026009921APending Publication Date: 2026-01-21テンパスエーアイインコーポレイテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025155037
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-02-28
Filing Date
2025-09-18
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Current cancer classification methods face challenges in accurately determining tumor origin, particularly for tumors of unknown origin, due to incomplete, inconsistent, and inaccurate data from pathology reports and sequencing, leading to misdiagnosis and limited access to personalized targeted therapies.

Method used

A multi-label classification approach utilizing RNA expression data, tumor genome sequencing, and pathology data, including digital images, to improve classification accuracy by iteratively refining classifiers and incorporating heterogeneous features from various data sources.

Benefits of technology

Enhances diagnostic accuracy for tumors of unknown origin, improving classification precision and recall rates, enabling personalized treatment recommendations and expanding access to targeted therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026009921000001_ABST
    Figure 2026009921000001_ABST
Patent Text Reader

Abstract

Systems and methods for identifying a diagnosis of a cancerous state for a somatic tumor specimen of a subject are provided.SOLUTION: The method receives sequencing information comprising an analysis of a plurality of nucleic acids from a somatic tumor specimen. The method identifies a plurality of features from the sequencing information including two or more of RNA, DNA, RNA splicing, viral, and copy number features. The method provides a first subset of features and a second subset of features from the identified plurality of features as input to a first classifier and a second classifier, respectively. The method generates two or more predictions of cancer status based at least in part on the identified plurality of features from the two or more classifiers. The method combines the two or more predictions with a final classifier to identify a diagnosis of cancer status for the subject's somatic tumor specimen.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is related to and claims priority to U.S. Provisional Patent Application No. 62 / 983,488, filed February 28, 2020, entitled "Systems and Methods for Multi-Label Cancer Classification," which is incorporated herein by reference in its entirety.

[0002] This application is related to and claims priority to U.S. Provisional Patent Application No. 62 / 855,750, filed May 31, 2019, entitled "Systems and Methods for Multi-Label Cancer Classification," which is incorporated herein by reference in its entirety.

[0003] This application is related to and claims priority to U.S. Provisional Patent Application No. 62 / 847,859, filed May 14, 2019, entitled "Systems and Methods for Multi-Label Cancer Classification," which is incorporated herein by reference in its entirety.

[0004] This application is related to and claims priority to U.S. Provisional Patent Application No. 62 / 902,950, filed September 19, 2019, entitled "System and Method for Expanding Clinical Options for Cancer Patients using Integrated Genomic Profiling," which is incorporated herein by reference in its entirety.

[0005] The present disclosure relates generally to classifying patients with respect to cancer status using nucleic acid sequencing data from cancer tissue and pathology reports. [Background technology]

[0006] With current advances in targeted cancer therapy, determining tumor mutational and transcriptional status is becoming increasingly useful when determining patient care. Molecularly targeted therapies, including immunotherapy, are already offering improved treatment options for cancer patients. To take advantage of these advances, patients need comprehensive molecular tumor profiling so that optimal personalized treatment can be selected. See Kumar-Sinha et al. 2018 Nat. Biotechnol. 36, 46–60. Therapies targeting specific molecular alterations are already standard of care in some tumor types (e.g., as suggested by the National Comprehensive Cancer Network (NCCN) guidelines for melanoma, colorectal cancer, and non-small cell lung cancer). These few known alterations in the NCCN guidelines can be addressed using individual assays or small next-generation sequencing (NGS) panels. However, to maximize the number of patients benefiting from personalized oncology, molecular alterations that can be targeted with off-label drug indications, combination therapies, or tissue-agnostic immunotherapies should be evaluated. See Schwaederle et al. 2016 JAMA Oncol. 2, 1452-1459, Schwaederle et al. 2015 J Clin Oncol. 32, 3817-3825, and Wheler et al. 2016 Cancer Res. 76, 3690-3701. Large-panel NGS assays have also cast a wider net for clinical trial enrollment. See Coyne et al. 2017 Curr. Probl. Cancer 41, 182-193, and Markman 2017 Oncology 31, 158, 168.

[0007] Genomic analysis of tumors is rapidly becoming routine clinical practice to provide patient-specific treatment and improve outcomes. See Fernandes et al. 2017 Clinics 72, 588-594. In fact, recent studies indicate that clinical care is guided by NGS assay results for 30-40% of patients who undergo such testing. See Hirshfield et al. 2016 Oncologist 21, 1315-1325; Groisberg et al. 2017 Oncotarget 8, 39254-39267; Ross et al. JAMA Oncol. 1, 40-49; and Ross et al. 2015 Arch. Pathol. Lab Med. 139, 642-649. There is growing evidence that patients who receive genetically guided therapy advice experience better outcomes. For example, using a matching score (e.g., a score based on the number of treatment-related and genomic aberrations per patient), Wheler et al. (2016 Cancer Res. 76, 3690-3701) showed that patients with higher matching scores had a higher frequency of stable disease, longer time to treatment failure, and higher overall survival. Such methods may be particularly useful for patients who have already failed multiple lines of therapy.

[0008] Genome analyses may include different genes as knowledge and accepted practices within the field of genome sequencing evolve. NCBI publishes a list of genes accepted and retained as part of the human genome based on the best evidence at that time in the NCBI Genebank. Each new iteration of NCBI Genebank includes removals and additions to the gene list. Removals may include replacing, canceling, or discontinuing genes that were once retained as part of the human genome but were later found to be discarded regions of nucleotides that are not coding for any gene function. As with most molecular biology databases, records are ongoing studies and are subject to change as scientists learn more about the genes. For example, some gene records are generated as a result of gene prediction during the analysis of an organism's genome. Sequence data and / or gene prediction algorithms may change over time; that is, some records may be discontinued as new data are added to subsequent genome builds or refinements are made to the gene prediction software. Other records, particularly those of known genes, persist from one genome build to another, while the information in the records continues to be updated as new knowledge is acquired. What is needed is a way to identify new genes and / or de-identify canceled or discontinued genes as these advances in understanding the human genome are made from the original NGS results. For example, gene CYorfl5A was replaced with gene TXLNGY (taxillin gamma pseudogene, Y-linked), and genes LOC388416 and LOC400951 were discontinued because new models did not predict a gene at the previously identified location. Current NGS sequencing may not target discontinued regions because they are not currently held by the sequencing community as predictive of gene-coding regions.

[0009] Targeted therapy has demonstrated significant improvements in patient outcomes, particularly with regard to progression-free survival. See Radovich et al. 2016 Oncotarget 7, 56491-56500. Furthermore, recent evidence reported from the IMPACT trial, which involved genetic testing of advanced tumors from 3,743 patients, with approximately 19% of patients receiving matched targeted therapy based on their tumor biology, showed a response rate of 16.2% for patients receiving matched therapy compared with 5.2% for patients receiving mismatched therapy. See Bankhead. “IMPACT Trial: Support for Targeted Cancer Tx Approaches.” MedPageToday. June 5, 2018. https: / / www.medpagetoday.com / 'meetingcoverage / asco / 73291. The IMPACT study also found that the 3-year overall survival rate for patients receiving molecularly matched therapy was more than double that of unmatched patients (15% vs. 7%). See the authors' references and ASCO Post. "2018 ASCO:IMPACT Trial Matches Treatment to Genetic Changes in the Tumor to Improve Survival Across Multiple Cancer Conditions." The ASCO POST. June 6, 2018. https: / / wmv.ascopost.com / News / 58897. Estimates of the percentage of patients whose care trajectory changes as a result of genetic testing vary from approximately 10% to over 50%. See Fernandes et al. 2017 Clinics 72, 588-594.

[0010] Despite the promise of matched targeted therapy, however, many patients still lack access to this type of treatment. For example, patients with tumors of unknown origin (e.g., metastatic tumors that remain unclassified even after physician analysis) cannot be offered targeted therapy until the location of the primary tumor is identified. See, e.g., Varadhachary 2007 Gastrointest Cancer Res 1(6):229-235. Without information about the primary tumor, it is difficult to offer targeted therapy and improve patient outcomes. In some cases, cancer origin classification can be performed using RNA sequencing data (e.g., RNA-Seq), which uses gene expression to identify tumor characteristics and can provide additional information. Most studies using RNA-Seq or microarrays are limited to differentiating between a small set of cancers. See, e.g., Bloom et al. 2004 Am J Pathol., 164(1):9-16; Tschentscher et al., 2003 Can. Res., 63(10), 2578-84; Young et al., 2001 Am J Pathol., 158(5), 1639-51; and Brueffer et al. 2018 ICO Precision Oncology 2, 1-18. Furthermore, it is not always possible to determine the origin of some tumors based solely on RNA sequencing data.

[0011] The use of incomplete and / or inaccurate data in classifier training sets can lead to the training of poorly performing classifiers, thus complicating the problem of determining tumor origin. This is particularly problematic when using pathology results to train, validate, and / or implement classifiers. Pathology reports can provide valuable information for classifying tumor origin. See, for example, Leong et al. 2011 Pathobiology 78, 99–14. Unfortunately, however, there is no standardized scheme for sample annotation during pathology review that all pathologists uniformly follow. This results in many unique inputs during pathology review, many of which may indicate the same diagnostic conclusion by different pathologists. Furthermore, the information reported may vary from one pathology report to another. For example, each pathology report typically contains a subset of information about the disease, stage, grade, pathology, and sample histology. However, the type of information included varies. Furthermore, the absence of a classification or label in any field of a pathology report does not necessarily indicate that the classification or label does not apply to the sample; rather, it may be that a particular pathologist did not consider the classification or label sufficiently relevant to include in the report. Beyond labeling confusion, pathology diagnoses may also be incorrect. While the rate of misdiagnosis is not fully known, any error in cancer diagnosis can have a serious impact on a patient's health and survival. See, for example, Kantola et al. 2001 British Journal of General Practice 51,106-11; Herreros-Villanueva et al. 2012 WorldJ Gastroenterol 18(23),2902-2908; Yang et al. 2015 Cancer 121,3113-3121; and Xie et al. 2015 IntJ Clin Exp Med 8(5),6762-6772.

[0012] Further concerns exist regarding the reliability and reproducibility of sequencing data used to predict tumor origin. Unfortunately, sample handling issues (e.g., mislabeling, exchange) are rampant in all laboratory environments. See, for example, Broman et al. 2015 Genes, Genomes, Genetics 5, 2177–2186, Toker et al. 2016 FlOOOResearch 5, 2103, and Lynch et al. 2012 PLos ONE 7(8), e4185. Furthermore, distinguishing sample problems from simple misdiagnosis can be difficult (e.g., when sequencing results differ from pathology data, calling into question the legitimacy of the pathology data). While multiple sample quality control methods have been proposed (see also the same authors and Pengelly et al. 2013 Genome Medicine 5, 89), ensuring accurate data provenance remains a key issue in using sequencing data for cancer classification. Summary of the Invention

[0013] Given the above background, there is a need for improved systems and methods for classifying cancers, particularly cancers of unknown origin, for example, to improve access to personalized therapy. Advantageously, the present disclosure provides solutions to these and other shortcomings in the art. For example, in some embodiments, the systems and methods described herein utilize multiple types of information from cancer patients, including RNA expression data, tumor genome sequencing, somatic genome sequencing, and / or pathology (including digital images of pathology slides with hematoxylin and eosin and / or immunohistochemical staining), to improve difficult classifications such as tumor origin. Similarly, in some embodiments, the use of multiple types of data further facilitates multi-label classification, the output of which provides additional information upon which personalized treatment decisions can be made. In combination, multiple types of data provide supporting evidence in resolving diagnoses and / or validating classification models. In some embodiments, the methods and systems described herein use classification streams. Advantageously, these streams iteratively improve poor classifier performance caused by training with incomplete, inconsistent, and / or inaccurate training data, such as is commonly found in pathology reports. A particular use of the methods and systems described herein is to determine the origin of the tumor in patients with two or more coexisting cancer diagnoses, where knowing the correct cancer to treat can improve survival rates.

[0014] One aspect of the present disclosure provides a method for determining a set of cancer states for a subject.The method is implemented in a computer system having one or more processors and a memory that stores one or more programs for execution by one or more processors.The method proceeds by obtaining one or more data structures that collectively comprise a first plurality of sequence reads in electronic form.The first plurality of sequence reads are obtained from a plurality of RNA molecules or derivatives of the above-mentioned plurality of RNA molecules (for example, derivatives such as cDNA, or proteins).The plurality of RNA molecules are derived from somatic biopsies obtained from a subject.

[0015] The method continues by determining a first set of sequence features for the subject from the first plurality of sequence reads, and applying at least the first set of sequence features to a trained classification model, thereby obtaining, for each respective cancer condition in the set of cancer conditions, a classifier result that provides a probability that the subject has or does not have the respective cancer condition.

[0016] In some embodiments, the plurality of RNA molecules is obtained by full transcriptome sequencing. In some embodiments, the one or more data structures further include a second plurality of sequence reads and a third plurality of sequence reads. The second plurality of sequence reads is obtained from the first plurality of DNA molecules or a derivative of the above-mentioned DNA molecules (e.g., a derivative from an amplification method). The third plurality of sequence reads is obtained from the second plurality of DNA molecules or a derivative of the above-mentioned DNA molecules. The first plurality of DNA molecules is derived from a somatic biopsy obtained from the subject, and the second plurality of DNA molecules is derived from a germline sample obtained from the subject or from a collection of normal controls not including the set of cancer conditions. In such embodiments, the method further includes determining a second set of sequence features for the subject from a comparison between the second plurality of sequence reads and the third plurality of sequence reads.

[0017] In some embodiments, the applying further comprises applying at least the first set of sequence features and the second set of sequence features to a trained classification model.

[0018] In some embodiments, the method further includes obtaining a pathology report for the subject. The pathology report includes at least one of a first estimate of tumor cellularity (e.g., of a somatic biopsy), an indication of whether the subject has metastatic or primary cancer, or a tissue site of origin of the somatic biopsy. The method includes extracting a plurality of pathological features from the pathology report for the subject, including the first estimate of tumor cellularity of the somatic biopsy.

[0019] In some embodiments, the trained classification model is selected based at least in part on a plurality of pathological features.

[0020] In some embodiments, applying the features to the trained classification model further comprises applying at least the plurality of pathological features, the first set of sequence features, and the second set of sequence features to the trained classification model.

[0021] In some embodiments, the trained classification model further provides one or more treatment recommendations to the subject, or to a healthcare professional caring for the subject, based on the likelihood that the subject has or does not have each respective cancer condition in the set of cancer conditions.

[0022] A further aspect of the present disclosure provides a method for classifying a subject into a cancer state. The method includes obtaining one or more data structures in an electronic format on a computer, collectively including a first plurality of sequence reads. The first plurality of sequence reads are obtained from a plurality of RNA molecules or derivatives of the above-mentioned plurality of RNA molecules (e.g., derivatives such as cDNA, or proteins). The plurality of RNA molecules are derived from a somatic biopsy obtained from the subject. The method continues by determining a first set of sequence features for the subject from the first plurality of sequence reads. The method includes applying at least the first set of sequence features to a trained classification model, thereby obtaining a classifier result that provides the probability that the subject has or does not have the cancer state.

[0023] In some embodiments, the trained classification model further provides one or more treatment recommendations to the subject, or to a healthcare professional caring for the subject, based on the likelihood that the subject has or does not have the cancer condition.

[0024] A further aspect of the present disclosure provides a method for classifying a subject into a predicted cancer state. The method includes obtaining one or more data structures, in electronic format, collectively including a first plurality of sequence reads and an indicator of the subject's predicted cancer state. The first plurality of sequence reads are obtained from a plurality of RNA molecules or derivatives of the above-mentioned plurality of RNA molecules (e.g., derivatives such as cDNA, or proteins). The plurality of RNA molecules are derived from a somatic biopsy obtained from the subject. The method includes determining a first set of sequence features for the subject from the first plurality of sequence reads. The method includes applying at least the first set of sequence features and the indicator of the subject's predicted cancer state to a trained classification model, thereby obtaining a classifier result of the predicted cancer state. The method further includes comparing the predicted cancer state with the predicted cancer state to provide a probability that the subject has or does not have the predicted cancer state.

[0025] In some embodiments, the trained classification model further provides one or more treatment recommendations to the subject, or to a healthcare professional caring for the subject, based on the likelihood that the subject has or does not have the predicted cancer condition.

[0026] Another aspect of the present disclosure provides a classification method. The method includes obtaining, in electronic form, for each respective subject in a plurality of subjects for each respective cancer condition in a set of cancer conditions, an indication of whether each subject has cancer, a first plurality of sequence reads, and a pathology report for each subject. The first plurality of sequence reads are obtained from a plurality of RNA molecules or derivatives of the above-mentioned plurality of RNA molecules (e.g., derivatives such as cDNA, or proteins). The pathology report includes at least one of a first estimate of tumor cellularity (e.g., from a somatic biopsy), an indication of whether each subject has metastatic or primary cancer, or the tissue site from which the somatic biopsy originated. The plurality of RNA molecules are derived from a somatic biopsy obtained from each subject. The method continues by determining, for each respective subject in the plurality of subjects, a corresponding set of first sequence features for each subject from the first plurality of sequence reads for each subject. The method includes extracting a plurality of pathological features from the pathology report for each subject, including the first estimate of tumor cellularity from the somatic biopsy. The method then includes inputting at least the first set of sequence features and a plurality of pathological features for each respective subject in the plurality of subjects into an untrained classification model, thereby training the untrained classification model for an indication of whether each respective subject in the plurality of subjects has each respective cancer condition in the set of cancer conditions, and obtaining a trained classification model configured to provide, for each respective cancer condition in the set of cancer conditions, a likelihood that the test subject has or does not have the respective cancer condition.

[0027] In some embodiments, the trained classification model comprises a trained classifier stream.

[0028] In some embodiments, the method further comprises obtaining a second plurality of sequence reads and a third plurality of sequence reads for each of the plurality of subjects for each of the cancer conditions in the set of cancer conditions. In some embodiments, the second plurality of sequence reads are obtained from the first plurality of DNA molecules or derivatives of the DNA molecules described above. In some embodiments, the third plurality of sequence reads are obtained from the second plurality of DNA molecules or derivatives of the DNA molecules described above. The first plurality of DNA molecules are derived from somatic biopsies obtained from each of the subjects. The second plurality of DNA molecules are derived from germline samples obtained from each of the subjects or from a collection of normal controls that do not include the set of cancer conditions. In some embodiments, the method also comprises determining a second set of sequence features for each of the plurality of subjects from a comparison of the second plurality of sequence reads and the third plurality of sequence reads for each of the subjects.

[0029] In some embodiments, inputting the features into the untrained classification model further comprises inputting at least a plurality of pathological features, the first set of sequence features, and the second set of sequence features for each respective subject in the plurality of subjects into the untrained classification model.

[0030] A method is provided for identifying a cancer state diagnosis for a subject's somatic tumor specimen (e.g., a somatic tumor specimen of unknown origin). The method includes receiving sequencing information including an analysis of multiple nucleic acids derived from the somatic tumor specimen. The method further includes identifying multiple features from the received sequencing information, the multiple features including an RNA feature, a DNA feature, an RNA splicing feature, a viral feature, and a copy number feature. Each RNA feature is associated with a respective target region of a first reference genome and represents the abundance of corresponding sequence reads that map to the respective target region encompassed by the sequencing information. Each DNA feature is associated with a respective target region of a second reference genome and represents the abundance of corresponding sequence reads that map to the respective target region encompassed by the sequencing information. Each RNA splicing feature is associated with a respective splicing event in the respective target region of the first reference genome and represents the abundance of corresponding sequence reads that map to the respective target region with the respective splicing event encompassed by the sequencing information. The first reference genome and the second reference genome may be the same reference genome or different reference genomes. Each viral feature is associated with a respective target region of the viral reference genome and represents the abundance of corresponding sequence reads mapping to the respective target region in the viral reference genome encompassed by the sequencing information. Each copy number feature is associated with a target region of a second reference genome and represents the abundance of corresponding sequence reads mapping to the respective target region in the second reference genome encompassed by the sequencing information. The method further includes providing a subset of first features from the identified plurality of features as input to a first classifier. The method further includes providing a subset of second features from the identified plurality of features as input to a second classifier. The method further includes generating two or more predictions of cancer status from the two or more classifiers based at least in part on the identified plurality of features. The two or more classifiers include at least a first classifier and a second classifier.The method further includes combining two or more predictions in a final classifier to identify a diagnosis of cancer status for the subject's somatic tumor specimen.

[0031] In some embodiments, combining the two or more predictions in the final classifier further includes scaling each prediction of the two or more predictions based at least in part on the confidence in each respective prediction, and generating a combined prediction based at least in part on each scaled prediction of the two or more predictions.

[0032] In some embodiments, the two or more predictions comprise a first prediction from a diagnostic classifier provided with (e.g., trained on) the RNA features, a second prediction from a cohort classifier provided with the RNA features, a third prediction from a tissue classifier provided with the RNA features, a fourth prediction from a diagnostic classifier provided with the RNA splicing features, a fifth prediction from a cohort classifier provided with the RNA splicing features, a sixth prediction from a diagnostic classifier provided with the CNV features, a seventh prediction from a cohort classifier provided with the CNV features, an eighth prediction from a diagnostic classifier provided with the DNA features, and a ninth prediction from a diagnostic classifier provided with the viral features.

[0033] Other embodiments are directed to systems, portable consumer devices, and computer-readable media associated with the methods described herein. Any embodiment disclosed herein may be applied to any aspect of the methods described herein, where applicable.

[0034] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description. Only exemplary embodiments of the present disclosure are shown and described herein. As will be realized, the present disclosure is capable of other and different embodiments, and its several details can be modified in various obvious aspects, all without departing from the present disclosure. Accordingly, the drawings and description should be regarded as illustrative in nature and not restrictive.

[0035] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. [Brief explanation of the drawings]

[0036] [Figure 1] 1 illustrates a block diagram of an example computing device according to some embodiments of the present disclosure. [Figure 2A] 10 summarizes and provides a flowchart and features of a process for classifying a subject to determine, for a set of cancer conditions, the likelihood that the subject has or does not have each respective cancer condition, according to some embodiments of the present disclosure, where optional blocks are indicated by dashed boxes. [Figure 2B] 10 summarizes and provides a flowchart and features of a process for classifying a subject to determine, for a set of cancer conditions, the likelihood that the subject has or does not have each respective cancer condition, according to some embodiments of the present disclosure, where optional blocks are indicated by dashed boxes. [Figure 2C] 10 summarizes and provides a flowchart and features of a process for classifying a subject to determine, for a set of cancer conditions, the likelihood that the subject has or does not have each respective cancer condition, according to some embodiments of the present disclosure, where optional blocks are indicated by dashed boxes. [Figure 3A] 1 summarizes and provides a flowchart and features of a process for training a classifier to estimate tumor cellularity according to some embodiments of the present disclosure, where optional blocks are indicated by dashed boxes. [Figure 3B] 1 summarizes and provides a flowchart and features of a process for training a classifier to estimate tumor cellularity according to some embodiments of the present disclosure, where optional blocks are indicated by dashed boxes. [Figure 4A]

[0023] Figure 4 summarizes the predicted mutation spectrum for a cohort of 500 patients, according to some embodiments of the present disclosure. Figure 4A shows the distribution of genomic alteration types for the most commonly mutated genes. Figure 4B shows a comparison of detection assays to the MSKCC IMPACT study plotted by the prevalence of altered genes that are common features of cancer. Figure 4C shows the predicted TCGA cancer status for each sample in an exemplary cohort of 500 records, according to some embodiments of the present disclosure. [Figure 4B]

[0023] Figure 4 summarizes the predicted mutation spectrum for a cohort of 500 patients, according to some embodiments of the present disclosure. Figure 4A shows the distribution of genomic alteration types for the most commonly mutated genes. Figure 4B shows a comparison of detection assays to the MSKCC IMPACT study plotted by the prevalence of altered genes that are common features of cancer. Figure 4C shows the predicted TCGA cancer status for each sample in an exemplary cohort of 500 records, according to some embodiments of the present disclosure. [Figure 4C]

[0023] Figure 4 summarizes the predicted mutation spectrum for a cohort of 500 patients, according to some embodiments of the present disclosure. Figure 4A shows the distribution of genomic alteration types for the most commonly mutated genes. Figure 4B shows a comparison of detection assays to the MSKCC IMPACT study plotted by the prevalence of altered genes that are common features of cancer. Figure 4C shows the predicted TCGA cancer status for each sample in an exemplary cohort of 500 records, according to some embodiments of the present disclosure. [Figure 5A] 1 summarizes treatment and clinical trial matching according to some embodiments of the present disclosure. [Figure 5B] 1 summarizes treatment and clinical trial matching according to some embodiments of the present disclosure. [Figure 5C] 1 summarizes treatment and clinical trial matching according to some embodiments of the present disclosure. [Figure 5D] 1 summarizes treatment and clinical trial matching according to some embodiments of the present disclosure. [Figure 5E]1 summarizes treatment and clinical trial matching according to some embodiments of the present disclosure. [Figure 5F] 1 summarizes treatment and clinical trial matching according to some embodiments of the present disclosure. [Figure 6A] Figure 6 summarizes a comparison of tumor-only and tumor-normal analyses according to some embodiments of the present disclosure. The use of paired tumor / normal samples is further described in Example 3. Figure 6A shows the proportion of somatic mutations detected in tumor-only analyses of 50 randomly selected patient samples, due to false positives and true positives. Figure 6B provides a breakdown of somatic mutation detection in tumor-normal matched DNA sequencing versus tumor-only sequencing. [Figure 6B] Figure 6 summarizes a comparison of tumor-only and tumor-normal analyses according to some embodiments of the present disclosure. The use of paired tumor / normal samples is further described in Example 3. Figure 6A shows the proportion of somatic mutations detected in tumor-only analyses of 50 randomly selected patient samples, due to false positives and true positives. Figure 6B provides a breakdown of somatic mutation detection in tumor-normal matched DNA sequencing versus tumor-only sequencing. [Figure 7A] 1 summarizes patient classification for cancer conditions according to some embodiments of the present disclosure. [Figure 7B] 1 summarizes patient classification for cancer conditions according to some embodiments of the present disclosure. [Figure 7C] 1 summarizes patient classification for cancer conditions according to some embodiments of the present disclosure. [Figure 7D] 1 summarizes patient classification for cancer conditions according to some embodiments of the present disclosure. [Figure 8A] 1 illustrates solid biopsy imaging according to some embodiments of the present disclosure. [Figure 8B] 1 illustrates solid biopsy imaging according to some embodiments of the present disclosure. [Figure 9]9 shows an example of a patient report generated in accordance with some embodiments of the present disclosure. Figure 9 includes a report of key findings and the different tests performed to generate the key findings. Additional information, such as suggested immunotherapy targets and potential resistance to various treatments, may also be included. [Figure 10A] 10 summarizes an example of a patient report for tumor origin prediction generated in accordance with some embodiments of the present disclosure. These figures are for illustrative purposes only, and no single figure 10 contains a complete patient report in itself. [Figure 10B] 10 summarizes an example of a patient report for tumor origin prediction generated in accordance with some embodiments of the present disclosure. These figures are for illustrative purposes only, and no single figure 10 contains a complete patient report in itself. [Figure 10C] 10 summarizes an example of a patient report for tumor origin prediction generated in accordance with some embodiments of the present disclosure. These figures are for illustrative purposes only, and no single figure 10 contains a complete patient report in itself. [Figure 10D] 10 summarizes an example of a patient report for tumor origin prediction generated in accordance with some embodiments of the present disclosure. These figures are for illustrative purposes only, and no single figure 10 contains a complete patient report in itself. [Figure 10E] 10 summarizes an example of a patient report for tumor origin prediction generated in accordance with some embodiments of the present disclosure. These figures are for illustrative purposes only, and no single figure 10 contains a complete patient report in itself. [Figure 10F] 10 summarizes an example of a patient report for tumor origin prediction generated in accordance with some embodiments of the present disclosure. These figures are for illustrative purposes only, and no single figure 10 contains a complete patient report in itself. [Figure 10G]10 summarizes an example of a patient report for tumor origin prediction generated in accordance with some embodiments of the present disclosure. These figures are for illustrative purposes only, and no single figure 10 contains a complete patient report in itself. [Figure 11A] Figures 11A and 11B show examples of transcriptionally distinct clusters of patient samples according to some embodiments. For example, Figure 11A shows clustering of RNA expression data from patient samples, where clusters identify the tissue origin of the patient sample (e.g., lung vs. oral cavity) and the general cancer status (e.g., adenocarcinoma vs. squamous). Figure 11B shows clusters for patients diagnosed with sarcoma, illustrating the heterogeneity of sarcomas. Figure 11C shows UMAP clusters from patients with testicular cancer. Figure 11D shows UMAP clustering by biopsy location for neuroendocrine cancer. [Figure 11B] Figures 11A and 11B show examples of transcriptionally distinct clusters of patient samples according to some embodiments. For example, Figure 11A shows clustering of RNA expression data from patient samples, where clusters identify the tissue origin of the patient sample (e.g., lung vs. oral cavity) and the general cancer status (e.g., adenocarcinoma vs. squamous). Figure 11B shows clusters for patients diagnosed with sarcoma, illustrating the heterogeneity of sarcomas. Figure 11C shows UMAP clusters from patients with testicular cancer. Figure 11D shows UMAP clustering by biopsy location for neuroendocrine cancer. [Figure 11C] Figures 11A and 11B show examples of transcriptionally distinct clusters of patient samples according to some embodiments. For example, Figure 11A shows clustering of RNA expression data from patient samples, where clusters identify the tissue origin of the patient sample (e.g., lung vs. oral cavity) and the general cancer status (e.g., adenocarcinoma vs. squamous). Figure 11B shows clusters for patients diagnosed with sarcoma, illustrating the heterogeneity of sarcomas. Figure 11C shows UMAP clusters from patients with testicular cancer. Figure 11D shows UMAP clustering by biopsy location for neuroendocrine cancer. [Figure 11D]Figures 11A and 11B show examples of transcriptionally distinct clusters of patient samples according to some embodiments. For example, Figure 11A shows clustering of RNA expression data from patient samples, where clusters identify the tissue origin of the patient sample (e.g., lung vs. oral cavity) and the general cancer status (e.g., adenocarcinoma vs. squamous). Figure 11B shows clusters for patients diagnosed with sarcoma, illustrating the heterogeneity of sarcomas. Figure 11C shows UMAP clusters from patients with testicular cancer. Figure 11D shows UMAP clustering by biopsy location for neuroendocrine cancer. [Figure 12A] 1 summarizes an example of using clustering of RNA expression data to determine which cancer labels represent transcriptionally relevant divisions of the data, according to some embodiments of the present disclosure. [Figure 12B] 1 summarizes an example of using clustering of RNA expression data to determine which cancer labels represent transcriptionally relevant divisions of the data, according to some embodiments of the present disclosure. [Figure 12C] 1 summarizes an example of using clustering of RNA expression data to determine which cancer labels represent transcriptionally relevant divisions of the data, according to some embodiments of the present disclosure. [Figure 13] 1 is an example confusion matrix illustrating the accuracy of an example classifier trained in accordance with some embodiments of the present disclosure. [Figure 14A] Figure 1 shows an example of gene frequency correlation between patients with known cancer status (light grey bars, "Actual") and patients with predicted tumor origin (dark grey bars, "Tumor of Unknown Origin (tuo) Predicted"), organized by actual or predicted cancer status, according to some embodiments of the present disclosure. These results demonstrate that the trained model generates cancer status predictions associated with DNA mutation profiles that mimic DNA mutation profiles associated with actual cancer status. [Figure 14B]Figure 1 shows an example of gene frequency correlation between patients with known cancer status (light grey bars, "Actual") and patients with predicted tumor origin (dark grey bars, "Tumor of Unknown Origin (tuo) Predicted"), organized by actual or predicted cancer status, according to some embodiments of the present disclosure. These results demonstrate that the trained model generates cancer status predictions associated with DNA mutation profiles that mimic DNA mutation profiles associated with actual cancer status. [Figure 15A] Figure 1 summarizes examples of correlations of RNA transcript levels between patients with tumors of known origin (tko) cancer state (tko_primary for primary cancers and tko_met for metastatic cancers) and patients with tumors of unknown origin organized by cancer state, according to some embodiments of the present disclosure. These results demonstrate that the trained model generates cancer state predictions (tuo) associated with RNA expression level profiles that mimic RNA expression level profiles associated with actual primary and / or metastatic cancer states. [Figure 15B] Figure 1 summarizes examples of correlations of RNA transcript levels between patients with tumors of known origin (tko) cancer state (tko_primary for primary cancers and tko_met for metastatic cancers) and patients with tumors of unknown origin organized by cancer state, according to some embodiments of the present disclosure. These results demonstrate that the trained model generates cancer state predictions (tuo) associated with RNA expression level profiles that mimic RNA expression level profiles associated with actual primary and / or metastatic cancer states. [Figure 15C]Figure 1 summarizes examples of correlations of RNA transcript levels between patients with tumors of known origin (tko) cancer state (tko_primary for primary cancers and tko_met for metastatic cancers) and patients with tumors of unknown origin organized by cancer state, according to some embodiments of the present disclosure. These results demonstrate that the trained model generates cancer state predictions (tuo) associated with RNA expression level profiles that mimic RNA expression level profiles associated with actual primary and / or metastatic cancer states. [Figure 16]

[0023] Figure 1 shows the performance of the classification model described herein, according to some embodiments of the present disclosure. For lymph node, liver, lung, and brain cancers, varying tumor cellularity within the range of 20-100% (in 10% increments) according to some embodiments of the present disclosure does not significantly affect classifier performance. The disclosed methods produce similar results for samples associated with other cancer conditions. [Figure 17A] 1 shows the error rate of the classification model described herein, according to some embodiments of the present disclosure. [Figure 17B] 1 shows the error rate of the classification model described herein, according to some embodiments of the present disclosure. [Figure 18] 1 illustrates the performance of an example classifier trained in accordance with some embodiments of the present disclosure. [Figure 19A] 1 shows a collection of anonymized case reports of individual patients classified with cancer status, according to some embodiments of the present disclosure. [Figure 19B] 1 shows a collection of anonymized case reports of individual patients classified with cancer status, according to some embodiments of the present disclosure. [Figure 20] 1 shows an example of viral variants (e.g., viral signatures) associated with different cancer and tumor cohorts according to some embodiments of the present disclosure. [Figure 21A] 1 summarizes examples of genomic variant patterns associated with different cancer conditions, according to some embodiments of the present disclosure. [Figure 21B]1 summarizes examples of genomic variant patterns associated with different cancer conditions, according to some embodiments of the present disclosure. [Figure 21C] 1 summarizes examples of genomic variant patterns associated with different cancer conditions, according to some embodiments of the present disclosure. [Figure 22] 1 illustrates an example of cancer status label determination according to some embodiments of the present disclosure. [Figure 23] 1 illustrates an artificial intelligence system for receiving a patient's health information and generating a prediction of the origin of the patient's tumor, according to some embodiments of the present disclosure. [Figure 24] FIG. 24 illustrates a stacked TUO classification using the artificial intelligence engine of FIG. 23 to predict cancer status in patients with tumors of unknown origin, according to some embodiments of the present disclosure. [Figure 25] 25 shows classification results from four of the lower-level model classifiers of FIG. 24 according to some embodiments of the present disclosure. [Figure 26] 26 illustrates meta-classification results combining the sum model classifiers of FIG. 25 according to some embodiments of the present disclosure. [Figure 27] 10 illustrates an example of a feature importance heatmap across each of the lower-level model classifiers, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0037] Like reference numerals refer to corresponding parts throughout the several views of the drawings.

[0038] To maximize the use of newly developed targeted therapies, it is essential to determine the specific cancer status affecting cancer patients. The present disclosure provides systems and methods useful for determining a patient's cancer status using RNA sequence features and features extracted from the patient's pathology report. In some embodiments, the present method employs a multi-label classification approach, annotating patient samples with a combination of genomic, pathological, and / or clinical features. The inclusion of these heterogeneous features determined from different attributes of the patient's medical history contributes to the clinically relevant accuracy of the classification disclosed herein across multiple tumor types. The present disclosure provides improved methods, particularly for classifying tumors of unknown origin.

[0039] In some embodiments, the systems and methods described herein employ a classification stream as the classification model. Advantageously, this facilitates refinement of the classifier over time and is particularly useful for initially training a classifier using unreliable data (e.g., data from pathology reports). In some embodiments, the systems and methods described herein employ an ensemble of adaptable classifiers as the classification model, for example, where the output of a first classifier helps define the structure of a downstream classification cascade (e.g., a classifier chain). Advantageously, these classifier ensembles improve performance when input test data, for example, from pathology reports, is inaccurate, inconsistent, and / or incomplete.

[0040] In one aspect, the present disclosure provides a method for training a classification model to determine the likelihood that a patient has or does not have a cancer condition. The present disclosure further provides systems and methods useful for predicting the type of treatment for a cancer patient based on whether the likelihood suggests that the patient has or does not have the respective cancer condition.

[0041] advantage In some embodiments, the present disclosure provides systems and methods for determining the cancer status of tumors of unknown origin using sequencing and pathology report data. Tumors of unknown origin account for an estimated 5% of cancer patients. See, for example, Fizazi et al. 2011 Annals of Oncology 22(6), pp. 664-668 and Example 4. As discussed in Example 4, the classification method disclosed herein enabled classification of cancer type for 867 subjects (7.6% of the sample set) who previously had only tumors of unknown origin. Advantageously, combining sequencing data and pathology report information to provide a diagnosis of tumors of unknown origin can also change patient diagnosis and clinical treatment recommendations (e.g., by providing improved recommendations over the initial diagnosis). For example, as described in the case report in Example 8, determining the origin of the tumor using the classification method described herein changed the treatment strategy for a patient with two pre-existing cancer diagnoses and a newly detected metastatic tumor.

[0042] Standard methods for molecular cancer classification only use sequencing data, which reduces diagnostic accuracy. For example, Sveen et al. (2017) developed an improved molecular classifier for colorectal cancer that demonstrated an accuracy rate of 85-92%, while classification methods trained according to embodiments described herein have precision and recall rates of 93% and 96% for colon cancer. See Clin Cancer Res 24(4), 794-806. Similarly, another 2019 study developed a molecular classifier for breast cancer that provided an average accuracy of 80%, while classification methods trained according to embodiments described herein have precision and recall rates of 95% and 96% for breast cancer. See Tao et al. (2019) Genes 10, 200. As described in Example 4, the methods described herein are applicable to a wide variety of patients with tumors or tumor origins of unknown origin.

[0043] Diagnostic information in pathology reports is typically recorded in free-form text boxes and requires some processing before it can be incorporated into classification models. As described in Example 5, the present disclosure advantageously presents methods for natural language processing of diagnostic values ​​from pathology reports. This allows for clustering of patient data into clinically and transcriptionally relevant diagnostic categories, as described in Example 6. Thus, embodiments of the present disclosure allow previously inaccessible data to be incorporated into training classification models, helping to improve the classification accuracy provided by these models.

[0044] In some embodiments, the present disclosure provides systems and methods for classifying cancer that utilize tumor and matched germline tissue sequencing data. For example, in some embodiments, the systems and methods provided herein use a plurality of sequence reads obtained from a somatic biopsy from a subject and another plurality of sequence reads obtained from a germline (non-cancerous) sample to classify the subject's cancer status. Advantageously, by using sequencing data from both the tumor sample (e.g., somatic) and the matched germline (non-cancerous) tissue, a more accurate picture of the patient's tumor biology is achieved because "false positive" somatic variants are identified (e.g., as discussed in Example 3, comparing somatic variants to germline variants filters out more than 20% of somatic variants, identifying them as false positives). The use of non-cancerous samples helps remove background mutations (e.g., mutations present in the subject but not associated with the subject's tumor). For example, as shown in Example 3 and Figure 6B, the use of sequencing data from both tumor samples and matched normal tissues reduced the false positive rate, provided more accurate classification results, and improved actionable outcomes. In particular, Example 3 shows that 16% of the subjects analyzed would have received a different clinical diagnosis if they had undergone tumor-only testing.

[0045] The methods described herein contrast with conventional methods used to classify a subject's cancer status: classifiers trained according to the embodiments described herein provide improved predictive results for tumors of unknown origin, and therefore, result in improved patient outcomes compared to other classification methods.

[0046] definition The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used in the description of the invention and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It will also be understood that the term "and / or," as used herein, refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that, as used herein, the terms "includes," "comprising," or any variation thereof, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, to the extent that the terms "including," "includes," "having," "has," "with," or variations thereof, are used in either the detailed description and / or claims, such terms are intended to be inclusive, similar to the term "comprising."

[0047] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "response to determining" or "response to detecting," depending on the context. Similarly, the phrases "when it is determined" or "when [a stated condition or event] is detected" may be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [a stated condition or event]" or "in response to detecting [a stated condition or event]," depending on the context.

[0048] It should also be understood that, although terms such as "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first subject may be referred to as a second subject, and similarly, a second subject may be referred to as a first subject, without departing from the scope of the present disclosure. The first subject and the second subject are both subjects, but are not the same subject. Furthermore, the terms "subject," "user," and "patient" are used interchangeably herein.

[0049] As used herein, the term "subject" or "patient" refers to any living or non-living human (e.g., a male human, a female human, a fetus, a pregnant woman, a child, etc.). In some embodiments, the subject is a male or female (e.g., a man, a woman, or a child) at any time.

[0050] As used herein, the terms "control," "control sample," "reference," "reference sample," "normal," and "normal sample" describe a sample from a subject without a particular condition or otherwise healthy. In one example, the methods disclosed herein can be performed on a subject with a tumor, and the reference sample is a sample taken from the subject's healthy tissue. The reference sample can be obtained from the subject or a database. The reference can be, for example, a reference genome used to map sequence reads obtained from sequencing a sample from the subject. The reference genome can refer to a haploid genome or a diploid genome to which sequence reads from a biological sample and a constitutional sample can be aligned and compared. An example of a constitutional sample can be the DNA of white blood cells obtained from a subject. For a haploid genome, only one nucleotide can be present at each locus. For a diploid genome, heterozygous loci can be identified, and each heterozygous locus can have two alleles, with either allele allowing for alignment to the locus.

[0051] As used herein, the term "locus" refers to a position (e.g., site) within a genome, e.g., on a particular chromosome. In some embodiments, a locus refers to a single nucleotide position within a genome, such as on a particular chromosome. In some embodiments, a locus refers to a small group of nucleotide positions within a genome, defined by mutations (e.g., substitutions, insertions, or deletions) of consecutive nucleotides, e.g., in a cancer genome. Because normal mammalian cells have diploid genomes, a normal mammalian genome (e.g., a human genome) generally has two copies of every locus in the genome, or at least two copies of every locus located on an autosome, e.g., one copy on the maternal autosome and one copy on the paternal autosome.

[0052] As used herein, the term "allele" refers to a particular sequence of one or more nucleotides at a chromosomal locus.

[0053] As used herein, the term "reference allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is either the dominant allele (e.g., the "wild-type sequence") that appears at that chromosomal locus within a collection of a species, or an allele that is predefined within a reference genome for that species.

[0054] As used herein, the term "variant allele" refers to a sequence of one or more nucleotides at a chromosomal locus that is not the dominant allele (e.g., not the "wild-type sequence") represented at that chromosomal locus within a collection of species or that is not a predefined allele within a reference genome for that species.

[0055] As used herein, the terms "single nucleotide variant," "SNV," "single nucleotide polymorphism," or "SNP" refer to a substitution of one nucleotide for a different nucleotide at a position (e.g., site) in a nucleotide sequence, e.g., a sequence read from an individual. A substitution of a first nucleobase X with a second nucleobase Y can be represented as "X>Y." For example, a cytosine to thymine SNP can be represented as "OT." The term "het-SNP" refers to a heterozygous SNP, where the genome is at least diploid and at least one (but not all) of two or more homologous sequences exhibits a particular SNP. Similarly, a "hom-SNP" is a homologous SNP, where each homologous sequence in a polyploid genome has an identical variant compared to a reference genome. As used herein, the term "structural variant" or "SV" refers to a large (e.g., greater than 1 kb) region of a genome that has undergone physical transformation, such as an inversion, insertion, deletion, or duplication. (See, for example, the overview of human genome SVs by Spielmann et al., 2018, Nat Rev Genetics 19:453-467).

[0056] As used herein, the term "indel" refers to an insertion and / or deletion event of a stretch of one or more nucleotides within a single locus or across multiple genes.

[0057] As used herein, the term " copy number variant ", " CNV " or " copy number variation " refers to the region of genome that is repeated. These can be categorized as short repeats or long repeats according to the number of nucleotides that are repeated across genome regions. Long repeats typically refer to the case where the entire gene or a large part of a gene is repeated one or more times.

[0058] As used herein, the term "mutation" refers to a detectable change in the genetic material of one or more cells. In certain instances, one or more mutations can be found in cancer cells and identify them (e.g., driver mutations and passenger mutations). Mutations can be transmitted from parent cells to daughter cells. Those skilled in the art will understand that a genetic mutation in a parent cell (e.g., a driver mutation) can induce additional, different mutations (e.g., passenger mutations) in daughter cells. Mutations generally occur in nucleic acids. In certain instances, a mutation can be a detectable change in one or more deoxyribonucleic acids or fragments thereof. Mutations generally refer to nucleotides that are added, deleted, substituted, inverted, or transposed to a new position in a nucleic acid. Mutations can be spontaneous or experimentally induced. A mutation in a specific tissue sequence is an example of a "tissue-specific allele." For example, a tumor can have a mutation that results in an allele at a locus that does not occur in normal cells. Another example of a "tissue-specific allele" is a fetal-specific allele that occurs in fetal tissue but not in maternal tissue.

[0059] As used herein, the term "genomic variant" can refer to one or more mutations, copy number variants, indels, single nucleotide variants, or variant alleles. Genomic variants can also refer to a combination of one or more of the above.

[0060] As used herein, the terms "cancer," "cancerous tissue," or "tumor" refer to an abnormal mass of tissue whose growth exceeds and is uncoordinated with the growth of normal tissue. In the case of blood cancer, this includes a volume of blood or other bodily fluid containing cancer cells. Cancers or tumors can be defined as "benign" or "malignant" depending on characteristics such as the degree of cellular differentiation, including morphology and functionality, growth rate, local invasion, and metastasis. "Benign" tumors are well differentiated, characteristically grow slower than malignant tumors, and remain localized at the site of origin. In addition, in some cases, benign tumors do not have the ability to infiltrate, invade, or metastasize to distant sites. "Malignant" tumors may be poorly differentiated (anaplastic) and characteristically grow rapidly with progressive infiltration, invasion, and destruction of surrounding tissue. Furthermore, malignant tumors may have the ability to metastasize to distant sites. Thus, cancer cells are cells found within an abnormal mass of tissue whose growth is uncoordinated with the growth of normal tissue. Thus, a "tumor sample" or "somatic biopsy" refers to a biological sample obtained from or derived from a tumor of a subject as described herein.

[0061] As used herein, the term "tumor cellularity" refers to the relative proportion of tumor cells (e.g., cancer cells) to normal cells in a sample. Normal cells may include normal tissue, normal stroma, and normal immune cells. A subject's tumor cellularity can be estimated from a subject's biological sample and included in the subject's pathology report.

[0062] As used herein, the term "somatic biopsy" refers to a biopsy of a subject. In some embodiments, the biopsy is a solid tissue biopsy. In some embodiments, the biopsy is a liquid biopsy.

[0063] As used herein, the terms "sequencing," "sequence determination," and the like generally refer to any and all biochemical processes that can be used to determine the order of biological macromolecules, such as nucleic acids or proteins. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule, such as an mRNA transcript or a genomic locus.

[0064] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence generated by any sequencing process described herein or known in the art. A read can be generated from one end of a nucleic acid fragment (a "single-end read"), or sometimes from both ends of the nucleic acid (e.g., paired-end read, dual-end read). The length of a sequence read is often associated with a particular sequencing technology. High-throughput methods provide sequence reads that can vary in size, for example, from tens to hundreds of base pairs (bp). In some embodiments, the sequence reads are sequence reads of an average, median, or representative length between about 15 bp and 900 bp (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, about 500 bp, etc.). In some embodiments, sequence reads have an average, median, or representative length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. Nanopore sequencing can provide sequence reads that can vary in size, for example, from tens to hundreds to thousands of base pairs. Illumina parallel sequencing can provide sequence reads that are less variable, for example, most of the sequence reads can be less than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides (about 20 to about 150) from a portion of a nucleic acid fragment, a string of nucleotides at one or both ends of a nucleic acid fragment, or the nucleotides of the entire nucleic acid fragment.Sequence reads can be obtained, for example, using sequencing techniques, or in a variety of ways using probes, for example, hybridization arrays or capture probes, or amplification techniques, for example, polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.

[0065] As used herein, the term "read segment" or "read" refers to any nucleotide sequence comprising a sequence read obtained from an individual and / or a nucleotide sequence derived from an initial sequence read from a sample obtained from an individual. For example, a read segment can refer to an aligned sequence read, a folded sequence read, or a stitched read. Additionally, a read segment can refer to an individual nucleotide base, such as a single nucleotide variant.

[0066] As used herein, the terms "read depth," "sequencing depth," or "depth" refer to the total number of read segments from a sample obtained from an individual at a given location, region, or locus. A locus can be as small as a nucleotide, as large as a chromosome arm, or as large as the entire genome. Sequencing depth can be expressed as "Y-fold," for example, 50-fold, 100-fold, etc., where "Y" refers to the number of times a locus is covered by sequence reads. In some embodiments, depth refers to the average sequencing depth across a genome, an exome, or a targeted sequencing panel. Sequencing depth can also apply to multiple loci, the whole genome, in which case Y can refer to the average number of times a locus or a haploid genome, a whole genome, or a whole exome is sequenced, respectively. When average depth is quoted, the actual depth for different loci included in a dataset can range across a range of values. Ultra-deep sequencing can refer to a sequencing depth of at least 100-fold at a locus.

[0067] As used herein, the term "sequencing breadth" refers to the analyzed fraction of a particular reference exome (e.g., a human reference exome), a particular reference genome (e.g., a human reference genome), or a portion of an exome or genome (e.g., as represented by the gene list in Table 2). The denominator of this fraction may be the repeat-masked genome, and thus 100% may correspond to the entire reference genome minus all masked portions. A repeat-masked exome or genome may refer to an exome or genome in which sequence repeats are masked (e.g., sequence reads aligned to unmasked portions of the exome or genome). Any portion of the exome or genome can be masked, and thus any specific portion of the reference exome or genome can be focused on. Broad-based sequencing may refer to sequencing and analyzing at least 0.1% of the exome or genome.

[0068] As used herein, the term "reference exome" refers to any particular known, sequenced, or characterized exome, whether partial or complete, of any tissue from any organism or pathogen that can be used to reference sequences identified from a subject. Exemplary reference exomes used for human subjects and many other organisms are provided in the online GENCODE database hosted by the GENCODE Consortium, e.g., Release 29 of the Human Exome Assembly (GRCh38.p12).

[0069] As used herein, the term "reference genome" refers to any specific known, sequenced, or characterized genome of any organism or pathogen, whether partial or complete, that can be used to reference sequences identified from a subject. Exemplary reference genomes used for human subjects and many other organisms are provided in online genome browsers hosted by the National Center for Biotechnology Information ("NCBI") or the University of California, Santa Cruz (UCSC). "Genome" refers to the complete genetic information of an organism or pathogen expressed in nucleic acid sequences. As used herein, a reference sequence or genome is often an assembled or partially assembled genome sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. A reference genome can be viewed as a representative set of genes or gene sequences for a species. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (UCSC equivalent: hgl6), NCBI build 35 (UCSC equivalent: hgl7), NCBI build 36.1 (UCSC equivalent: hgl8), GRCh37 (UCSC equivalent: hgl9), and GRCh38 (UCSC equivalent: hg38).

[0070] As used herein, the term "assay" refers to a technique for determining the characteristics of a substance, such as a nucleic acid, a protein, a cell, a tissue, or an organ. An assay (e.g., a first assay or a second assay) can include a technique for determining the copy number variation of a nucleic acid in a sample, the methylation status of a nucleic acid in a sample, the fragment size distribution of a nucleic acid in a sample, the mutation status of a nucleic acid in a sample, or the fragmentation pattern of a nucleic acid in a sample. Any assay known to those skilled in the art can be used to detect any of the nucleic acid characteristics mentioned herein. Nucleic acid characteristics can include sequence, genomic identity, copy number, methylation status at one or more nucleotide positions, nucleic acid size, the presence or absence of mutations in a nucleic acid at one or more nucleotide positions, and the fragmentation pattern of a nucleic acid (e.g., the nucleotide positions at which the nucleic acid fragments). Assays or methods can have a particular sensitivity and / or specificity, and their relative usefulness as diagnostic tools can be measured using the ROC-AUC statistic.

[0071] The term "classification" can refer to any number or other feature associated with a particular characteristic of a sample. For example, a "+" symbol (or the word "positive") can indicate that the sample is classified as having a deletion or amplification. In another example, the term "classification" can refer to the oncogenic pathogen infection status, the amount of tumor tissue in the subject and / or sample, the size of the tumor in the subject and / or sample, the stage of the tumor in the subject, the tumor burden in the subject and / or sample, and the presence of tumor metastasis in the subject. Classifications can be binary (e.g., positive or negative) or can have more classification levels (e.g., a scale of 1 to 10 or 0 to 1). The terms "cutoff" and "threshold" can refer to a predetermined number used in a given operation. For example, a cutoff size can refer to the size above which fragments are excluded. A threshold can be a value above or below which a particular classification applies. Any of these terms can be used in any of these contexts.

[0072] As used herein, the term "relative abundance" can refer to the ratio of a first amount of nucleic acid fragments with a particular characteristic (e.g., aligning with a particular region of the exome) to a second amount of nucleic acid fragments with a particular characteristic (e.g., aligning with a particular region of the exome).In one example, relative abundance can refer to the ratio of the number of mRNA transcripts that encode a particular gene in a sample (e.g., aligning with a particular region of the exome) to the total number of mRNA transcripts in the sample.

[0073] As used herein, the term "untrained classifier" refers to a classifier that has not been trained on a training dataset or that has been partially trained on a training dataset.

[0074] As used herein, an "effective amount" or a "therapeutically effective amount" is an amount sufficient to produce beneficial or desired clinical results during treatment. An effective amount can be administered to a subject in one or more doses. In terms of treatment, an effective amount is an amount sufficient to slow, alleviate, stabilize, reverse, or delay the progression of a disease, or otherwise reduce the pathological consequences of a disease. An effective amount is generally determined by a medical professional on a case-by-case basis and is within the skill of a person skilled in the art. When determining the appropriate dosage to achieve an effective amount, several factors are typically taken into consideration. These factors include the age, sex, and weight of the subject, the condition being treated, the severity of the condition, and the form and effective concentration of the therapeutic agent being administered.

[0075] As used herein, the term "tumor mutation burden" (TMB) refers to the level of mutations present in a patient's tumor cells. Herein, TMB was calculated by dividing the number of nonsynonymous mutations by the size of the gene panel (e.g., 2.4 Mb). See, for example, Beaubier et al. 2019 Oncotarget 10, 2384-2396. All non-silent somatic coding mutations with greater than 100x coverage and greater than 5% allelic fragmentation, including missense, insertion, or deletion, and stop-loss variants, were included in the count of nonsynonymous mutations. A hypermutated tumor was considered TMB-high if its TMB was at least 9 mutations per Mb. This threshold was established by examining the enrichment of tumors with orthogonally defined hypermutations (MSI-H) in the Tempus clinical database.

[0076] Some aspects are described below with reference to illustrative applications. It should be understood that numerous specific details, relationships, and methods are described to provide a thorough understanding of the features described herein. However, it will be readily apparent to those skilled in the relevant art that the features described herein can be implemented without one or more of the specific details or using other methods. The features described herein are not limited by the order of acts or events illustrated, and some acts may occur in a different order and / or simultaneously with other acts or events. Furthermore, not all illustrated acts or events are required to implement a methodology in accordance with the features described herein.

[0077] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0078] Example System Implementation An overview of some aspects of the present disclosure and some definitions used in the disclosure are provided, and details of an exemplary system are described in combination with FIG. 1. FIG. 1 is a block diagram illustrating a system 100 according to some implementations. In some implementations, the system 100 includes one or more processing units CPU 102 (also referred to as a processor), one or more network interfaces 104, a user interface 106 including (optionally) a display 108 and an input system 110, non-persistent memory 111, persistent memory 112, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communication between the system components. Non-persistent memory 111 typically includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while persistent memory 112 typically includes CD-ROM, digital versatile disk (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage, magnetic disk storage, optical disk storage, flash memory device, or other non-volatile solid-state storage. Persistent memory 112 optionally includes one or more storage devices located remotely from CPU 102. Persistent memory 112, and the non-volatile memory device(s) 112 within non-persistent memory, include non-transitory computer-readable storage media. In some implementations, non-persistent memory 111, or alternatively, non-transitory computer-readable storage media, store the following programs, modules, and data structures, or a subset thereof, sometimes in combination with persistent memory 112: an optional operating system 116 that handles various basic system services and includes procedures for performing hardware-dependent tasks; an optional network communication module (or instructions) 118 for connecting the system 100 with other devices and / or the communication network 104; an optional classifier training module 120 for training a classifier to determine a set of cancer states, the classifier training module including a dataset for one or more reference subjects 122, the dataset for each reference subject including at least a first plurality of sequence reads 124, (optionally) a second plurality of sequence reads 128, (optionally) a third plurality of sequence reads 130, (optionally) a pathology report 134, and an indicator of a diagnosed cancer state 138 for the respective reference subject, the dataset for each reference subject further including one or more reference features 126 derived from the first plurality of sequence reads 124, one or more reference features 132 derived from a comparison of the second plurality of sequence reads 128 and the third plurality of sequence reads 130, and one or more reference features 136 derived from the pathology report 134; a patient classification module 140 for classifying test subjects into a particular set of cancer states using a classifier (e.g., as trained using the classifier training module 120) using DNA sequence information, RNA sequence information, and pathology report information; and The patient classification module 140 further includes a dataset including, for each test subject, a first plurality of sequence reads 144 including one or more features 146 derived from the first plurality of sequence reads for each test subject 142, a second plurality of sequence reads 148, a third plurality of sequence reads 150 including one or more features 152 from comparing the second and third plurality of sequence reads, and (optionally) a pathology report 154 including one or more features 156 derived from the pathology report.

[0079] In various implementations, one or more of the above-identified elements are stored in one or more of the aforementioned memory devices and correspond to sets of instructions for performing the functions described above. The above-identified modules, data, or programs (e.g., sets of instructions) need not be implemented as separate software programs, procedures, data sets, or modules; thus, various subsets of these modules and data may be combined or otherwise rearranged in various implementations. In some implementations, non-persistent memory 111 optionally stores a subset of the above-identified modules and data structures. Additionally, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the above-identified elements are stored in a computer system other than that of visualization system 100 and are addressable by visualization system 100, such that visualization system 100 may retrieve all or a portion of such data when needed.

[0080] While FIG. 1 illustrates “system 100,” this diagram is intended more as a functional description of various features that may be present in a computer system than as a structural schematic of the implementations described herein. In practice, items shown separately may be combined and some items may be separated, as will be recognized by those skilled in the art. Furthermore, while FIG. 1 depicts certain data and modules in non-persistent memory 111, some or all of these data and modules may instead be stored in persistent memory 112 or in more than one memory. For example, in some embodiments, at least dataset 122 is stored on a remote storage device that may be part of a cloud-based infrastructure. In some embodiments, at least dataset 122 is stored on a cloud-based infrastructure. In some embodiments, dataset 122, classifier training module 120, and patient classification module 140 may also be stored on a remote storage device.

[0081] Patient classification A system according to the present disclosure is disclosed with reference to FIG. 1, and a method according to the present disclosure is described in detail with reference to FIGS.

[0082] Determining a set of cancer conditions for a subject Block 202. Referring to block 202 of FIG. 2A , the method determines a set of cancer conditions for the subject. Referring to block 204, in some embodiments, the set of cancer conditions consists of a single cancer condition (e.g., to determine whether the subject has a particular cancer condition). In some embodiments, the one cancer condition is selected from the subject's pathology report or other medical record. In some embodiments, the set of cancer conditions consists of two, three, or four different cancer conditions. Referring to block 206, in some embodiments, the set of cancer conditions includes five or more different cancer conditions. Referring to block 208, in some embodiments, the set of cancer conditions includes the likelihood of cancer origin from each respective tissue of a plurality of tissues (e.g., the set of cancer conditions provides information about the tissue of origin). In some embodiments, a cancer condition in the set of cancer conditions is the likelihood that the subject has metastatic cancer. In some embodiments, a cancer condition in the set of cancer conditions is the likelihood that the subject has primary cancer.

[0083] In some embodiments, the method classifies the subject to a cancer condition, hi some embodiments, the cancer condition is selected from a set of cancer conditions.

[0084] In some embodiments, the method classifies the subject into a predicted cancer state. In some embodiments, the predicted cancer state is selected from a set of cancer states. In some embodiments, the predicted cancer state (e.g., a prediction or determination made by a pathologist) is determined from the subject's pathology report. In some embodiments, the predicted cancer state is determined from one or more cancer states from the subject's pathology report.

[0085] In some embodiments, varying tumor cellularity between 20 and 100% does not significantly affect classification performance, as shown in Figure 16. All data shown in Figure 16 pertains to liver metastatic samples. The liver is a representative type of cancer because many tumors of unknown origin can be found in that organ. This analysis illustrates that the classification model performs well in the setting of low metastatic purity. Classification of tumors of unknown origin is described in further detail in Example 4.

[0086] Block 210. Referring to block 210 of FIG. 2A, one or more data structures of interest are obtained in electronic format. The one or more data structures collectively comprise a first plurality of sequence reads. The first plurality of sequence reads are obtained from a plurality of RNA molecules or derivatives of the above-mentioned plurality of RNA molecules (e.g., derivatives such as cDNA). In some embodiments, the plurality of RNA molecules are obtained by full transcriptome sequencing. In some embodiments, these sequence reads are derived from RNA isolated from a solid tumor or a hematological tumor (e.g., a solid biopsy).

[0087] Figures 15A-15C show expression data (e.g., the amount of sequence reads obtained for a particular RNA, e.g., from an RNA molecule, for each) for a patient with a primary tumor of known origin (tko_primary) and a patient with a metastatic tumor of known origin (tko_met) compared with the expression profile of a patient with a tumor of unknown origin (tuo). Each patient in these figures has one of the following cancers (as indicated along the x-axis): colorectal, non-small cell lung, pancreatic, esophageal, gastric, bladder, or bile duct cancer. Each figure shows the patient's expression level for one gene (e.g., a gene known to be associated with cancer). The tuo patients were classified into cancer statuses by a classification model trained as described herein. As can be seen from the expression profiles, there is a general correlation between RNA expression in patients with tumors of known origin (both metastatic and primary tumors) and patients with tumors of unknown origin. These figures demonstrate that RNA data can be useful for classifying patients with tumors of unknown origin.

[0088] Referring to block 214, in some embodiments, the one or more data structures further include a second plurality of sequence reads and a third plurality of sequence reads. In some embodiments, referring to block 215, the second plurality of sequence reads are obtained from the first plurality of DNA molecules or derivatives of the above-mentioned DNA molecules, and the third plurality of sequence reads are obtained from the second plurality of DNA molecules or derivatives of the above-mentioned DNA molecules. Referring to block 216, in some embodiments, the first plurality of DNA molecules are derived from a somatic biopsy obtained from the subject, and the second plurality of DNA molecules are derived from a germline sample obtained from the subject or from a collection of normal controls not including the set of cancer conditions. Referring to block 217, in some embodiments, the first plurality of DNA molecules and the second plurality of DNA molecules are obtained by whole exome sequencing.

[0089] Referring to block 218, in some embodiments, the second and third plurality of sequence reads are generated by next-generation sequencing. In some embodiments, the second and third plurality of sequence reads are generated by short-read paired-end next-generation sequencing. Referring to block 220, in some embodiments, the second and third plurality of sequence reads are obtained by targeted panel sequencing using multiple probes. In such embodiments, each respective probe in the plurality of probes uniquely represents a different portion of the reference genome. In such embodiments, each sequence read in the second and third plurality of sequence reads corresponds to at least one probe in the plurality of probes.

[0090] Figures 14A and 14B show that DNA expression data can provide information useful for classifying subjects into cancer states. Figures 14A and 14B each show the expression levels of specific genes. The amount of sequence reads obtained from DNA molecules often correlates between patients with known cancer states and patients with the same predicted tumor origin (e.g., particularly for patients with known or predicted bladder cancer, endocrine cancer, endometrial cancer, esophageal cancer, non-small cell lung cancer, ovarian cancer, and pancreatic cancer in Figure 14A, and patients with known or predicted colorectal cancer, non-small cell lung cancer, ovarian cancer, and pancreatic cancer in Figure 14B).

[0091] Referring to block 222, in some embodiments, the one or more data structures (e.g., from a data store described in more detail below) further include a pathology report for the subject (e.g., a pathology report for the subject is obtained). In some embodiments, the pathology report includes one or more of the subject's IHC protein level, the subject's age, the subject's sex, a disease diagnosis, a treatment category, a type of treatment, and a treatment outcome.

[0092] In some embodiments, the pathology report further includes image files. In some embodiments, the pathology report further includes one or more image features extracted from one or more image files of a somatic biopsy (e.g., a tumor biological sample) from the subject. In some embodiments, the extracted image features include tumor size, tumor stage, tumor grade, tumor purity, degree of invasiveness, degree of immune infiltration in the tumor, cancer stage, anatomical site of origin of the tumor, etc. In some embodiments, one or more of these extracted image features are incorporated into the pathology report. In some embodiments, the image features are extracted according to the methods described in U.S. Patent Application No. 62 / 824,039, filed March 26, 2019, and entitled "PD-L1 Prediction Using H&E Slide Images."

[0093] In some embodiments, a digital pathology image of a somatic biopsy (e.g., a somatic biopsy image file) provides essential clues about the cellular composition of a subject and is then sequenced to obtain a first plurality of sequence reads (e.g., obtained from a plurality of RNA molecules) and a second plurality of sequence reads (e.g., obtained from a first plurality of DNA molecules). Somatic biopsies are often a heterogeneous mixture of necrotic tissue, lymphocytes and other immune cells, stromal cells, and tumor cells. The imaging itself provides essential information about the cellular composition of the sample being sequenced. Further analysis of the image file (e.g., via a convolutional neural network, such as that described in U.S. Patent Application No. 16 / 732,242, filed December 31, 2019, entitled "Artificial Intelligence Segmentation of Tissue Images") can capture even higher levels of information about the tissue morphology of the biopsy location. A deep ranking neural network performs a search task and can be used to find other images in the dataset that share common features, providing information about the identity of the tumor.

[0094] In some embodiments, the one or more data structures further include an indicator of viral status (e.g., as described in U.S. Patent Application No. 62 / 810,849, filed February 26, 2019, entitled "Systems and Methods for Using Sequencing Data for Pathogen Detection," which is incorporated by reference in its entirety) (e.g., an indicator of viral status of a subject is obtained). In some embodiments, the indicator of viral status includes a count of virus-associated sequence reads (see, e.g., FIG. 20 and Example 9). In such embodiments, the method further includes applying the indicator of viral status to a trained classification model.

[0095] In some embodiments, both DNA expression data and RNA expression data are used to train a classification model. In some embodiments, one of DNA expression data or RNA expression data is used to train a classification model. In some embodiments, both DNA expression data and RNA expression data are used to determine a set of cancer states and / or cancer states for patients with one or more tumors of unknown origin. In some embodiments, one of DNA expression data or RNA expression data is used to determine a set of cancer states and / or cancer states for patients with one or more tumors of unknown origin.

[0096] Referring to block 224, in some embodiments, the somatic biopsy includes a microdissected formalin-fixed, paraffin-embedded (FFPE) tissue section, a surgical biopsy, a skin biopsy, a punch biopsy, a prostate biopsy, a bone biopsy, a bone marrow biopsy, a needle biopsy, a CT-guided biopsy, an ultrasound-guided biopsy, a fine needle aspiration, a suction biopsy, a fresh tissue, or a blood sample. In some embodiments, the germline sample includes blood or saliva from the subject. This serves to separate the tumor sample from a normal sample (e.g., the patient's own control sample). In some embodiments, the somatic biopsy is a somatic biopsy of a breast tumor, a glioblastoma, a prostate tumor, a pancreatic tumor, a kidney tumor, a colorectal tumor, an ovarian tumor, an endometrial tumor, a breast tumor, or a combination thereof. A biopsy is typically performed after one or more minimally invasive clinical tests suggest that the patient has or may have one or more tumors. The type of biopsy often depends on the location of the tumor. For example, biopsies of kidney tumors are frequently performed endoscopically, while biopsies of ovarian tumors frequently involve tissue scrapings.

[0097] Referring to block 226, in some embodiments, the first plurality of sequence reads are generated by next-generation sequencing with one or more spike-in controls. In some embodiments, the first plurality of sequence reads are generated from short-read paired-end next-generation sequencing. In some embodiments, the second plurality of sequence reads and / or the third plurality of sequence reads are generated by next-generation sequencing with one or more spike-in controls. In some embodiments, the first, second, and / or third plurality of sequence reads are generated from short-read paired-end next-generation sequencing.

[0098] Next-generation sequencing methods for use in the methods described herein are disclosed in Shendure 2008 Nat. Biotechnology 26:1135-1145 and Fullwood et al. 2009 Genome Res. 19:521-532, each of which is incorporated herein by reference. Next-generation sequencing methods well known in the art include synthesis technology (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), sequencing by ligation (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using sequencing by synthesis using reversible dye terminators.

[0099] Methods for mRNA sequencing are also well known in the art. In some embodiments, mRNA is reverse transcribed into cDNA before sequencing. For example, RNA-seq methods for use in accordance with block 210 are disclosed in Nagalakshmi et al., 2008, Science 320, 1344-1349, and Finotell and Camillo, 2014, Briefings in Functional Genomics 14(2), 130-142, each of which is incorporated herein by reference. In some embodiments, mRNA sequencing is carried out by whole exome sequencing (WES). In some embodiments, WES is carried out by isolating RNA from tissue samples, optionally selecting desired sequences and / or depleting undesired RNA molecules, generating a cDNA library, and then sequencing the cDNA library using, for example, next-generation sequencing technology. For an overview of the use of whole-exome sequencing technology in cancer diagnosis, see Serrati et al., 2016, Onco Targets Ther. 9, 7355-7365 and Cieslik, M. et al., 2015, Genome Res. 25, 1372-81, the contents of each of which are incorporated herein by reference in their entirety for all purposes. In some embodiments, mRNA sequencing is performed by nanopore sequencing. A summary of the use of nanopore sequencing technology for the human genome can be found in Jain et al., 2018, Nature 36(4), 338-345. This list is not exhaustive of RNA sequencing methods that can be used in accordance with the methods described herein. In some embodiments, RNA sequencing is performed according to one or more sequencing methods known in the art. See, for example, Kukurba et al., 2015, Cold Spring Harb Protoc. 11:951-969, an overview of RNA sequencing methods.

[0100] RNA-seq is a next-generation sequencing-based methodology for RNA profiling that enables the measurement and comparison of gene expression patterns across multiple subjects. In some embodiments, millions of short strings, called "sequence reads," are generated by sequencing random positions in cDNA prepared from input RNA obtained from a subject's tumor tissue. In some embodiments, RNA-seq gene expression data is generated from formalin-fixed, paraffin-embedded tumor samples using an exome capture-based RNA-seq protocol. These reads can then be computationally mapped to a reference genome to reveal a "transcription map," and the number of sequence reads aligned to each gene provides a measure of its expression level (e.g., abundance). In some embodiments, RNA-seq expression levels (e.g., raw read counts) are normalized (e.g., to correct for GC content, sequencing depth, and / or gene length). In some embodiments, methods for mapping raw RNA-seq reads to the transcriptome, methods for quantifying gene counts, and normalization are performed as described in U.S. Patent Application No. 62 / 735,349, filed September 24, 2018, entitled "Methods of Normalizing and Correcting RNA Expression Data."

[0101] In some alternative embodiments, rather than using RNA-seq, RNA profiling is performed using microarrays. Such microarrays are disclosed in Wang et al., 2009, Nat Rev Genet 10, 57-63; Roy et al., 2011, Brief Funct Genomic 10:135-150; Shendure, 2008 Nat Methods 5, 585-587; Cloonan et al., 2008, "Stem cell transcriptome profiling via massive-scale mRNA sequencing," Nat Methods 5, 613-619; Mortazavi et al., 2008, "Mapping and quantifying mammalian transcriptomes by RNA-Seq," Nat Methods 5, 621-628; and Bullard et al., 2010, "Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments," BMC Bioinformatics 11, p. 94, each of which is incorporated herein by reference.

[0102] The first computational step in the RNA-seq data analysis pipeline is read mapping, which aligns reads to a reference genome or transcriptome by identifying genomic regions that match the read sequence. Any of a variety of alignment tools can be used for this task. See, for example, Hatem et al., 2013 BMC Bioinformatics 14, 184, and Engstrom et al., 2013 Nat Methods 10, 1185-1191, each of which is incorporated herein by reference. In some embodiments, the mapping process begins by constructing an index for either the reference genome or the read, which is then used to search for a set of positions in the reference sequence where the read is more likely to align. Once this subset of possible mapping positions is identified, alignment is performed within these candidate regions using slower and more sensitive algorithms. See, for example, Flicek and Birney, 2009, Nat Methods 6 (Suppl. 11), S6-S12, which is incorporated herein by reference. In some embodiments, the mapping tool is a methodology that utilizes a hash table or the Burrows-Wheeler transform (BWT). See, e.g., Li and Homer, 2010 Brief Bioinformatics 11, 473-483, which is incorporated herein by reference.

[0103] After mapping, read counts are calculated using reads aligned to each coding unit, such as an exon, transcript, or gene, to provide an estimate of its abundance (e.g., expression) level. In some embodiments, only the coding region of the genome is available for mapping, thus preventing the mapping of discontinued or canceled genes from previous iterations of the human genome. In some embodiments, such counts take into account the total number of reads that overlap the exons of a gene. However, in some cases, because some of the sequence reads map outside the boundaries of known exons, alternative embodiments consider the entire length of the gene and also count reads from introns. Furthermore, in some embodiments, spliced ​​reads are used to model the abundance of different splicing isoforms of a gene. See, for example, Trapnell et al., 2010 Nat Biotechnol 28, 511-515, and Gatto et al., 2014 Nucleic Acids Res 42, p.e71, each of which is incorporated herein by reference.

[0104] As explained above, quantifying transcript abundance from RNA-seq data is typically implemented in an analysis pipeline through two computational steps: aligning reads to a reference genome or transcriptome, followed by estimating transcript and isoform abundance based on the aligned reads. Unfortunately, the reads generated by most commonly used RNA-seq techniques are generally much shorter than the transcripts from which they are sampled. As a result, it is not always possible to uniquely assign short sequence reads to specific genes in the presence of transcripts with similar sequences. Such sequence reads are referred to as "multiple reads" because they are homologous to more than one region of the reference genome. In some embodiments, such multiple reads are discarded, i.e., they do not contribute to gene abundance counts. In some embodiments, programs such as MMSEQ or RSEM are used to resolve ambiguities. See examples of methodologies used to resolve multiple reads in Turro et al., 2011 Genome Biol 12, p. R13, and Nicolae et al., Algorithms Mol Biol 6, 9, each incorporated herein by reference.

[0105] Another aspect of RNA-seq is the normalization of sequence read count.In some embodiments, this includes normalization to take into account different sequencing depths.For example, see Lin et al., 2011 Bioinformatics 27, 2031-2037, Robinson Oshlack, 2010 Genome Biol 11, R25, and Li et al., 2012 Biostatistics 13, 523-538, each of which is incorporated herein by reference.In some embodiments, sequence read count is normalized to take into account gene length bias.See Finotell and Camillo, 2014 Briefings in Functional Genomics 14(2), 130-142, each of which is incorporated herein by reference.

[0106] In some embodiments, the fourth plurality of sequence reads is obtained from an additional plurality of RNA molecules, the additional plurality of RNA molecules being isolated from normal, healthy tissue (e.g., the use of paired tumor / normal analysis is described in Example 3). In some embodiments, the amount of each sequence read in the first plurality of sequence reads is compared to the amount of a corresponding sequence read from the fourth plurality of sequence reads (e.g., essentially normalizing the amount of RNA sequence reads in the subject).

[0107] In some embodiments, the second plurality of sequence reads and the third plurality of sequence reads are obtained by targeted panel sequencing using multiple probes.Each probe in the plurality of probes is uniquely targeted to a different part of reference genome (for example, human reference genome).Each sequence read in the second plurality of sequence reads and each sequence read in the third plurality of sequence reads corresponds to at least one probe in the plurality of probes.In some embodiments, for example, whole genome sequencing is used instead of targeted panel sequencing.

[0108] In some embodiments, the second plurality of sequence reads has an average depth across the plurality of probes of at least 50x. In some embodiments, the second plurality of sequence reads has an average depth across the plurality of probes of at least 400x. In other embodiments, the second plurality of sequence reads has an average depth of at least 10x, 15x, 20x, 25x, 30x, 40x, 50x, 75x, 100x, 150x, 200x, 250x, 300x, 400x, 500x, or more.

[0109] In some embodiments, the plurality of probes includes probes for at least 300 different genes. In some embodiments, the plurality of probes includes probes for at least 500 different genes. In still other embodiments, the plurality of probes includes at least 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000, 4000, 5000, or more different genes. In some embodiments, the plurality of probes includes probes for at least 50, 100, 150, 200, 250, 300, 400, 500, or more different genes selected from the Targeted Gene List (e.g., Table 2).

[0110] In some embodiments, the plurality of probes comprises probes for at least 500 different genes selected from the Targeted Gene List. The Targeted Gene List is derived from the IDT xGen Exome Research Panel, which includes approximately 19,400 exons from the human genome. See, e.g., https: / / www.idtdna.com / pages / products / next-generation-sequencing / hybridization-capture / lockdown-panels / xgen-exome-research-panel.

[0111] In some embodiments, whole-exome sequencing of cDNA libraries is performed using Integrated DNA Technologies (IDT) XGEN® LOCKDOWN® technology with the xGen Exome Research Panel. Briefly, the xGen Exome Research Panel covers 51 Mb of end-to-end tile probe space of the human genome, providing deep and uniform coverage for whole-exome target capture. The cDNA library was hybridized to biotinylated DNA capture probes covering the reference human exome. The hybridized probes were recovered by binding to streptavidin beads. Post-capture PCR was performed to enrich the captured sequences. The amplified products were then sequenced using sequencing by synthesis (SBS) technology (Bently et al., 2008, Nature 456(7218), 53-59, the contents of which are incorporated herein by reference in their entirety for all purposes).

[0112] In some embodiments, metastatic or primary cancers (e.g., in some embodiments, defined as somatic biopsies) include tumors from a common primary site of origin (e.g., metastatic or primary cancers originate from one tumor). In some embodiments, metastatic or primary cancers include tumors originating from two or more different organs (e.g., tumors originate from multiple organs and / or tumors originate from any of several possible organs).

[0113] In some embodiments, the metastatic or primary cancer (e.g., in some embodiments, defined as a somatic biopsy) comprises a tumor of a predetermined stage of brain cancer, a predetermined stage of glioblastoma, a predetermined stage of prostate cancer, a predetermined stage of pancreatic cancer, a predetermined stage of kidney cancer, a predetermined stage of colorectal cancer, a predetermined stage of ovarian cancer, a predetermined stage of endometrial cancer, or a predetermined stage of breast cancer. Figure 7B shows the results of classifying samples from different brain cancers into broad categories (e.g., different tumor grades), as discussed in Example 4 below.

[0114] Referring to block 226, in some embodiments, the first plurality of sequence reads are generated from short-read next-generation sequencing using one or more spike-in controls. In some embodiments, the one or more spike-in controls calibrate for variation in sequence reads across a population of cells (e.g., the volume of RNA reads obtained from each cell may vary significantly, and spiking serves to normalize reads across a set of cells).

[0115] Next, in block 230 of FIG. 2B , the method continues by determining a first set of sequence features for the subject from the first plurality of sequence reads. In some embodiments, the first set of sequence features includes between 15,000 and 22,000 features. In some embodiments, the first set of sequence features includes 20,000 RNA variables (e.g., transcriptome). In some embodiments, the features for the subject further include any of the features described herein (e.g., RNA features, DNA features, CNV features, viral features, or any combination thereof). In some embodiments, methods for generating features in the plurality of features may include one or more of the methods disclosed in U.S. Patent Application No. 16 / 657,804, entitled “Data Based Cancer Research and Treatment Systems and Methods,” filed October 18, 2019 (hereinafter, the '804 patent), which is incorporated herein by reference in its entirety.

[0116] In some embodiments, determining the first set of sequence features further comprises deconvolving the first plurality of sequence reads by comparing the first plurality of sequence reads to a deconvolved RNA expression model comprising at least one cluster identified as corresponding to a cancer condition. In some embodiments, the deconvolution process is performed as described by U.S. Patent Application No. 16 / 732,229, filed December 31, 2018, and entitled "Transcriptome Deconvolution of Metastatic Tissue Samples," which is incorporated herein in its entirety.

[0117] In some embodiments, as shown in block 232 of FIG. 2B , the first set of sequence features derived from the first plurality of sequence reads includes one or more gene fusions, one or more copy number variations, one or more somatic mutations, one or more germline mutations, one or more gene fusions, tumor mutational burden (TMB), one or more microsatellite instability indices (MSI), an indicator of pathogen load, an indicator of immune infiltration, or an indicator of tumor cellularity.

[0118] In some embodiments, the determining step includes aligning each respective sequence read in the first plurality of sequence reads with a reference genome to determine a set of first sequence features of interest. In some embodiments, one or more gene fusions are determined as discussed in McPherson et al. 2011 PLoS Comput Biol 7(5):elOOl 138. In some embodiments, one or more copy number variations are determined as described in Shilien and Malkin 2009 Genome Med 1,62. In some embodiments, one or more somatic mutations and / or one or more germline mutations are discovered by comparing the second plurality of sequence reads and the third plurality of sequence reads with a reference genome, respectively. In some embodiments, one or more microsatellite instability indices are determined as described by Buhard et al. 2006 J Clinical Onco 24(2),241. In some embodiments, tumor mutation burden is determined as described in Chalmers et al. 2017 Genome Med 9,34. In some embodiments, the index of pathogen burden and / or the index of immune infiltration are determined, for example, as described by Barber et al. 2015 PLoS Pathog 11(1):el004558 and Pages et al. 2010 Oncogene 29,1093-1102. In some embodiments, the index of tumor cellularity is determined from a somatic biopsy by comparing the number of cancerous cells and the number of normal cells obtained in the somatic biopsy. In some embodiments, the index of tumor cellularity is determined from one or more images of the somatic biopsy (e.g., by counting and identifying cancerous cells versus non-cancerous cells in one or more images).

[0119] Next, at block 234 of FIG. 2B , in some embodiments, the method continues by determining a second set of sequence features for the subject from a comparison of the second plurality of sequence reads to the third plurality of sequence reads. In some embodiments, the second set of sequence features includes between 400 features and 2,000 features. In some embodiments, the second set of sequence features includes 500 DNA variables. In other embodiments, the second set of sequence features includes at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000, 7500, 10,000, or more sequence features.

[0120] In some embodiments, as shown in block 236 of FIG. 2B , the second set of sequence features derived from comparing the second plurality of sequence reads to the third plurality of sequence reads comprises one or more copy number variations, one or more somatic mutations, one or more germline mutations, tumor mutational burden (TMB), one or more microsatellite instability indices (MSI), an indicator of pathogen load, an indicator of immune infiltration, or an indicator of tumor cellularity. In some embodiments, the second set of sequence features is derived as described above with respect to FIG. 232. In some embodiments, the second set of sequence features comprises one or more DNA variant patterns (e.g., as discussed in Example 9 and illustrated by FIGS. 21A-21C). In some embodiments, the one or more DNA variant patterns are determined for the subject by comparing the second and third plurality of sequence reads to a reference set of DNA variant patterns, each DNA variant pattern in the reference set for a respective cancer state in the set of cancer states.

[0121] In some embodiments, determining the set of second sequence features comprises aligning each respective sequence read in the second plurality of sequence reads and the third plurality of sequence reads with a reference genome to determine the set of second sequence features of interest. The second plurality of sequence reads and the third plurality of sequence reads must be aligned in advance so that they can be compared with each other.

[0122] At block 238 of FIG. 2C , in some embodiments, the method continues by extracting a plurality of pathological features from the pathology report for the subject, including a first estimate of tumor cellularity from the somatic biopsy and an indication of whether the subject has metastatic or primary cancer. Referring to block 240, in some embodiments, the plurality of pathological features include one or more of immunohistochemistry (IHC) protein levels, tissue site, tumor cellularity, extent of tumor infiltration by lymphocytes, tumor mutation burden (TMB), microsatellite status (e.g., MSI), viral status (e.g., HPV+ / -), subject age, subject gender, disease diagnosis (e.g., including cancer diagnosis, cancer stage, and / or cancer subtype), treatment category (e.g., treated or untreated), type of treatment, or treatment outcome. In some embodiments, other sections of the pathology report known in the art may also be considered. In some embodiments, a subset of possible pathological features is considered in classifying a subject into a set of cancer states.

[0123] In some embodiments, the method further includes complementing the first estimate of tumor cellularity of the somatic biopsy with a second estimate of tumor cellularity from one or more images of the somatic biopsy (e.g., analyzing one or more somatic biopsy images to determine the extent of tumor growth and / or progression). The images of the somatic biopsy may include images of histological slides generated from the somatic biopsy, or radiological scans of a solid tumor or somatic biopsy. In some embodiments, the method further includes complementing the first estimate of tumor cellularity of the somatic biopsy with a second estimate of tumor cellularity from the abundance of one or more mutations in a second plurality of sequence reads (e.g., from sequence reads derived from the somatic biopsy).

[0124] Pathology reports typically require data cleaning (which may include, for example, natural language processing as described in Example 5, or manual abstraction) before meaningful features can be extracted. Natural language processing in pathology reports is, in some embodiments, performed as described in U.S. Application No. 16 / 702,510, entitled "Clinical Concept Identification, Extraction, and Prediction System and Related Methods," filed December 3, 2019, which is incorporated herein in its entirety. Terminology in pathology reports is not necessarily standardized, and natural language processing determines which terms are synonymous and may be collapsed together for downstream analysis.

[0125] Referring back to block 242, in some embodiments, the pathology report further includes one or more image features extracted from one or more images of the test subject's somatic biopsy. In some embodiments, the pathology features are extracted from the pathology report (and any associated files, such as images) itself. In alternative embodiments, the pathology features are determined or extracted from an alternative source (e.g., without the need for a pathology report). For example, in some embodiments, an electronic medical record (EMR) that focuses on pathology needs and allows pathologists to input features (e.g., tumor purity, etc.) is available in a structured format, and the pathology features are parsed from the structured report (e.g., requiring less data cleaning than a largely unstructured pathology report).

[0126] In some embodiments, the plurality of pathological features comprises at least 200 pathological features. In some embodiments, the plurality of pathological features comprises between 200 and 500 features in the pathology record. In some embodiments, the plurality of pathological features comprises 400 pathological features in the medical record. In other embodiments, the plurality of pathological features comprises at least 10, 15, 20, 25, 30, 40, 50, 75, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, or more pathological features. A "pathological feature" is a feature derived from information typically present in a pathology record or pathology image.

[0127] Next, in block 244 of Figure 2C, the method continues by applying at least the first set of sequence features to the trained classification method, thereby obtaining a classification result that provides, for each respective cancer condition in the set of cancer conditions, the likelihood that the subject has or does not have the respective cancer condition.

[0128] Referring to block 246, in some embodiments, the trained classification method includes a trained classifier stream. Referring to block 248, in some embodiments, as a non-limiting example, the trained classifier stream is a decision tree. Decision tree algorithms suitable for use as the classification model of block 244 are described, for example, in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, 395-396, which is incorporated herein by reference. Tree-based methods partition the feature space into a set of rectangles and then fit a model (as if constant) to each. In some embodiments, the decision tree is a random forest regression. One particular algorithm that can be used as the classification model of block 244 is a classification and regression tree (CART). Other examples of particular decision tree algorithms that can be used as classifiers in block 244 include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and 411-412, which are incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forests are described in Breiman, 1999, "Random Forests—Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is incorporated herein by reference in its entirety. In some embodiments, xgboost and / or lightgbm are additional decision tree methods that can be used as trained classifier streams.See, e.g., Chen et al. 2016 KDD'16: Proc 22nd ACM SIGKDD Int Conf Knowledge Disc. Data Mining, 785-794, and Wang et al. 2017 ICCBB: Proc 2017 Int Conf Comp Biol and Bioinform, 7-11.

[0129] In some embodiments, by way of non-limiting example, the trained classifier stream includes regression. The regression algorithm can be any type of regression. For example, in some embodiments, the regression algorithm is logistic regression. Logistic regression algorithms are disclosed in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103-144, John Wiley & Son, New York, which is incorporated herein by reference. In some embodiments, the regression algorithm is logistic regression with lasso, L2, or elastic net regularization.

[0130] In some embodiments, by way of non-limiting example, the trained classifier stream comprises a neural network. Examples of neural network algorithms, including convolutional neural network algorithms, are disclosed, for example, in Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference.

[0131] In some embodiments, by way of non-limiting example, the trained classifier stream comprises a support vector machine (SVM). Examples of SVM algorithms are found in, for example, Cristianini and Shawe-Taylor, 2000, "An Introduction to Support Vector Machines," Cambridge University Press, Cambridge; Boser et al., 1992, "A training algorithm for optimal margin classifiers," in Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, "Statistical Learning Theory," Wiley, New York; Mount, 2001, "Bioinformatics: sequence and genome analysis," Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Duda, "Pattern Classification," Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265; and Hastie, 2001, "The Elements of Statistical Learning," Springer, New York; and Furey et al. al., 2000, Bioinformatics 16, 906-914, each of which is incorporated herein by reference in its entirety. When used for classification, SVMs separate a given set of binary labeled data training set with a hyperplane that is maximally separated from the labeled data. When linear separation is not possible, SVMs can work in conjunction with the technique of "kernels" that automatically achieve a nonlinear mapping to the feature space. The hyperplane found by the SVM in the feature space corresponds to a nonlinear decision boundary in the input space.

[0132] In some embodiments, the trained classifier stream includes a first classifier, a second classifier, and a third classifier. In such embodiments, applying includes inputting all or a portion of the plurality of pathological features, the first set of sequence features, and the second set of sequence features into the first classifier, thereby obtaining an intermediate result. In such embodiments, applying further includes, if the intermediate result satisfies a first predetermined threshold or range, inputting the intermediate result into the second classifier but not into the third classifier, thereby obtaining a probability that the subject has or does not have a first cancer condition in the set of cancer conditions. In such embodiments, applying further includes, if the intermediate result fails to satisfy the first predetermined threshold or range, inputting the intermediate result into the third classifier but not into the second classifier, thereby obtaining a probability that the subject has or does not have the first cancer condition.

[0133] In some embodiments, the first classifier, the second classifier, and the third classifier each comprise a classifier of a respective classifier type (e.g., a respective cancer state classifier). In some embodiments, the first classifier comprises a classifier of the first classifier type, and the second and third classifiers each comprise a classifier of the second classifier type. In some embodiments, the first classifier comprises a classifier of the first classifier type, the second classifier comprises a classifier of the second classifier type, and the third classifier comprises a classifier of the third classifier type.

[0134] In some embodiments, the trained classifier stream includes a first classifier and a second classifier. In such embodiments, the applying includes inputting all or a portion of the plurality of pathological features, the first set of sequence features, and the second set of sequence features into the first classifier, thereby obtaining an intermediate result. In such embodiments, the applying further includes, if the intermediate result satisfies a first predetermined threshold or range, inputting the intermediate result into a second classifier, thereby obtaining a probability that the subject has or does not have a first cancer condition in the set of cancer conditions.

[0135] In some embodiments, the first classifier and the second classifier each comprise a classifier of a respective classifier type, hi some embodiments, the first classifier comprises a classifier of the first classifier type and the second classifier comprises a classifier of the second classifier type.

[0136] In some embodiments, the trained classifier stream used in block 244 includes a K-nearest neighbor model, a random forest model, a logistic regression, a support vector machine, or a neural network.

[0137] A nearest neighbor algorithm suitable for use as the classifier in block 244 is described below. For nearest neighbors, given a query point xo (object), a set of k training points xo that are closest to xo is found. (r) ,r,...,k (here training objects) are identified, and then the point xo is classified using its k nearest neighbors, where the distance to these neighbors is a function of the expression of the discriminant gene set. In some embodiments, the distance is calculated using Euclidean distance in feature space.

number

[0138] Neural network algorithms, including multilayer neural network algorithms, suitable for use as the classifier in block 244 are disclosed, for example, in Vincent et al., 2010 J Mach Learn Res 11, 3371-3408; Larochelle et al., 2009 J Mach Learn Res 10, 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference. A neural network has a layered structure, including a layer of input units (and biases) connected to a layer of output units by a layer of weights. For regression, the layer of output units typically includes only one output unit. However, some neural networks can handle multiple quantitative responses in a seamless manner. In a multilayer neural network, there are input units (input layer), hidden units (hidden layer), and output units (output layer). Additionally, there is a single bias unit connected to each unit other than the input units. Additional examples of neural networks suitable for use as the classifier in block 244 are disclosed in Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, Inc., New York, and Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, each of which is incorporated herein by reference in its entirety.Additional examples of neural networks suitable for use as the classifier in block 244 are also described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC, and Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is incorporated herein by reference in its entirety.

[0139] Referring to block 250, in some embodiments, the trained classifier stream includes a plurality of classifiers (e.g., a combination of classifiers). The plurality of classifiers includes a first subset of classifiers and a second subset of classifiers. Each classifier in the second subset of classifiers receives as input at least the output of at least one classifier in the first subset of classifiers. Each classifier in the first subset of classifiers receives as input at least a plurality of pathological features, a first set of sequence features, and all or a portion of a second set of sequence features. The output of the second subset of classifiers collectively provides, for each respective cancer condition in the set of cancer conditions, a probability that the subject has or does not have the respective cancer condition.

[0140] In some embodiments, the trained classifier stream includes multiple classifiers (e.g., a chain of classifiers). A first classifier in the multiple classifiers is used to determine the likelihood that a subject has or does not have a first cancer condition in the set of cancer conditions when the tumor cellularity meets a predetermined threshold (e.g., when the tumor cellularity is of high purity). A second classifier in the multiple classifiers is used to determine the likelihood that a subject has or does not have a first cancer condition in the set of cancer conditions when the tumor cellularity fails to meet the predetermined threshold (e.g., when the tumor cellularity is of low purity).

[0141] In some embodiments, each individual classifier in a classifier chain performs binary classification on a subset of the subject's features. In such embodiments, the classification results from an upstream classifier can be input to a downstream classifier. In some embodiments, a hyperparameter search for an optimal sequence of classifiers can be performed. In some embodiments, an ensemble model including one or more chains of classifiers classifies the subject by majority voting (e.g., each chain of classifiers receives one vote).

[0142] In some embodiments, each classifier in the plurality of classifiers is a classifier of a respective classifier type, hi some embodiments, one or more classifiers in the plurality of classifiers are classifiers of a first classifier type and one or more classifiers in the plurality of classifiers are classifiers of a second classifier type.

[0143] Referring to block 252, in some embodiments, the method further includes applying the second set of sequence features and the plurality of pathological features (e.g., as obtained above as described with respect to blocks 234 and 238, respectively) to the trained classification model.

[0144] In some embodiments, applying further includes applying one or more image features extracted from one or more images of a somatic biopsy from the test subject (e.g., tumor imaging data separate from the pathology report). FIG. 8A is an example of a tissue image (e.g., biopsy) used to predict tumor versus healthy tissue, where region 802 corresponds to a tissue region likely to be tumor. FIG. 8B is another image of the same example tissue sample, where region 804 corresponds to predicted lymphocytes. Tissue images such as those shown in FIGS. 8A and 8B may, in some embodiments, be used according to the methods disclosed herein to estimate tumor cellularity. For example, in some embodiments, tumor cellularity may be calculated as the ratio of the area of ​​the region predicted to be tumor to the total area of ​​tissue in the image. In alternative embodiments, tumor cellularity may be calculated as the ratio of the cell count for the region predicted to be tumor to the cell count for all tissue in the respective image. In some embodiments, a respective tumor cellularity value is determined for each image in the one or more images of the somatic biopsy. In such embodiments, an overall tumor cellularity value for the subject is determined by averaging each respective tumor cellularity value.

[0145] In some embodiments, the applying step further comprises applying one or more epigenetic or metabolomic features of the subject obtained from the subject's somatic cell sample to the trained classification model to obtain a classifier result. Epigenetic modifications are known to contribute to cancer progression in some cases. See Sharma el al. 2010 Carcinogenesis 31:27-36. Similarly, metabolic reprogramming has also been correlated with cancer diagnosis. See, e.g., Yang et al. 2017 Scientific Reports 7:43353. In some embodiments, the applying step further comprises applying one or more microbiota features of the subject to the trained classification model to obtain a classifier result. In recent years, the gut microbiota has been recognized as contributing, inter alia, to patient response to cancer therapy. See Guglielmi 2018 https: / / www.nature.eom / articles / d41 586-018-05208-8 and Gopalakrishnan et al.2018 Cell 33:570-580.

[0146] In some embodiments, the classifier results are further used (e.g., by a pathologist) to provide one or more treatment recommendations (e.g., as described below for Example 2) to the subject or a healthcare professional caring for the subject based on the likelihood that the subject has or does not have each respective cancer condition in the set of cancer conditions. In some embodiments, the trained classification model further provides one or more treatment recommendations to the subject or a healthcare professional caring for the subject based on the likelihood that the subject has or does not have the cancer condition. In some embodiments, the trained classification model further provides one or more treatment recommendations to the subject or a healthcare professional caring for the subject based on the likelihood that the subject has or does not have the predicted cancer condition.

[0147] In some embodiments, the classification model changes the tumor origin diagnosis for the subject (e.g., as described in Example 8). In some embodiments, this change also changes the recommended course of treatment for the subject.

[0148] In some embodiments, results from the trained classification model are used (e.g., from a pathologist or other healthcare provider) to provide a patient report (e.g., as illustrated in FIGS. 10A-10G and described in Example 7) as part of providing one or more treatment recommendations. In some embodiments, the patient report includes detailed information regarding the classification results (e.g., the likelihood that the subject has or does not have each respective cancer condition in a set of cancer conditions, the likelihood that the subject has or does not have a cancer condition, or the likelihood that the subject has or does not have a predicted cancer condition) and / or treatment recommendations. An example layout of a patient report 900 is illustrated by FIG. 9, and FIGS. 10A-10G show specific examples of particular sections of the patient report (e.g., 902-920 in example patient report 900).

[0149] Detailed examples of specific patient reports are provided in Example 7 below. Briefly, these sections provide patients and healthcare professionals with more information about their diagnosis. This can help improve patient care (e.g., by suggesting specific clinical trials, as shown in FIG. 10E, or by identifying relevant FDA-approved therapies, as shown in FIG. 10G) and give patients a sense of information and control over their diagnosis. Research has shown that there are significant limitations in physician-patient communication in effectively communicating information about cancer diagnoses. See, for example, Cartwright et al. 2015 J. Cancer Educ. 29, 311-317 or Nord et al. 2003 J. Public Health 25(5), 313-317. Through the classification methods described herein, more accurate information about a patient's cancer is obtained, and this information is provided to patients and healthcare professionals in a clear manner through the corresponding patient report.

[0150] Diagnosing the cancer status of a subject Each and every embodiment described with respect to Figures 2A, 2B, and 2C can also be applied to additional methods of diagnosing a cancerous condition in a subject, as described below, which provide for identifying a diagnosis of a cancerous condition in a somatic tumor specimen of a subject.

[0151] The method further includes receiving sequencing information comprising an analysis of a plurality of nucleic acids from a somatic tumor specimen. In some embodiments, the sequencing information comprises a plurality of DNA sequence reads (e.g., from circulating tumor DNA or a tumor biopsy). In some embodiments, the sequencing information comprises a plurality of RNA sequence reads (e.g., from circulating RNA or a tumor biopsy). In some embodiments, the sequencing information comprises both DNA sequencing reads and RNA sequencing reads.

[0152] The method further includes identifying a plurality of features from the received sequencing information, the plurality of features including two or more of an RNA feature (e.g., RNA feature 2341), a DNA feature (e.g., DNA feature 2342), an RNA splicing feature (e.g., RNA splicing feature 2349a), a viral feature, and a copy number feature (e.g., copy number feature 2349b). In some embodiments, methods for generating one or more of the features described herein may include one or more of the methods of the '804 patent.

[0153] Each RNA feature (e.g., from RNA features 2341) is associated with a respective target region of the first reference genome and represents the abundance of corresponding sequence reads mapping to the respective target region encompassed by the sequencing information. In some embodiments, each RNA feature in the RNA features is associated with a coding region of a gene. In some embodiments, the RNA features are obtained from cDNA sequencing.

[0154] Each DNA feature (e.g., from DNA features 2342) is associated with a respective target region of the second reference genome and represents the abundance of corresponding sequence reads that map to the respective target region encompassed by the sequencing information.

[0155] Each RNA splicing feature (e.g., from RNA splicing feature 2349a) is associated with a respective splicing event in a respective target region of the first reference genome, and represents the abundance of corresponding sequence reads that map to each target region with each splicing event contained in the sequencing information. In some embodiments, the RNA splicing feature is associated with a predetermined exome skipping event. In some embodiments, each RNA splicing feature is associated with two or more target regions (e.g., two or more non-contiguous exons).

[0156] Each viral feature is associated with a respective target region of the viral reference genome and represents the abundance of corresponding sequence reads mapping to each target region in the viral reference genome encompassed by the sequencing information. In some embodiments, viral features are used when a subject has an indication of viral status (e.g., as described above with respect to block 222). In some embodiments, the target region in the viral reference genome comprises virus-associated sequence reads (see, e.g., Figure 20 and Example 9). In some embodiments, viral features are identified by DNA sequencing or RNA sequencing of a sample from a subject. In some embodiments, the viral reference genome comprises one or more viral genomes (e.g., the viral reference genome represents multiple viral genomes). For example, a viral feature may comprise features from one or more viruses of interest.

[0157] Each copy number feature (e.g., from copy number variation 2349b) is associated with a target region of the second reference genome and represents the abundance of corresponding sequence reads that map to each target region of the second reference genome encompassed by the sequencing information. In some embodiments, the copy number feature is associated with a structural variant of the second reference genome.

[0158] In some embodiments, the first reference genome associated with the RNA features (and / or RNA splicing features) is the same as the second reference genome associated with the DNA features (and / or copy number features). In some embodiments, either the first or second reference genome comprises a human genome (e.g., NCBI build 34 (UCSC equivalent: hgl6), NCBI build 35 (UCSC equivalent: hgl7), NCBI build 36.1 (UCSC equivalent: hgl8), GRCh37 (UCSC equivalent: hgl9), and GRCh38 (UCSC equivalent: hg38)). In some embodiments, both the first reference genome and the second reference genome comprise a human genome.

[0159] In some embodiments, the target regions are predetermined regions of a genome (e.g., a first or second reference genome). In some embodiments, the predetermined regions of a genome represent regions known to be associated with a particular disease (e.g., a particular type of cancer). In some embodiments, the predetermined regions are genes. In some embodiments, each respective target region is a coding region (e.g., a gene) in the reference genome. In some embodiments, the target regions are non-coding regions (e.g., introns) in the reference genome. In some embodiments, the target regions are a combination of coding and non-coding genomic regions in the reference genome. In some embodiments, the target regions correspond to a group of genomic regions in the reference genome. In some embodiments, the target regions are at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten regions (e.g., coding and / or non-coding) in the reference genome.

[0160] In some embodiments, each target region corresponds to a respective feature in the plurality of features. In some embodiments, each feature in the plurality of features corresponds to a respective target region. In some embodiments, each target region corresponds to two or more features in the plurality of features. In some embodiments, each feature in the plurality of features corresponds to two or more target regions (e.g., two or more target regions may be operably linked or have similar expression patterns).

[0161] In some embodiments, the target regions can be approximately the same length. In some embodiments, the target regions can be different lengths. In some embodiments, the target regions have approximately equal lengths. In some embodiments, the target regions are at least 100 nucleobases, at least 200 nucleobases, at least 300 nucleobases, at least 400 nucleobases, at least 500 nucleobases, at least 600 nucleobases, at least 700 nucleobases, at least 800 nucleobases, at least 900 nucleobases, or at least 1,000 nucleobases in length. In some embodiments (e.g., particularly for RNA splicing features and / or structural variant copy number features), the target region is at least 1 Kb, at least 2 Kb, at least 3 Kb, at least 4 Kb, at least 5 Kb, at least 6 Kb, at least 7 Kb, at least 8 Kb, at least 9 Kb, at least 10 Kb, at least 15 Kb, at least 20 Kb, at least 25 Kb, at least 30 Kb, at least 40 Kb, at least 50 Kb, at least 60 Kb, at least 70 Kb, at least 80 Kb, at least 90 Kb, at least 100 Kb, at least 150 Kb, at least 200 Kb in length. In some embodiments (e.g., particularly for target regions corresponding to two or more genomic regions), each genomic region in a respective target region is between 100 and 500 nucleic acid bases, between 200 and 500 nucleic acid bases, between 200 and 400 nucleic acid bases, between 100 and 1000 nucleic acid bases, or between 500 and 1000 nucleic acid bases, between 100 and 10,000 nucleic acid bases, between 100 and 100,000 nucleic acid bases, between 5000 and 10,000 nucleic acid bases, between 10,000 and 50,000 nucleic acid bases, between 10,000 and 100,000 nucleic acid bases, between 50,000 and 100,000 nucleic acid bases, or between 50,000 and 150,000 nucleic acid bases in length.

[0162] In some embodiments, a predetermined number of target regions are evaluated for identifying a plurality of features. In some embodiments, the predetermined number of target regions is at least 100 target regions, at least 200 target regions, at least 300 target regions, at least 400 target regions, at least 500 target regions, at least 600 target regions, at least 700 target regions, at least 800 target regions, at least 900 target regions, at least 1,000 target regions, at least 2,500 target regions, at least 5,000 target regions, at least 7,500 target regions, at least 10,000 target regions, at least 20,000 target regions, or at least 50,000 target regions.

[0163] In some embodiments, for each feature in the plurality of features, the abundance of each sequence read corresponds to the raw number of sequence reads associated with each target region of the first or second reference genome. In some embodiments, for each feature in the plurality of features, the abundance of each sequence read corresponds to a normalized number of sequence reads. In some embodiments, normalization of sequence reads is performed as described above with respect to block 226.

[0164] In some embodiments, the sequencing information is deconvolved before feature identification. In some embodiments, deconvolution includes identifying sequence reads in the sequencing information originating from healthy tissue (e.g., sequence reads from normal, non-tumor cells) and removing these sequence reads from the sequencing information (e.g., to reduce background noise). In some embodiments, the deconvolution model includes a supervised machine learning model, a semi-supervised machine learning model, or an unsupervised machine learning model, as described in U.S. Patent Application No. 16 / 732,229, filed December 31, 2018, entitled "Transcriptome Deconvolution of Metastatic Tissue Samples," which is incorporated herein in its entirety. In some embodiments, a different deconvolution model is determined for each type of cancer or set of types of cancer (e.g., a liver cancer deconvolution model). In some embodiments, the deconvolution model removes expression data from cell populations that are not the cell type of interest (e.g., tumor or other type of cancer tissue). In some embodiments, the deconvolution model uses machine learning algorithms, such as unsupervised or supervised clustering techniques, to examine gene expression data and quantify the level of tumor relative to normal cell populations present in the data. In some embodiments, training the deconvolution model involves identifying common expression features shared across sequence reads from tissue normal samples, primary samples, and metastatic samples, such that the deconvolution model can predict the ratio of metastatic tumor to background tissue and identify which portion of sequence reads are attributable to tumor and which portion are attributable to background tissue. In some embodiments, sequence reads attributable to background tissue are removed from the sequencing information.

[0165] In some embodiments, the plurality of features are obtained by low-pass whole genome sequencing. In some embodiments, low-pass sequencing refers to the average coverage of a reference genome (e.g., a first, second, or viral reference genome) by a plurality of DNA or RNA sequencing reads (e.g., sequencing information). In some embodiments, the average coverage of the plurality of sequence reads (either DNA or RNA sequence reads) is less than 0.25x, less than 0.5x, less than 1x, less than 2x, less than 3x, less than 4x, less than 5x, less than 6x, less than 7x, less than 8x, less than 9x, or less than 10x across the reference genome. In some embodiments, the average coverage of the plurality of sequence reads is between 0.1x and 1x across the reference genome. In some embodiments, the average coverage of the plurality of sequence reads is between 0.1x and 5x across the reference genome. In some embodiments, the average coverage of the plurality of sequence reads is between 0.1x and 10x across the reference genome. In some embodiments, the average coverage rate of the multiple sequence reads is 1x to 5x across the reference genome.

[0166] The method further includes providing a first subset of features from the identified plurality of features as input to the first classifier. The method further includes providing a second subset of features from the identified plurality of features as input to the second classifier. In some embodiments, the first subset of features includes all or a portion of the identified plurality of features. In some embodiments, the second subset of features includes all or a portion of the identified plurality of features.

[0167] In some embodiments, providing the first subset of features to the first classifier and the second subset of features to the second classifier comprises providing the same subset of features to both the first classifier and the second classifier. In some embodiments, the same subset of features are RNA features. In some embodiments, the same subset of features are DNA features. In some embodiments, the same subset of features are RNA splicing features. In some embodiments, the same subset of features are viral features. In some embodiments, the same subset of features are copy number features.

[0168] In some embodiments, each feature in the plurality of features is associated with a respective target region (e.g., for RNA, DNA, copy number, and / or RNA splicing features), the plurality of features collectively represent a plurality of target regions, each region in the plurality of target regions is a gene, and the plurality of target regions include any of the following: GPM6A, CDX1, SOX2, NAPSA, CDX2, MUC12, SLAMF7, HNF4A, ANXA10, TRPS1, GATA3, SLC34A2, NKX2-1, SLC22A31, ATP10B, STEAP2, CLDN3, SPATA6, NRCAM, USH1C, SOX17, TMPRSS2, MECOM, WT1, CDHR1, HOXA13, SOX10, SALL1, CPE, NPR1, CLRN3, THSD4, ARL14, SFTPB, COL17A1, KLHL1 4, EPS8L3, NXPE4, FOXA2, SYT11, SPDEF, GRHL2, GBP6, PAX8, ANOl, KRT7, HOXA9, TYR, DCT, LYPDl, MSLN, TP63, CDH1, ESR1, HNF1B, HOXA10, TJP3, NRG3, TMC5, PRLR, GATA2, DCDC2, INS, NDUFA4L2, TBX5, ABCC3, FOLH1, HIST1 and two or more (e.g., in some embodiments, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or fifty or more) of H3G, S100A1, PTHLH, ACER2, RBBP8NL, TACSTD2, C19orf77, PTPRZ1, BHLHE41, FAM155A, MYCN, DDX3Y, FMN1, HIST1H3F, UPK3B, TRIM29, TXNDC5, BCAM, FAM83A, TCF21, MIA, RNF220, AFAP1, KRT5, SOX21, KANK2, GPM6B, Clorfl 16, FOXF1, MEIS1, EFHD1, and XKRX.

[0169] In some embodiments, each feature in the plurality of features is associated with a respective target region (e.g., for RNA, DNA, copy number, and / or RNA splicing features), the plurality of features collectively represent a plurality of target regions, each region in the plurality of target regions is a gene, and the plurality of target regions include: ENSG00000150625, ENSG00000 113722, ENSG00000181449, ENSG00000131400, ENSG00000165556, ENSG00000205277, ENSG00000026751, ENSG00000101076, ENSG00000 10951 1, ENSG00000 104447, ENSG00000 107485, ENSG00000 157765, ENSG00000136352, ENSG00000259803, ENSG000001 18322, ENSG00000157214, ENSG00000165215, ENSG00000132122, ENSG00000091 129, ENSG0000000661 1, ENSG00000164736, ENSG00000 184012, ENSG00000085276, ENSG00000184937, ENSG00000148600, ENSG00000 106031, ENSG00000100146, ENSG00000 103449, ENSG00000 109472, ENSG00000 169418, ENSG00000 180745, ENSG00000 187720, ENSG00000 179674, ENSG00000 168878, ENSG00000065618, ENSG00000 197705, ENSG00000198758, ENSG00000 137634, ENSG00000125798, ENSG00000132718, ENSG00000 124664, ENSG00000083307, ENSG00000183347, ENSG00000125618, ENSG00000131620, E NSG00000135480, ENSG00000078399, ENSG00000077498, ENSG00000080166, ENSG00000 150551, ENSG00000102854, ENSG00000073282, ENSG00000039068, ENSG0000009 1831, ENSG00000108753, ENSG00000253293, ENSG00000 105289, ENSG00000 185737, ENSG00000103534, ENSG000001 13494, ENSGOOOOO 179348, ENSG00000146038, ENSG00000254647, ENSG00000185633, ENSG00000089225, ENSGOOOOO 108846, ENSG00000086205, ENSG00000256018, ENSGOOOOO 160678, ENSG00000087494, ENSGOOOOO 177076, ENSGOOOOO 130701, ENSGOOOOO 184292, ENSG00000095932, ENSGOOOOO 106278, ENSG00000123095, ENSG00000204442, ENSGOOOOO 134323, ENSG00000067048, ENSG00000248905, ENSG00000256316, ENSG00000243566, ENSGOOOOO 137699, ENSG00000239264, ENSGOOOOO 187244, ENSGOOOOO 147689, ENSGOOOOO 118526, ENSG00000261857, ENSG00000187147, ENSGOOOOO 196526, ENSGOOOOO 186081, ENSG00000125285, ENSGOOOOO 197256, ENSG00000046653, ENSGOOOOO 182795, ENSGOOOOO 103241, ENSG00000143995, ENSGOOOOO 115468, and ENSGOOOOO 182489 (e.g., in some embodiments, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, or 50 or more).

[0170] In some embodiments, providing the subset of first features to the first classifier and providing the subset of second features to the second classifier comprises providing RNA features to both the first classifier and the second classifier. In some embodiments, providing the subset of first features to the first classifier and providing the subset of second features to the second classifier comprises providing RNA features to the first classifier and DNA features to the second classifier.

[0171] In some embodiments, the first classifier is a diagnosis classifier (e.g., classifier 2382a) and the second classifier is a cohort classifier (e.g., classifier 2382b). In some embodiments, the first classifier is a diagnosis classifier and the second classifier is a tissue classifier (e.g., tissue classifier 2382c).

[0172] The method further includes generating two or more predictions of the cancer status based at least in part on the identified plurality of features from two or more classifiers, wherein the two or more classifiers include at least a first classifier and a second classifier.

[0173] In some embodiments, the two or more predictions comprise a first prediction from a diagnostic classifier provided with RNA features, a second prediction from a cohort classifier provided with RNA features, a third prediction from a tissue classifier provided with RNA features, a fourth prediction from a diagnostic classifier provided with RNA splicing features, a fifth prediction from a cohort classifier provided with RNA splicing features, a sixth prediction from a diagnostic classifier provided with CNV features, a seventh prediction from a cohort classifier provided with CNV features, an eighth prediction from a diagnostic classifier provided with DNA features, and a ninth prediction from a diagnostic classifier provided with viral features.

[0174] The method further includes combining the two or more predictions in a final classifier to identify a cancer status diagnosis for the somatic tumor specimen (e.g., identify a TUO classification 2382d). In some embodiments, combining the two or more predictions in the final classifier further includes scaling each prediction of the two or more predictions based at least in part on a respective confidence level in each respective prediction, and generating a combined prediction based at least in part on each scaled prediction.

[0175] In some embodiments, the corresponding confidence level for each prediction of the two or more predictions is at least 0.5, at least 0.6, at least 0.7, at least 0.8, or at least 0.9. In some embodiments, the confidence level for the prediction is at least 0.9, at least 0.95, or at least 0.99. In some embodiments, the scaling of each prediction is linear (e.g., the scaling includes a linear combination of each prediction and the corresponding confidence level in the respective prediction). In some embodiments, the scaling of each prediction is non-linear (e.g., the scaling includes a non-linear combination of each prediction and the corresponding confidence level in the respective prediction). In some embodiments, scaling each prediction of the two or more predictions is performed as described below with respect to features module 2340.

[0176] In some embodiments, multiple classifiers can be used to generate a prediction of the cancer status, with each classifier provided with a respective subset of features from the identified plurality of features, and in some such embodiments, each prediction from each classifier in the plurality of classifiers is combined by a final classifier to determine a final diagnosis of the cancer status of the subject's somatic tumor specimen.

[0177] In some embodiments, the method further comprises providing the RNA splicing feature to a third classifier. In some such embodiments, generating comprises generating three or more predictions of the cancer status from three or more classifiers based at least in part on the identified plurality of features, the three or more classifiers comprising at least a first classifier, a second classifier, and a third classifier. In some such embodiments, combining comprises combining the three or more predictions to identify a diagnosis of the cancer status for the somatic tumor specimen.

[0178] In some embodiments, the method further includes providing the viral signature to a third classifier. In some such embodiments, generating comprises generating three or more predictions of the cancerous status from three or more classifiers based at least in part on the identified plurality of signatures, the three or more classifiers comprising at least a first classifier, a second classifier, and a third classifier. In some such embodiments, combining comprises combining the three or more predictions to identify a diagnosis of the cancerous status for the somatic tumor specimen.

[0179] In some embodiments, the method further comprises providing the copy number feature to a third classifier. In some such embodiments, generating comprises generating three or more predictions of the cancer status from three or more classifiers based at least in part on the identified plurality of features, the three or more classifiers comprising at least a first classifier, a second classifier, and a third classifier. In some such embodiments, combining comprises combining the three or more predictions to identify a diagnosis of the cancer status for the somatic tumor specimen.

[0180] In some embodiments, the method further includes providing the RNA feature to a first classifier, the copy number feature to a second classifier, and the RNA splicing feature to a third classifier. In some such embodiments, generating comprises generating three or more predictions of the cancer status based at least in part on the identified plurality of features from the three or more classifiers, the three or more classifiers comprising at least the first classifier, the second classifier, and a third classifier. In some such embodiments, combining comprises combining the three or more predictions to identify a diagnosis of the cancer status for the somatic tumor specimen.

[0181] In some embodiments, the method further includes providing the RNA features to a first classifier, the first classifier being a diagnostic classifier; providing the RNA features to a second classifier, the second classifier being a cohort classifier; and providing the RNA features to a third classifier, the third classifier being a tissue classifier. In some such embodiments, generating comprises generating three or more predictions of a cancer status based at least in part on the identified plurality of features from the three or more classifiers, the three or more classifiers including at least the first classifier, the second classifier, and the third classifier. In some such embodiments, combining comprises combining the three or more predictions to identify a diagnosis of a cancer status for the somatic tumor specimen.

[0182] In some such embodiments, the method further includes providing the DNA feature to a fourth classifier, the fourth classifier being a diagnostic classifier; providing the RNA splicing feature to a fifth classifier, the fifth classifier being a diagnostic classifier; and providing the RNA splicing feature to a sixth classifier, the sixth classifier being a cohort classifier. In some such embodiments, generating includes generating six or more predictions of the cancer status based at least in part on the identified plurality of features from the six or more classifiers, the six or more classifiers including at least a first classifier, a second classifier, a third classifier, a fourth classifier, a fifth classifier, and a sixth classifier. In some such embodiments, combining includes combining the six or more predictions to identify a diagnosis of the cancer status for the somatic tumor specimen.

[0183] In some embodiments, the final classification diagnosis differentiates between cancers of the same type (e.g., between types of sarcoma). In some embodiments, the final classification diagnosis differentiates between cancers based on location of origin (e.g., to identify the origin of metastases). In some embodiments, the final classification diagnosis differentiates between two or more cancer types, three or more cancer types, four or more cancer types, five or more cancer types, six or more cancer types, seven or more cancer types, eight or more cancer types, nine or more cancer types, or ten or more cancer types.

[0184] In some embodiments, the final classification of the cancerous condition comprises differentiating between lung adenocarcinoma, lung squamous cell carcinoma, oral adenocarcinoma, and oral adenocarcinoma. In some embodiments, the final classification of the cancerous condition comprises differentiating between systemic sarcoma, epithelioma, Ewing's sarcoma, gliosarcoma, leiomyosarcoma, meningioma, mesothelioma, and Rosai-Dorfman. In some embodiments, the final classification of the cancerous condition comprises differentiating between liver metastases of pancreatic origin, upper gastrointestinal origin, and biliary origin. In some embodiments, the final classification of the cancerous condition comprises differentiating between brain metastases of glioblastoma, oligodendroglioma, astrocytoma, and medulloblastoma. In some embodiments, the final classification of the cancerous condition comprises differentiating between non-small cell lung carcinoma squamous cell and adenocarcinoma. In some embodiments, the final classification of the cancerous condition comprises differentiating between one or more sarcomas with morphological features or protein expression of a carcinoma and one or more carcinomas with morphological features or protein expression of a sarcoma. In some embodiments, the final classification diagnosis of the cancerous condition comprises differentiating between one or more neuroendocrine, one or more carcinomas, and one or more sarcomas.

[0185] In some embodiments, the method further includes receiving final classifier diagnoses of cancer status for the somatic tumor specimens of the plurality of subjects. In some such embodiments, the method further includes calculating an entropy score for each subject in the plurality of subjects based at least in part on the respective final classifier diagnoses for each subject in the plurality of subjects. In some such embodiments, the method further includes identifying an entropy threshold based at least in part on the accuracy of the entropy score for each subject in the plurality of subjects. In some such embodiments, the method further includes training the final classifier using subjects from the subjects whose entropy scores satisfy the entropy threshold.

[0186] In some embodiments, the entropy score provides a basis for weighting (e.g., scaling) the respective contributions from each classifier from two or more classifiers in a final classifier. In some embodiments, the entropy score is used to remove subjects with low accuracy (e.g., high uncertainty) predictions from a plurality of subjects for purposes of training a final classifier (e.g., as part of improving the performance of the final classifier). In some embodiments, the entropy score is between 0 and 1 (e.g., at least 0, at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or at least 1). In some embodiments, the entropy score is between -1 and 1. In some embodiments, the entropy score is between -1 and 0. In some embodiments, the entropy score is in the range of -5 to 5.

[0187] In some embodiments, each entropy score is associated with a classification accuracy (e.g., the accuracy of the prediction determined from the final classifier). In some embodiments, the entropy scores are used to bin subjects from multiple subjects (e.g., subjects with the same entropy score are evaluated together). In some such embodiments, an average classification accuracy is determined for each entropy score (e.g., for each bin), and an entropy threshold is used to discard subjects with classification accuracies below the accuracy percentile. In some embodiments, identifying the entropy threshold includes identifying an accuracy percentile for the final classifier across multiple subjects. In some embodiments, subjects with entropy scores associated with the accuracy percentile are retained to train the final classifier. In some embodiments, the accuracy percentile is at least 0.75, at least 0.8, at least 0.85, at least 0.9, at least 0.925, at least 0.95, at least 0.975, or at least 0.99 accuracy (e.g., subjects with predictions at the lowest accuracy percentile are used to train the final classifier).

[0188] In some embodiments, identifying a diagnosis of the cancer condition further includes receiving subject information including one or more clinical events and differentiating the cancer condition between a new tumor and a recurrence of a previous tumor based at least in part on the one or more clinical events. In some embodiments, the one or more clinical events are received from a pathology report 134 (e.g., the pathology report is obtained as described with respect to block 222 above). In some embodiments, the one or more clinical events include at least one previous disease diagnosis. In some embodiments, the one or more clinical events include at least one previous treatment for the disease.

[0189] In some embodiments, the subject is being treated with a cancer drug, and the method further comprises using the diagnosis to evaluate the subject's response to the cancer drug. In some embodiments, the cancer drug is a hormone, immunotherapy, radiography, or cancer drug. In some embodiments, the cancer drug is lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, tetravalent human papillomavirus (types 6, 11, 16, and 18) vaccine, pertuzumab, pemetrexed, nilotinib, nilotinib, denosumab, abiraterone acetate, Promacta, imatinib, everolimus, palbociclib, erlotinib, bortezomib, or bortezomib.

[0190] In some embodiments, the method further includes providing the subject with an identified diagnosis of the cancerous condition for the somatic tumor specimen. In some embodiments, the identified diagnosis is provided to the subject as part of a patient report (e.g., patient report 900 as described with respect to Figures 9 and 10A-10G).

[0191] In some embodiments, the method further comprises administering a therapeutic regimen to the subject based at least in part on the diagnosis. In some such embodiments, the therapeutic regimen comprises administering to the subject a drug for cancer (e.g., one or more diagnosed cancers). In some embodiments, the drug for cancer is a hormone, immunotherapy, radiography, or cancer drug. In some embodiments, the drug for cancer is lenalidomide, pembrolizumab, trastuzumab, bevacizumab, rituximab, ibrutinib, tetravalent human papillomavirus (types 6, 11, 16, and 18) vaccine, pertuzumab, pemetrexed, nilotinib, nilotinib, denosumab, abiraterone acetate, Promacta, imatinib, everolimus, palbociclib, erlotinib, bortezomib, or bortezomib.

[0192] Training a classifier A classification model may be trained to perform the classification method described above and with reference to Figures 2A, 2B, and 2C: i) determine a set of cancer states for a test subject (e.g., a probability generated for each cancer state in the set of cancer states), and / or ii) classify the subject into a cancer state. Each and every embodiment described above and with reference to Figures 2A, 2B, and 2C may also be applied to the method of training a classification model as described below with reference to Figures 3A and 3B. Training a classification model may further require a dataset of data of reference subjects with known cancer states. A method for providing a trained classification model is described below, and a specific example of developing a trained classifier is detailed in Example 1.

[0193] In one embodiment, pre-training assessment can be performed across the entire training dataset to identify inputs that are best suited for training and those that are not. In one example, pre-training assessment can include calculating an entropy score for subjects and including patients whose scores meet a threshold while excluding subjects who fail to meet the threshold. The entropy score can function as a function that takes a probability vector and maps it to a single number that characterizes how "uncertain" the outcome is, where the uncertainty can be rooted in the prediction from the classifier. For a probability vector with components pi, entropy can be defined as: -Σpi log(pi). As an example, consider the case of a six-sided die. In that case, pi = [1 / 4, 1 / 6, 1 / 6, 1 / 6, 1 / 6], so the entropy is log(6), approximately 1.79. Now, if the die were rigged, it would always give the same answer, i.e., pi = [1, 0, 0, 0, 0]. In this case, the entropy is 0. The entropy score for a particular model can vary over a large range of values. In some cases, a high entropy score may be associated with an increased number of model errors. Therefore, identifying an entropy cutoff for a particular model may include evaluating model performance at each cutoff for a range of values ​​and selecting the cutoff with the best performance as measured by model accuracy. In one example, selecting a range of values ​​may include grouping all holdout model predictions from model training interactions by entropy score using various cutoffs in the entropy score, for example, in the negative range of [-4, -3.5, -3, -2.5, -2, -1.5, -1, -0.5, 0], and observing the model accuracy for each cutoff from [0.925, 0.925, 0.927, 0.930, 0.938, 0.950, 0.972, 0.969, 0.965]. The best overall model accuracy may be obtained by filtering any entropy scores above -1. Different entropy scores may be identified for other model training iterations.For each final training model, the best score for that training set can be used.Once identified, the entropy score of the subject can be used to determine how reliable the model is in predicting the origin site of the subject's tumor.In one example, if the entropy score is too high, the subject's TUO result may be invalid and not reported.

[0194] Block 302. Referring to block 302 of Figure 3A, a method for training a classifier stream to determine a set of cancer conditions is provided. As shown in Figures 16 and 18, the trained classification model achieves high accuracy for many cancer conditions.

[0195] Block 304. Referring to block 304 of FIG. 3A, the method obtains, in an electronic format, for each respective subject in a plurality of subjects (e.g., reference subjects), for each respective cancer condition in a set of cancer conditions, an indication of cancer, a first plurality of sequence reads, and an indication of whether the respective subject has a pathology report for the respective subject. The corresponding first plurality of sequence reads for each subject in the plurality of subjects are obtained from a respective plurality of RNA molecules or derivatives of the above-mentioned plurality of RNA molecules (e.g., derivatives such as cDNA). Each respective plurality of RNA molecules is derived from a corresponding somatic biopsy obtained from the respective subject. The pathology report for each subject includes at least one of a first estimate of tumor cellularity, an indication of whether the respective subject has metastatic or primary cancer, or a tissue site from which the somatic biopsy originated. In one example, the subjects in the plurality of subjects can be filtered based on their entropy scores to remove subjects with poor training.

[0196] In some embodiments, the plurality of RNA molecules is obtained by whole transcriptome sequencing (eg, as described above with reference to block 226 of Figure 2B).

[0197] Referring to block 306, in some embodiments, the method further includes obtaining a second plurality of sequence reads and a third plurality of sequence reads for each subject. The second plurality of sequence reads are obtained from the first plurality of DNA molecules or derivatives of the aforementioned DNA molecules. The third plurality of sequence reads are obtained from the second plurality of DNA molecules or derivatives of the aforementioned DNA molecules. The first plurality of DNA molecules are derived from a somatic biopsy obtained from each subject. The second plurality of DNA molecules are derived from a germline sample obtained from each subject or from a collection of normal controls not including the set of cancer conditions.

[0198] In some embodiments, the first and / or second plurality of DNA molecules are obtained by whole genome sequencing (e.g., as described above with reference to block 226 of FIG. 2B). In some embodiments, the first and / or second plurality of DNA molecules are obtained by targeted panel sequencing or panel sequencing.

[0199] In some embodiments, each subject in the plurality of subjects is a human. In some embodiments, the plurality of subjects comprises at least 50 subjects, at least 100 subjects, at least 150 subjects, at least 200 subjects, at least 250 subjects, at least 300 subjects, at least 400 subjects, at least 500 subjects, at least 750 subjects, at least 1000 subjects, at least 1500 subjects, at least 2000 subjects, at least 3000 subjects, at least 4000 subjects, or at least 5000 subjects.

[0200] Block 308. With reference to block 308 of FIG. 3A, the method continues by determining, for each respective subject in the plurality of subjects, a corresponding set of first sequence features for the respective subject from the first plurality of sequence reads for the respective subject (e.g., as described above with reference to block 230 of FIG. 2B).

[0201] Block 310. Referring to block 310 of FIG. 3A, in some embodiments, the method continues by determining, for each respective subject in the plurality of subjects, a set of second sequence features for the respective subject from a comparison of the second plurality of sequence reads and the third plurality of sequence reads for the respective subject (e.g., as described above with respect to blocks 234 and 236 of FIG. 2B).

[0202] Block 312. With reference to block 312 of FIG. 3B, the method continues by extracting, for each respective subject in the plurality of subjects, a plurality of pathological features from the pathology report for each subject, including a first estimate of tumor cellularity from the somatic biopsy and an indication of whether the respective subject has metastatic or primary cancer (e.g., as described above with reference to blocks 238-242 of FIG. 2C).

[0203] Referring to block 314, in some embodiments, extracting multiple pathological features from the pathology report further includes normalizing the pathology report. In some embodiments, normalizing the pathology report includes one or more data cleaning steps to enable comparison between pathology reports of different subjects. Various components of the pathology report are useful for determining the origin of cancer. Of particular use are diagnostic labels, which provide valuable information regarding cancer classification, such as the patient's disease state, disease stage and grade, pathology, and histology. In some embodiments, normalizing the pathology report includes natural language processing (NLP), which may include relabeling of the medical professional's diagnosis entry. Because there is no standardized scheme for sample annotation during pathology review, some processing of the diagnostic label in the pathology report is often required. Instead, the pathology report includes a "Diagnosis" field, which is an open-text box, allowing the medical professional to enter any value of their choice (see, for example, the Diagnosis field entry column in Table 1, as discussed in Example 5 below).

[0204] 12A-12B illustrate the accuracy of NLP re-labeling of diagnostic entries by comparing clustering performed according to different sets of labels determined by NLP. FIG. 12A provides a summary of the data, showing that each of the data points included in the analysis is from a respective patient with the blanket diagnosis or label of "sarcoma." As shown in FIG. 12B, using different, more specific labels for each data point results in clusters that are each more closely associated with a single label. In some embodiments, as discussed in more detail below in Example 6, there may be a loss of information in the resulting clusters when the labels are highly specific (e.g., overly specific).

[0205] Block 316. Referring to block 316 of Figure 3B, the method continues by inputting at least the first set of sequence features and a plurality of pathological features for each respective subject in the plurality of subjects into an untrained classification model. The method continues by training the untrained classification model for an indication of whether each respective subject in the plurality of subjects has each respective cancer condition in the set of cancer conditions, thereby obtaining a trained classification model. The trained classification model is configured to provide: i) for each respective cancer condition in the set of cancer conditions, a likelihood that the test subject has or does not have the respective cancer condition; ii) a likelihood that the test subject has or does not have the cancer condition; or iii) a likelihood that the test subject has or does not have the predicted cancer condition.

[0206] Referring to block 318, in some embodiments, the inputting further includes inputting the second set of sequence features for each respective subject in the plurality of subjects (e.g., along with the first set of sequence features and the plurality of pathological features) into an untrained classification model to obtain a trained classification model.

[0207] With reference to block 320, in some embodiments, the trained classification model includes a trained classifier stream. With reference to block 322 (and as described further below), in some embodiments, the trained classifier stream includes, by way of non-limiting example, a hierarchical model, a deep neural network, a multi-task multi-kernel learning engine, or a nearest neighbor engine. Exemplary nearest neighbor and neural network algorithms suitable for use in block 316 are described above with respect to block 244 of FIG. 2C.

[0208] A multi-task, multi-kernel learning engine suitable for use as the classifier in block 316 is described, for example, in Widmer et al., 2015, Framework for Multi-Task Multiple Kernel Learning and Applications in Genome Analysis. arXiv:1506.09153vl, which is incorporated herein by reference in its entirety. The goal of multi-task, multi-kernel learning methods is to identify one or more subsets of similar features in input data, enabling the discovery of underlying structure in the input data. One specific algorithm that can be used to identify data subsets for multi-task, multi-kernel learning is the least absolute shrinkage and selection operator (Lasso). Additional algorithms are described in detail, for example, in Yousefi et al., 2017, Multi-Task Learning Using Neighborhood Kernels. arXiv:1707.03426vl, which is incorporated herein by reference.

[0209] Hierarchical algorithms suitable for use as the classification model in block 316 are described, for example, in Galea et al., 2017 Scientific Reports 7:14981 and Silla et al., 2011 Data Mining and Knowledge Discovery 22:31-72, each of which is incorporated herein by reference. The result of hierarchical classification is typically layered or branched, e.g., a directed acyclic graph.

[0210] Additional embodiments directed to retrieving patient data from a patient data store In some embodiments, the artificial intelligence system 2300 retrieves features associated with a patient from a patient data store. In some embodiments, the patient data store includes one or more feature modules 2340 that contain a collection of features available for all patients in the system. In some embodiments, these features are used to generate a prediction of the origin of the patient's tumor. While the feature range across all patients is informationally dense, an individual patient's Feature Set, in some embodiments, is sparsely populated across the aggregate feature range of all features across all patients. For example, the feature range across all patients may extend to tens of thousands of features, while a patient's unique Feature Set may include only a subset of hundreds or thousands of the aggregate feature range based on the records available for that patient.

[0211] In some embodiments, the feature set may include a diverse set of fields available within a patient's health record. In some embodiments, clinical information, such as information in the health record 2344, is based on fields entered into an electronic medical record (EMR) or electronic health record (EHR) 2346 by a physician, nurse, or other medical professional or representative. In some embodiments, other clinical information is curated from other sources, such as molecular fields from gene sequencing reports (2345). In some embodiments, sequencing may include next-generation sequencing (NGS), including long-read, short-read, paired-end, or other forms of sequencing the patient's somatic and / or normal genome. In some embodiments, the comprehensive set of features in the additional feature module combines together various features across various medical disciplines, which may include diagnosis, response to treatment regimens, genetic profiles, clinical and phenotypic features, and / or other medical, geographic, demographic, clinical, molecular, or genetic features. For example, the subset of features may include molecular data features, such as features from the RNA feature module 2341 or the DNA feature module 2342, that include sequencing results of a patient's germline or somatic specimen.

[0212] In some embodiments, another subset of features (imaging features from imaging feature module 2347) includes features identified, for example, through a pathologist's review of a specimen, e.g., review of stained H&E or IHC slides. As another example, a subset of features may include derived features 2349 obtained from analysis of the individual and combined results of such feature sets. Features derived from DNA sequencing and RNA sequencing may include genetic variants from variant science module 2348 present in the sequenced tissue. Further analysis of genetic variants may include additional steps, such as identifying single or multiple nucleotide polymorphisms, identifying whether the variation is an insertion or deletion event, identifying loss or gain of function, identifying fusions, identifying splicing, calculating copy number variation (CNV), calculating microsatellite instability, calculating tumor mutation burden (TMB), or other structural variations in DNA and RNA. Analysis of slides for H&E staining or IHC staining may reveal characteristics such as tumor infiltration, programmed cell death-ligand 1 (PD-L1) status, human leukocyte antigen (HLA) antigen status, or other immunological features.

[0213] In some embodiments, features derived from structured, curated, or electronic medical or health records may include clinical features such as diagnosis, symptoms, therapy, outcome; patient name, date of birth, sex, ethnicity, date of death, address, smoking status, date of diagnosis of cancer, disease, disorder, diabetes, depression, other physical or mental illness, patient demographics such as personal medical history, family medical history; clinical diagnosis such as date of initial diagnosis, date of metastatic diagnosis, cancer staging, tumor characterization, tissue of origin; treatment and outcome such as line of therapy, therapy group, clinical trial, medications prescribed or taken, surgery, radiation therapy, imaging, adverse effects, associated outcomes, genetic testing; performance score, lab test, pathology results, prognostic indicators; laboratory information such as date of genetic testing, testing provider used, testing method used, e.g., gene sequencing method or gene panel; genetic results, e.g., genes included, variants, expression levels / status, or corresponding dates for any of the above. Clinical features may also include imaging features.

[0214] In some embodiments, the omics feature module 2343 includes features derived from information from additional medical or research-based omics fields, including proteomics, transcriptomics, epigenomics, metabolomics, microbiomics, and other multi-omics fields. In some embodiments, features derived from organoid modeling laboratories include DNA and RNA sequencing information closely linked to each organoid, as well as results from treatments applied to those organoids. In some embodiments, features derived from imaging data further include stained slides, reports related to tumor size, changes in tumor size over time, including treatments during changes, and machine learning approaches for classifying PDL1 status, HLA status, or other features from imaging data. In some embodiments, the other features include additional derived feature sets from other machine learning approaches based at least in part on any new features and / or combinations of those listed above. For example, imaging results may need to be combined with MSI calculations derived from RNA expression to determine additional imaging features. In some embodiments, the machine learning model may generate a likelihood that the patient's cancer will metastasize to a particular organ, or a probability of metastasis to further organs in the patient's body in the future. In some embodiments, other features that can be extracted from medical information are also used. There are thousands of features, and the above list of feature types is merely representative and should not be construed as an exhaustive or exclusive list of features.

[0215] In some embodiments, the variance module 2350 includes one or more microservices, servers, scripts, or other executable algorithms that generate variance features associated with de-identified patient features from the feature set 2305. In some embodiments, the variance module may retrieve input from the feature set and provide the variance for the storage device 2310. Example variance modules 2352a-n may include one or more of the following variances as a set of variance modules 2353a-n:

[0216] In some embodiments, the IHC (immunohistochemistry) module identifies antigens (proteins) in cells of tissue sections by utilizing the principle of antibodies specifically binding to antigens in biological tissues. IHC staining is widely used in the diagnosis of abnormal cells, such as those found in cancerous tumors. Specific molecular markers are characteristic of specific cellular events, such as proliferation or cell death (apoptosis). IHC is also widely used in basic research to understand the distribution and localization of biomarkers and differentially expressed proteins in different parts of biological tissues. Visualizing antibody-antigen interactions can be achieved in several ways. In the most common case, antibodies are conjugated to enzymes (e.g., peroxidase) that can catalyze a color reaction during immunoperoxidase staining. Alternatively, antibodies can be tagged with fluorophores such as fluorescein or rhodamine in immunofluorescence. In some embodiments, approximations are generated from RNA expression data, H&E slide imaging data, or other data.

[0217] In some embodiments, the therapy module identifies differences in how cancer cells (or other cells nearby) grow and thrive, as well as drugs that "target" these differences. Treatment with these drugs is called targeted therapy. For example, many targeted drugs are lethal to cancer cells because their internal "programming" makes them different from normal, healthy cells, while sparing most healthy cells. Targeted drugs may block or turn off chemical signals that tell cancer cells to grow and divide rapidly without affecting normal cells; change proteins in cancer cells so they die; stop the creation of new blood vessels to nourish them; trigger the patient's immune system to kill the cancer cells; or deliver toxins to the cancer cells to kill them. Some targeted drugs are more "targeted" than others. Some may target only a single change in cancer cells, while others may affect several different changes. Others enhance the way the patient's body fights cancer cells. This may affect where these drugs work and what side effects they cause. In some embodiments, matching targeted therapy may involve identifying a therapy target in a patient and meeting any other inclusion or exclusion criteria that may identify patients for whom the therapy is likely to be effective.

[0218] In some embodiments, the clinical trials module identifies and tests hypotheses for treating cancers with specific characteristics by matching patient characteristics to clinical trials. These trials have inclusion and exclusion criteria that must be matched to enroll patients and may be captured and structured from publications, clinical trial reports, or other documents.

[0219] In some embodiments, the amplification module identifies genes that are disproportionately increased in count (e.g., the number of gene products present in a specimen) relative to other genes. Amplification may cause the increased count genes to go dormant, become overactive, or operate in another unexpected manner. In some embodiments, amplification may be detected at the gene level, variant level, RNA transcript or expression level, or protein level. In some embodiments, detection is performed across all different detection mechanisms or levels and validated against each other.

[0220] In some embodiments, the isoform module specifies alternative splicing (AS), a biological process in which more than one mRNA type (isoform) is generated from the same gene transcript through different combinations of exons and introns. Large-scale genomic studies estimate that 30-60% of mammalian genes are alternatively spliced. The possible alternative splicing patterns for a gene can be highly complex, and this complexity increases rapidly as the number of introns in the gene increases. In silico, alternative splicing prediction can identify genomic loci through searching mRNA sequences against genomic sequences, extracting sequences for the genomic loci, extending the sequences up to 20 kb on both ends, searching the genomic sequence (repeated sequences were masked), extracting splicing pairs (GT-AG consensus or two boundaries of an alignment gap with more than two expressed sequence tags aligned on both ends of the gap), assembling the splicing pairs according to their coordinates, determining gene boundaries (splice pair predictions are generated at this point), generating predicted gene structures by aligning the mRNA sequences to the genomic template, and comparing the splice pair predictions and gene structure predictions to find alternatively spliced ​​isoforms, which may find large insertions or deletions within a set of mRNAs that share a large portion of the aligned sequences.

[0221] In some embodiments, a SNP (single nucleotide polymorphism) module identifies a single nucleotide substitution occurring at a specific position in the genome, where each variation is present to some discernible extent in the population (e.g., >1%). For example, at a particular base position, or locus, in the human genome, a C nucleotide may appear in most individuals, but in a minority of individuals, the position is occupied by an A. This means that an SNP exists at this particular position, and the two possible nucleotide variations (C or A) are said to be alleles at this position. SNPs underlie differences in human susceptibility to a wide range of diseases (e.g., sickle cell anemia, β-thalassemia, and cystic fibrosis resulting from SNPs). Disease severity and how the body responds to treatment are also manifestations of genetic variation. For example, a single base mutation in the APOE (apolipoprotein E) gene is associated with a lower risk of Alzheimer's disease. A single nucleotide variant (SNV) is a single nucleotide variation that can occur in somatic cells without frequency limitations. Somatic single nucleotide variations (e.g., caused by cancer) may also be referred to as single nucleotide changes. In some embodiments, an MNP (multiple nucleotide polymorphism) module identifies substitutions of consecutive nucleotides at specific positions in a genome.

[0222] In some embodiments, the indel module can identify insertions or deletions of bases in the genome of organisms classified within small genetic variations. Indels typically measure 1-10,000 base pairs in length, while microindels are defined as indels resulting in a net change of 1-50 nucleotides. Indels can be constructed with SNPs or point mutations. Indels insert and / or delete nucleotides from a sequence, while point mutations are a form of substitution that replaces one of the nucleotides without changing the overall number in DNA. Indels, which are insertions and / or deletions, can be used as genetic markers in natural populations, particularly in phylogenetic studies. Indel frequencies tend to be significantly lower than those of single nucleotide polymorphisms (SNPs), except near highly repetitive regions, including homopolymers and microsatellites.

[0223] In some embodiments, the MSI (microsatellite instability) module can identify genetic hypermutability (predisposition to mutations) resulting from impaired DNA mismatch repair (MMR). The presence of MSI represents phenotypic evidence of impaired MMR. MMR corrects errors that naturally occur during DNA replication, such as single-base mismatches or short insertions and deletions. Proteins involved in MMR correct polymerase errors by binding to mismatched segments of DNA, excising the errors, and inserting the correct sequence in their place. Cells with abnormally functioning MMR are unable to correct errors that occur during DNA replication, which causes the cells to accumulate errors in their DNA. This generates novel microsatellite fragments. Polymerase chain reaction-based assays can reveal these novel microsatellites and provide evidence for the presence of MSI. Microsatellites are repetitive sequences of DNA. These sequences can be made up of repeating units ranging from 1 to 6 base pairs in length. The length of these microsatellites varies greatly from person to person, contributing to an individual's DNA "fingerprint," but each individual has a set length of microsatellites. The most common microsatellites in humans are dinucleotide repeats of the nucleotides C and A, which occur tens of thousands of times throughout the genome. Microsatellites are also known as simple sequence repeats (SSRs).

[0224] In some embodiments, the TMB (tumor mutation burden) module may identify a measure of mutations carried by tumor cells and is a predictive biomarker being studied to assess its association with response to immuno-oncology (IO) therapy. Tumor cells with high TMB may harbor more neoantigens, along with an associated increase in cancer-fighting T cells in the tumor microenvironment and periphery. These neoantigens can be recognized by T cells and induce anti-tumor responses. TMB has recently emerged as a quantitative marker that can help predict potential response to immunotherapy across different cancers, including melanoma, lung cancer, and bladder cancer. TMB is defined as the total number of mutations per coding region of the tumor genome. Importantly, TMB is consistently reproducible. It provides a quantitative measurement that can be used to better inform treatment decisions, such as the selection of targeted therapy or immunotherapy, or enrollment in clinical trials.

[0225] In some embodiments, the CNV (copy number variation) module can specifically identify variations from the normal genome in the copy number of genes, parts of genes, or other parts of the genome that are not defined by genes, and any subsequent effects from analyzing the sequence of genes, variants, alleles, or nucleotides. CNV is a phenomenon in which structural variations can occur in segments of nucleotides, or base pairs, including repeats, deletions, or inversions.

[0226] In some embodiments, a fusion module may identify a hybrid gene formed from two previously separated genes. Hybrid genes can be the result of translocations, interstitial deletions, or chromosomal inversions. Gene fusions can play an important role in tumorigenesis. Fusion genes can contribute to tumorigenesis because they can produce abnormal proteins that are much more active than non-fusion genes. Many fusion genes are cancer-causing oncogenes. Fusion genes include BCR-ABL, TEL-AML1 (ALL with t(12;21)), AML1-ETO (M2 AML with t(8;21)), and TMPRSS2-ERG, which has an interstitial deletion on chromosome 21 and is commonly found in prostate cancer. In the case of TMPRSS2-ERG, the fusion product regulates prostate cancer by disrupting androgen receptor (AR) signaling and inhibiting AR expression by oncogenic ETS transcription factors. Most fusion genes are found in hematological cancers, sarcomas, and prostate cancer. BCAM-AKT2 is a fusion gene specific and unique to high-grade serous ovarian cancer. Oncogenic fusion genes can result in gene products with new or distinct functions from the two fusion partners. Alternatively, proto-oncogenes can be fused to strong promoters, whereby oncogenic functions are set up through upregulation caused by the strong promoter of the upstream fusion partner. The latter is common in lymphomas, where oncogenes are juxtaposed to the promoters of immunoglobulin genes. Oncogenic fusion transcripts can also be caused by trans-splicing or read-through events. Because chromosomal translocations play such an important role in neoplasia, a specialized database of chromosome aberrations and gene fusions in cancer has been created. This database is called the Mitelman Database of Chromosome Aberrations and Gene Fusions in Cancer.

[0227] In some embodiments, the VUS (Variant of Unknown Significance) module may identify variants that are detected in a patient's genome (particularly in a patient's cancer specimen) but cannot be classified as pathogenic or benign at the time of detection. VUS are cataloged from publications to identify whether they are classified as benign or pathogenic.

[0228] In some embodiments, the DNA pathway module identifies defects in DNA repair pathways that allow cancer cells to accumulate genomic alterations that contribute to their aggressive phenotype. Cancerous tumors depend on residual DNA repair capacity to survive damage induced by genotoxic stress, resulting in isolated DNA repair pathways being inactivated in cancer cells. DNA repair pathways are generally thought of as mutually exclusive mechanistic units that deal with different types of lesions in different cell cycle phases. However, recent preclinical studies provide strong evidence that multifunctional DNA repair hubs, which are involved in multiple conventional DNA repair pathways, are frequently altered in cancer. Identifying potentially affected pathways may lead to important patient treatment considerations.

[0229] In some embodiments, the raw count module determines the count of variants detected from sequencing data. For DNA, in some embodiments, this includes the number of reads from sequencing that correspond to a particular variant in a gene. For RNA, in some embodiments, this includes gene expression counts or transcriptome counts from sequencing.

[0230] In some embodiments, classification involves classification by one or more trained models to generate predictions, while other structural variant classifications may involve evaluating features from the feature set, variations from the variation module, and other classifications from one or more classification modules. The structural variant classification may provide the classification in a stored classification storage device. An exemplary classification module may include classification of CNVs, such that "reportable" may mean that the CNV has been identified in one or more reference databases as affecting tumor cancer characterization, disease state, or pharmacogenetics; "non-reportable" may mean that the CNV has not been so identified; and "conflicting evidence" may mean that the CNV has both evidence suggesting "reportable" and "non-reportable." Furthermore, therapy-related classifications may be similarly confirmed by reference to any reference datasets of therapies that may be affected by the detection (or non-detection) of the CNV. Other classifications may involve the application of machine learning algorithms, neural networks, regression techniques, graphing techniques, inductive reasoning approaches, or other artificial intelligence assessments within the module. In some embodiments, the clinical trial classifier may include evaluating variants identified from variation modules identified as significant or reportable, evaluating all available clinical trials to identify inclusion and exclusion criteria, mapping patient variants and other information to the inclusion and exclusion criteria, and classifying the clinical trial as applicable or not applicable to the patient. In some embodiments, similar classifications are performed for therapy, loss of function, gain of function, diagnostic, microsatellite instability, tumor mutation burden, indels, SNPs, MNPs, fusions, CNVs, splicing, and other variations that can be classified based on the results of the variation modules. Additionally, in some embodiments, models trained to classify tumor type for patients with tumors of unknown origin are generated according to the disclosure herein. In some embodiments, the classifications are generated and stored as part of the feature set 2305 in the stored classification database 2330.

[0231] In some embodiments, each of the feature sets, variation modules, structural variants, and feature stores are communicatively coupled to a data bus to transfer data between each module for processing and / or storage. In some embodiments, each of the feature sets, variation modules, and classifications may be communicatively coupled to one another for independent communication without sharing a data bus.

[0232] In addition to the above features and listed modules, in some embodiments, the feature modules may further include one or more of the following modules within the respective modules, either as sub-modules or as standalone modules:

[0233] In some embodiments, the germline / somatic DNA feature module includes a set of features associated with DNA origin information for the patient or the patient's tumor. These features may include raw sequencing results, genes, mutations, variant calls, and variant characterizations, such as those stored in FASTQ, BAM, VCF, or other sequencing file types known in the art. In some embodiments, genomic information from the patient's normal sample is stored as germline, and genomic information from the patient's tumor sample is stored as somatic.

[0234] In some embodiments, the RNA feature module includes a feature set associated with a patient's RNA-derived information, such as transcriptome information. These features may include raw sequencing results, transcriptome expression, genes, mutations, variant calls, and variant characterization.

[0235] In some embodiments, the metadata module includes a set of features associated with the human genome, protein structures, and their effects, for example, changes in energy stability based on protein structure.

[0236] In some embodiments, the clinical module includes feature sets associated with information derived from the patient's clinical records and records from the patient's family. These may be extracted from unstructured clinical documents, EMRs, EHRs, or other sources of the patient's medical history. The information may include the patient's symptoms, diagnoses, treatments, medications, therapies, hospice, response to treatment, laboratory test results, medical history, respective geographic location, demographics, or other characteristics of the patient that may be found in the patient's medical record. Information regarding treatments, medications, therapies, etc. may be obtained as recommendations or prescriptions and / or as confirmation that such treatments, medications, therapies, etc. have been administered or taken.

[0237] In some embodiments, the imaging module includes feature sets associated with information derived from patient imaging records. The imaging records may include H&E slides, IHC slides, radiology images, and other medical imaging that physicians may prescribe in the course of diagnosing and treating various diseases and disorders. These features may include TMB, ploidy, purity, nuclear-cytoplasmic ratio, macronuclei, cell state changes, biological pathway activation, hormone receptor changes, immune cell infiltration, immune biomarkers for MMR, MSI, PDL1, CD3, FOXP3, HRD, PTEN, PIK3CA, collagen or stromal composition, appearance, density, or characterization, tumor budding, size, aggressiveness, metastasis, immune status, chromatin morphology, and other characterizations of cells, tissues, or tumors for prognostic purposes.

[0238] In some embodiments, an epigenomic module, e.g., an omics-derived epigenomic module, includes a set of features associated with information derived from DNA modifications that regulate gene expression rather than changes to DNA sequence. These modifications are often the result of environmental factors based on what a patient may breathe, eat, or drink. These features may include DNA methylation, histone modifications, or other factors that inactivate genes or cause changes in gene function without changing the sequence of nucleotides in the gene.

[0239] In some embodiments, a microbiome module, e.g., an omics-derived microbiome module, includes a feature set associated with information derived from a patient's viruses and bacteria. Viral genomes may be generated to identify which viruses are present in a patient's specimen based on genomic features that map to viral DNA or RNA (e.g., viral reference genomes) instead of the human genome. These features may include viral infections, which may affect the treatment and diagnosis of certain diseases, as well as bacteria present in the patient's gastrointestinal tract, which may affect the effectiveness of medications taken by the patient.

[0240] In some embodiments, a proteome module, e.g., an omics-derived proteome module, includes a feature set associated with information derived from proteins produced in a patient. These features may include protein composition, structure, and activity, when and where the protein is expressed, protein production, rate of degradation, and steady-state abundance, how the protein is modified, e.g., post-translational modifications such as phosphorylation, trafficking of the protein between subcellular compartments, participation of the protein in metabolic pathways, how proteins interact with each other, or post-translational modifications of the protein from RNA, such as phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, or nitrosylation.

[0241] In some embodiments, the additional omics modules are included in feature sets associated with all different branches of omics, including: cognitive genomics, a feature set that includes the study of changes in cognitive processes associated with genetic profiles; comparative genomics, a feature set that includes the study of the association of genome structure and function across different species or strains; functional genomics, a feature set that includes the study of gene and protein function and interactions, including transcriptomics; interactomics, a feature set that includes studies related to large-scale analysis of gene-gene, protein-protein, or protein-ligand interactions; metagenomics, a feature set that includes the study of the metagenomic, i.e., genetic material recovered directly from environmental samples; neurogenomics, a feature set that includes the study of genetic influences on nervous system development and function; pangenomics, a feature set that includes the study of the entire set of gene families found within a given species; and sequencing and analysis of individual genomes, where genotypes are known, to compare individual genotypes with published literature to determine trait expression and potential disease risk to enhance personalized medicine suggestions. personal genomics, a set of features including the study of genomics related to the genome; epigenomics, a set of features including the study of the structure of the genome, including proteins and RNA binders, alternative DNA structures, and chemical modifications of DNA; nucleomics, a set of features including the study of the complete set of genomic components that form the cell nucleus as a complex and dynamic biological system; lipidomics, a set of features including the study of cellular lipids, including modifications made to any specific set of lipids produced by the patient; proteomics, a set of features including the study of proteins, including modifications made to any specific set of proteins produced by the patient; immunoproteomics, a set of features including the study of large sets of proteins involved in the immune response; phosphoproteomics, a set of features including the study of protein phosphorylation patterns, including modifications made to any specific set of proteins produced by the patient, including the use of proteomics mass spectrometry data for protein expression studies.Nutriproteomics, a set of features that involves the study of molecular targets of dietary nutritional and non-nutritional components; proteogenomics, a set of features that involves the study of biological research at the intersection of proteomics and genomics, including data identifying gene annotations; structural genomics, a set of features that involves the study of the three-dimensional structure of all proteins encoded by a given genome using a combination of modeling approaches; glycomics, a set of features that involves the study of sugars and carbohydrates and their effects in patients; foodomics, a set of features that involves the study of the intersection between the domains of food and nutrition through the application and integration of technologies to improve consumer well-being, health, and knowledge; transcriptomics, a set of features that involves the study of RNA molecules, including mRNA, rRNA, tRNA, and other non-coding RNAs produced in cells; and metabolomics, a set of features that involves the study of chemical processes, including metabolites, or the unique chemical fingerprints left by specific cellular processes, and their small molecule metabolite profiles. pharmacogenomics, a set of features that involves the study of the sum of variation within the human genome and its effect on drugs; pharmacomicrobiomics, a set of features that involves the study of variation within the human microbiota and its effect on drugs; toxicogenomics, a set of features that involves the study of gene and protein activity within specific cells or tissues of an organism in response to a toxicant; and the mitointeractome, a set of features that involves the study of processes by which mitochondrial proteins interact with the biological substrates of normal behavior, including the application of psychogenomics to the study of drug addiction to develop more effective treatments and objective diagnostic tools, preventative measures, and cures for these disorders.These include psychogenomics, a set of features that involves the study of the process of applying the powerful tools of genomics and proteomics to achieve a better understanding of the biological substrates of brain diseases that manifest as behavioral abnormalities; stem cell genomics, a set of features that involves the study of stem cell biology to establish stem cells as a model system for understanding human biology and disease states; connectomics, a set of features that involves the study of neural connections within the brain; microbiomics, a set of features that involves the study of the genomes of microbial communities inhabiting the gastrointestinal tract; cellomics, a set of features that involves the study of quantitative cellular analysis and the use of bioimaging methods and bioinformatics; tomomics, a set of features that involves the study of tomography and omics methods to understand the biochemistry of tissues or cells with high spatial resolution from imaging mass spectrometry data; ethomics, a set of features that involves the study of high-throughput machine measurements of patient behavior; and videoomics, a set of features that involves the study of video analysis paradigms inspired by genomics principles, where sequential image sequences, or videos, can be interpreted as the capture of a single image evolving over time to reveal patient insights into mutations. ,

[0242] In some embodiments, a sufficiently robust feature set includes all of the features disclosed above. However, models and predictions based on available features include models optimized and trained from a much more restricted selection of features than an exhaustive feature set. In some embodiments, such a constrained feature set includes tens to hundreds of features. For example, a model's constrained feature set may include the genomic results of sequencing a patient's tumor, derived features based on the genomic results, the patient's tumor origin, the patient's age at diagnosis, the patient's gender and race, and symptoms that the patient brought to a physician's attention during a routine checkup.

[0243] In some embodiments, the feature store may enrich a patient's feature set through the application of machine learning and analytics by selecting any features, variations, or calculated outputs derived from the patient's features and adding variations to those features. In some embodiments, such a feature store may generate new features from the original features found in the feature module, or identify and store key insights or analyses based on the features. In some embodiments, feature selection is based at least on the variations or calculations generated, including calculations of insertions or deletions of single or multiple nucleotide polymorphisms in the genome, tumor mutation burden, microsatellite instability, copy number variations, fusions, or other such calculations. In some embodiments, exemplary outputs of generated variations or calculations that may inform future variations or calculations include findings of hypertrophic cardiomyopathy (HCM) and variant mMYH7. In some embodiments, previously classified variants may be identified in a patient's genome that may inform classification of novel variants or indicate additional risk for disease. An exemplary approach involves enrichment of variants and their respective classifications to identify regions in MYH7 associated with HCM. Novel variants detected from patient sequencing localized in this region would increase the patient's risk of HCM. In some embodiments, features that can be utilized for such alteration detection include the structure of MYH7 and the classification of variants therein. In some embodiments, enrichment-focused models can isolate such variants. An exemplary output of the generated alterations or calculations that can inform future alterations or calculations includes the finding of lung cancer and variants in EGFR, an epidermal growth factor receptor gene mutated in approximately 10% of non-small cell lung cancers and approximately 50% of lung cancers from non-smokers. In some embodiments, previously classified variants can be identified in the patient's genome that can inform the classification of novel variants or indicate additional risk for disease. Exemplary approaches can include identifying nearby regions or enriching variants and their respective classifications for those that interact with EGFR and have evidence of association with cancer.Novel variants detected from patient sequencing that are localized in this region or interact with this region would increase the patient's risk. In some embodiments, features that can be utilized in such alteration detection include the structure of EGFR and the classification of variants therein. In some embodiments, enrichment-focused models can isolate such variants.

[0244] In some embodiments, the above-referenced classification models may include one or more classification models 2382a-n, which may be implemented as artificial intelligence engines and may include gradient boosting models, random forest models, neural networks (NNs), regression models, naive Bayes models, or machine learning algorithms (MLAs). The MLAs or NNs may be trained from a training dataset. In an exemplary predictive profile, the training dataset may include imaging, pathology, clinical, and / or molecular reports, as well as patient details, e.g., curated from EHRs or gene sequencing reports. MLAs include supervised algorithms (e.g., algorithms where features / classifications in the data set are annotated) using linear regression, logistic regression, decision trees, classification and regression trees, naive Bayes, and nearest neighbor clustering; unsupervised algorithms (e.g., algorithms where features / classifications in the data set are not annotated) using Apriori, average clustering, principal component analysis, random forests, and adaptive boosting; and semi-supervised algorithms (e.g., algorithms where an incomplete number of features / classifications in the data set are annotated) using generative approaches (e.g., Gaussian mixtures, multinomial mixtures, hidden Markov models), sparse separation, graph-based approaches (e.g., minimum cut, harmonic functions, manifold regularization), heuristic approaches, or support vector machines. NNs include conditional random fields, convolutional neural networks, attention-based neural networks, deep learning, long-short-term memory networks, or other neural models. The training dataset includes pathology reports covering multiple tumor samples, RNA expression data for each sample, and imaging data for each sample. Although MLA and neural network identify different approaches to machine learning, these terms may be used interchangeably herein. Thus, unless expressly stated otherwise, a reference to an MLA may include a corresponding NN, or a reference to a NN may include a corresponding MLA.Training may include providing an optimized dataset, labeling these traits as they occur in patient records, and training an MLA to predict or classify based on new inputs. Artificial neural networks (NNs) are efficient computational models that have demonstrated their strength in solving difficult problems in artificial intelligence. They have also been shown to be universal approximators (capable of representing a wide range of functions given appropriate parameters). In some embodiments, some MLAs may identify important features and assign coefficients or weights to them. The coefficients may be multiplied by the frequency of occurrence of the feature to generate a score, and when the score of one or more features exceeds a threshold, a particular classification may be predicted by the MLA. In some embodiments, the coefficient scheme may be combined with a rule-based scheme to generate more complex predictions, such as predictions based on multiple features. For example, 10 key features may be identified across different classifications. In some embodiments, a list of coefficients may exist for the key features, and a rule set may exist for the classification. In some embodiments, the rule set may be based on the number of occurrences of a feature, its scaled weight, or other qualitative and quantitative evaluations of the feature encoded in logic known to those skilled in the art. In other MLAs, features may be organized in a binary tree structure. For example, the key feature that best identifies the classification may be present at the root of the binary tree and each subsequent branch in the tree until the classification is assigned based on reaching the terminal node of the tree. For example, the binary tree may have a root node that tests for a first feature. The occurrence or non-occurrence of this feature must be present (a binary decision), and the logic may traverse the branch that is true for the item being classified. Additional rules may be based on thresholds, ranges, or other qualitative and quantitative tests. Supervised methods are useful when the training dataset has many known values ​​or annotations, but the nature of EMR / EHR documents is such that annotations may not be many. When considering large amounts of unlabeled data, unsupervised methods are useful for binning / bucketing the cases in the dataset.A single instance of the above model, or two or more such instances combined, may constitute a model for purposes of the models, artificial intelligence, neural network, or machine learning algorithms herein.

[0245] In some embodiments, the stacked TUO classifier 2400 may receive one or more features from the artificial intelligence engine 2300 of FIG. 23 and predict cancer status in the TUO classification 2382 using one or more classifiers 2382a-n.

[0246] In some embodiments, the set of cancer conditions is selected from the group consisting of acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), adolescent cancer, adrenocortical carcinoma, AIDS-related cancer, Kaposi's sarcoma (soft tissue sarcoma), AIDS-related lymphoma (lymphoma), primary CNS lymphoma (lymphoma), anal cancer, appendix cancer, astrocytoma, childhood (brain cancer), atypical teratoid / rhabdoid tumor, childhood, central nervous system (brain cancer), basal cell carcinoma of the skin, bile duct cancer, bladder cancer, bone cancer (including Ewing's sarcoma and osteosarcoma and malignant fibrous histiocytoma), brain tumor, breast cancer, bronchial tumor ( Lung Cancer, Burkitt's Lymphoma, Carcinoid Tumor (Gastrointestinal), Carcinoma of Unknown Primary, Cardiac (Heart) Tumor, Pediatric, Central Nervous System, Atypical Teratoid / Rhabdoid Tumor, Pediatric (Brain Cancer), Medulloblastoma and Other CNS Embryonal Tumors, Pediatric (Brain Cancer), Germ Cell Tumors, Pediatric (Brain Cancer), Primary CNS Lymphoma, Cervical Cancer, Pediatric Cancer, Pediatric Cancer, Rare Bile Duct Carcinoma, Chordoma, Pediatric (Bone Cancer), Chronic Lymphocytic Leukemia (CLL), Chronic Myeloid Leukemia (CML), Chronic Myeloproliferative Neoplasms, Colorectal Cancer, Craniopharyngioma, Pediatric (Brain Cancer), Skin T Cell lymphoma, ductal carcinoma in situ (DCIS), pediatric (brain cancer), endometrial cancer (uterine cancer), ependymoma, pediatric (brain cancer), esophageal cancer, nasal neuroblastoma (head and neck cancer), Ewing's sarcoma (bone cancer), extracranial germ cell tumor, pediatric, extragonadal germ cell tumor, eye cancer, intraocular melanoma, retinoblastoma, fallopian tube cancer, fibrous histiocytoma of bone, malignant and osteosarcoma, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumor, gastrointestinal stromal tumor (GIST) (soft tissue sarcoma), germ cell tumor, pediatric central nervous system germ cell tumor (brain cancer), pediatric head and neck Extracerebrospinal germ cell tumors, extragonadal germ cell tumors, ovarian germ cell tumors, testicular cancer, gestational trophoblastic disease, hairy cell leukemia, head and neck cancer, cardiac tumors, children, hepatocellular (liver) cancer, histiocytosis, Langerhans cell, Hodgkin's lymphoma, hypopharyngeal cancer (head and neck cancer), intraocular melanoma, pancreatic islet cell tumors, pancreatic neuroendocrine tumors, Kaposi's sarcoma (soft tissue sarcoma), kidney (renal cell) cancer, Langerhans cell histiocytosis, laryngeal cancer (head and neck cancer), leukemia, oral cavity and lip cancer (head and neck cancer), liver cancer, lung cancer (non-small cell, small cell, pleuropulmonary blastoma, and tracheobronchial tumors), lymphoma, male breast cancer,Malignant fibrous histiocytoma and osteosarcoma of bone, melanoma, melanoma, intraocular (eye), Merkel cell carcinoma (skin cancer), mesothelioma, malignant, metastatic cancer, metastatic squamous cell carcinoma of occult primary (head and neck cancer), midline duct carcinoma with NUT gene alterations, oral cancer (head and neck cancer), multiple endocrine neoplasia syndrome, multiple myeloma / plasma cell neoplasm, mycosis fungoides (lymphoma), myelodysplastic syndrome, myelodysplastic / myeloproliferative neoplasm, chronic myeloid leukemia (CML), acute myeloid leukemia (AML), chronic myeloproliferative neoplasm, cancer of the nasal cavity and paranasal sinuses (head and neck cancer) Neck cancer), nasopharyngeal cancer (head and neck cancer), neuroblastoma, non-Hodgkin's lymphoma, non-small cell lung cancer, oral cavity cancer, cancer of the lip and oral cavity and oropharyngeal cancer (head and neck cancer), osteosarcoma and malignant fibrous histiocytoma of bone, ovarian cancer, pancreatic cancer, pancreatic neuroendocrine tumors (pancreatic islet cell tumors), papillomatosis (pediatric larynx), paraganglioma, cancer of the paranasal sinuses and nasal cavity (head and neck cancer), parathyroid cancer, penile cancer, pharyngeal cancer (head and neck cancer), pheochromocytoma, pituitary tumor, plasma cell neoplasm / multiple myeloma, pleuropulmonary blastoma (lung cancer), breast cancer during pregnancy, primary central nervous system Central nervous system (CNS) lymphoma, primary peritoneal cancer, prostate cancer, rectal cancer, recurrent cancer, renal cell (kidney) cancer, retinoblastoma, rhabdomyosarcoma, childhood (soft tissue sarcoma), salivary gland cancer (head and neck cancer), childhood rhabdomyosarcoma (soft tissue sarcoma), childhood vascular tumor (soft tissue sarcoma), Ewing's sarcoma (bone cancer), Kaposi's sarcoma (soft tissue sarcoma), osteosarcoma (bone cancer), soft tissue sarcoma, uterine sarcoma, Sézary syndrome (lymphoma), skin cancer, small cell lung cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma of the skin, squamous cell neck cancer of unknown primary, metastatic (head and neck cancer), gastric Diagnoses include cancer (Stomach (Gastric) Cancer), T-cell lymphoma, lymphoma (mycosis fungoides and Sézary syndrome), testicular cancer, laryngeal cancer (head and neck cancer), nasopharyngeal cancer, oropharyngeal cancer, hypopharyngeal cancer, thymoma and thymic carcinoma, thyroid cancer, tracheobronchial tumor (lung cancer), transitional cell carcinoma of the renal pelvis and ureter (renal (renal cell) cancer), ureter and renal pelvis, transitional cell carcinoma of the kidney (renal (renal cell) cancer), urethra cancer, cervical cancer, endometrium, uterine sarcoma, vaginal cancer, vascular tumor (soft tissue sarcoma), or vulvar cancer.

[0247] In some embodiments, the diagnosis is brain non-glioma (epithelioma, hemangioblastoma, medulloblastoma, meningioma), breast (ductal, lobular), colon, endometrium (endometrium, serous endometrium, endometrial stromal sarcoma), gastroesophageal (esophageal adenocarcinoma, stomach), gastrointestinal stromal tumor, glioma (glioma, oligodendroglioma), head and neck adenocarcinoma, hematologic (acute lymphoblastic leukemia, acute myeloid leukemia, B-cell lymphoma, chronic lymphocytic leukemia, chronic myeloid leukemia, Rosai-Dorfman, T-cell lymphoma), hepatobiliary (cholangiocarcinoma, gallbladder, liver), lung adenocarcinoma, melanoma, mesothelioma tumors of the following organs may include: tumors, neuroendocrine (gastrointestinal neuroendocrine, high-grade neuroendocrine lung, low-grade neuroendocrine lung, pancreatic neuroendocrine, cutaneous neuroendocrine), ovary (ovarian clear cell, ovarian granulosa, ovarian serous), pancreas, prostate, kidney (chromophobe kidney, renal clear cell, renal papillary), sarcoma (chondrosarcoma, chordoma, Ewing's sarcoma, fibrous sarcoma, leiomyosarcoma, liposarcoma, osteosarcoma, rhabdomyosarcoma, synovial sarcoma, angiosarcoma), squamous (cervical, esophageal squamous, head and neck squamous, lung squamous, cutaneous squamous / basal), thymus, thyroid, or urothelium.

[0248] In some embodiments, the diagnosis may include one or more entries in the ICD-10-CM, or International Classification of Disease. The ICD provides a method for classifying diseases, injuries, and causes of death. The World Health Organization (WHO) published the ICD to standardize how diagnosed cases of disease, including cancer, are recorded and tracked. For example, a classification from any chapter of the ICD or Chapter 2, C, and D codes for cancer. C codes include neoplasms of the lips, oral cavity, and pharynx (C00-C14), neoplasms of the digestive tract (C15-C26), neoplasms of the respiratory system and intrathoracic organs (C30-C39), neoplasms of mesothelium and soft tissue (C45), neoplasms of bone, joints, and articular cartilage (C40-C41), neoplasms of the skin (melanoma, Merkel cell, and other cutaneous histology) (C43, C44, C4a), Kaposi's sarcoma (C46), neoplasms of the peripheral and autonomic nervous system, retroperitoneum, peritoneum, and soft tissue (C47, C48, C49), neoplasms of the breast and female genitalia (C50-C58), neoplasms of the male genitalia (C60-C63). ), neoplasms of the urinary tract (C64-C68), neoplasms of the eye, brain, and other parts of the central nervous system (C69-C72), neoplasms of the thyroid, other endocrine glands, and unspecified local sites (C73-C76), malignant neuroendocrine tumors (C7a._), secondary neuroendocrine tumors (C7B), neoplasms of other unspecified local sites (C76-80), secondary and unspecified malignant neoplasms of lymph nodes (C77), secondary cancers of the respiratory and digestive tract, other unspecified sites (C78-80), malignant neoplasms of unspecified site (C80), lymphoma, or malignant neoplasms of hematopoietic and related tissues (C81-C96).

[0249] In some embodiments, the cancer status may include categorization into broadly construed cohort classes, which may include blood cancer, bone cancer, brain cancer, bladder cancer, breast cancer, rectal and colon cancer, endometrial cancer, kidney cancer, leukemia, liver cancer, lung cancer, melanoma, non-Hodgkin's lymphoma, pancreatic cancer, prostate cancer, thyroid cancer, or other tissue / organ-based classifications.

[0250] In some embodiments, the cancerous condition is cancer of the lips, base of the tongue, tongue (excluding base of the tongue), gums, floor of the mouth and other endocervical areas, salivary glands, oropharynx, nasopharynx (excluding posterior wall), posterior wall of the nasopharynx, hypopharynx, pharynx, esophagus, stomach, small intestine, large intestine (excluding appendix), appendix, rectum, anal canal and anus, liver, intrahepatic bile duct, gallbladder and extrahepatic bile duct, pancreas, unspecified digestive tract, nasal cavity (including nasal cartilage), middle ear, paranasal sinuses, sinus sinuses, nose, larynx, trachea, lungs and bronchi, thymus, heart, mediastinum, pleura, respiratory tract, bones and joints (skull and face, excluding mandible), bones of the skull and face, mandible, blood, bone marrow, and hematopoietic system, spleen, reticuloendothelium, skin, peripheral nerves, retroperitoneum and peritoneum, kidney, kidneys ... The biopsy site for the biopsy specimen may include one or more ICD-03 codes, including: complex and soft tissue, breast, vagina and labia, vulva, cervix, uterine corpus, uterus, ovaries, fallopian tubes, other female genitalia (excluding fallopian tubes), placenta, penis, prostate, testicles, epididymis, spermatic cord, male genitalia, scrotum, kidneys, renal pelvis, ureters, bladder, other urinary organs, orbit and lacrimal gland (excluding retina, eyes, nose), retina, eyeball, eye, nose, meninges (e.g., cerebrum and spine), brain and cranial nerves, and spinal cord (excluding ventricles, cerebellum), ventricles, cerebellum, other nervous system, thyroid, adrenal glands, parathyroid glands, pituitary gland, head and pharyngeal duct, pineal gland, other endocrine glands, unspecified lymph nodes, and unknown.

[0251] In some embodiments, diagnostic classifier 2382a may be trained with labels corresponding to one or more of the above diagnostic cancer classifications. The input to the model is a feature matrix having multiple patient feature vectors. For each model, the patient feature vector may include features from an increasing number of feature modules 2340, stored features in feature set 2305, variance modules 2350, or classification 2380. For each patient, the supervision signal may identify which of the diagnostic cancer classifications the patient feature vector is labeled with.

[0252] In some embodiments, cohort classifier 2383b may be trained with labels corresponding to one or more of the above cohort class cancer classifications. In some embodiments, the input to the model is a feature matrix having multiple patient feature vectors. For each model, the patient feature vector may include features from an increasing number of feature modules 2340, stored features in feature set 2305, variation modules 2350, or classification 2380. For each patient, the supervision signal may identify which of the cohort cancer classifications the patient feature vector is labeled with.

[0253] In some embodiments, tissue classifier 2382c may be trained with labels corresponding to one or more of the above site of biopsy class cancer classifications. In some embodiments, the input to the model is a feature matrix having multiple patient feature vectors. For each model, the patient feature vector may include one or more of features from feature module 2340, stored features in feature set 2305, variation module 2350, or classification 2380. For each patient, the supervised signal may identify which site of biopsy cancer classification the patient feature vector is labeled with.

[0254] In some embodiments, the stacked TUO classifier 2382d (also referred to as the final classifier) ​​may include one or more classifiers 2382a-n.

[0255] In some embodiments, a set of classifications may be trained and provided to the classifier. In other embodiments, multiple characteristic classifications may be available for classification by the classifier. In one example, characteristic classification may be performed between distinct tumor / tissue types sharing a common cell lineage: one or more sarcomas and one or more carcinomas, one or more squamous cell carcinomas and one or more carcinomas, or one or more neuroendocrine carcinomas and one or more carcinomas. In one example, differentiation may occur between lung adenocarcinoma, lung squamous cell carcinoma, oral adenocarcinoma, and oral adenocarcinoma. In one example, differentiation may occur between systemic sarcoma, epithelioma, Ewing's sarcoma, gliosarcoma, leiomyosarcoma, meningioma, mesothelioma, and Rosai-Dorfman. In addition to differentiation based on cell lineage, differentiation between metastatic sites and the site of origin may be performed when tumor tissue has spread widely but differentiation is insufficient. Examples may include distinguishing between liver metastases of pancreatic, upper gastrointestinal, or biliary origin, breast metastases of salivary gland, squamous, or ductal origin, brain metastases of glioblastoma, oligodendroglioma, astrocytoma, or medulloblastoma (including Wnt, Whh, Group 3, and Group 4), and lung metastases of NSCLC adenocarcinoma or squamous, between gynecological organs of the endometrium, ovary, or fallopian tube, and between endometrial, serous, and clear cell carcinoma. In one example, differentiation may be made between one or more sarcomas with morphological features or protein expression of a carcinoma and one or more carcinomas with morphological features or protein expression of a sarcoma.

[0256] In some embodiments, only a single RNA classifier may be implemented to generate a tissue classification for diagnosis classification 2382a, cohort classification 2382b, or TUO classification 2382d. In some embodiments, the input to the RNA classifier may include 20,000+ transcripts from whole-exome RNA sequencing, or a subset of transcripts (100, 500, 1000, 2000, 5000, etc.) may be selected based on correlation with outcome variables or supervised signals. In some embodiments, the RNA transcripts may be deconvolved or normalized. In some embodiments, two or more RNA classifiers, such as a combination of diagnosis classification 2382a, cohort classification 2382b, or tissue classification, may be combined to generate a diagnosis classification, cohort classification, and tissue classification 2382a-c based on RNA features 2341. In some embodiments, two or more classifiers based on one or more feature modules 2340 may be combined for TUO classification 2382d. For example, RNA features 2341 and DNA features 2342 may be received and combined to generate a tissue classification for a diagnostic classification 2382a, a cohort classification 2382b, or a TUO classification 2382d. Input to a DNA classifier may include genes, genes and their variants, as represented by protein (P-dot) notation. In some embodiments, the classifier may begin operation when input features are available to the system, and refined TUO classifications may be generated as each additional classification becomes available.

[0257] In some embodiments, RNA features 2341 may be normalized, such as by any of the methods disclosed in U.S. patent application Ser. No. 16 / 581,706, filed Sep. 24, 2019, entitled "Methods of Normalizing and Correcting RNA Expression Data," and deconvoluted, such as by any of the methods disclosed in U.S. patent application Ser. No. 16 / 732,229, filed Dec. 31, 2019, entitled "Transcriptome Deconvolution of Metastatic Tissue Samples," both of which are incorporated by reference herein in their entireties. Normalized and / or deconvoluted RNA features may be expressed as transcripts per million. In some embodiments, RNA features 2341 may be represented as expression, such as gene expression data quantified by Kallisto and quantile-normalized for GC content and length at the transcript level, followed by a library depth normalization step in which a scaling factor is calculated as the median ratio of sample expression to its geometric mean across all reference samples. After normalization, quality control using principal component analysis may be used to filter samples with aberrant expression. RNA features may be represented as a matrix of n patients by 19,147 genes, or feature selection may be performed, for example, by applying a variance threshold, where hyperparameter search may identify an optimal threshold of variance to optimize performance on the test data. In some embodiments, feature selection may reduce the number of transcripts required to train and apply classification models 2382a-n from 19,147 transcripts to approximately 7,000 transcripts. In some embodiments, feature selection methods may select the best 250 transcripts, 1,000 transcripts, or 10,000 transcripts given different selection criteria or hyperparameters. In some embodiments, methods for generating RNA signatures may include one or more of the methods of the '804 patent.

[0258] In some embodiments, the inhibitors of GPM6A, CDX1, SOX2, NAPSA, CDX2, MUC12, SLAMF7, HNF4A, ANXA10, TRPS1, GATA3, SLC34A2, NKX2-1, SLC22A31, ATP10B, STEAP2, CLDN3, SPATA6, NRCAM, USH1C, SOX17, TMPRSS2, MECOM, WT1, CDHR1, HOXA13, SOX10, SALL1, CPE, NPR1, CLRN3, THSD4, ARL14, SFTPB, COL17A1, KLHL14, EPS8L3, NXPE4, FOXA2, SYT11, SPDEF, GRHL2, GBP6, PAX8, ANO1, KRT7, HOX A9, TYR, DCT, LYPDl, MSLN, TP63, CDH1, ESR1, HNF1B, HOXA10, TJP3, NRG3, TMC5, PRLR, GATA2, DCDC2, INS, NDUFA4L2, TBX5, ABCC3, FOLH1, HIST1H3G, S100A1, PTHLH, ACER2, RBBP8NL, TACSTD2, C19orf77, PTPRZ1, BHLHE41, FAM155A, MYCN, DDX3Y, FMN1, HIST1H3F, UPK3B, TRIM29, TXNDC5, BCAM, FAM83A, TCF21, MIA, RNF220, AFAPl, KRT5, SOX21, KANK2, GPM6B, Clorfl 16, FOXF1, MEIS1, EFHD1, or XKRX. In some embodiments, features from genes identified by the following Ensembl gene IDs may be used: ENSG00000150625(GPM6A), ENSG00000 113722(CDX1), ENSG00000181449(SOX2), ENSG00000131400(NAPSA), ENSG00000 165556(CDX2), ENSG00000205277(MUC12), ENSG00000026751(SLAMF7), ENSG00000 101076(HNF4A), ENSG00000 10951 1(ANXA10), ENSG00000 104447(TRPS1), ENSG00000 107485(GATA3), ENSG00000157765(SLC34A2)、ENSG00000136352(NKX2-1)、ENSG00000259803(SLC22A31)、ENSG000001 18322(ATPlOB)、ENSG00000157214(STEAP2)、ENSG00000165215(CLDN3)、ENSG00000132122(SPATA6)、ENSG00000091 129(NRCAM)、ENSG0000000661 1(USH1C)、ENSG00000164736(SOX17)、ENSG00000184012(TMPRSS2)、ENSG00000085276(MECOM)、ENSG00000184937(WT1)、ENSG00000148600(CDHR1)、ENSG00000 106031(HOXA13)、ENSG00000 100146(SOXIO)、ENSG00000 103449(SALLl)、ENSG00000109472(CPE)、ENSG00000169418(NPRl)、ENSG00000 180745(CLRN3)、ENSG00000 187720(THSD4)、ENSG00000 179674(ARL14)、ENSG00000 168878(SFTPB)、ENSG00000065618(COL17A1)、ENSG00000 197705(KLHL14)、ENSG00000198758(EPS8L3)、ENSG00000137634(NXPE4)、ENSG00000 125798(FOXA2)、ENSG00000132718(SYT11)、ENSG00000 124664(SPDEF)、ENSG00000083307(GRHL2)、ENSG00000183347(GBP6)、ENSG00000125618(PAX8)、ENSG00000131620(ANOl)、ENSG00000135480(KRT7)、ENSG00000078399(HOXA9)、ENSG00000077498(TYR)、ENSG00000080166(DCT)、ENSG00000150551(LYPDl)、ENSG00000102854(MSLN)、ENSG00000073282(TP63)、ENSG00000039068(CDH1)、ENSG00000091831(ESR1)、ENSG00000 108753(HNF1B)、ENSG00000253293(HOXAIO)、ENSG00000 105289(TJP3)、ENSG00000185737(NRG3)、ENSG00000103534(TMC5)、ENSG00000 113494(PRLR)、ENSG00000 179348(GATA2)、ENSG00000146038(DCDC2)、ENSG00000254647(INS)、ENSG00000185633(NDUFA4L2)、ENSG00000089225(TBX5)、ENSG00000 108846(ABCC3)、ENSG00000086205(FOLH1)、ENSG00000256018(HIST1H3G)、ENSG00000 160678(S100A1)、ENSG00000087494(PTHLH)、ENSG00000 177076(ACER2)、ENSG00000130701(RBBP8NL)、ENSG00000 184292(TACSTD2)、ENSG00000095932(C19orf77)、ENSG00000 106278(PTPRZ1)、ENSG00000123095(BHLHE41)、ENSG00000204442(F AMI 55A)、ENSG00000 134323(MYCN)、ENSG00000067048(DDX3Y)、ENSG00000248905(FMN1)、ENSG00000256316(HIST1H3F)、ENSG00000243566(UPK3B)、ENSG00000137699(TRIM29)、ENSG00000239264(TXNDC5)、ENSG00000 187244(BCAM)、ENSG00000147689(FAM83A)、ENSG00000 118526(TCF21)、ENSG00000261857(MIA)、ENSG00000187147(RNF220)、ENSG00000 196526(AFAPl)、ENSG00000 186081(KRT5)、ENSG00000125285(SOX21)、ENSG00000197256(KANK2), ENSG00000046653(GPM6B), ENSG00000 182795(Clorfl 16), ENSG00000 103241(FOXF1), ENSG00000143995(MEIS1), ENSG000001 15468(EFHD1), and ENSG00000 182489(XKRX).

[0259] The transcript isoform information associated with these genes can be selected as input features of the sequencing results. For example, for each gene, GPM6A has transcript isoforms GPM6A-201, GPM6A-202, GPM6A-203, GPM6A-204, GPM6A-205, GPM6A-206, GPM6A-207, GPM6A-208, GPM6A-209, GPM6A-210, GPM6A-211, GPM6A-212, GPM6A-213, GPM6A-214, GPM6A-215, GPM6A-216, GPM6A-217, GPM6A-218, GPM6A-219, GPM6A-220, GPM6A-221, GPM6A-222, GPM6A-223, GPM6A-224, GPM6A-225, GPM6A-226, GPM6A-227, GPM6A-228, GPM6A-229, GPM6A-230, GPM6A-231, GPM6A-232, GPM6A-233, GPM6A-234, GPM6A-235, GPM6A-236, GPM6A-237, GPM6A-238, GPM6A-239, GPM6A-240, GPM6A-241, GPM6A-242, GPM6A-243, GPM6A-244, GPM6A-245, GPM6A-246, GPM6A-247, GPM6A-248, GPM6A-249, GPM6A-250, GPM6A-251, GPM6A-252, GPM6A- 1, GPM6A-212, GPM6A-213, GPM6A-214, GPM6A-215, GPM6A-216, GPM6A-217, GPM6A-218, GPM6A-219, GPM6A-220, GPM6A-221; CDX1 may be associated with transcript isoforms CDXl-201 and CDXl-202; SOX2 may be associated with transcript isoform SOX2-201; NAPSA may be associated with transcript isoforms NAPSA-201, NAPSA-202, NAPSA-203, NAPSA-204, NAPSA-205, NAPSA-206, and NAPSA-207; In some embodiments, transcripts may be selected at the transcript level, so that for each gene, only the feature-selected transcript is included, instead of each gene with all of its transcripts. A "transcript" of a gene is an mRNA molecule associated with the gene.

[0260] In some embodiments, RNA splicing features 2349a can be generated from RNA alternative splicing, such as alternative splicing scores. Alternative splicing scores may be calculated for 1,500 common exon skipping events in the human genome. In some embodiments, spliced ​​transcript alignment to reference (STAR), a rapid RNA-Seq read mapper with support for splice junctions, and fusion read detection can be applied. STAR aligns reads by finding the Maximally Mappable Prefix (MMP) between a read (or read pair) and the genome using a Suffix Array index. Different portions of a read can be mapped to different genomic locations corresponding to splicing or RNA fusions. The genome index includes known splice junctions from annotated gene models, enabling sensitive detection of spliced ​​reads. STAR performs local alignments and automatically soft-clipping the ends of reads with high mismatches. STAR or similar splicing identifiers can be used to generate splice junction indexes for each RNA sample.The splice junction index can then be normalized to calculate the percentage of spliced ​​in common alternative splicing events (PSI) score, and can be represented as a matrix for n patients, with -5000 alternatively spliced ​​transcripts.Transcript splicing can be detected in any of the RNA transcripts associated with each gene.In some embodiments, the method for generating RNA splicing features can include one or more of the methods of the '804 patent.

[0261] In some embodiments, copy number variation (CNV) 2349b (e.g., copy number features) can be generated from raw sequencing read data corresponding to each probe in a sequencing assay. DNA CNV, or copy number data, can be generated using a bioinformatics pipeline that identifies structural variants from DNA sequencing by comparing the read depth of a sample to a pool of normal samples. The raw sequencing data can first be normalized to account for variance introduced through different sequencing methods, bioinformatics procedures, or other bias-introducing factors. Normalization can include depth normalization to the normal pool, GC correction across GC percentiles for all target regions, principal component noise correction to the normal pool, log ratio calculations for both the normal pool and matched normal samples, and / or site band level imputation to account for probe-target mismatch. After normalization, the sequencing data can be used to identify copy number data expressed as the average log odds ratio for each probe within a site band. The log odds ratio is log((observed reads) / (expected reads)). In some embodiments, methods for generating CNV features may include one or more of the methods of the '804 patent. CNVs may be generated using a fixed-width sliding window and mapped to chromosomes, genes, variants, or cytobands for each sequenced sample, and may be represented as a matrix of n patients by 550 cytobands, or alternatively, a matrix of n patients by approximately 600 genes. In some embodiments, other sequencing panels include different numbers of genes, e.g., 100 genes, 300 genes, 1000 genes, or 20,000 genes.

[0262] In some embodiments, the cytobands for each gene include: 1Op11.1, 1Op11.21, 1Op13, 10p14, 1Op15.1, 10p15.3, 1Oq11.21, 1Oq11.23, 10q21.2, 10q22.1, 10q22.3, 10q23.2, 10q23.31, 10q23.33, 10q24.2, 10q24.31, 10q24.32, 10q24.33, 10q25.2, 10q25.3, 10q26.1, 10q26.13, 10q26.2, 10q26.3, 1lpl l.2,1lpl3,1lpM.l,1lpl4.3,1lpl5.1,1lpl5.2,1lpl5.4,I1p15.5,1lql2.1,1lpl2.2,1lpl2.3,1lpl3.1,1lpl3.2,1lpl3.3,1lpl3.4,1lpl3.5,1lqM.l,I Iq2 1, 1lq22.2, 1lq22.3, 1lq23.1, 1lq23.2, 1l1q23.3, 1lq24.1, 1lq24.2, 1lq24.3, 1lq25, 12p 11.21, 12pl2.1, 12pl3.1, 12pl3.2, 12pl3.31, 12pl3.32, 12pl3.33, 12ql2, 12ql3.12, 12ql3 .13, 12ql3.2, 12ql3.3, 12ql4.1, 12ql4.3, 12ql5, 12q21.31, 12q21.33, 12q23.1, 12q23.2, 12q23.3, 12q24.12, 12q24.13, 12q24.21, 12q24.31, 12q24.33, 13ql2.1 1, 13ql2.13, 13ql2.2, 13ql2.3, 13ql3.1, 13ql3.3, 13ql4.1 1, 13ql4.2, 13ql4.3, 13q21.1, 13q22.1, 13q31.1, 13q32.1, 13q33.1, 13q34, 14ql 1.2, 14ql2, 14ql3.2, 14ql3.3, 14q21.1, 14q21.2, 14q22.1, 14q23.2, 14q23.3, 14q24.1 , 14q24.3, 14q31.1, 14q32.12, 14q32.13, 14q32.2, 14q32.31, 14q32.32, 14q32.33, 15ql 1.2, 15ql3.3, 15ql4, 15ql5.1, 15q21.1, 15q21.2, 15q22.2, 15q22.31, 15q22.33、15q24.1、15q24.3、15q25.1、15q25.3、15q26.1、15q26.3、16pl 1.2、16pl2.1、16pl2.2、16pl3.1 1、16p 13.12、16pl3.13、16pl3.2、16pl3.3、16ql2.1、16q21、16q22.1、16q 22.2、16q22.3、16q23.1、16q23.2、16q23.3、16q24.1、16q24.3、17pl 1.2、17pl2、17pl3.1、17pl3.2、17pl3.3、17ql 1.2、17ql2、17q21.1、17q21.2、17q21.31、17q21.32、17q21.33、17q22、17q2 3.1、17q23.2、17q23.3、17q24.1、17q24.2、17q24.3、17q25.1、17q25.3、18pl 1.21、18pl 1.32、18ql 1.2, 18ql2.3, 18q21.1, 18q21.2, 18q21.32, 18q21.33, 18q22.3, 18q23, 19pl3.1 1, 19pl3.12, 19pl3.2, 19pl3.3, 19ql2, 19ql3.11, 19ql3.12, 19ql3.2, 19ql3.31, 19ql3.32, 19ql3.33, 19ql3.41, 19ql3.42, 19ql3.43, lpl l.2, 1p12, 1p 13.1, lpl3.2, lpl3.3, lp21.3, lp22.1, lp22.2, lp22.3, lp31.1, lp31.3, lp32.1, lp32.3, lp33, lp34.1, lp34.2, lp34.3, lp35.1, lp36.1 1, lp36.12, lp36.13, lp36.21, lp36.22, lp36.23, lp36.31, lp36.32, lp36.33, lq21.1, lq21.2, lq21.3, lq22, lq23.1, lq23.3, lq24.2, lq24.3, lq25.2, lq31.2, lq32.1, lq32.3, lq41, lq42.12, lq42.13, lq42.2, lq43, lq44, 20pl l.21, 20pl 1.22, 20pl 1.23, 20pl2.1, 20pl3, 20ql l.21, 20ql l.23, 20ql2, 20ql3.12, 20ql3.13, 20ql3.2, 20ql3.32, 20ql3.33, 21ql 1.2, 21q21.1, 21q21.3, 21q22.1 1, 21q22.12, 21q22.2, 21q22.3, 22ql 1.21, 22ql l.22, 22ql l.23, 22ql2.1, 22ql2.2, 22ql2.3, 22ql3.1, 22ql3.2, 22ql3.31, 22ql3.33, 2pl l.2、2pl3.1、2pl3.2、2pl3.3、2pl5、2pl6.1、2pl6.3、2p21、2p22.2、2 p23.1、2p23.2、2p23.3、2p24.1、2p24.2、2p24.3、2p25.1、2p25.3、2ql 11.1、2q1 l.2, 2ql2.2, 2ql2.3, 2ql3, 2ql4.2, 2ql4.3, 2q22.1, 2q22.2, 2q22.3, 2q23.3, 2q24.2, 2q31.1, 2q31.2, 2q31.3, 2q32.2, 2q32.3, 2q33.1, 2q33.2, 2q34, 2q35, 2q36.1, 2q36.3, 2q37.1, 2q37.3, 3pl 1.1, 3pl2.1, 3pl3, 3pl4.1, 3pl4.2, 3pl4.3, 3p2 1.1, 3p21.2, 3p21.31, 3p22.1, 3p22.2, 3p24.1, 3p25.1, 3p25.2, 3p25.3, 3p26.1, 3p26.3, 3ql 1.1, 3ql3.1 1, 3ql3.2, 3ql3.31, 3q21.1, 3q21.2, 3q21.3, 3q22.1, 3q22.2, 3q22.3, 3q23, 3q26.1, 3q26.2, 3q26.32, 3q26.33, 3q27.1, 3q27.2, 3q27.3, 3q28, 3q29, 4pl l, 4pl3, 4pl4, 4p 15.31, 4pl5.33, 4pl6.1, 4pl6.3, 4ql 1、4ql2、4ql3.2、4ql3.3、4q21.21、4q21.22、4q21.23、 4q21.3、4q24、4q25、4q27、4q28.1、4q31.1、4q31.21、4q 31.3,4q32.1,4q32.3,4q34.3,4q35.1,4q35.2,5pl2,5pl3.1 1.1、5ql 1.2, 5ql2.3, 5ql3.1, 5ql3.2, 5ql4.1, 5ql4.2, 5ql4.3, 5ql5, 5q22.2, 5q23.2, 5q23.3, 5q31.1, 5q31.2, 5q31.3, 5q32, 5q33.1, 5q33.3, 5q34, 5q35.1, 5q35.2, 5q35.3, 6pl l.2, 6p21.1, 6p21.2, 6p21.31, 6p21.32, 6p21.33, 6p22.2, 6p22.3, 6p24.1, 6p25.3, 6ql 1.1, 6ql3, 6ql5, 6ql6.1, 6ql6.2, 6q21, 6q22.1, 6q22.31, 6q22.33, 6q23.2, 6q23.3, 6q24.1, 6q24.2, 6q25.1, 6q25.3, 6q26, 6q27, 7pl l.2, 7pl2.2, 7pl4.1, 7pl4.3, 7p 15.1, 7pl5.2, 7p21.1, 7p21.2, 7p22.1, 7p22.2, 7p22.3, 7ql 1.21, 7q21.11, 7q21.12, 7q21.2, 7q21.3, 7q22.1, 7q22.3, 7q31.1, 7q31.2, 7q31.31, 7q31.33, 7q32.1, 7q34, 7q36.1, 7q36.3, 8pl 1.21, 8pl l.22, 8pl 1.23, 8pl2, 8p21.2, 8p21.3, 8p22, 8p23.1, 8p23.3, 8ql 1.21, 8ql 1.23, 8ql2.1, 8ql3.1, 8ql3.2, 8ql3.3, 8q21.1 1, 8q21.12, 8q21.3, 8q22.2, 8q22.3, 8q23.1, 8q24.1 1, 8q24.13, 8q24.21, 8q24.22, 8q24.3, 9p 13.1 9pl3.2, 9pl3.3, 9p21.1, 9p21.3, 9p24.1, 9p24.3, 9q21.11, 9q21.2, 9q21.32, 9q21.33, 9q22.1, 9q22.2, 9q22.32, 9q22.33, 9q31.2, 9q32, 9q33.1, 9q33.2, 9q33.3, 9q34.11, 9q34.12, 9q34.13, 9q34.2, 9q34.3, Xpl l.21, Xpl l.22, Xpl l.23、Xpl l.3、Xpl l.4、Xp21.2、Xp21.3、Xp22.2、Xp22.33、Xql l.2、Xql2、Xql3.1、Xql3.2、Xq21.1、Xq22.1、Xq22.3"1, and Xq28.

[0263] In some embodiments, germline / somatic DNA features 2342 may be represented as genes or variants that are detected or not detected in a sample. DNA features, such as DNA variants, may be detected using one or more variant callers, e.g., FreeBay and Pindel. Ensemble methods may enable improved variant detection. In some embodiments, tumor specimen sequencing results may be evaluated for variants to identify variants present in the sample. In some embodiments, tumor specimens may be sequenced alongside and compared to normal specimens from the same patient to identify somatic and germline alterations. For example, if a patient has a variant in both the tumor and normal specimens, that variant may be removed from further evaluation because it is unlikely to contribute to tumor growth. In some embodiments, a variant reference set or a database of all variant classifications may be used to annotate the pathogenicity of each change detected in the specimen sequencing results. In some embodiments, changes may be represented either as a gene and amino acid change (i.e., KRAS G12V) or as a gene and functional effect (BRAF loss of function). In some embodiments, pathogenic changes may be one-hot encoded for representation during modeling. Selecting gene and amino acid change representations instead of nucleotide change representations may improve performance by reducing the number of variants, since some nucleotide representations are semantically identical changes. In some embodiments, methods for generating DNA features may include one or more of the methods of the '804 patent. DNA features may be represented as a matrix of patient n names with 20,000 variants. In some embodiments, feature selection may reduce the number of variants required to train and apply classification models 2382a-n from 20,000 variants to approximately 7,000 variants. In some embodiments, feature selection methods may select the best 250, 1,000, or 10,000 variants given different selection criteria or hyperparameters.An exemplary feature-selected gene list is provided above for the RNA feature set.

[0264] In some embodiments, the viral genomic signature 2343a can be expressed as the presence or absence of a virus in the specimen result. In some embodiments, the viral genomic signature is determined as described in U.S. Application No. 62 / 978,067, filed February 18, 2020, entitled "Systems and Methods for Detecting Viral DNA from Sequencing," which is incorporated herein by reference in its entirety. The sequencing results may be matched to a human reference genome, which leaves some portions of the sequencing results unmatched. In some examples, the unmatched portions may be compared to a bacterial or viral reference genome. The match identifies the presence of bacteria or viruses in the specimen, which may affect the specific cancer status for each patient specimen. In some embodiments, identifying bacteria may include detecting Salmonella typhi, Streptococcus bovis, Chlamydia pneumoniae, Mycoplasma, or Helicobacter pylori, and identifying the presence of viruses may include detecting Hepatitis B (HBV), Hepatitis C (HCV), Human T-lymphotropic Virus (HTLV), Human Papillomavirus (HPV), Kaposi's Sarcoma-Associated Herpesvirus (HHV-8), Merkel Cell Polyomavirus (MCV), or Epstein-Barr Virus (EBV). A viral genomic signature may be represented as a matrix of n patients and 13 bacteria and viruses, x bacteria, or y viruses (e.g., 3 viruses). In some embodiments, a method for generating a viral signature may include one or more of the methods of the '804 patent.

[0265] In some embodiments, the feature module 2340 may include other features, such as clinical features. Clinical features may include patient information from the patient's electronic health record, test results, diagnosis, and treatment. In one example, clinical information may include the patient's history of breast cancer diagnosis and notes of subsequent remission. A classifier trained to identify a diagnostic cancer status may first identify a diagnosis using one or more of the features introduced above as RNA features, DNA features, RNA fusions, viral features, or copy number features, and as a second step, may further identify whether the patient had a previous diagnosis with reference to clinical features. A patient with a previous diagnosis of breast cancer in remission and a cancer status classification associated with breast cancer may further be identified as having a recurrence of breast cancer noted in the cancer status.

[0266] In some embodiments, any feature in feature module 2340 may be provided to one or more classifiers 2382a-n to generate a TUO classification. Feature combinations may include DNA features only, RNA features only, a combination of DNA and RNA features, any combination of DNA features and other features, any combination of RNA features and other features, any combination of DNA and RNA features with other features, including combinations of RNA features, DNA features, splicing features, CNV features, and viral / bacterial genomic features. It should be understood that one or more combinations of models may be trained and selected for each new patient based on the combination of features available and associated with the patient. For example, a patient sequenced only for DNA may have DNA features, CNV features, and viral features, but no RNA features. One or more models may receive DNA features, CNV features, and viral features. In some embodiments, the TUO classifier may receive predicted outputs from one or more classifiers 2382a-n and combine them to generate a TUO classification to identify a diagnosis of a cancerous condition for the patient. In some embodiments, each of the classifiers may be a diagnostic classifier 2382a that uses linear regression on the RNA feature set, the DNA feature set, the splicing feature set, the CNV feature set, and the viral feature set. In some embodiments, each of the classifiers may be a cohort / subtype classifier 2382b that uses linear regression on the RNA feature set, the DNA feature set, the splicing feature set, the CNV feature set, and the viral feature set. In some embodiments, each of the classifiers may be a tissue classifier 2382c that uses linear regression on the RNA feature set, the DNA feature set, the splicing feature set, the CNV feature set, and the viral feature set.In some embodiments, each of the classifiers may be one or more, two or more, or three or more of the diagnostic classifier 2382a, the cohort classifier 2382b, and the tissue classifier 2382c, which use linear regression on the RNA feature set, the DNA feature set, the splicing feature set, the CNV feature set, and the viral feature set. In some embodiments, a boosting algorithm may be used to improve the classifiers for the RNA feature set, the DNA feature set, the splicing feature set, the CNV feature set, and the viral feature set. The boosting algorithm may identify a subset of genes that produces better similarity between subjects given the labels selected for boosting. Classifiers may be boosted for RNA labels, imaging labels, DNA labels, clinical information labels, tumor grade labels, tumor staging labels, or other labels in the dataset provided for classification. In one embodiment, boosting may be implemented as described in Skurichina, M., Duin, R. Bagging, Boosting and the Random Subspace Method for Linear Classifiers. Pattern Anal Appl 5, 121-135 (2002), which is incorporated herein by reference in its entirety.

[0267] In some embodiments, classifiers 2382a-n generate a classification of diagnoses, cohorts, or tissues from the feature sets received as input at each classifier. In some embodiments, classification may also be referred to as a prediction, based on the received features. Sub-models, which are models that provide predictions to meta-classifier 2382d, may be viewed through a viewer 2500, such as a web page, application, or other display device capable of displaying graphs. In some embodiments, the web page may be accessed from a web address. A user, such as a physician, may access a patient's TUO classification results by the patient's unique, anonymized ID or by identifying information such as the patient's name or medical record number. In some embodiments, the unique ID may include a combination of letters and numbers, such as "20form," as shown in FIGS. 25 and 26. In some embodiments, viewer 2500 may include one or more graphs 2510, 2520, 2530, and 2540 corresponding to the classifiers and feature sets.

[0268] In some embodiments, the sub-predicted CNVs illustrated in graph 2510 visually depict the predicted / classified results of a specimen's sequencing results using copy number analysis. The top 10 classification results are arranged in ranked order. The Y-axis identifies the cancer state and the X-axis identifies the likelihood of the presence of the cancer state generated in the sequencing results. For the specimen associated with ID "20form," a lung adenocarcinoma classification is predicted to be approximately 48% likely, a pancreatic classification is predicted to be approximately 18% likely, and a biliary classification is predicted to be approximately 17% likely.

[0269] In some embodiments, the sub-predicted RNA illustrated in graph 2520 visually depicts the prediction / classification results of the sequencing results of a specimen using RNA transcripts. The top 10 classification results are arranged in ranked order. However, only two results are associated with the likelihood of the presence of a cancerous condition generated in this sequencing result. For the specimen associated with ID "20form," a classification of lung adenocarcinoma is predicted with approximately an 88% likelihood, and a classification of lung squamous cell carcinoma is predicted with approximately a 6% likelihood.

[0270] In some embodiments, the sub-prediction DNA illustrated in graph 2530 visually depicts the prediction / classification results of the sequencing results of the specimen using DNA variants. The top 10 classification results are arranged in ranked order. For the specimen associated with ID "20form", a classification of lung adenocarcinoma is predicted to be approximately 65% ​​likely, a classification of pancreatic is predicted to be approximately 20% likely, and a classification of biliary tract is predicted to be approximately 3% likely.

[0271] In some embodiments, the sub-predicted RNA splicing graph 2540 illustrates a visual depiction of the predicted / classified results of a specimen's sequencing results using RNA splicing analysis. The top 10 classification results are listed in ranked order. For the specimen associated with ID "20form," a classification of acute lymphoblastic leukemia is predicted to be approximately 0.025% likely, a classification of acute myeloid leukemia is predicted to be approximately 0.019% likely, and a classification of B-cell lymphoma is predicted to be approximately 0.017% likely.

[0272] In some embodiments, physician review of each sub-prediction outcome may allow additional insight into the drivers of TUO meta-classifier classification.

[0273] In some embodiments, meta-classifier 2382d may combine results from one or more classifiers 2382a-n to generate a TUO classification that can be viewed through viewer 2600, such as a web page, application, or other display device capable of displaying graphs. In some embodiments, the web page may be accessed from a web address. A user, such as a physician, may access a patient's TUO classification results by the patient's unique, anonymized ID. Viewer 2600 may include one or more graphs 2610, 2620, and 2630 for displaying the combined results of graphs 2510, 2520, 2530, and 2540 corresponding to classifiers and Feature Sets for CNVs, DNA variants, RNA transcripts, and RNA splicing.

[0274] In some embodiments, the rollup predictions of CNV, DNA variant, RNA transcript, and RNA splicing illustrated in graph 2610 visually depict the sum of sub-prediction probabilities for each cancer classification cohort. The highest sum result is listed on the Y-axis, and the cumulative probability is represented along the X-axis. For the specimen associated with ID "20form," the classification of lung across all sub-prediction classifiers is approximately 97% likely, while the next closest classifications of neuroendocrine and squamous are approximately less than 2% likely, indicating that the meta-classifier is certain that the TUO originates from the lung.

[0275] In some embodiments, the rollup subtype predictions of CNVs, DNA variants, RNA transcripts, and RNA splicing illustrated in graph 2620 visually depict the sum of sub-prediction probabilities for each cancer classification diagnosis. The highest sum result is listed on the Y-axis, and the cumulative probability is represented along the X-axis. For the specimen associated with ID "20form," the classification of lung adenocarcinoma across all sub-prediction classifiers is approximately 95% likely, while the next closest classifications for high-grade neuroendocrine lung cancer and lung squamous cell carcinoma are approximately less than 2% likely, indicating that the meta-classifier is confident that the TUO should be diagnosed as lung adenocarcinoma.

[0276] In some embodiments, selection of any bar graph in graph 2610 or any bar graph in graph 2620 automatically adds a Shapley Additive Explanation (SHAP) feature importance plot to graph 2630, visually showing how each sub-prediction from CNVs, DNA variants, RNA transcripts, and RNA splicing contributes to the rolled-up cancer classification diagnosis, cohort, or tissue classification (based on which bar graph is selected). For example, if a user selects the lung adenocarcinoma bar graph in graph 2620, the contribution probability from each sub-prediction from graphs 2510, 2520, 2530, and 2540 is mapped to graph 2630, where the Y-axis corresponds to the likelihood of how much the sub-prediction contributes to the total likelihood, and the X-axis corresponds to whether the likelihood increases or decreases the total likelihood.

[0277] In some embodiments, predicting multiple target variables in a stacked setting improves overall model performance by enabling the meta-classifier to understand the semantic associations between cohorts, tissues, and diagnoses. For example, a cancer status cohort RNA model may predict "sarcoma" with high confidence, while a cancer status diagnostic RNA model may be split between lung adenocarcinoma and osteosarcoma. In another example, the meta-classifier may favor osteosarcoma because additional cohort-level evidence is weighted in favor of sarcoma. In another example, a cancer status tissue model may predict colon tissue at the biopsy site, and the cancer status diagnostic RNA model may predict a diagnosis of colon cancer and liver cancer with fairly equal likelihood. In another example, the stacked model may underweight the likelihood of a colon cancer diagnosis given the "contamination" of the underlying colon tissue as identified by the biopsy site, and the meta-classifier favors a cancer status diagnosis of liver cancer.

[0278] In some embodiments, meta-classifier 2382d may receive classifications from nine separate classifiers: RNA diagnosis, RNA cohort, RNA splicing diagnosis, RNA splicing cohort, CNV diagnosis, CNV cohort, DNA variant diagnosis, DNA variant cohort, and viral diagnosis. In some embodiments, a heat map of feature importance is shown according to which features contribute to the different classifiers, providing additional clarification regarding the scaling of the meta-classifier's importance factor. While some features and their respective importance may be immediately discernible, many importance scores determined from performance may not be readily discernible.

[0279] The examples provided herein are illustrative and are not intended to limit the feature importance scaling factors to only the possible examples provided. The y-axis identifies the classifier, the x-axis identifies the cancer classification diagnosis, and the cells where the x-axis and y-axis meet are color-coded using the importance of the classifier to accurately predict the diagnosis from the classifier.

[0280] In some embodiments, the RNA classification of the diagnosis is weighted across the majority of cancer state diagnoses, as shown in the heatmap, the RNA classification of the cohort class is weighted across the majority of cancer state cohorts, and the RNA classification of the biopsy tissue or site is also weighted across the majority of cancer state sites in the biopsy. In some embodiments, because the majority of granulosa ovarian diagnoses include the presence of FOXL2 alterations, the DNA classification of the diagnosis is weighted only if the classification includes granulosa ovarian cancer. In some embodiments, the viral classification of the diagnosis is weighted for a small number of classes. For example, HPV affects anal squamous, head and neck squamous, and cervical cancers, and polyomavirus affects most Merkel cell carcinomas. The presence of these viral reads is a highly informative feature in some classes, but may provide less diagnostic value for samples that do not have tumors affected by the virus. Therefore, the significance score for any one diagnosis is naturally lower. In some embodiments, the diagnostic copy number classification is weighted for the diagnosis of glioma, prostate, ovarian serous, melanoma, oligodendroglioma and leiomyosarcoma. In some embodiments, the diagnostic RNA splicing classification is weighted when classifying prostate cancer or breast cancer diagnosis.

[0281] Additional illustrative examples: In one example, a patient visits a physician with concerns about breast pain. The physician identifies a breast lump and sends it for imaging, resulting in a CT scan or MRI, which identifies additional tumors in the patient's bone and liver. The physician biopsies the liver tumor and sends it for identification and sequencing. The pathologist is unable to confirm the tumor's liver origin and labels the specimen as a tumor of unknown origin (TUO). The sequencing laboratory sequences the patient's DNA and RNA from the tumor and DNA from the patient's blood. The physician receives notification that the specimen is a TUO and, in addition to the initial sequencing order, orders an additional TUO classification from the laboratory. In response to the TUO classification order, the laboratory provides the sequencing results and other results derived from the sequencing results to an artificial intelligence engine. The artificial intelligence engine identifies the liver tumor as originating from the breast based on a multimodal model combining RNA, DNA, CNV, and splicing modal classifiers. A report is generated identifying the TUO classification and supporting information provided by the classifier results as to why the classification is a reasonable prediction of tissue origin. Based on the identification of breast as the tissue origin for a liver tumor, a physician may select a line of treatment for the patient using FDA-approved drugs / therapy to target breast cancer tumors rather than simply platinum chemotherapy, which is offered to all patients with TUO.

[0282] In another example, a sequencing laboratory includes an ordering system that provides a comprehensive breakdown of sequencing assays, reports, classifications, and predictions that a physician can order. In some embodiments, a physician has one or more assay options to choose from: tumor-only or matched tumor-normal sequencing; reporting TMB, MSI, CNV, fusions, splicing, and other sequencing alterations; H&E and / or IHC staining; predicted IHC staining from H&E staining; predicted PD-L1 or other biomarker status from H&E staining; predicted metastasis to one or more organs; predicted origin for tumors of unknown origin; and other sequencing-related tests, predictions, or reported ordering items. Physicians can specify their preferred ordering by selecting one or more of the available options and paying the associated fees for each. After receiving somatic and / or germline specimens from a patient, performing sequencing according to the ordered assays, and completing all ordered items, the laboratory may generate and return to the physician a report summarizing the sequencing results and therapies, treatments, clinical trials, and other insights that may influence the physician's treatment choices for the patient. "Matched tumor-normal," "tumor-normal matched," and "tumor-normal sequencing" refer to processing genomic information from a subject's normal, non-cancerous germline sample (e.g., saliva, blood, urine, stool, hair, healthy tissue, or other collection of cells or fluids from a subject), and from the subject's tumor, somatic sample, e.g., a smear, biopsy, or other collection of cells or fluids from a subject containing tumor tissue, cells, or DNA (particularly circulating tumor DNA, ctDNA). DNA and RNA features identified from next-generation sequencing (NGS) of a subject's tumor or normal specimen can be cross-referenced to remove genomic mutations and / or variants that appear as part of the subject's germline from somatic analysis. The use of somatic and germline datasets leads to significantly improved mutation identification and reduced false positive rates. "Tumor-normal matched sequencing" provides more accurate variant calling due to improved germline mutation filtering.For example, generating somatic variant calls based at least in part on germline and somatic specimens can include identifying and removing common mutations. In this manner, variant calls from germline samples are removed from variant calls from somatic samples as non-driver mutations. Variant calls that occur in both germline and somatic specimens can be presumed to be normal for the patient and removed from further bioinformatics calculations. [Example]

[0283] Example 1 - Classification of Exemplary Patient Cohorts Using the methods described herein, we developed a classifier for a targeted oncology panel using hybrid capture next-generation sequencing. The classifier combines whole-transcriptome RNA-seq and targeted DNA tiling probes for comprehensive gene rearrangement and microsatellite instability (MSI) detection. In addition to the clinical testing capabilities of the classifier, the DNA-seq and RNA-seq assay components support tools for assessing tumor immune status, including HLA typing, neoantigen prediction, DNA repair gene analysis, MSI status, tumor mutation burden, and immune cell typing and expression.

[0284] Referring to Figures 4A-4C, a cohort of subjects was analyzed to investigate the effectiveness of using genome-wide expression patterns for cancer status classification. A cohort of 500 patients with tumors of known or unknown origin was examined. Analysis from tumor-normal matched sequencing of DNA mutational spectra across multiple cancer states, whole transcriptome profiling, genomic rearrangement detection, and immunogenicity landscape based on immunotherapy biomarkers in the patient cohort is described below.

[0285] Patients in the 500-patient cohort were randomly selected from a larger patient set. To be eligible for inclusion in the cohort, each patient was required to have complete data elements for tumor-normal matched sequencing and clinical data. After filtering for eligibility, the patient set was randomly sampled via a pseudorandom number generator. Based on pathological diagnosis, patients were divided into eight cancer conditions, including 50 patients per condition: brain, breast, colorectal, lung, ovarian, endometrial, pancreatic, and prostate. In addition, 50 tumors from a combined set of rare malignancies and 50 tumors of unknown origin were included in the cohort, for a combined total of 500 patients.

[0286] We first examined the mutational spectrum of a 500-patient cohort and compared it to the broad patterns of genomic alterations observed in larger studies across multiple cancer conditions. As shown in Figure 4A, we identified gene-by-gene genomic alterations for all 500 patients. Genomic alterations included single nucleotide polymorphisms (SNPs / indels), fusions (FUS), and a subset of copy number variants (CNY), amplifications (AMP), and deletions (DEL). The most commonly mutated genes were well-known driver mutations in solid tumors, including TP53, KRAS, PIK3CA, CDKN2A, PTEN, ARID 1A, APC, ERBB2 (HER2), EGFR, IDH1, and CDKN2B. Of these, CDKN2A, CDKN2B, and PTEN were found to be most commonly homozygous deletions, as expected for tumor suppressor genes. Alterations are grouped by type and those occurring in at least five patients (eg, at least 1% prevalence in the population) are plotted.

[0287] We next compared the mutational spectrum data shown in Figure 4A to a previously published pan-cancer analysis using the Memorial Sloan Kettering Cancer Center (MSKCC) IMPACT panel. See Zehir et al. 2017 Nat. Med. 23, 703-713. As shown in Figure 4B, both the IMPACT panel and the 500-patient cohort exhibited the same highly mutated genes at similar relative frequencies, indicating that the mutational spectrum of the 500-patient cohort is representative of a broad collection of tumors sequenced in previously published large-scale studies.

[0288] Each sample in the 500-patient cohort was further investigated by RNA-seq whole-transcriptome profiling. The trained classification model was then used to predict cancer status from each transcriptome. As shown in Figure 4C, this classification was particularly successful in predicting breast, prostate, brain, colorectal, pancreatic, and lung cancer. The bubbles in Figure 4C indicate the percentage of samples from each cohort type that were predicted to have a given TCGA cancer status. In some embodiments, the accuracy of each prediction is quantified using bootstrapping.

[0289] The Cancer Genome Atlas (TCGA) dataset referred to herein is a publicly available dataset containing over 2 petabytes of genomic data for over 11,000 cancer patients, including clinical information about the cancer patients, metadata about the samples collected from such patients (e.g., sample portion weight, etc.), histopathology slide images from the sample portions, and molecular information derived from the samples (e.g., mRNA / miRNA expression, protein expression, copy number, etc.). The TCGA dataset contains data on 33 different cancers: breast (ductal carcinoma, lobular carcinoma), central nervous system (glioblastoma multiforme, low-grade glioma), endocrine (adrenocortical carcinoma, papillary thyroid carcinoma, paraganglioma, and pheochromocytoma), gastrointestinal tract (bile duct carcinoma, colorectal adenocarcinoma, esophageal carcinoma, hepatocellular carcinoma of the liver, pancreatic ductal adenocarcinoma, and gastric carcinoma), gynecological system (cervical carcinoma, ovarian serous cystadenoma, uterine carcinosarcoma, and endometrial carcinoma), head and neck (head and neck squamous cell carcinoma, uveal melanoma), hematology (acute myeloid leukemia, thymoma), skin (cutaneous melanoma), soft tissue (sarcoma), chest (lung adenocarcinoma, lung squamous cell carcinoma, and mesothelioma), and urinary tract (chromophobe renal cell carcinoma, clear cell renal carcinoma, papillary renal carcinoma, prostate adenocarcinoma, testicular germ cell carcinoma, and urinary bladder carcinoma).

[0290] TCGA cancer conditions in Figure 4C include: ACC: adrenocortical carcinoma, BRCA: invasive breast carcinoma, COAD: colon adenocarcinoma, GBM: glioblastoma multiforme, HNSCC: head and neck squamous cell carcinoma, LGG: brain low-grade glioma, LIHC: liver hepatocellular carcinoma, LUAD: lung adenocarcinoma, LUSC: lung squamous cell carcinoma, MESO: mesothelioma, OV: ovarian serous cystadenoma, PAAD: pancreatic adenocarcinoma, PCPG: pheochromocytoma and paraganglioma, PRAD: prostate adenocarcinoma, SARC: sarcoma, SKCM: skin cutaneous melanoma, STAD: gastric adenocarcinoma, THYM: thymoma, UCEC: endometrial carcinoma, and UCS: uterine carcinosarcoma.

[0291] Example 2 – Matching Therapies to Clinical Trials We investigated the extent to which extensive molecular profiling can help match patients to therapy. Factors considered included consensus clinical guidelines for case reports of response and / or resistance to therapy. A knowledge database of therapy and prognostic evidence was compiled from sources including the National Comprehensive Cancer Network (NCCN), CIViC (see Griffith et al. 2017 Nat. Genet. 49, 170-174), and DGIdb (see Finan et al. 2017 Sci. Transl. Med. 9, eaagll66).

[0292] Clinically actionable entries in the knowledge database are structured by both associated disease and evidence level (e.g., tier). Binning of somatic evidence is performed as described by the ASCO / AMP / CAP working group. See Li et al. 2017 J. Mol. Diagnostics 19, 4-23, incorporated herein by reference. Patients are then matched to clinically actionable entries by gene, specific variant, diagnosis, and level of evidence.

[0293] Across all cancer conditions, 90.8% of patients were matched to a therapy option based on evidence of response to treatment, and 22.6% were matched based on evidence of resistance to treatment (see Figure 5A). As shown in Figure 5B, the maximum tier of matched therapy evidence varied significantly across cancer conditions. For example, 58.0% of colorectal cancer patients were matched to Tier IA evidence, the majority of which was for resistance to therapy based on detected KRAS mutations. On the other hand, no pancreatic cancer patients were matched to Tier IA evidence. This result was expected because, in contrast to pancreatic cancer, several molecular-based consensus guidelines exist for colorectal cancer.

[0294] We next determined the contribution of each molecular assay component to patient-therapy matching. First, we examined therapy evidence matches based on copy number variants (CNVs), single nucleotide variants (SNVs), and indels. Overall, 140 patients (28%) were matched to accurate medication options using Tier IA or Tier IB evidence, including FDA-approved, well-potent consensus therapies.

[0295] Next, we examined the contribution of therapy matching based on RNA-seq gene expression profiles of clinically relevant genes. Genes included in this analysis were selected based on their relevance to disease diagnosis, prognosis, and / or possible therapeutic intervention. In this example, up to 43 genes were evaluated for each cancer condition based on the sample's specific cancer status. To make expression calls, percentiles of expression for new patients were calculated relative to all cancer samples, all normal samples, matched cancer samples, and matched normal samples in the processed TCGA and GTEx databases. For example, tumor expression for breast cancer patients was compared to all cancer samples, all normal samples, all breast cancer samples, and all breast normal tissue samples in the reference database. The specific thresholds used for gene expression calling evolved over the course of the assay's use. Criteria specific to each gene and cancer condition at the time of reporting were used to determine gene expression calls. Therefore, the thresholds applied to specific genes may vary across this dataset.

[0296] Over- or under-expressed gene calls in 136 patients (28.7%) were examined for 16 genes using clinical trials, case reports, or preclinical evidence reported in the literature (results shown in Figure 5C). Metastatic cases were similarly more likely to have at least one reportable expression call compared with non-metastatic tumors. The most commonly reported gene was overexpression of NGR1, observed in 35 cases (7.3% of tumor samples) across the cohort.

[0297] Based on the immunotherapy biomarkers identified by this assay, we assessed the proportion of the cohort eligible for immunotherapy. As shown in Figure 5D, 52 patients (10.4%) would have been considered potential candidates for immunotherapy based on TMB, MSI status, and PD-L1 IHC results alone. The number of MSI-H and TMB-high cases was distributed across multiple cancer conditions, with 22 patients (4.4%) positive for both biomarkers. PD-L1-positive IHC alone was measured in 15 patients (3%), the highest among lung cancer patients. In 13 patients (2.6%), primarily lung and breast cancer cases, TMB-high status alone was measured. Finally, the combination of PD-L1-positive IHC and TMB-high status was observed in a minority of cases, measured in only 2 patients (0.4%).

[0298] Combining the above results, comprehensive molecular profiling was used to match therapy options for 455 patients (91%) (see Figure 5E). Additionally, 1,996 clinical trial matches were reported for the 500-patient cohort. At least one clinical trial was matched to 481 patients (96.2%). Of these patients, 77.2% were matched to at least one biomarker-based clinical trial for genetic variants in their final report.

[0299] As shown in Figure 5F, the frequency of biomarker-based clinical trial matching varied by diagnosis and exceeded that of disease-based clinical trial matching. For example, gynecological cancers and pancreatic cancers typically matched to biomarker-based clinical trials, while rare cancers had the fewest biomarker-based clinical trial matches, with the ratio of biomarker-based and disease-based trial matching being roughly equal. The difference between biomarker-based and disease-based trial matching is likely due to the frequency and heterogeneity of targetable changes in these cancer conditions.

[0300] The classification method described herein is unique in its use of matched tumor-normal DNA and whole-transcriptome RNA-seq to provide a comprehensive view of somatic genomic alterations, including MSI status, for targeted cancer therapy, immuno-oncology, and clinical trial enrollment. As described above, the method has been validated in multiple study formats.

[0301] Example 3 - Comparison of paired tumor / normal samples with tumor-only samples in cancer studies Cancer testing assays most commonly use only tumor samples.However, there is a potential advantage in using paired tumor samples and normal samples for diagnosis.In particular, this allows comparison of each patient's own germline mutation with the mutation in each patient's tumor (for example, the patient's somatic mutation).

[0302] For this example, 50 cases were randomly selected from a cohort of 500 patients with a range of tumor mutation burden (TMB) profiles (e.g., as shown in Figure 6A) and then re-evaluated using a tumor-only analysis pipeline. After filtering using a publicly available aggregate database, 8,557 coding variants were identified. Further filtering with an internally developed list of technical artifacts, an internal pool of normal samples, and classification criteria reduced the number of variants to 642, while retaining all true somatic alterations (72.3%). Among the 642 filtered tumor-only variants, 27.7% of these variants were classified as somatic false positives (e.g., actually germline variants or artifacts).

[0303] To assess the impact of tumor-only testing on therapy and to compare it with the therapy insights derived from the Tempus platform (e.g., from tumor-normal testing and RNA-seq and immuno-oncology (IO) analysis), we determined the therapy that would be proposed to each of these 50 patients in both scenarios. Eight of the 50 patients (16%) would have received a different clinical recommendation if they had received a tumor-only test instead of a full Tempus test. Of these eight patients, four had different recommendations due to information obtained via RNA-seq or due to tumors with low clonality somatic mutations, a feature that is difficult to detect with tumor-only testing. For example, in a prostate cancer patient, DNA-seq did not indicate any contraindications to the antiandrogen therapy the patient was receiving, but RNA-seq showed androgen receptor (AR) overexpression, indicating potential resistance. The other four patients had different therapies and potentially would not have received genetic counseling due to the tumor-only test reporting germline mutations as somatic.

[0304] Finally, we compared therapies recommended for all DNA variants detected by the Tempus platform with those recommended by the patient-facing website My Cancer Genome (MCG). While 43 cases received a recommended therapy via Tempus' tumor-only testing, therapy was found for only 5 cases via MCG.

[0305] In the aforementioned analyses, the use of tumor-normal matched sequencing in clinically reported results can filter variants and more accurately classify true somatic positives, leading to differential therapy recommendations for patients. This is illustrated in Figure 6B, where tumor-only analysis showed an increased number of false positives compared to the number of false positives generated by tumor / normal matched analysis.

[0306] Example 4 - Cancer Classification and Detection of Tumors of Unknown Origin Classification of cancer status is clinically essential for providing effective treatment to patients. Identifying the tissue of origin is an important preliminary step to determining optimal treatment strategies. In the test set of samples analyzed according to the methods described herein, 419 were tumors of unknown origin cohort. In addition, more than 41 samples had unknown gynecological origin, more than 166 samples had unknown gastrointestinal origin, and more than 241 samples were described as poorly differentiated. Collectively, these 867 samples comprise 7.6% of cancer patients in the entire sample set. The aforementioned methods are not sufficient to determine the cancer status of samples in these cohorts. See, e.g., Bloom et al. 2004 Am J Pathol., 164(1):9-16; Tschentscher et al. 2003 Can. Res., 63(10), 2578-84; Young et al. 2001 Am J Pathok, 158(5), 1639-51; and Amar et al. 2015 Nuc. Acids Res. 43, 7779-7789. Some of the methods disclosed herein use a combination of genomic, pathological, and clinical features, either to train a classification model or as input to a classification model, to enable classification of many tumors of unknown origin, thus providing patient information and improving patient outcomes (e.g., by enabling treatment appropriate to the identified cancer state).

[0307] 7A-7D show examples of model classification and demonstrate the accuracy of the methods disclosed herein. Diagnosis, cancer subtype, tissue site, and histology can all be correctly classified. FIG. 7A shows that there is high accuracy in predicting cancer status using the classification model described herein. Samples from different cohorts (e.g., samples with known cancer categories) were correctly categorized into their predicted labels most of the time (e.g., as indicated by the size of the circles for each category).

[0308] Figure 7B shows that tumor grade can be predicted even beyond cancer status. This example shows prediction results for brain cancer grade and subtype labels for two different cohorts of brain cancer samples. Samples identified as "brain cancer" in the pathology report (e.g., a set of "non-glioblastoma brain cancers") were classified into nine different categories using a classification model trained as described herein. Similar classification refinement was observed for samples designated "glioblastoma." Glioblastoma is a grade IV astrocytoma (e.g., a specific type of brain cancer originating from astrocytes that is locally highly invasive but does not usually metastasize beyond the brain and spinal cord). See, for example, Giese et al. 1996 Int. J. Cancer 67, 275-282. As shown in Figure 7B, samples within the pathologist-curated glioblastoma cohort, many of which fall into World Health Organization (WHO) grade IV, can be further differentiated by classification.

[0309] Beyond cancer status and subtype, the classification model described herein also provides information about tissue site and histology (e.g., thus increasing the resolution of classifier results and providing important information for informing treatment options). Figure 7C shows that tissue type can be accurately predicted. For example, a breast cancer sample with a tumor in the breast is classified as breast tissue and breast cancer. In addition, a pancreatic cancer sample with metastasis to the liver is further classified as "pancreas" and "tissue liver."

[0310] Figure 7D shows that using the classification method, squamous and adenocarcinoma phenotypes are summarized with reasonable accuracy. The x-axis indicates the TCGA type selected during pathological review. HNSC, CESC, and LUSC samples should be classified under the squamous label. LUAD and COAD samples should be classified under the adenocarcinoma label. KIRC samples should be classified under the carcinoma label.

[0311] Example 5 - Natural Language Processing of Diagnostic Values ​​from Pathology Reports Predictive models (e.g., classifiers) trained on RNA expression levels are likely to perform better when the labels / classes they predict are specific enough that they distinguish different RNA expression profiles (e.g., when the labels sufficiently define different clusters). Both broad and very specific labels can be problematic. For example, if a label is too general and many different RNA profiles are grouped together by being associated with the same label, the model may not recognize meaningful patterns that may be associated with that label. If a label is too specific and similar RNA profiles are arbitrarily separated by such a label, the model may encounter the same obstacle. In either case, inaccurate labels lead to a loss of information.

[0312] Thus, providing a classification algorithm that uses training labels corresponding to a relatively homogeneous set of tumors facilitates the identification of strong expression signatures that are robust to known confounders of RNA expression, including stage, tissue site, and immune infiltration. At the same time, the robustness of the signature may be reduced if the labels are overly narrowly defined, thus arbitrarily dividing an otherwise homogeneous cohort of tumors into two or more labels.

[0313] Cancer RNA expression represents the combined signal of the heterogeneous mixture of cell types present in a collected sample. While tissue specificity for RNA expression exists (e.g., genes with expression unique to certain cell types), the observed tumor expression profile can be highly confounded by the distribution of cell types present in the sample. For example, in Figure 11A, which shows clustered RNA expression data, each sample is labeled by both the tissue of origin of the sample (lung vs. oral cavity) and the cohort / general cancer status (adenocarcinoma vs. squamous cell) associated with the sample. Figure 11A shows three distinct clusters: 1102, 1104, and 1106. Due to a shared transcriptional signature for some patients (e.g., as seen in cluster 1102), oral squamous cell tumors appear more similar to lung squamous cell tumors than oral adenocarcinoma, regardless of biopsy tissue location.

[0314] The present disclosure includes a series of different tools for identifying optimal cohorts (e.g., classes) for cancer classification, avoiding classes defined by confounding factors. These tools combine natural language processing, unsupervised clustering of expression data, cross-validation, and the incorporation of known biology to identify and apply an optimal set of labels to the training data. An example of clustering a dataset is shown in Figures 12A-12C, described below. The goal of clustering is for all partitions (e.g., all labels) to identify some biologically relevant pathological subtype of the disease, providing further information to patients and / or healthcare professionals. In some embodiments, multiple iterations of clustering are required to obtain clusters that accurately describe the patient data and reveal actionable information.

[0315] Natural Language Processing before clustering: As discussed above with respect to block 314, the pathology diagnosis field typically allows for unstructured entry by medical personnel, such as a free text box containing the pathology assessment diagnosis (e.g., in addition to histology and stage data from abstracted clinical records). Methods are described herein for defining a label set from multiple pathology diagnosis field entries, which can then be used to analyze the biological relevance of algorithmically defined clusters of RNA expression profile data and annotate training data. In some embodiments, these methods also include identifying distinct text patterns associated with the label set (e.g., renaming diagnostic labels). [Table 1]

[0316] Table 1 illustrates various possible diagnostic entries found in a pathology report, the corresponding known cancer conditions associated with each diagnostic entry, and the corresponding NLP-assigned labels. After natural language processing, diagnostic entries from multiple different pathology reports (e.g., "metastatic prostate cancer" and "prostate adenocarcinoma") are mapped to the same cancer diagnostic category (e.g., "prostate"). Such normalization and data cleaning facilitate the extraction of a set of common features from the pathology report and the subsequent use of these features to refine a classifier for predicting the status / presence of a cancer condition.

[0317] In some embodiments, all labels are defined by one or more of the following criteria: Label text (e.g., "Lung Adenocarcinoma"). A set of regular expressions that can be selected and customized to search natural language text for the presence or absence of a label (e.g., ['lung.*adeno,''adeno.*lung,''^luad$'). A set of database text fields (e.g., cohort, diagnosis, histology, tissue site, etc.). Prioritization levels to establish a hierarchy to resolve rare situations where a single sample may be erroneously tagged with two mutually exclusive tags, such as lung adenocarcinoma and squamous lung. In some embodiments, rarer labels receive higher priority than more common or more abundant labels. In some embodiments, more abundant labels receive higher priority than rarer labels.

[0318] In some embodiments, the labels ...

Claims

1. 1. A method for determining a set of cancer conditions for a subject, comprising:

1. A computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, (A) in electronic form, at least: obtaining one or more data structures that collectively comprise a first plurality of sequence reads, the first plurality of sequence reads being obtained from a plurality of RNA molecules, the plurality of RNA molecules being from a somatic biopsy obtained from the subject; (B) determining a set of first sequence features for the subject from the first plurality of sequence reads; and (C) applying at least the first set of sequence features to a trained classification model, thereby obtaining, for each respective cancer condition in the set of cancer conditions, a classifier result that provides a likelihood that the subject has or does not have the respective cancer condition.

2. 1. A method for classifying a subject into cancer status, comprising: a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, the method comprising: (A) in electronic form, at least: obtaining one or more data structures that collectively comprise a first plurality of sequence reads, the first plurality of sequence reads being obtained from a plurality of RNA molecules, the plurality of RNA molecules being from a somatic biopsy obtained from the subject; (B) determining a set of first sequence features for the subject from the first plurality of sequence reads; and (C) applying at least the first set of sequence features to a trained classification model, thereby obtaining a classifier result providing a likelihood that the subject has or does not have the cancer condition.

3. 1. A method for classifying a subject into a probable cancer status, comprising: a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, the method comprising: (A) in electronic form, at least: a first plurality of sequence reads obtained from a plurality of RNA molecules, the plurality of RNA molecules being derived from a somatic biopsy obtained from the subject; and and an indicator of the probable cancer status of the subject. (B) determining a set of first sequence features for the subject from the first plurality of sequence reads; and (C) applying at least the first set of sequence features and the indicator of the predicted cancer status of the subject to a trained classification model, thereby obtaining a classifier result of a predicted cancer status; (D) comparing the predicted cancer status to the expected cancer status to provide a likelihood that the subject has or does not have the expected cancer status.

4. 4. The method of claim 3, wherein the predicted cancer state and the predicted cancer state are each independently selected from the set of: non-cancer, breast cancer, colorectal cancer, esophageal cancer, head / neck cancer, lung cancer, lymphoma, ovarian cancer, pancreatic cancer, prostate cancer, renal cancer, and uterine cancer.

5. the one or more data structures further comprising: a second plurality of sequence reads, the second plurality of sequence reads being obtained from the first plurality of DNA molecules; and a third plurality of sequence reads, the third plurality of sequence reads being obtained from the second plurality of DNA molecules; the first plurality of DNA molecules is derived from a somatic biopsy obtained from the subject; the second plurality of DNA molecules is derived from a germline sample obtained from the subject or from a collection of normal controls not including the set of cancer conditions; 4. The method of claim 1, further comprising determining a second set of sequence features for the subject from a comparison of the second plurality of sequence reads to the third plurality of sequence reads.

6. 6. The method of claim 5, wherein said applying (C) further comprises applying at least the first set of sequence features and the second set of sequence features to a trained classification model.

7. The method further comprises: obtaining a pathology report for the subject, the pathology report including at least one of a first estimate of tumor cellularity, an indication of whether the subject has metastatic or primary cancer, or a tissue site of origin of the somatic biopsy; and extracting a plurality of pathological features from the pathology report for the subject, the pathological features comprising a first estimate of the tumor cellularity of the somatic biopsy.

8. The method of claim 7 , wherein the trained classification model is selected based at least in part on the plurality of pathological features.

9. 8. The method of claim 7, wherein said applying (C) further comprises applying at least said plurality of pathological features, said first set of sequence features and said second set of sequence features to said trained classification model.

10. The method of any one of claims 1 to 3, wherein the set of cancer conditions consists of a single cancer condition.

11. The method of any one of claims 1 to 3, wherein the set of cancer conditions consists of two different cancer conditions.

12. The method of any one of claims 1 to 3, wherein the set of cancer conditions comprises five or more different cancer conditions.

13. 13. The method of any one of claims 1 and 4-12, wherein the set of cancer states provides a likelihood of origin of the cancer from each respective tissue of a plurality of tissues.

14. 13. The method of any one of claims 5 to 12, wherein the first plurality of sequence reads, the second plurality of sequence reads, and the third plurality of sequence reads are generated by next-generation sequencing.

15. 13. The method of any one of claims 5 to 12, wherein the first plurality of sequence reads, the second plurality of sequence reads, and the third plurality of sequence reads are generated by short-read paired-end next-generation sequencing.

16. the second plurality of sequence reads and the third plurality of sequence reads are obtained by targeted panel sequencing using a plurality of probes; each probe in said plurality of probes uniquely targets a respective portion of a reference genome; 13. The method of any one of claims 5 to 12, wherein each sequence read in the second plurality of sequence reads and each sequence read in the third plurality of sequence reads corresponds to at least one probe in the plurality of probes.

17. 17. The method of claim 16, wherein the second plurality of sequence reads has an average depth across the plurality of probes of at least 50-fold.

18. 17. The method of Claim 16, wherein the second plurality of sequence reads has an average depth across the plurality of probes of at least 400-fold.

19. 17. The method of claim 16, wherein the plurality of probes comprises probes for at least 300 different genes.

20. 17. The method of claim 16, wherein the plurality of probes comprises probes for at least 500 different genes.

21. 17. The method of claim 16, wherein the plurality of probes comprises probes for at least 500 different genes selected from a Targeted Gene Listing.

22. The method of any one of claims 7 to 18, wherein the plurality of pathological features comprises at least 200 pathological features.

23. 6. The method of claim 5, wherein the second plurality of sequence reads and the third plurality of sequence reads are obtained by whole exome sequencing.

24. the somatic biopsy comprises a microdissected formalin-fixed paraffin-embedded (FFPE) tissue section, a surgical biopsy, a skin biopsy, a punch biopsy, a prostate biopsy, a bone biopsy, a bone marrow biopsy, a needle biopsy, a CT-guided biopsy, an ultrasound-guided biopsy, a fine needle aspiration, aspiration biopsy, a fresh tissue or a blood sample; 24. The method of any one of claims 1 to 23, wherein the germline sample comprises blood or saliva from the subject.

25. The method of any one of claims 1 to 24, wherein the trained classification model comprises a trained classifier stream.

26. the trained classifier stream includes a first classifier, a second classifier, and a third classifier, and the applying (C) comprises: inputting all or a portion of the plurality of pathological features, the first set of sequence features, and the second set of sequence features into the first classifier, thereby obtaining an intermediate result; if the intermediate result satisfies a first predetermined threshold or range, inputting the intermediate result into the second classifier but not into the third classifier, thereby obtaining a probability that the subject has or does not have a first cancer condition in the set of cancer conditions; and if the intermediate result fails to meet a first predetermined threshold or range, entering the intermediate result into the third classifier but not the second classifier, thereby obtaining a probability that the subject has or does not have the first cancer condition.

27. the trained classifier stream includes a first classifier and a second classifier, and the applying (C) comprises: inputting all or a portion of the plurality of pathological features, the first set of sequence features, and the second set of sequence features into the first classifier, thereby obtaining an intermediate result; and if the intermediate result meets a first predetermined threshold or range, inputting the intermediate result into the second classifier, thereby obtaining a likelihood that the subject has or does not have a first cancer condition in the set of cancer conditions.

28. 27. The method of claim 25 or 26, wherein the trained classifier stream is a decision tree.

29. the trained classifier stream comprises a plurality of classifiers; the plurality of classifiers includes a subset of first classifiers and a subset of second classifiers; each classifier in the second subset of classifiers receives as input at least the output of at least one classifier in the first subset of classifiers; each classifier in the subset of first classifiers receives as input at least a plurality of pathological features, a first set of sequence features, and all or a portion of a second set of sequence features; 27. The method of claim 25 or 26, wherein the output of the second classifier subset collectively provides, for each respective cancer condition in the set of cancer conditions, a likelihood that the subject has or does not have the respective cancer condition.

30. the trained classifier stream comprises a plurality of classifiers; using a first classifier in the plurality of classifiers to determine the likelihood that the subject has or does not have a first cancer condition in a set of cancer conditions if the tumor cellularity meets a predetermined threshold; 27. The method of claim 25 or 26, wherein a second classifier in the plurality of classifiers is used to determine the likelihood that the subject has or does not have the first cancer condition in the set of cancer conditions if tumor cellularity fails to meet a predetermined threshold.

31. 31. The method of any one of claims 7 to 30, wherein the method further comprises complementing the first estimate of tumor cellularity of the somatic biopsy with a second estimate of tumor cellularity from one or more images of the somatic biopsy.

32. 32. The method of any one of claims 7-31, wherein the method further comprises complementing the first estimate of tumor cellularity of the somatic biopsy with a second estimate of tumor cellularity from the abundance of one or more mutations in the second plurality of sequence reads.

33. the first set of sequence features comprises between 15,000 features and 22,000 features; the second set of sequence features comprises between 400 features and 2,000 features; The method of claim 7, wherein the plurality of pathological features comprises between 200 features and 500 features.

34. 8. The method of claim 7, wherein the first plurality of sequence reads, the second plurality of sequence reads, and the third plurality of sequence reads are generated by short-read next-generation sequencing using one or more spike-in controls.

35. 35. The method of claim 34, wherein the one or more spike-in controls calibrate for variation in sequence reads across a population of cells.

36. 8. The method of claim 7, wherein the pathology report further comprises one or more image features extracted from one or more images of a somatic biopsy from the test subject.

37. 37. The method of claim 36, wherein said applying (C) further comprises applying one or more image features extracted from one or more images of the somatic biopsy from the test subject.

38. 4. The method of claim 1, wherein the first set of sequence features derived from the first plurality of sequence reads comprises one or more gene fusions, one or more copy number variations, one or more somatic mutations, one or more germline mutations, tumor mutational burden, one or more indices of microsatellite instability, indices of pathogen burden, indices of immune infiltration, or indices of tumor cellularity.

39. 6. The method of Claim 5, wherein the second set of sequence features derived from the second plurality of sequence reads comprises one or more gene fusions, one or more single nucleotide variants, one or more copy number variations, one or more somatic mutations, one or more germline mutations, tumor mutational burden, one or more indices of microsatellite instability, indices of pathogen burden, indices of immune infiltration, or indices of tumor cellularity.

40. 8. The method of claim 7, wherein the plurality of pathological features comprises one or more of an IHC protein level, an age of the test subject, a gender of the test subject, a disease diagnosis, a treatment category, a type of treatment, or a treatment outcome.

41. 41. The method of any one of claims 1 to 40, wherein the somatic biopsy is a somatic biopsy of a breast tumor, a glioblastoma, a prostate tumor, a pancreatic tumor, a kidney tumor, a colorectal tumor, an ovarian tumor, an endometrial tumor, a breast tumor, or a combination thereof.

42. 42. The method of any one of claims 7 to 41, wherein the metastatic cancer or the primary cancer comprise tumors of a common primary site of origin.

43. 42. The method of any one of claims 7 to 41, wherein the metastatic cancer or the primary cancer comprises tumors originating from two or more different organs.

44. 42. The method of any one of claims 7 to 41, wherein the metastatic cancer or the primary cancer comprises a tumor of a predetermined stage of brain cancer, a predetermined stage of glioblastoma, a predetermined stage of prostate cancer, a predetermined stage of pancreatic cancer, a predetermined stage of renal cancer, a predetermined stage of colorectal cancer, a predetermined stage of ovarian cancer, a predetermined stage of endometrial cancer, or a predetermined stage of breast cancer.

45. 10. The method of claim 1, wherein a cancer condition in the set of cancer conditions is the likelihood that the subject has metastatic cancer.

46. 46. ​​The method of any one of claims 1-45, wherein said applying (C) further comprises applying one or more epigenetic or metabolomic features of the subject obtained from the germline sample of the subject to the trained classification model to obtain the classifier result.

47. 10. The method of claim 1, wherein the trained classification model further provides one or more treatment recommendations to the subject or a healthcare professional caring for the subject based on the likelihood that the subject has or does not have each respective cancer condition in the set of cancer conditions.

48. 3. The method of claim 2, wherein the trained classification model further provides one or more treatment recommendations to the subject or a healthcare professional caring for the subject based on the likelihood the subject has or does not have the cancerous condition.

49. 4. The method of claim 3, wherein the trained classification model further provides one or more treatment recommendations to the subject or a healthcare professional caring for the subject based on the likelihood the subject has or does not have the predicted cancer condition.

50. wherein said determining (B) comprises aligning each respective sequence read in the first plurality of sequence reads to a reference genome to determine the first set of sequence features for the subject; 50. The method of any one of claims 5-49, wherein said determining (C) comprises aligning each respective sequence read in the second plurality of sequence reads and the third plurality of sequence reads to a reference genome to determine the second set of sequence features for the subject.

51. 47. The method of any one of claims 5 to 46, wherein the trained classifier stream comprises a logistic regression, a K-nearest neighbor model, a random forest model, or a neural network.

52. 52. The method of claim 51 , wherein a boosting algorithm is applied to the trained classifier stream.

53. The method of any one of claims 1 to 3, wherein the somatic biopsy comprises one of a solid biopsy of the subject or a liquid biopsy of the subject.

54. 4. The method of claim 1, wherein (B) determining a first set of sequence features further comprises deconvoluting the first plurality of sequence reads by comparing the first plurality of sequence reads to a deconvoluted RNA expression model comprising at least one cluster identified as corresponding to a cancer state.

55. 1. A non-transitory computer-readable storage medium having stored thereon program code instructions that, when executed by a processor, cause the processor to perform a method for classifying a subject into a cancer status, the method comprising: (A) in electronic form, at least: obtaining one or more data structures that collectively comprise a first plurality of sequence reads, the first plurality of sequence reads being obtained from a plurality of RNA molecules, the plurality of RNA molecules being from a somatic biopsy obtained from the subject; (B) determining a set of first sequence features for the subject from the first plurality of sequence reads; and (C) applying at least the first set of sequence features to a trained classification model, thereby obtaining a classifier result, wherein the classifier result provides, for each respective cancer condition in a set of cancer conditions, a likelihood that the subject has or does not have the respective cancer condition.

56. 1. A computer system for determining a set of cancer conditions for a subject, the computer system comprising: at least one processor; and a memory storing at least one program for execution by said at least one processor, said at least one program comprising: (A) in electronic form, at least: obtaining one or more data structures that collectively comprise a first plurality of sequence reads, the first plurality of sequence reads being obtained from a plurality of RNA molecules, the plurality of RNA molecules being from a somatic biopsy obtained from the subject; (B) determining a set of first sequence features for the subject from the first plurality of sequence reads; and (C) instructions for applying at least the first set of sequence features to a trained classification model, thereby obtaining a classifier result, the classifier result providing, for each respective cancer condition in a set of cancer conditions, a likelihood that the subject has or does not have the respective cancer condition.

57. 1. A non-transitory computer-readable storage medium having stored thereon program code instructions that, when executed by a processor, cause the processor to perform a method for classifying a subject into a cancer status, the method comprising: (A) in electronic form, at least: obtaining one or more data structures that collectively comprise a first plurality of sequence reads, the first plurality of sequence reads being obtained from a plurality of RNA molecules, the plurality of RNA molecules being from a somatic biopsy obtained from the subject; (B) determining a set of first sequence features for the subject from the first plurality of sequence reads; and (C) applying at least the first set of sequence features to a trained classification model, thereby obtaining a classifier result that provides a likelihood that the subject has or does not have the cancer condition.

58. 1. A computer system for classifying a subject into a cancer status, said computer system comprising: at least one processor; and a memory storing at least one program for execution by said at least one processor, said at least one program comprising:

1. A computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, (A) in electronic form, at least: obtaining one or more data structures that collectively comprise a first plurality of sequence reads, the first plurality of sequence reads being obtained from a plurality of RNA molecules, the plurality of RNA molecules being from a somatic biopsy obtained from the subject; (B) determining a set of first sequence features for the subject from the first plurality of sequence reads; and (C) applying at least the first set of sequence features to a trained classification model, thereby obtaining a classifier result that provides a likelihood that the subject has or does not have the cancer condition.

59. 1. A non-transitory computer-readable storage medium having stored thereon program code instructions that, when executed by a processor, cause the processor to perform a method for classifying a subject into a probable cancer status, the method comprising: (A) in electronic form, at least: a first plurality of sequence reads obtained from a plurality of RNA molecules, the plurality of RNA molecules being derived from a somatic biopsy obtained from the subject; and obtaining one or more data structures that collectively comprise an indicator of a probable cancer status of the subject; (B) determining a set of first sequence features for the subject from the first plurality of sequence reads; and (C) applying at least the first set of sequence features and the indicator of the predicted cancer status of the subject to a trained classification model, thereby obtaining a classifier result of a predicted cancer status; (D) comparing the predicted cancer status to the expected cancer status to provide a likelihood that the subject has or does not have the expected cancer status.

60. 1. A computer system for classifying a subject into a probable cancer status, said computer system comprising: at least one processor; and a memory storing at least one program for execution by said at least one processor, said at least one program comprising:

1. A computer system having one or more processors and a memory storing one or more programs for execution by said one or more processors, (A) in electronic form, at least: a first plurality of sequence reads obtained from a plurality of RNA molecules, the plurality of RNA molecules being derived from a somatic biopsy obtained from the subject; and obtaining one or more data structures that collectively comprise an indicator of a probable cancer status of the subject; (B) determining a set of first sequence features for the subject from the first plurality of sequence reads; and (C) applying at least the first set of sequence features and the indicator of the predicted cancer status of the subject to a trained classification model, thereby obtaining a classifier result of a predicted cancer status; (D) a computer system comprising instructions for comparing the predicted cancer status to the expected cancer status to provide a likelihood that the subject has or does not have the expected cancer status.

61. 1. A method for training a classifier, comprising: a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors, the method comprising: (A) in electronic form, for each respective object in the plurality of objects: For each respective cancer condition in the set of cancer conditions, an indication of whether the respective subject has the indication of cancer; and a first plurality of sequence reads obtained from a plurality of RNA molecules, the plurality of RNA molecules being derived from a somatic biopsy obtained from the respective subject; and obtaining a pathology report for each of the subjects, the pathology report including at least one of a first estimate of tumor cellularity, an indication of whether each of the subjects has metastatic or primary cancer, or a tissue site of origin of the somatic biopsy; (B) for each respective subject in the plurality of subjects, determining a corresponding set of first sequence features for the respective subject from the first plurality of sequence reads for the respective subject; (C) extracting, for each respective subject in the plurality of subjects, a plurality of pathological features from the pathology report for the respective subject, the pathological features including a first estimate of the tumor cellularity of the somatic biopsy and an indication of whether the respective subject has metastatic cancer or primary cancer; (D) inputting at least the first set of sequence features and the plurality of pathological features for each respective subject in the plurality of subjects into an untrained classification model, thereby training the untrained classification model for an indication of whether each respective subject in the plurality of subjects has each respective cancer condition in the set of cancer conditions, thereby obtaining a trained classification model configured to provide, for each respective cancer condition in the set of cancer conditions, a likelihood that a test subject has or does not have the respective cancer condition.

62. The method further comprises: For each of the multiple objects, a second plurality of sequence reads, the second plurality of sequence reads being obtained from the first plurality of DNA molecules; and obtaining a third plurality of sequence reads, the third plurality of sequence reads being obtained from a second plurality of DNA molecules; obtaining the first plurality of DNA molecules from a somatic biopsy obtained from the subject, and the second plurality of DNA molecules from a germline sample obtained from the subject or from a collection of normal controls not including the set of cancer conditions; and determining a second set of sequence features for the subject from a comparison of the second plurality of sequence reads to the third plurality of sequence reads.

63. 63. The method of Claim 62, wherein said inputting (D) further comprises applying at least said first set of sequence features and said second set of sequence features to a trained classification model.

64. 62. The method of claim 61 , wherein the trained classification model comprises a trained classifier stream.

65. 65. The method of claim 64, wherein the trained classifier stream comprises a logistic regression, a hierarchical model, a deep neural network, a multi-task multi-kernel learning engine, or a nearest neighbor engine.

66. 66. The method of claim 65, wherein a boosting algorithm is applied to the trained classifier stream.

67. 62. The method of claim 61 , wherein extracting the plurality of pathological features from the pathology report further comprises normalizing the pathology report.

68. 1. A non-transitory computer-readable storage medium having stored thereon program code instructions that, when executed by a processor, cause the processor to perform a method for training a classifier, the method comprising: (A) in electronic form, for each respective object in the plurality of objects: For each respective cancer condition in the set of cancer conditions, an indication of whether the respective subject has the indication of cancer; and obtaining a first plurality of sequence reads, the first plurality of sequence reads being obtained from a plurality of RNA molecules, the plurality of RNA molecules being derived from a somatic biopsy obtained from the respective subject; (B) for each respective subject in the plurality of subjects, determining a corresponding set of first sequence features for the respective subject from the first plurality of sequence reads for the respective subject; (C) extracting, for each respective subject in the plurality of subjects, a plurality of pathological features from the pathology report for the respective subject, the pathological features including a first estimate of the tumor cellularity of the somatic biopsy and an indication of whether the respective subject has metastatic cancer or primary cancer; (D) inputting at least the first set of sequence features and the plurality of pathological features of each respective subject in the plurality of subjects into an untrained classification model, thereby training the untrained classification model for an indication of whether each respective subject in the plurality of subjects has each respective cancer condition in the set of cancer conditions, thereby obtaining a trained classification model configured to provide, for each respective cancer condition in the set of cancer conditions, a likelihood that a test subject has or does not have the respective cancer condition.

69. 1. A computer system for training a classifier, the computer system comprising: at least one processor; and a memory storing at least one program for execution by said at least one processor, said at least one program comprising: (A) in electronic form, for each respective object in the plurality of objects: For each respective cancer condition in the set of cancer conditions, an indication of whether the respective subject has the indication of cancer; and obtaining a first plurality of sequence reads, the first plurality of sequence reads being obtained from a plurality of RNA molecules, the plurality of RNA molecules being derived from a somatic biopsy obtained from the respective subject; (B) for each respective subject in the plurality of subjects, determining a corresponding set of first sequence features for the respective subject from the first plurality of sequence reads for the respective subject; (C) extracting, for each respective subject in the plurality of subjects, a plurality of pathological features from the pathology report for the respective subject, the pathological features including a first estimate of the tumor cellularity of the somatic biopsy and an indication of whether the respective subject has metastatic cancer or primary cancer; (D) inputting at least the first set of sequence features and the plurality of pathological features of each respective subject in the plurality of subjects into an untrained classification model, thereby training the untrained classification model for an indication of whether each respective subject in the plurality of subjects has each respective cancer condition in the set of cancer conditions, thereby obtaining a trained classification model configured to provide, for each respective cancer condition in the set of cancer conditions, a likelihood that a test subject has or does not have the respective cancer condition.

70. 1. A method for identifying a diagnosis of cancer status in a patient somatic tumor specimen of unknown origin, said method comprising: receiving sequencing information comprising an analysis of a plurality of nucleic acids from the somatic tumor specimen; Identifying a plurality of features from the received sequencing information, wherein the plurality of features includes an RNA feature, a DNA feature, an RNA splicing feature, a viral feature, and a copy number feature; generating two or more predictions of cancer status based at least in part on the identified features from two or more classifiers; and combining the two or more predictions to identify a diagnosis of the cancer condition for the somatic tumor specimen of the patient.

71. Combining the two or more predictions further comprises: scaling each prediction of the two or more predictions based at least in part on a confidence level in each respective prediction; and generating a combined prediction based at least in part on each prediction of the two or more predictions.

72. 71. The method of claim 70, wherein the two or more classifiers are selected from a diagnostic classifier, a cohort classifier, or a tissue classifier.

73. The two or more predictions are:

71. The method of claim 70, comprising a first prediction from a diagnostic classifier for RNA features, a second prediction from a cohort classifier for RNA features, a third prediction from a tissue classifier for RNA features, a fourth prediction from a diagnostic classifier for RNA splicing features, a fifth prediction from a cohort classifier for RNA splicing features, a sixth prediction from a diagnostic classifier for CNV features, a seventh prediction from a cohort classifier for CNV features, an eighth prediction from a diagnostic classifier for DNA features, and a ninth prediction from a diagnostic classifier for viral features.

74. The plurality of features may include GPM6A, CDX1, SOX2, NAPSA, CDX2, MUC12, SLAMF7, HNF4A, ANXA10, TRPS1, GATA3, SLC34A2, NKX2-1, SLC22A31, ATP10B, STEAP2, CLDN3, SPATA6, NRCAM, USH1C, SOX17, TMPRSS2, MECOM, WT1, CDHR1, HOXA13, SOX10, SALL1, CPE, NPR1, CLRN3, THSD4, ARL14, SFTPB, COL17A1, KLHL14, EPS8L3, NXPE4, FOXA2, SYT1 1, SPDEF, GRHL2, GBP6, PAX8, ANOl, KRT7, HOXA9, TYR, DCT, LYPDl, MSLN, TP63, CDH1, ESR1, HNFIB, HOX A10, TJP3, NRG3, TMC5, PRLR, GATA2, DCDC2, INS, NDUFA4L2, TBX5, ABCC3, FOLH1, HIST1H3G, S100A1, PT 71. The method of claim 70, comprising one or more of HLH, ACER2, RBBP8NL, TACSTD2, C19orf77, PTPRZ1, BHLHE41, FAM155A, MYCN, DDX3Y, FMN1, HIST1H3F, UPK3B, TRIM29, TXNDC5, BCAM, FAM83A, TCF21, MIA, RNF220, AFAP1, KRT5, SOX21, KANK2, GPM6B, Clorfl 16, FOXF1, MEIS1, EFHD1, and XKRX.

75. 1. A method for identifying a diagnosis of a cancerous state in a somatic tumor specimen of a subject, said method comprising: receiving sequencing information comprising an analysis of a plurality of nucleic acids from the somatic tumor specimen; identifying a plurality of features from the received sequencing information, the plurality of features comprising two or more of an RNA feature, a DNA feature, an RNA splicing feature, a viral feature, and a copy number feature; each RNA feature is associated with a respective target region of the first reference genome and represents the abundance of corresponding sequence reads mapping to said respective target region encompassed by said sequencing information; each DNA feature is associated with a respective target region of a second reference genome and represents an abundance of corresponding sequence reads mapping to said respective target region encompassed by said sequencing information; each RNA splicing feature is associated with a respective splicing event in a respective target region of the first reference genome and represents an abundance of corresponding sequence reads mapping to the respective target region with the respective splicing event encompassed by the sequencing information; each viral feature is associated with a respective target region of a viral reference genome and represents an abundance of corresponding sequence reads mapping to the respective target region in the viral reference genome encompassed by the sequencing information; Identifying each copy number feature associated with a target region of the second reference genome and representing an abundance of corresponding sequence reads mapping to the respective target region of the second reference genome encompassed by the sequencing information; providing a first subset of features from the identified plurality of features as input to a first classifier; providing a second subset of features from the identified plurality of features as input to a second classifier; generating two or more predictions of cancer status based at least in part on the identified plurality of features from two or more classifiers, the two or more classifiers including at least the first classifier and the second classifier; and combining the two or more predictions in a final classifier to identify a diagnosis of the cancer status for the somatic tumor specimen of the subject.

76. Combining the two or more predictions in the final classifier further comprises: scaling each prediction of the two or more predictions based at least in part on a respective confidence level in each respective prediction; and generating a combined prediction based at least in part on each scaled prediction.

77. 76. The method of claim 75, wherein providing the first subset of features to the first classifier and providing the second subset of features to the second classifier comprises providing the same subset of features to both the first classifier and the second classifier.

78. 78. The method of claim 77, wherein the subset of the same features are RNA features.

79. 76. The method of claim 75, wherein the first classifier is a diagnostic classifier and the second classifier is a cohort classifier.

80. 76. The method of claim 75, wherein the first classifier is a diagnostic classifier and the second classifier is a tissue classifier.

81. 80. The method of Claim 79, wherein providing the subset of first features to the first classifier and providing the subset of second features to the second classifier comprises providing RNA features to both the first classifier and the second classifier.

82. 76. The method of Claim 75, wherein providing the subset of first features to the first classifier and providing the subset of second features to the second classifier comprises providing RNA features to the first classifier and providing DNA features to the second classifier.

83. moreover, providing an RNA splicing signature to a third classifier; generating three or more predictions of cancer status based at least in part on the identified plurality of features from three or more classifiers, the three or more classifiers including at least the first classifier, the second classifier, and the third classifier; 83. The method of claim 82, wherein said combining combines said three or more predictions to identify a diagnosis of said cancer status for said somatic tumor specimen.

84. moreover, providing virus signatures to a third classifier; generating three or more predictions of cancer status based at least in part on the identified plurality of features from three or more classifiers, the three or more classifiers including at least the first classifier, the second classifier, and the third classifier; 83. The method of claim 82, wherein said combining combines said three or more predictions to identify a diagnosis of said cancer status for said somatic tumor specimen.

85. moreover, providing copy number features to a third classifier; generating three or more predictions of cancer status based at least in part on the identified plurality of features from three or more classifiers, the three or more classifiers including at least the first classifier, the second classifier, and the third classifier; 83. The method of claim 82, wherein said combining combines said three or more predictions to identify a diagnosis of said cancer status for said somatic tumor specimen.

86. moreover, providing RNA features to the first classifier; providing copy number features to said second classifier; providing an RNA splicing signature to a third classifier; generating three or more predictions of cancer status based at least in part on the identified plurality of features from three or more classifiers, the three or more classifiers including at least the first classifier, the second classifier, and the third classifier; 76. The method of claim 75, wherein said combining combines said three or more predictions to identify a diagnosis of said cancer status for said somatic tumor specimen.

87. moreover, providing RNA features to the first classifier, wherein the first classifier is a diagnostic classifier; providing RNA features to the second classifier, wherein the second classifier is a cohort classifier; providing the RNA features to a third classifier, the third classifier being a tissue classifier; generating three or more predictions of cancer status based at least in part on the identified plurality of features from three or more classifiers, the three or more classifiers including at least the first classifier, the second classifier, and the third classifier; 76. The method of claim 75, wherein said combining combines said three or more predictions to identify a diagnosis of said cancer status for said somatic tumor specimen.

88. moreover, providing the DNA signature to a fourth classifier, the fourth classifier being a diagnostic classifier; providing the RNA splicing signature to a fifth classifier, wherein the fifth classifier is a diagnostic classifier; providing the RNA splicing signature to a sixth classifier, wherein the sixth classifier is a cohort classifier; wherein said generating generates six or more predictions of cancer status based at least in part on the identified plurality of features from six or more classifiers, wherein the six or more classifiers include at least the first classifier, the second classifier, the third classifier, the fourth classifier, the fifth classifier, and the sixth classifier; 88. The method of claim 87, wherein said combining combines said six or more predictions to identify a diagnosis of said cancer status for said somatic tumor specimen.

89. The two or more predictions are:

76. The method of claim 75, comprising a first prediction from a diagnostic classifier provided with RNA features, a second prediction from a cohort classifier provided with RNA features, a third prediction from a tissue classifier provided with RNA features, a fourth prediction from a diagnostic classifier provided with RNA splicing features, a fifth prediction from a cohort classifier provided with RNA splicing features, a sixth prediction from a diagnostic classifier provided with CNV features, a seventh prediction from a cohort classifier provided with CNV features, an eighth prediction from a diagnostic classifier provided with DNA features, and a ninth prediction from a diagnostic classifier provided with viral features.

90. Each feature in the plurality of features is associated with a respective target region, the plurality of features collectively represent a plurality of target regions, each region in the plurality of target regions is a gene, and the plurality of target regions include GPM6A, CDX1, SOX2, NAPSA, CDX2, MUC12, SLAMF7, HNF4A, ANXA10, TRPS1, GATA3, SLC34A2, NKX2- 1, SLC22A31, ATP10B, STEAP2, CLDN3, SPATA6, NRCAM, USH1C, S0X17, TMPRSS2, MECOM, WT1, CDHR1, HOXA 13, SOXIO, SALLl, CPE, NPR1, CLRN3, THSD4, ARL14, SFTPB, COL17A1, KLHL14, EPS8L3, NXPE4, FOXA2, SY T11, SPDEF, GRHL2, GBP6, PAX8, ANOl, KRT7, HOXA9, TYR, DCT, LYPDl, MSLN, TP63, CDH1, ESR1, HNF1B, HO XAIO, TJP3, NRG3, TMC5, PRLR, GATA2, DCDC2, INS, NDUFA4L2, TBX5, ABCC3, FOLH1, HIST1H3G, S100A1, P 76. The method of claim 75, comprising ten or more of THLH, ACER2, RBBP8NL, TACSTD2, C19orf77, PTPRZ1, BHLHE41, FAM155A, MYCN, DDX3Y, FMN1, HIST1H3F, UPK3B, TRIM29, TXNDC5, BCAM, FAM83A, TCF21, MIA, RNF220, AFAP1, KRT5, SOX21, KANK2, GPM6B, Clorfl 16, FOXF1, MEIS1, EFHD1, and XKRX.

91. Each feature in the plurality of features is associated with a respective target region, the plurality of features collectively represent a plurality of target regions, each region in the plurality of target regions is a gene, and the plurality of target regions are selected from the group consisting of ENSG00000 150625, ENSG000001 13722, ENSG00000181449, ENSG00000 131400, ENSG00000165556, ENSG00000205277, ENSG00000026751, ENSG00000 101076, ENSG00000 10951 1, ENSG00000 104447, ENSG00000 107485, ENSG00000 157765, ENSG00000136352, ENSG00000259803, ENSG000001 18322, ENSG00000157214, ENSG00000165215, ENSG00000132122, ENSG00000091 129, ENSG0000000661 1, ENSG00000164736, ENSG00000184012, ENSG00000085276, ENSG00000 184937, ENSG00000148600, ENSG00000 106031, ENSG00000100146, ENSG00000 103449, ENSG00000 109472, ENSG00000169418, ENSG00000 180745, ENSG00000 187720, ENSG00000 179674, ENSG00000 168878, ENSG00000065618, ENSG00000 197705, ENSG00000198758, ENSG00000137634, ENSG00000125798, ENSG00000132718, ENSG00000 124664, ENSG00000083307, ENSG00000183347, ENSG00000125618, ENSG00000131620, ENSG00000 135480, ENSG00000078399, ENSG00000077498, ENSG00000080166, ENSG00000150551, ENSG00000102854, ENSG00000073282, ENSG00000039068, ENSG00000091831, ENSG00000108753, ENSG00000253293, ENSG00000 105289, ENSG00000185737, ENSG00000103534, ENSG00000 113494, ENSG00000 179348, ENSG00000146038, ENSG00000254647, ENSG00000 185633, ENSG00000089225, ENSG00000 108846, ENSG00000086205, ENSG00000256018, ENSG00000 160678, ENSG00000087494, ENSG00000 177076, ENSG00000 130701, ENSG00000 184292, ENSG00000095932, ENSG00000 106278, ENSG00000 123095, ENSG00000204442, ENSG00000134323, ENSG00000067048, ENSG000002 48905, ENSG00000256316, ENSG00000243566, ENSG00000137699, ENSG0000023 9264, ENSG00000 187244, ENSG00000147689, ENSG000001 18526, ENSG00000261857, ENSG00000187147, ENSG00000 196526, ENSG00000 76. The method of claim 75, comprising ten or more of ENSG00000 125285, ENSGOOOOO 197256, ENSG00000046653, ENSGOOOOO 182795, ENSGOOOOO 103241, ENSG00000143995, ENSGOOOOO 115468, and ENSGOOOOO 182489.

92. 76. The method of claim 75, wherein the plurality of features is obtained by low-pass whole genome sequencing.

93. 76. The method of claim 75, wherein the RNA signature is obtained from cDNA sequencing.

94. 76. The method of claim 75, wherein the RNA signature is associated with a coding region of a gene.

95. moreover, receiving the final classifier diagnosis of the cancer status for the somatic tumor specimens of a plurality of subjects; calculating an entropy score for each subject based at least in part on the respective final classifier diagnosis for each subject in the plurality of subjects; identifying an entropy threshold based at least in part on the accuracy of the entropy score for each subject in the plurality of subjects; and training the final classifier using subjects from among those whose entropy scores satisfy the entropy threshold.

96. 96. The method of claim 95, wherein specifying an entropy threshold comprises specifying a percentile of the accuracy of the final classifier across the plurality of subjects.

97. 76. The method of claim 75, wherein the final classification diagnosis of the cancer condition comprises differentiating between lung adenocarcinoma, lung squamous cell carcinoma, oral adenocarcinoma, and oral adenocarcinoma.

98. 76. The method of claim 75, wherein said final classification diagnosis of said cancerous condition comprises differentiating between systemic sarcoma, epithelioma, Ewing's sarcoma, gliosarcoma, leiomyosarcoma, meningioma, mesothelioma, and Rosai-Dorfman.

99. 76. The method of claim 75, wherein the final classification diagnosis of the cancer condition comprises differentiating between liver metastases of pancreatic origin, upper gastrointestinal origin, and biliary origin.

100. 76. The method of claim 75, wherein said final classification diagnosis of said cancer condition comprises differentiating between glioblastoma, oligodendroglioma, astrocytoma, and medulloblastoma brain metastasis.

101. 76. The method of claim 75, wherein said final classification diagnosis of said cancer condition comprises differentiating between non-small cell lung cancer squamous cell and adenocarcinoma.

102. 76. The method of claim 75, wherein the final classification diagnosis of the cancerous condition comprises distinguishing between one or more sarcomas with morphological characteristics or protein expression of a carcinoma and one or more carcinomas with morphological characteristics or protein expression of a sarcoma.

103. 76. The method of claim 75, wherein said final classification diagnosis of said cancer condition comprises distinguishing between one or more neuroendocrine, one or more carcinoma, and one or more sarcoma.

104. identifying said diagnosis of said cancerous condition further comprises: receiving subject information including one or more clinical events; and differentiating the cancer status between a new tumor and a recurrence of a previous tumor based at least in part on said one or more clinical events.