Dynamic tissue typing

By employing AI systems with machine learning models to analyze biological data and navigate hierarchical type trees, the challenge of accurately identifying tumor tissue of origin is addressed, leading to improved diagnostic accuracy and treatment outcomes.

WO2025137139A1PCT designated stage expired Publication Date: 2025-06-26CARIS MPI INC

Patent Information

Application Number
PCT/US2024/060823
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-23
Filing Date
2024-12-18
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current methods for identifying the primary tissue/organ of origin for tumors, especially in cases of metastatic cancer of unknown primary origin (CUP), are often inaccurate, leading to suboptimal or ineffective treatment.

Method used

The use of artificial intelligence (AI) systems that employ machine learning models to predict tumor tissue of origin by analyzing biological data, such as DNA mutations and RNA expression levels, and navigating a hierarchical type tree to provide accurate assessments.

Benefits of technology

This approach enables accurate identification of tumor tissue of origin, improves diagnostic accuracy, and aids in selecting optimal treatment regimens, thereby enhancing patient outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024060823_26062025_PF_FP_ABST
    Figure US2024060823_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Systems, apparatuses, and methods as described herein can provide in part a validated AI model integrated with tumor profiling that enhances diagnostic accuracy, including resolution of CUP cases, and prompts clinically relevant therapeutic recommendation changes without requiring additional specimen. Machine learning models in a hierarchal sample type tree can be used, e.g., to determine a tumor type of a cancer.
Need to check novelty before this filing date? Find Prior Art

Description

PATENT Attorney Docket No.: 110588-0886WO1-1478774 Client Reference No.: CMI 886.601 DYNAMIC TISSUE TYPING BACKGROUND

[0001] When a cancer is identified for treatment, it is important to know the primary origin of the cancer, such as whether the cancer originated in the current tissue / organ at which the tumor was identified or some other tissue / organ. In some situations, a doctor might identify that the cancer originated from some other unknown tissue / organ but not know which one (e.g., a metastatic cancer of unknown primary origin, or CUP). In other situations, a pathologist might misidentify the cancer as originating from the current tissue / organ, when the cancer actually originated in another tissue / organ. Information about the tumor type may play a role in selecting optimal treatment regimens. Therefore, misidentification or non- identification of the primary tissue / organ of the tumor may lead to suboptimal to ineffective treatment of the patient. In the laboratory setting, identification of a tumor type also can be used as a quality safeguard. BRIEF SUMMARY

[0002] Provided herein are systems and methods that employ artificial intelligence (AI) to predict tumor tissue of origin. Applications thereof include identifying origin of CUP and flagging potential misdiagnoses for additional workup during routine molecular testing.

[0003] Embodiments can use biological data measured from a sample and a set of machine learning models to navigate a hierarchal type tree to provide accurate assessments of a sample type (e.g., a tumor type). For example, biological data (e.g., DNA mutations and / or RNA expression) can be measured from a biological sample. The biological data can form a sample vector that can be compared to reference vectors corresponding to different sample types. Such a comparison can occur by training respective machine learning models with different subsets of training samples (e.g., all samples for top level and portions for lower levels).

[0004] The hierarchy can provide as much specificity as possible or desired about the sample type, as well as ensure that an accurate type (e.g., location) is identified. The hierarchy can also include multiple sample subtypes for the same top-level type (e.g., same tumor origin). The hierarchy can form a label tree, where each node in the label tree corresponds to a different samples type, e.g., tumor origin or tumor type for a given origin.

[0005] In some embodiments, a top-level machine learning model can determine a probability (likelihood) that a sample is from each of N different origins, which can correspond to N top-level nodes of the sample tree (e.g., a tumor tree). The top-level model can be trained using top-level training samples, each having a label corresponding to one of the N different origins. A first origin node with the highest probability that is greater than a threshold may have multiple (M) sample types that further specify the type of sample. In such a situation, a second-level machine learning model can be used. The second-level machine learning model can be trained using second-level training samples, each having the first origin (e.g., ovarian epithelial tumor) and having a respective second-level label of the M sample types indicating a subtype for the first origin (e.g., serous ovarian cancer, endometoid ovarian cancer, mucinous ovarian cancer, clear cell ovarian cancer, and ovarian carcinosarcoma with respect to first origin ovarian epithelial tumor). The second-level training samples can be a subset of the top-level training samples. Each top-level node having multiple sample types can have a separately trained ML model that selects among the M second level subtypes for that top-level node. When multiple top-level origins each have multiple subtypes, there can be a plurality of second-level machine learning models. Further levels (e.g., a third-level model or a fourth-level model) can be employed for each node / type that has further subtypes.

[0006] These and other embodiments of the disclosure are described in detail below. For example, other embodiments are directed to systems, devices, and computer readable media associated with methods described herein.

[0007] A better understanding of the nature and advantages of embodiments of the present disclosure may be gained with reference to the following detailed description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] FIG.1 shows a sample type tree 100 according to embodiments of the present disclosure.

[0009] FIG.2 shows a diagram of a top-level ML model classifying among a set of top- level types according to embodiments of the present disclosure.

[0010] FIG.3 illustrates a selection and use of a second-level ML model according to embodiments of the present disclosure.

[0011] FIG.4 is a block diagram of a system for training a machine learning model to predict a sample type (e.g., origin) of a biological sample based on biological data obtained from the biological sample.

[0012] FIG.5 is a block diagram of a system for using a trained machine learning model 570 to predict a sample origin of sample data from a subject.

[0013] FIG.6 shows an example top level of a hierarchal tumor type tree according to embodiments of the present disclosure.

[0014] FIG.7 shows an example top level and several lower levels of a hierarchal tumor type tree according to embodiments of the present disclosure.

[0015] FIG.8 shows results of a subject identified as having cholangiocarcinoma according to embodiments of the present disclosure.

[0016] FIGS.9A-C show an example hierarchical tree according to embodiments of the present disclosure.

[0017] FIG.10A and FIG.10B show performance of the trained tissue typing system in a retrospective and prospective analysis, respectively.

[0018] FIG.11A is a flowchart illustrating a method for determining a tumor type using a hierarchal tumor type tree according to embodiments of the present disclosure.

[0019] FIG.11B shows method 1105 of training machine learning models forming a hierarchal tumor type tree according to embodiments of the present disclosure.

[0020] FIG.11C shows a method of determining a tumor type using a hierarchal tumor type tree according to embodiments of the present disclosure

[0021] FIG.12 illustrates a measurement system according to an embodiment of the present disclosure.

[0022] FIG.13 shows a block diagram of an example computer system usable with systems and methods according to embodiments of the present disclosure.

[0023] FIGs.14A-14B show cohort baseline demographics for a study using an example tissue typing system as provided herein, termed GPSai. See Examples for further details. Shown are age (FIG.14A) and sex (FIG.14B) of all patients in training, validation, and CUP cohorts. Horizontal midline represents the median in FIG.14A.

[0024] FIGs.15A-15D show calculation of hierarchical positive predictive value (hPPV) and hierarchical sensitivity (hSens). Hierarchical metrics were calculated per sample with values ranging from 0-1.0, and then averaged across all samples. The schematic shows level predictions where the “root node” is the broadest category, and not counted toward the calculation. FIG.15A illustrates 100% hPPV and hSens; FIG.15B illustrates 100% hPPV and 67% hSens; FIG.15C illustrates 67% hPPV and 67% hSens; and FIG.15D illustrates 100% hPPV and 100% hSens, where the model made a more granular prediction than “truth” diagnosis. For example, a newly received specimen may have a diagnosis of “metastatic carcinoma, consistent with known breast primary” without specifying the subtype. Most likely the cancer was subtyped, but the information was not provided by the ordering physician.

[0025] FIG.16 illustrates model scoring using the example GPSai system. A representative example of scoring is shown. Subcategories (e.g., lung adenocarcinoma, lung squamous cell carcinoma) can sum to unity of the major category score (e.g., non-small cell lung carcinoma).

[0026] FIG.17 illustrates threshold selection for the example GPSai system provided herein. The dashed line at 0.55 indicates the threshold determined to optimize both CUP call rate (indicated) and hierarchical positive predictive value (hPPV, indicated) for metastatic cases. A GPS score ≥0.55 indicates high confidence in the tumor type prediction and will be included on the molecular report provided to the ordering physician.

[0027] FIG.18 illustrates cases where the GPSai call differed from outside submitted diagnosis. Confirmation of diagnosis by select orthogonal methods are shown.

[0028] FIGs.19A-B illustrate diagnosis and Level 1 targeted-therapy recommendation changes prompted by the example GPSai system. FIG.19A shows percentage of cases with changes in Level 1 targeted therapy recommendations among all cases with diagnosis change prompted by GPSai. Although some cases were eligible / ineligible for more than one drug, no more than one drug per case was included in this analysis. FIG.19B shows a breakdown of the technologies used to identify the biomarker for Level 1 targeted-therapy treatment associations. If drug eligibility could be driven by whole exome sequencing (WES) / copy number alterations (CNA) or whole transcriptome sequencing (WTS), that drug was chosen first, then immunohistochemistry (IHC)-driven drugs. If no biomarker-driven drug association changes were present, diagnosis-only changes were recorded.

[0029] FIG.20 shows changes in drug eligibility based on example GPSai results. Based on diagnosis + biomarker or diagnosis-only, shown are all cases with added eligibility or ineligibility for the indicated drugs.

[0030] FIG.21 shows the example GPSai system as part of physician-in-the-loop paradigm for management of CUP and diagnostically ambiguous metastatic tumors. TERMS

[0031] Nucleic acids include deoxyribonucleotides or ribonucleotides and polymers thereof in either single- or double-stranded form, or complements thereof. Nucleic acids can contain known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, and non-naturally occurring, which have similar binding properties as the reference nucleic acid, and which are metabolized in a manner similar to the reference nucleotides. Examples of such analogs include, without limitation, phosphorothioates, phosphoramidates, methyl phosphonates, chiral-methyl phosphonates, 2-O-methyl ribonucleotides, peptide-nucleic acids (PNAs). Nucleic acid sequence can encompass conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res.19:5081 (1991); Ohtsuka et al., J. Biol. Chem.260:2605-2608 (1985); Rossolini et al., Mol. Cell Probes 8:91-98 (1994)). The term nucleic acid is sometimes used interchangeably with oligonucleotide, and polynucleotide. Examples of nucleic acids include without limitation genomic DNA (gDNA), genes, copy DNA (cDNA), RNA, messenger RNA (mRNA) transcripts, circulating or cell- free nucleic acids (cfNA, including total (cfTNA), DNA (cfDNA) and RNA (cfRNA)).

[0032] A sample as used herein includes any relevant biological sample that can be used for molecular profiling, e.g., sections of tissues such as biopsy or tissue removed during surgical or other procedures, bodily fluids, autopsy samples, and frozen sections taken for histological purposes. Any biopsy technique known in the art can be applied. The biopsy technique applied can depend on the tissue type to be evaluated (e.g., colon, prostate, kidney, bladder, lymph node, liver, bone marrow, blood cell, lung, breast, etc.), the size and type of the tumor (e.g., solid or suspended, blood or ascites), among other factors. Representative biopsy techniques include, but are not limited to, excisional biopsy, incisional biopsy, needlebiopsy (e.g., core-needle or fine-needle), surgical biopsy, and bone marrow biopsy. Such samples can also include blood and blood fractions or products (e.g., serum, buffy coat, plasma, platelets, red blood cells, and the like), sputum, malignant effusion, cheek cells tissue, cultured cells (e.g., primary cultures, explants, and transformed cells), stool, urine, other biological or bodily fluids (e.g., prostatic fluid, gastric fluid, intestinal fluid, renal fluid, lung fluid, cerebrospinal fluid, and the like), etc. The sample can comprise biological material that is a fresh frozen & formalin fixed paraffin embedded (FFPE) block, formalin-fixed paraffin embedded, or is within an RNA preservative + formalin fixative. The sample can comprise a fixed tumor sample. Unless otherwise noted, a “sample” as referred to herein for molecular profiling of a patient may comprise more than one physical specimen. As still another non-limiting example, a molecular profile may be generated for a subject using a “sample” comprising a solid tumor specimen and a bodily fluid specimen. In some embodiments, a sample is a unitary sample, i.e., a single physical specimen. More than one sample of more than one type can be used for each patient.

[0033] The term “sequence analysis” as used herein refers to determining a nucleotide sequence. The entire sequence or a partial sequence of a polynucleotide, e.g., DNA or mRNA, can be determined, and the determined nucleotide sequence can be referred to as a “read” or “sequence read.” Any suitable sequencing method can be used to detect, and determine the amount of, nucleotide sequence species, amplified nucleic acid species, or detectable products generated from the foregoing. Examples of sequencing platforms include, without limitation, the 454 platform (Roche) (Margulies, M. et al.2005 Nature 437, 376- 380), Illumina Genomic Analyzer (or Solexa platform), MiSeq, HiSeq, NextSeq and Novaseq (Illumina, Inc, San Diego, CA), SOLID System (Applied Biosystems; see PCT patent application publications WO 06 / 084132 entitled “Reagents, Methods, and Libraries For Bead-Based Sequencing” and WO07 / 121,489 entitled “Reagents, Methods, and Libraries for Gel-Free Bead-Based Sequencing”), the Helicos True Single Molecule DNA sequencing technology (Harris TD et al.2008 Science, 320, 106-109), the single molecule, real-time (SMRT™) technology of Pacific Biosciences, and nanopore sequencing (Soni G V and Meller A.2007 Clin Chem 53: 1996-2001), Ion semiconductor sequencing (Ion Torrent Systems, Inc, San Francisco, CA), or DNA nanoball sequencing (Complete Genomics, Mountain View, CA), VisiGen Biotechnologies approach (Invitrogen) and polony sequencing. Such platforms allow sequencing of many nucleic acid molecules isolated from a specimen at high orders of multiplexing in a parallel manner (Dear Brief Funct Genomic Proteomic 2003;1: 397-416; Haimovich, Methods, challenges, and promise of next-generation sequencing in cancer biology. Yale J Biol Med.2011 Dec;84(4):439-46). These non-Sanger-based sequencing technologies are commonly referred to as next-generation sequencing, NGS, NextGen sequencing, next generation sequencing, and variations thereof. These platforms can allow sequencing of clonally expanded or non-amplified single molecules of nucleic acid fragments. Certain platforms involve, for example, sequencing by ligation of dye-modified probes (including cyclic ligation and cleavage), pyrosequencing, and single-molecule sequencing. See, e.g., Satam et al., Next-Generation Sequencing Technology: Current Trends and Advancements, Biology (Basel).2023 Jul 13;12(7):997; Satam et al., Correction: Satam et al. Next-Generation Sequencing Technology: Current Trends and Advancements. Biology 2023, 12, 997, Biology (Basel).2024 Apr 24;13(5):286.

[0034] In various embodiments, at least 1,000, 5,000, 10,000 or 50,000 or 100,000 or 500,000 or 1,000,000 or 5,000,000 nucleic acid molecules, or more, can be analyzed. At least a same number of sequence reads can be analyzed. Examples sizes of a sample used for sequencing can include 30, 50, 100, 200, 300, 500, 1,000, 5,000, or 10,000 or more nanograms, or 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 ml.

[0035] A molecular profile as used herein can refer to molecular characteristics of a biological sample, including without limitation genetic, epigenetic, proteomic, and / or expression characteristics of molecules in the biological sample. The biological sample may be a biopsy of a tumor. A summary and / or analysis of a patient / disease state based on a molecular profile may be included in a molecular profiling report, as further described herein.

[0036] The term “phenotype” as used herein can mean any trait or characteristic that can be identified in part or in whole by using the systems and / or methods provided herein. In some embodiments, the systems can include one or more computer programs on one or more computers in one or more locations, e.g., configured for use in a method described herein. Phenotypes may be determined by analyzing a biological sample obtained from a subject. Phenotypes to be characterized can be any phenotype of interest, including without limitation a tissue, tissue origin, tumor type, anatomical origin of a sample (or sample origin), medical condition, ailment, disease, disorder, or useful combinations thereof. A phenotype can be any observable characteristic or trait of, such as a disease or disorder, a stage of a disease or disorder, susceptibility to a disease or disorder, prognosis of a disease stage or disorder, a physiological state, or response / potential response (or lack thereof) to interventions such astherapeutics, e.g., a stage or grade of a tumor, or a tumor type / origin (e.g., tissue origin). A phenotype can result from a subject’s genetic makeup as well as the influence of environmental factors and the interactions between the two, as well as from epigenetic modifications to nucleic acid sequences. In various embodiments, a phenotype in a subject is characterized by obtaining a biological sample from a subject and analyzing the sample using the systems and / or methods provided herein. For example, characterizing a phenotype for a subject or individual can include detecting a disease or disorder (including pre-symptomatic early-stage detection), determining a prognosis, diagnosis, or theranosis of a disease or disorder, or determining the stage or progression of a disease or disorder. Characterizing a phenotype can include identifying appropriate treatments or treatment efficacy for specific diseases, conditions, disease stages and condition stages, predictions and likelihood analysis of disease progression, particularly disease recurrence, metastatic spread or disease relapse. A phenotype can also be a clinically distinct type or subtype of a condition or disease, such as a cancer or tumor. Phenotype determination can also be a determination of a physiological condition, or an assessment of organ distress or organ rejection, such as post-transplantation. The compositions and methods described herein allow assessment of a subject on an individual basis, which can provide benefits of more efficient and economical decisions in treatment.

[0037] A medical condition as used herein can refer to a disease or disorder of a subject, as well as a stage or severity of the medical condition.

[0038] A “sample type” as used herein includes a tissue or anatomical origin of a biological sample, also referred to as an origin location of a sample, e.g., a particular organ or tissue of origin (TOO). A sample type may be referred to as the sample origin in some contexts. Examples of such sample types include without limitation tissue types including muscle, epithelial, connective tissue, nervous tissue, or any combination thereof. As further examples, the anatomical origin can be the stomach, liver, small intestine, large intestine, rectum, anus, lungs, nose, bronchi, kidneys, urinary bladder, urethra, pituitary gland, pineal gland, adrenal gland, thyroid, pancreas, parathyroid, prostate, cervix / uterine (such as cervix / uterine carcinoma), heart, blood vessels, lymph node, bone marrow, thymus, spleen, skin, tongue, nose, eyes, ears, teeth, uterus, vagina, testis, penis, ovaries, breast, mammary glands, brain, spinal cord, nerve, bone, ligament, tendon, or any combination thereof. A sample type can include its histological type (also referred to as histology herein), which refers to tissue appearance, organization, and function at a microscopic level.

[0039] A “tumor type” refers to the sample type of a tumor sample and may also be referred to as the tumor origin and the like. A tumor type may include subtypes, e.g., a specific organ in a grouping (e.g., stomach), as well as its histology (e.g., carcinoma, adenocarcinoma, Pheochromocytom, atypical, benign, lymphoma (Hodgkin or non-hodgkin), types of cell (such as natural killer cells, T-cells, B-cells, mature, early, etc.), hyperplasia, and neoplasm. The histological type of a cancer refers to the type of tissue in which the cancer originates. There are hundreds of different types of cancers when histology is considered. Major categories include: 1) carcinoma, which originates from the epithelial tissue; 2) sarcoma, which originates in supportive and connective tissue; 3) myeloma, which originates in plasma cells in bone marrow; 4) leukemia, liquid / blood cancers originating in the bone marrow; 5) lymphoma, which develops in glands and nodes of the lymphatic system; and 6) mixed types (e.g., adenosquamous, carcinoma, mixed mesodermal tumor, or carcinosarcoma teratocarcinoma). See, e.g., FIGs.9A-9C and related discussion for a classification of tumor types provided herein. A tumor origin may include a grouping of organs, e.g., Esophagus / Stomach, as such cancers may share similar features (e.g., DNA mutations and / or RNA expression levels). For purposes of the invention, a tumor type can include blood cancers that may not have formed a mass (tumor) unless otherwise clear in context.

[0040] Sample types can be classified using a hierarchical tree, e.g., wherein top-level categories may correspond to general organ systems, groups or organ types, and sub- categories are more specific organ systems, groups or organ types, or portions thereof. A hierarchical tree for classifying tumor types, which may be referred to as a “hierarchal tumor type tree” or “tumor tree,” with tumor origin as a first level in the tree. A first-level or top- level tumor type can have multiple subtypes, referred to as second-level or lower-level tumor types. Tumor types may include histology. For example, ovarian epithelial tumor can be a first level node in the tree, where second-level types can include serous ovarian / fallopian tube / peritoneal, clear cell ovarian cancer, endometrioid ovarian cancer, mucinous ovarian cancer, and / or ovarian carcinosarcoma / malignant mixed mesodermal tumor. A tree can have additional levels, e.g., a third level and a fourth level. For instance, serous ovarian / fallopian tube / peritoneal can include third-level types of high-grade serous ovarian / fallopian tube / peritoneal cancer and low-grade serous ovarian / fallopian tube / peritoneal cancer. An exemplary hierarchal tumor type tree is shown in FIGs.9A-9C. The system and methods provided can be applied to alternate classification schemes, e.g., having different more or less granulated top-levels and / or lower levels.

[0041] A subject (individual, patient, or the like) can be any animal which may benefit from the methods described herein. A subject or individual can be any animal which may benefit from the methods described herein, including, e.g., humans and non-human mammals, such as primates, rodents, horses, dogs and cats. Subjects include without limitation eukaryotic organisms, most preferably a mammal such as a primate, e.g., chimpanzee or human, cow; dog; cat; a rodent, e.g., guinea pig, rat, mouse; rabbit; or a bird; reptile; or fish. Subjects specifically intended for treatment using the methods described herein include humans. A subject may also be referred to herein as an individual or a patient. The subject can have a pre-existing disease or disorder, including without limitation cancer. Alternatively, the subject may not have any known pre-existing condition. The subject may also be non-responsive to an existing or past treatment, such as a treatment for cancer.

[0042] Theranostics as used herein can include therapy-related diagnostic testing that provides the ability to affect therapy or treatment of a medical condition such as a disease or disease state. Theranostics testing provides a theranosis in a similar manner that diagnostics or prognostic testing provides a diagnosis or prognosis, respectively. As used herein, theranostics encompasses any desired form of therapy related testing, including predictive medicine, personalized medicine, precision medicine, integrated medicine, pharmacodiagnostics and Dx / Rx partnering. Therapy related tests can be used to predict and assess drug response in individual subjects, thereby providing personalized medical recommendations. Predicting a likelihood of response can be determining whether a subject is a likely responder or a likely non-responder to a candidate therapeutic agent, e.g., before the subject has been exposed or otherwise treated with the treatment. Assessing a therapeutic response can be monitoring a response to a treatment, e.g., monitoring the subject’s improvement or lack thereof over a time course after initiating the treatment. Therapy related tests are useful to select a subject for treatment who is particularly likely to benefit or lack benefit from the treatment or to provide an early and objective indication of treatment efficacy in an individual subject. Characterization using the systems and methods provided herein may indicate that treatment should be altered to select a more promising treatment, thereby avoiding the expense of delaying beneficial treatment and avoiding the financial and morbidity costs of less efficacious or ineffective treatment(s).

[0043] Theranosis can comprise predicting a treatment efficacy or lack thereof, classifying a patient as a responder or non-responder to treatment. A predicted “responder” can refer to a patient likely to receive a benefit from a treatment whereas a predicted “non-responder” canbe a patient unlikely to receive a benefit from the treatment. Unless specified otherwise, a benefit can be any clinical benefit of interest, including without limitation cure in whole or in part, remission, or any improvement, reduction or decline in progression of the condition or symptoms. The theranosis can be directed to any appropriate treatment, e.g., the treatment may comprise at least one of chemotherapy, immunotherapy, targeted cancer therapy, a monoclonal antibody, small molecule, surgery, radiation, or any useful combinations thereof.

[0044] A classification of a medical condition can include a diagnosis, prognosis, or theranosis of the subject, including a tumor type, e.g., tumor origin. The term “classification” as used herein refers to any number(s) or other characters(s) that are associated with a particular property of a sample, e.g., a medicinal condition of a subject from whom the sample was obtained. For example, a “-“ symbol (or the word “negative”) or a “+” symbol (or the word “positive”) coulda sample is classified as having deletions or amplifications, respectively. The classification can be binary (e.g., positive or negative) or have more levels of classification (e.g., a scale from 1 to 10 or 0 to 1), including probabilities. Different techniques for determining a classification can be combined to obtain a final classification from the initial or intermediate classification for each of the different techniques, e.g., by majority vote or a requirement that all initial / intermediate classifications are the same (e.g., positive).

[0045] The terms “genetic variant,” “nucleotide variant,” and “mutation” may be used herein interchangeably to refer to changes or alterations to the reference human gene sequence at a particular locus, including, but not limited to, nucleotide base point mutations, polymorphisms, copy number variations (e.g., deletions and amplifications, such as duplications), insertions, inversions, substitutions, translocations, rearrangements, fusions, breaks, or repeats (e.g., microsatellite instability) in the coding and non-coding regions, or other types of sequence variants. Certain mutations may be detectable on a given molecule, e.g., using long read techniques, such as single molecule sequencing such as Single-molecule real-time (SMRT) sequencing and nanopore sequencing. Deletions may be of a single nucleotide base, a portion or a region of the nucleotide sequence of the gene, or of the entire gene sequence. Insertions may be of one or more nucleotide bases. The variants may occur in transcriptional regulatory regions, untranslated regions of mRNA, exons, introns, exon / intron junctions, etc. Variants can potentially result in stop codons, frame shifts, deletions of amino acids, altered gene transcript splice forms or altered amino acid sequence.

[0046] An allele or gene allele comprises generally a naturally occurring gene having a reference sequence or a gene containing a specific nucleotide variant.

[0047] A haplotype refers to a combination of genetic (nucleotide) variants in a region of an mRNA or a genomic DNA on a chromosome found in an individual. Thus, a haplotype includes a number of genetically linked polymorphic variants which are typically inherited together as a unit.

[0048] The term “locus” refers to a specific position or site in a gene sequence or protein. Thus, there may be one or more contiguous nucleotides in a particular gene locus, or one or more amino acids at a particular locus in a polypeptide. Moreover, a locus may refer to a particular position in a gene where one or more nucleotides have been deleted, inserted, or inverted.

[0049] A “biomarker” or “marker” can comprise a gene and / or gene product depending on the context. “Biomarkers” or “sets of biomarkers” can be used to train and test machine learning models and classify samples. Particular biomarkers may be used, such as particular nucleic acids (e.g., DNA or RNA, such as mRNA and microRNA), proteins, lipids, carbohydrates and metabolites and optionally also include a state of such molecules. Examples of the state of a biomarker include various aspects that can be queried such as presence, level (quantity, concentration, etc., such as an expression level), sequence, location, activity, structure, modifications, covalent or non-covalent binding partners, and the like. As a non-limiting examples, a set of biomarkers may include a gene or gene product (i.e., mRNA or protein) having a specified sequence (e.g., KRAS mutant), and / or a gene or gene product and a level thereof (e.g., amplified ERBB2 gene or overexpressed HER2 protein), all of which may be circulating biomarkers that are detectable in body fluids, such as blood, plasma, and serum. Various numbers of biomarkers can be used to generate an input data structure (also referred to as an input feature vector), e.g., at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, or 100,000 biomarkers can be used.

[0050] Unless specified otherwise or understood by one of skill in art, the terms “polypeptide,” “protein,” and “peptide” are used interchangeably herein to refer to an amino acid chain in which the amino acid residues are linked by covalent peptide bonds. The amino acid chain can be of any length of at least two amino acids, including full-length proteins. Unless otherwise specified, polypeptide, protein, and peptide also encompass variousmodified forms thereof, including but not limited to glycosylated forms, phosphorylated forms, etc. A polypeptide, protein or peptide can also be referred to as a gene product.

[0051] The terms “label” and “detectable label” in the physical context can refer to any composition detectable by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical, chemical or similar methods. Such labels include biotin for staining with labeled streptavidin conjugate, magnetic beads (e.g., DYNABEADS™), fluorescent dyes (e.g., fluorescein, Texas red, rhodamine, green fluorescent protein, and the like), radiolabels (e.g., 3H, 125I, 35S, 14C, or 32P), enzymes (e.g., horse radish peroxidase, alkaline phosphatase and others commonly used in an ELISA), and calorimetric labels such as colloidal gold or colored glass or plastic (e.g., polystyrene, polypropylene, latex, etc) beads. Patents teaching the use of such labels include U.S. Pat. Nos.3,817,837; 3,850,752; 3,939,350; 3,996,345; 4,277,437; 4,275,149; and 4,366,241. Means of detecting such labels are well known to those of skill in the art. Thus, for example, radiolabels may be detected using photographic film or scintillation counters, fluorescent markers may be detected using a photodetector to detect emitted light. Enzymatic labels are typically detected by providing the enzyme with a substrate and detecting the reaction product produced by the action of the enzyme on the substrate, and calorimetric labels are detected by simply visualizing the colored label. Labels can include, e.g., ligands that bind to labeled antibodies, fluorophores, chemiluminescent agents, enzymes, and antibodies which can serve as specific binding pair members for a labeled ligand. An introduction to labels, labeling procedures and detection of labels is found in Polak and Van Noorden Introduction to Immunocytochemistry, 2nd ed., Springer Verlag, NY (1997); and in Haugland Handbook of Fluorescent Probes and Research Chemicals, a combined handbook and catalogue Published by Molecular Probes, Inc. (1996).

[0052] Detectable labels include, but are not limited to, nucleotides (labeled or unlabelled), compomers, sugars, peptides, proteins, antibodies, chemical compounds, conducting polymers, binding moieties such as biotin, mass tags, calorimetric agents, light emitting agents, chemiluminescent agents, light scattering agents, fluorescent tags, radioactive tags, charge tags (electrical or magnetic charge), volatile tags and hydrophobic tags, biomolecules (e.g., members of a binding pair antibody / antigen, antibody / antibody, antibody / antibody fragment, antibody / antibody receptor, antibody / protein A or protein G, hapten / anti-hapten, biotin / avidin, biotin / streptavidin, folic acid / folate binding protein, vitamin B12 / intrinsic factor, chemical reactive group / complementary chemical reactive group (e.g.,sulfhydryl / maleimide, sulfhydryl / haloacetyl derivative, amine / isotriocyanate, amine / succinimidyl ester, and amine / sulfonyl halides) and the like.

[0053] In the machine learning context, a “label” can refer to a known classification of a training sample. For example, a label can include a first-level classification (such as a tumor origin) and potentially one or more additional classifications for one or more additional layers when a tumor type has multiple subtypes.

[0054] The terms “primer”, “probe,” and “oligonucleotide” may be used herein interchangeably to refer to a relatively short nucleic acid fragment or sequence. They can comprise DNA, RNA, or a hybrid thereof, or chemically modified analog or derivatives thereof. Typically, they are single-stranded. However, they can also be double stranded having two complementing strands which can be separated by denaturation. Normally, primers, probes and oligonucleotides have a length of from about 8 nucleotides to about 200 nucleotides, preferably from about 12 nucleotides to about 100 nucleotides, and more preferably about 18 to about 50 nucleotides. They can be labeled with detectable markers or modified using conventional manners for various molecular biological applications.

[0055] The term “isolated” when used in reference to nucleic acids (e.g., genomic DNAs, cDNAs, mRNAs, or fragments thereof) is intended to mean that a nucleic acid molecule is present in a form that is substantially separated from other naturally occurring nucleic acids and / or biological materials that are normally associated with the molecule. Because a naturally existing chromosome (or a viral equivalent thereof) includes a long nucleic acid sequence, an isolated nucleic acid can be a nucleic acid molecule having only a portion of the nucleic acid sequence in the chromosome but not one or more other portions present on the same chromosome. An isolated nucleic acid can include naturally occurring nucleic acid sequences that flank the nucleic acid in the naturally existing chromosome (or a viral equivalent thereof). An isolated nucleic acid can be substantially separated from other naturally occurring nucleic acids that are on a different chromosome of the same organism. An isolated nucleic acid can be from extracellular material, such as nucleic acid fragments in a bodily fluid. An isolated nucleic acid can also be a composition in which the specified nucleic acid molecule is significantly enriched so as to constitute at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or at least 99% of the total nucleic acids in the composition.

[0056] An isolated nucleic acid can be a hybrid nucleic acid having the specified nucleic acid molecule covalently linked to one or more nucleic acid molecules that are not the nucleic acids naturally flanking the specified nucleic acid. For example, an isolated nucleic acid can be in a vector. In addition, the specified nucleic acid may have a nucleotide sequence that is identical to a naturally occurring nucleic acid or a modified form or mutein thereof having one or more mutations such as nucleotide substitution, deletion / insertion, inversion, and the like.

[0057] An isolated nucleic acid can be extracted from a tissue sample or bodily fluid, prepared from a recombinant host cell (in which the nucleic acids have been recombinantly amplified and / or expressed), or can be a chemically synthesized nucleic acid having a naturally occurring nucleotide sequence or an artificially modified form thereof.

[0058] For the purpose of comparing two different nucleic acid or polypeptide sequences, one sequence (test sequence) may be described to be a specific percentage identical to another sequence (comparison sequence). The percentage identity can be determined by the algorithm of Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 90:5873-5877 (1993), which is incorporated into various BLAST programs. The percentage identity can be determined by the “BLAST 2 Sequences” tool, which is available at the National Center for Biotechnology Information (NCBI) website. See Tatusova and Madden, FEMS Microbiol. Lett., 174(2):247- 250 (1999). For pairwise DNA-DNA comparison, the BLASTN program is used with default parameters (e.g., Match: 1; Mismatch: -2; Open gap: 5 penalties; extension gap: 2 penalties; gap x_dropoff: 50; expect: 10; and word size: 11, with filter). For pairwise protein-protein sequence comparison, the BLASTP program can be employed using default parameters (e.g., Matrix: BLOSUM62; gap open: 11; gap extension: 1; x_dropoff: 15; expect: 10.0; and wordsize: 3, with filter). Percent identity of two sequences is calculated by aligning a test sequence with a comparison sequence using BLAST, determining the number of amino acids or nucleotides in the aligned test sequence that are identical to amino acids or nucleotides in the same position of the comparison sequence, and dividing the number of identical amino acids or nucleotides by the number of amino acids or nucleotides in the comparison sequence. When BLAST is used to compare two sequences, it aligns the sequences and yields the percent identity over defined, aligned regions. If the two sequences are aligned across their entire length, the percent identity yielded by the BLAST is the percent identity of the two sequences. If BLAST does not align the two sequences over their entire length, then the number of identical amino acids or nucleotides in the unaligned regions of the test sequenceand comparison sequence is considered to be zero and the percent identity is calculated by adding the number of identical amino acids or nucleotides in the aligned regions and dividing that number by the length of the comparison sequence. Various versions of the BLAST programs can be used to compare sequences, e.g., BLAST 2.1.2 or BLAST+ 2.2.22.

[0059] The term “mapping” or “aligning” refers to a process that relates a sequence to a location or coordinate (e.g., a genomic coordinate) in a reference (e.g., a reference genome) having a known reference sequence, where the sequence is similar to the known reference sequence at the location in the reference. The degree of similarity can be measured or reported in terms of a “mapping quality.” In one example of a mapping quality used herein, a mapping quality of X for a sequence with respect to a reported location or coordinate in a reference indicates that the probability of the sequence mapping to a different location is no greater than 10^(-X / 10). For instance, a mapping quality of 30 indicates a less than 0.1% probability of the sequence mapping to an alternate location.

[0060] A “reference genome” or “reference sequence” may be an entire genome sequence of a reference organism, one or more portions of a reference genome that may or may not be contiguous, a consensus sequence of many reference organisms, a compilation sequence based on different components of different organisms, or any other appropriate reference sequence. As examples, a reference genome / sequence can be at least 1,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 5,000,000, 10,000,000, 50,000,000, 100,000,000, 500,000,000, one billion, or 3 billion nucleotides long, e.g., a full human genome or a repeat masked human genome. A reference may also include information regarding variations of the reference known to be found in a population of organisms.

[0061] A “machine learning model” (ML model) can refer to a software module configured to be run on one or more processors to provide a classification or numerical value of a property of one or more samples. An ML model can include various parameters (e.g., for coefficients, weights, thresholds, functional properties of function, such as activation functions). As examples, an ML model can include at least 10, 100, 1,000, 5,000, 10,000, 50,000, 100,000, or one million parameters. An ML model can be generated using sample data (e.g., training samples) to make predictions on test data. Various number of training samples can be used, e.g., at least 10, 100, 1,000, 5,000, 10,000, 50,000, 100,000, or at least 200,000 training samples. One example is an unsupervised learning model such as hidden Markov model (HMM),clustering (e.g., hierarchical clustering, k-means, mixture models,model-based clustering, density-based spatial clustering of applications with noise (DBSCAN), and OPTICS algorithm), approaches for learning latent variable models such as Expectation–maximization algorithm (EM), method of moments, and blind signal separation techniques (e.g., principal component analysis, independent component analysis, non- negative matrix factorization, singular value decomposition), and anomaly detection (e.g., local outlier factor and isolation forest). Another example type of model is supervised learning that can be used with embodiments of the present disclosure. Example supervised learning models may include different approaches and algorithms including analytical learning, statistical models, artificial neural network (e.g. including convolutional and / or transformer layers) that may have 1-10 layers as examples, recurrent neural network (e.g., long short term memory, LSTM), boosting (meta-algorithm), bootstrap aggregating (bagging) such as random forests, support vector machine (SVM), support vector (SVR), Bayesian statistics, case-based reasoning, decision tree learning, inductive logic programming, linear regression, logistic regression, Gaussian process regression, genetic programming, group method of data handling, kernel estimators, learning automata, learning classifier systems, minimum message length (decision trees, decision graphs, etc.), multilinear subspace learning, naive Bayes classifier, maximum entropy classifier, conditional random field, nearest neighbor algorithm, probably approximately correct learning (PAC) learning, ripple down rules, a knowledge acquisition methodology, symbolic machine learning algorithms, subsymbolic machine learning algorithms, minimum complexity machines (MCM), ordinal classification, data pre-processing, handling imbalanced datasets, statistical relational learning, or Proaftn (a multicriteria classification algorithm), or an ensemble of any of these types. Supervised learning models can be trained in various ways using various cost / loss functions that define the error from the known label (e.g., least squares and absolute difference from known classification) and various optimization techniques, e.g., using backpropagation, steepest descent, conjugate gradient, and Newton and quasi-Newton techniques.

[0062] The term “about” or “approximately” can mean within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For purposes of the disclosure, “about” refers to ±10% unless otherwise stated.

[0063] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, betweenthe upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within embodiments of the present disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded in the range (e.g., range can be greater than or less than specified number), and each range where either, neither, or both limits are included in the smaller ranges is also encompassed within the present disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the present disclosure. DETAILED DESCRIPTION

[0064] Although a majority of metastatic tumors are classified based on standard pathological and clinical evaluation, a subset is more challenging to diagnose and may present with unclear or potentially incorrect primary organ and histological diagnoses. Among these are ‘cancers of unknown primary’ (CUP), which compose ~2% of cancer cases and are associated with lack of effective treatment options and poor outcomes. Metastatic tumors may also be assigned a diagnostic label but present with diagnostic discrepancies upon further evaluation. The rate of pathology diagnostic discrepancies, including tumor lineage, is estimated to be between 6% and 9%, and a change in treatment plan may be required for over 1 / 3 of cancers with a discrepant diagnosis. As such, accurate diagnosis of pathologically ambiguous cancers is critical to effective treatment planning. Provided herein is an artificial intelligence (AI) system, GPSai, that predicts tumor tissue of origin in CUP and flags potential misdiagnoses for additional workup during routine molecular testing. Such a prediction can also be useful for identifying misdiagnoses either through data entry or other processing errors.

[0065] Biological data (including without limitation, e.g., DNA mutations and / or RNA expression levels) measured from a biological sample can be used to determine an origin of a sample. However, a challenge is the vast number of different tissue types. In the case of a tumor sample, unique combinations of primary tumor site and histology can be in the hundreds or more. Differentiating among such a high number of different tumor types is challenging.

[0066] To address such a problem, embodiments provided herein can use a hierarchical labeling system. A top level in a hierarchal tumor type tree can have labels corresponding toanatomical origin / site / organ or groupings of such, with each node in the tree corresponding to a different origin / site / organ. A top-level type can have multiple subtypes (also referred to as types). A second level type can also have subtypes, corresponding to third level types, and so on.

[0067] A set of machine learning models (ML models or just models) can be used to navigate such a tree to determine which labels apply. The samples can be labeled in such a manner that each machine learning model is trained with a different set of samples, some of which may be subsets of a higher level. Each sample can be labeled with multiple tumor types, if applicable. For example, a sample can be labeled with a top-level type and a second level type, where the second level type has no further subtypes.

[0068] A top-level machine learning model can be trained using all the training samples, using only labels for the top-level types. This top-level machine learning model can provide an output that identifies a most likely top-level type (e.g., a most likely tumor origin). The most likely top-level type (e.g., a first tumor origin) can have multiple subtypes (e.g., greater than or less than 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 50, 60, 70, 80, 90 or 100). The subset of samples labeled with the first tumor origin can then be used to train a second-level machine learning model corresponding to the first tumor origin. More specifically, the second-level machine learning model can be trained using samples labeled with multiple subtypes. This second-level machine learning model can provide high accuracy given it has a defined task to differentiate among a limited number of different tumor types. Such is also true for the top level. Additional machine learning models can be trained for each node that has multiple subtypes.

[0069] Each level for which a type can be determined (e.g., that a probability / likelihood is greater than a threshold) and provided in a report. In some instances, only a top level may be provided, e.g., if a subtype cannot be identified with sufficient accuracy. In this manner, embodiments can provide a clinician at least some useful information even if not the most detailed information. For instance, a broad category of pancreatobiliary is useful, even without defining the specific subtype, e.g., as gall bladder cancer, pancreatic adenocarcinoma, or cholangiocarcinoma, because identifying the tumor type as pancreatobiliary alone can provide useful clinical information such as helping guide the treatment.

[0070] In some embodiments, a sum of probabilities for each classification of a particular ML model can be enforced to sum to 1. And to provide a true probability for each subtype, the determined probabilities for subtypes (child nodes) can be scaled using the probability of a parent node.

[0071] The systems and methods provided herein for identifying tumor type can be applied to predicting sample types more broadly, e.g., healthy samples or those from diseases other than cancer. The different settings require appropriate training samples for use in training the models. I. DYNAMIC TUMOR TYPING USING HIERARCHAL TREE & ML MODELS

[0072] Labels for different tumor types (e.g., 90 diagnostic labels) are arranged in a nested or hierarchical manner to form a tree. A top-level can correspond to tumor origins (e.g., 26 broad diagnostic labels) and more granular diagnostic labels in lower levels. Each node in the tree corresponds to a different label. Each node that has child nodes can have a corresponding ML model for determining which child node (label for tumor type) applies.

[0073] FIG.1 shows a sample type tree 100 according to embodiments of the present disclosure. Sample type tree 100 can be a hierarchal tumor type tree. Each node in sample type tree 100 corresponds to a different sample type, except for a sample node 110. Each solid node has a corresponding ML model. Thus, in the example shown, six ML models would be used to navigate sample type tree 100.

[0074] Each node (except sample node 110) has a corresponding label. A given sample has as many labels (types) as levels that are applicable. For example, a given training sample might be identified as having the 4th-level tumor type corresponding to node 152 with sufficiently high probability (i.e., greater than or equal to a threshold). Such a training sample would have four labels, one for each level. Example values for the threshold can be 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, and 95%.

[0075] A top-level model 115 can determine which of the five possible types in the top level is most likely. If the classification output by the top-level model indicates the label type for a node 121, then a second-level model 125 can determine which 2nd-level type out of the two shown is most likely. If the classification output by the top-level model indicates the label type for node 122, then a different second-level model 127 can determine which 2nd- level type out of the three shown is most likely. Various numbers of total ML models can beused, e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, and 25. Each ML model can have the same architecture (e.g., a neural network) or different architecture type, e.g., one or more of each of SVM, decision tree, and neural network. Each ML model of a same type can have the same or different parameters, e.g., different numbers of nodes and layers; different types of layers, such as convolutional or transformer layers; different learning rates, max epochs, and the like.

[0076] Each of the models is trained using a different set of training samples. Top-level model 115 is trained with all samples having the five top-level labels. Second-level model 125 is trained using the subset of training samples with the label (tumor type) corresponding to node 121. More specifically, second-level model 125 can be trained using samples labeled with types corresponding to nodes 132 and 134. Whereas second-level model 127 is trained using the subset of training samples with the label (tumor type) corresponding to node 122. Thus, a ML model is trained for each parent node that has child nodes and is trained using the training samples having the label corresponding to the parent node. A. First level (tumor origin)

[0077] FIG.2 shows a diagram of a top-level ML model classifying among a set of top- level types according to embodiments of the present disclosure. Biological data 205 is provided to a top-level ML model 210. The biological data can correspond to any of the biomarkers described herein, e.g., including one or more DNA mutations (e.g., sequence variants and / or copy number variants) and / or one or more RNA expression levels. Other data may also be used, e.g., demographic data of the subject and personal characteristics, such as sex, height, weight, and age. Such data can form an input feature vector that is used for each ML model. Different models can have different input feature vectors. An input feature vector can be referred to as a biosignature.

[0078] Top-level ML model 210 can output probabilities corresponding to each of the possible N top-level types, e.g., tumor origins 1-N. In other embodiments, a model can output only the most likely tumor type for a given level. If a particular tumor origin has subtypes, then a 2nd-level model can be used. As shown, tumor origin 2 is assigned 2nd-level ML model A, and tumor origin 3 is assigned 2nd-level ML model B. An ML model selector can receive a given label and select the corresponding ML model. For instance, the ML model selector could receive the label for tumor origin 2 and select 2nd-level ML model A from a set of ML models.

[0079] Example numbers of nodes in the top-level layer include at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 50, 60, 70, 80, 90 or 100. Such are example values for N sample (e.g., tumor) origins.

[0080] Non-limiting examples of top-level types include prostate, bladder, endocervix, peritoneum, stomach, esophagus, ovary, parietal lobe, cervix, endometrium, liver, sigmoid colon, upper-outer quadrant of breast, uterus, pancreas, head of pancreas, rectum, colon, breast, intrahepatic bile duct, cecum, gastroesophageal junction, frontal lobe, kidney, tail of pancreas, ascending colon, descending colon, gallbladder, appendix, rectosigmoid colon, fallopian tube, brain, lung, temporal lobe, lower third of esophagus, upper-inner quadrant of breast, transverse colon, and skin. B. Additional levels

[0081] FIG.3 illustrates a selection and use of a second-level ML model according to embodiments of the present disclosure.

[0082] A top-level ML model 310 processes biological data 305 to obtain an output classification of tumor origin 320, as a top-level type. An ML model selector 330 uses the output classification to retrieve, from a database 340, an ML model corresponding to the tumor origin 320.

[0083] Database 340 can store each of set of ML models with an association of a label (type) for a particular level that corresponds to the ML model. Then, when a particular output classification (type) is obtained with sufficient confidence (e.g., probability is greater than a threshold), then the type can be used to retrieve the corresponding ML model that can indicate which subtype is most likely.

[0084] In the example shown, second-level ML model 350 is identified as having the highest likelihood. Second-level ML model 350 can process biological data 305 (or a portion of such data and / or other biomarkers) to obtain probabilities for M tumor types that are possible in the tree for tumor origin 320.

[0085] Additional levels can be analyzed in a similar manner. For example, tumor type M may have multiple child nodes (i.e., further subtypes), and thus tumor type M can be used by ML model selector 330 to select a corresponding third-level ML model to determine which further subtype is most likely.

[0086] Example numbers of nodes in a lower level layer include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 50, 60, 70, 80, 90 or 100. Such are example values for M second-level sample (e.g., tumor) types. Non-limiting examples of top-level types and second level types include those shown in FIGS.9A-C. C. Adjustments for additional levels

[0087] In some implementations, the probabilities output by a particular ML model are constrained to sum to 1. Thus, the sum of the probabilities at the top-level sum to 1. And each second-level model outputs probabilities that sum to 1. This may occur when neural networks are used.

[0088] A constraint of summing to 1 can introduce problems, e.g., when all ML models are run. Suppose that a patient has a pancreatobiliary cancer in truth, but when the second-level model for bladder / urinary tract (which should have a probability of 0) is run, these zero or near-zero probabilities could be forced to add up to one. Embodiments can reinterpret the output probabilities within the larger scope of the entire tree, e.g., by post-processing the model outputs. For instance, the probabilities of a subtype can be adjusted using the probability of a top-level type. For instance, the probabilities of the subtypes can be scaled by the probability of the top-level type or normalized to sum to the top-level type probability. Alternatively, if a particular top-level probability is below a threshold, a second-level machine learning model may not be executed, even though multiple subtypes are possible.

[0089] Some embodiments can connect different models, even though each one of the nodes with children corresponds to an independent model. For the connections, models can be integrated together by propagating probability adjustments down the tree. For example, this can allow the probabilities presented to pathologists to be interpretable in a wider scope and setting to help clinical decisions.

[0090] For instance, the top-level type of bladder / urinary tract cancer can have two child nodes for urothelial carcinoma and bladder adenocarcinoma. The output of the second-level model can be Bladder=.7, Urothelial=.3, even though the bladder / urinary tract has a zero probability. To fix this, embodiments can adjust the raw model scores to produce a final adjust model score. The adjustment can scale the raw model scores using the probability of the parent node so that the sum of all subtype model scores always equals 1 across all modelsfor a given level. In this manner, model scores from each model are adjusted so that model scores of children always sum to the model score of the children’s’ parent label / node. As an alternative fix, if the probability of bladder / urinary tract is below a threshold, the corresponding second-level machine learning model may not be executed.

[0091] As an example, suppose the top-level probabilities are: Lung 90%, Bladder 10%, all other major labels = 0%. Then, suppose running this case through the lung second-level model gives the following raw model scores: Lung Adenocarcinoma (LUAD): 80%, Lung Squamous Cell Carcinoma (LUSC): 20%. And suppose running through the Bladder model gives the raw model scores: Bladder Adenocarcinoma: 70%, Urothelial 30%. A score adjustment procedure can perform the following adjustments. Since top level Lung was 90%, multiply by 90% the raw score of each Lung subtype. Since top level Bladder was 10%, multiply by 10% the raw score of each Bladder subtype. This can be depicted as LUAD = 90% * 80%, LUSC = 90% * 20%, Bladder Adeno=10% * 70%, Urothelial=10% *30%. This gives Lung 90%, LUAD 72%, LUSC 18%, Bladder 10%, Bladder Adeno 7%, and Urothelial 3%. This has the property that all 26 major labels have scores which sum to 1 and the most granular / specific label for each tumor type sum to 1. Specifically, 72 + 18 + 7 + 3 = 100, and 90+10=100. II. SYSTEM FOR DETERMINING TUMOR ORIGIN / TYPE

[0092] Systems can perform the techniques for determining a sample type (e.g., a tumor type) using a hierarchal sample type tree. The description below relates to a single classifier (e.g., a single model or a group of models acting as a single classifier). Such description is applicable to each of the ML models usable for each parent note that has multiple child nodes in the hierarchal sample type tree (e.g., a hierarchal tumor type tree).

[0093] A system can generate a set of one or more training data structures (also referred to as training vectors) that can be used to train a machine learning model to provide various classifications, such as characterizing a phenotype (e.g., sample type) of a biological sample. Besides sample type, characterizing a phenotype can include providing a diagnosis, prognosis, theranosis or other relevant classification. For example, the classification may include a disease state, a predicted efficacy of a treatment for a disease or disorder of a subject, or the anatomical origin (e.g., a tumor type) of a sample based on a molecular profile, e.g., including a set of biomarkers in an input feature vector.

[0094] Once trained, the trained machine learning model can then be used to process input data provided by the system and make predictions based on the processed input data. The input data may include a set of features related to a subject such as data representing a molecular profile including one or more biomarkers, such as DNA mutation(s) and RNA expression level(s), and data representing a phenotype of interest, e.g., a disease and / or anatomical origin. The prediction may include output data based on the machine learning model’s processing of a specific input set of features. The input data (e.g., a profile) may include, without limitation, data representing a molecular profile including one or more subject biomarkers, data representing a disease or anatomical origin (e.g., a tumor type), and data representing a proposed treatment type as desired.

[0095] The training data for a model that determines a tumor type can comprise biological data representing one or more biomarkers and label data representing sample type, e.g., an origin / type of the cells in a sample, such as a tumor sample. The label data can include a known training label of a type (e.g., an origin) of the disease / disorder, such as a primary site for cancer. Other data that may be used to train a model include data representing a disease or disorder, data representing a sample (e.g., manner of preparation, type of sample, etc.), or any combination thereof. The system (e.g., a combination of trained models) may then predict an anatomical origin of a biological sample, e.g., a tumor type of a tumor sample. In some implementations, the disease or disorder may include a type of cancer, and the anatomical origins can include various tissues and organs, as well as a particular type of cancer at a particular anatomical origin.

[0096] In some implementations, the output data generated includes a probability of the classification at any level of the tree, e.g., a probability that the biological sample is derived from tissue from a particular organ. A. Basics of training and use of a model

[0097] A system may include a plurality of such ML models. A machine learning model training system may be implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.

[0098] The ML model training system can train a machine learning model using training data items from a database (or data set) of training data items, e.g., comprising trainingvectors (also referred to as reference vectors). The training data items may include a plurality of feature vectors. Each training vector may include a plurality of values that each correspond to a particular feature of a training sample that the training vector represents. The training features may be referred to as independent variables. In addition, the system can maintain a respective weight for each feature that is included in the feature vectors. The database can store a set of multiple training data items, with each training data item in the set of multiple training items being associated with a respective label.

[0099] A ML model (or just model) can receive an input training data item (or training sample) that is processed to generate an output. The input training data item may include a plurality of features, e.g., as vector or data structure (or independent variables “X”), and a training label (or dependent variable “Y”) that can correspond to a tumor type at any level in a tumor type tree. Other examples of training labels include detection, prognosis, and theranosis of disease / disorder (e.g., cancer). The label identifies a correct classification (or prediction) for the training data item, i.e., the classification that should be identified as the classification of the training data item by the output values generated by the model. A trained model can predict Y for a given X, Y = f(X).

[0100] To enable a model to generate accurate outputs for received data items, the system may train the model to adjust the values of the parameters of the model, e.g., to determine trained values of the parameters from initial values. These parameters derived from the training steps may include weights that can be used during the prediction stage using the fully trained model.

[0101] The system can train the machine learning model to optimize an objective function. Optimizing an objective function may include, for example, minimizing a loss function. Generally, the loss function is a function that depends on the (i) output generated by the model by processing a given training data item and (ii) the label for the training data item, i.e., the target output that the model should have generated by processing the training data item.

[0102] The system can train the model to minimize the (cumulative) loss function by performing multiple iterations of conventional machine learning model training techniques on training data items from the database, e.g., hinge loss, stochastic gradient methods, stochastic gradient descent with backpropagation, or the like, to iteratively adjust the values of theparameters of the model. A fully trained machine learning model may then be deployed as a predicting model that can be used to make predictions based on input data that is not labeled.

[0103] Various classification methodologies can be applied to the chosen attributes as desired, including without limitation a neural network model, a linear regression model, a random forest model or other decision tree model, a logistic regression model, a naive Bayes model, a quadratic discriminant analysis model, a K-nearest neighbor model, a support vector machine, clustering, or various forms of or combinations thereof. In some embodiments, the machine learning approach comprises an XGBoost multi-class classification. XGBoost is a decision-tree-based ensemble machine learning algorithm that uses a gradient boosting framework. Combinations of classification methods can be employed.

[0104] For embodiments using a neural network, examples numbers of layers include at least 3, 4, 5, 6, 7, 8, 9, and 10. Example numbers of nodes for a give layer include at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 500, 1,000, 1,500, 2,000, 5,000, 10,000, 20,000, and 50,000. Various activation functions can be used, such as logistic, tanh, ReLU, softmax, and sigmoid.

[0105] Once trained, the ML model can process new inputs (e.g., a sample vector) and provide an output, e.g., a predicted tumor type. A sample vector can be compared to a reference vector by training the ML models. Such a comparison is implicit in the training and input of a new sample vector into the ML model. In other embodiments, the comparison can be more explicit. For example, a representative molecular profile (or representative vector) can be determined from all the training samples having a particular label (e.g., sample type). For instance, an average, median, or majority vote for each value can be used to generate the representative vector that is to act as a reference vector for comparison. A sample vector can be compared to such a representative vector, e.g., in clustering or SVM techniques, to determine which representative vector is closest. For example, centroids of clusters can be used to generate the representative vectors.

[0106] In some embodiments, a system can be trained to employ a hierarchical tree to predict a sample type (e.g., tumor type), wherein the system comprises multiple levels and models as described above. When determining a tumor type, input data can be generated based on the biopsied tumor (e.g., NGS data, as further described herein) and provided as an input to the system. The model can be trained using biological signatures (reference vectors) of training samples having different cancer types. Based on the output generated by thesystem, the computer can determine whether biopsied tumor represented by the input data originated in the biopsy location (e.g., liver) or in some other portion of the subject’s body such as the pancreas. One or more treatments can then be determined based on the predicted tumor type as opposed to basing treatments on the anatomical origin of biopsied tumor, if different from the prediction.

[0107] Optionally the machine learning classification system can comprise a voting module, e.g., where multiple top-level machine learning models are used. For each of the tumor types, the probability can be compared to a threshold, which can be used to determine whether the classification of the tumor type is likely, unlikely, or indeterminate. B. Training system

[0108] FIG.4 is a block diagram of a system for training a machine learning model to predict a sample type (e.g., origin) of a biological sample based on biological data obtained from the biological sample.

[0109] The system 400 includes two or more distributed computers 410, 411, a network 430, and an application server 440. The application server 440 includes an extraction unit 442, a memory unit 444, a vector generation unit 450, and a machine learning model 470. Machine learning model 470 can represent a top-level machine learning model or any one of the lower-level machine learning models to determine a subtype of a sample type (e.g., a tumor type). When machine learning model 470 is a lower-level model, the training samples can only be for samples having a corresponding higher-level type. For instance, a second- level model for a first sample type is trained using samples verified to have the first sample type.

[0110] Each distributed computer 410, 411 may include a smartphone, a tablet computer, laptop computer, or a desktop computer, or the like. Alternatively, the distributed computers 410, 411 may include server computers that receive data input by one or more terminals 405, 406, respectively. The terminal computers 405, 406 may include any user device including a smartphone, a tablet computer, a laptop computer, a desktop computer or the like. The network 430 may include one or more networks 430 such as a LAN, a WAN, a wired ethernet network, a wireless network, a cellular network, the internet, or any combination thereof.

[0111] The application server 440 is configured to obtain, or otherwise receive, biomarker data records 420, 422, 424, provided by one or more distributed computers such as the first distributed computer 410 and the second distributed computer 411 using the network 430. In some embodiments, each respective distributed computer 410, 411 may provide different types of data records 420, 422, 424. For example, the first distributed computer 410 may provide biomarker data records 420, 422, 424 representing biomarkers for a biological sample from a subject and the second distributed computer 411 may provide sample record 414 representing anatomical origin or other sample data for a subject obtained from the sample database 412. However, the present disclosure need not be limited to two computers 410, 411 providing data records 420, 422, 424, 414, e.g., only one computer can be used or more than two.

[0112] The biomarker data records 420, 422, 424 may include any type of biomarker data that describes relevant biological attributes of a biological sample. By way of example, the example of FIG.4 shows the biomarker data records as including data records representing DNA 420, protein 422, and RNA 424. These biomarker data records may each include data structures having fields that structure information 420, 422, 424 describing biomarkers of a subject such as a subject’s DNA 420, protein 422, or RNA 424. As a non-limiting example, the DNA data records 420 may include next generation sequencing data including without limitation single variants / point mutations, insertions and deletions, substitutions, translocations, fusions, breaks, duplications, amplifications, loss, copy numbers, repeats, total mutational burden, microsatellite instability, or the like. Alternatively, or in addition, the RNA records 424 may include RNA data such as gene expression levels or the presence of gene fusions, and / or may include without limitation data derived from whole transcriptome sequencing. Alternatively, or in addition, the protein data records 422 may include protein expression data (e.g., amount and / or localization) such as obtained using immunohistochemistry (IHC). In various embodiments, the biomarker data may be obtained by whole exome sequencing, whole transcriptome sequencing, whole genome sequencing, or a combination thereof. Biomarker data may also be obtained by NGS panels specific for more limited sets of genes of interest. Such panels may be obtained by various NGS library preparation techniques such as targeted amplification or bait capture. In some embodiments, the biomarker data comprises boosted gene panels, wherein the depth of coverage for certain biomarkers of interest (e.g., cancer genes) is boosted relative to other biomarkers. As a non- limiting example, higher bait coverage of such certain biomarkers of interest can be used.

[0113] The patient sample data records 414 may describe various aspects of a biological sample, e.g., a tissue and / or organ from which the sample is derived. For example, the sample data records 414 obtained from the sample database 412 may include one or more data structures having fields that structure data attributes of a biological sample such as a medical condition 414a-1 (e.g., a disease or disorder), a tissue or organ 414a-2 where the sample was obtained, a sample format 414a-3 (e.g., bodily fluid or tissue biopsy), a verified sample type label 414a-4 (e.g., top level or lower-level type(s)), or any combination thereof. The sample record 414 can include up to n data records describing a sample, where n is any positive integer greater than 0. For example, though the example of FIG.4 trains the machine learning model using patient sample data describing medical condition 414a-1, tissue / organ 414a-2 where sample was obtained, and sample format 414a-3, the present disclosure is not so limited. For example, in some implementations, the machine learning model 370 can be trained to predict the sample type (e.g., origin of sample) using patient sample information that includes the tissue or organ 414a-2 where the sample was obtained and sample format 414a-3without including the medical condition 414a-1.

[0114] Alternatively, or in addition, the sample data records 414 may also include fields that structure data attributes describing details of the biological sample, including attributes of a subject from which the sample is derived. An example of a disease or disorder may include, for example, a type of cancer. A tissue or organ may include, for example, a type of tissue (e.g., muscle tissue, epithelial tissue, connective tissue, nervous tissue, etc.) or organ (e.g., colon, lung, brain, etc.). The verified sample type can be obtained by pathology review, and may be confirmed using orthogonal methods (e.g., IHCs). In some embodiments, sample format may be included, such as data representing the format of the patient sample, such as tumor sample, bodily fluid, fresh or frozen, biopsy, FFPE, or the like. In some implementations, attributes of a subject from which the sample is derived include clinical attributes such as pathology details of the sample, subject age and / or sex, prior subject treatments, or the like. If the sample is a metastatic sample of unknown primary origin (i.e., a cancer of unknown primary (CUP)), the attributes may include the location from which the sample was taken. As a non-limiting example, a metastatic lesion of unknown primary origin may be found in the liver or brain. Accordingly, though the example shows that sample data may include a disease or disorder, a tissue or organ, and a sample type, the sample data may include other types of information, as described herein. Moreover, there is no requirements that the sample data be limited to human “patients.” Instead, the biomarker data records 420,422, 424 and sample data records 414 may be associated with any desired subject including any non-human organism.

[0115] In some implementations, each of the data records 420, 422, 424, 414 may include keyed data that enables the data records from each respective distributed computer to be correlated by application server 440. The keyed data may include, for example, data representing a subject identifier. The subject identifier may include any form of data that identifies a subject and that can associate biomarker data for the subject with sample data for the subject.

[0116] The first distributed computer 410 may provide 208 the biomarker data records 420, 422, 424 to the application server 440. The second distributed computer 411 may provide the sample data records 414 to the application server 440. The application server 440 can provide the biomarker data records 420, 422, 424 and the sample data records 414 to the extraction unit 442.

[0117] The extraction unit 442 can process the received biomarker data 420, 422, 424 and sample data records 414 in order to extract data 420a, 422a, 424a, 414a-1, 414a-2, 414a-3 that can be used to train the machine learning model. For example, the extraction unit 442 can obtain data structured by fields of the data structures of the biomarker data records 420, 422, 424, obtain data structured by fields of the data structures of the sample data records 414, or a combination thereof. The extraction unit 442 may perform one or more information extraction algorithms such as keyed data extraction, pattern matching, natural language processing, or the like to identify and obtain data 420a, 422a, 424a, 414a-1, 414a-2, 414a-3 from the biomarker data records 420, 422, 424 and sample data records 414, respectively. The extraction unit 442 may provide the extracted data to the memory unit 444. The extracted data unit may be stored in the memory unit 444 such as flash memory (as opposed to a hard disk) to improve data access times and reduce latency in accessing the extracted data to improve system performance. In some implementations, the extracted data may be stored in the memory unit 444 as an in-memory data grid.

[0118] In more detail, the extraction unit 442 may be configured to filter a portion of the biomarker data records 420, 422, 424 and sample data records 414 such as 420a, 422a, 424a, 414a-1, 414a-2, 414a-3 that will be used to generate an input data structure 460 (also referred to as a sample vector) for processing by the machine learning model 470 from the portion of the sample data 414a that will be used as a label for the generated input data structure 460.Such filtering includes the extraction unit 442 separating the biomarker data and a first portion of the sample data that includes a medical condition 414a-1, tissue / organ 414a-2 where sample was obtained (e.g., biopsied), sample format 414a-3 details, or any combination thereof, from the verified sample type 414a-4 (e.g., origin). The verified sample origin of the sample may be a different tissue / organ or the same tissue / organ than the sample was obtained from such as the case of biopsied metastatic lesions which spread from the primary tumor (i.e., the true origin). The application server 440 can then use the biomarker data 420a, 422a, 424a, and the first portion of the sample data that includes the medical condition 414a-1, tissue or organ 414a-2 (e.g., location of biopsy), sample format 414a-3, or a combination thereof, to generate the input data structure 460. In addition, the application server 440 can use the second portion of the sample data describing the verified sample type 414a-4 oas the label for the generated data structure.

[0119] The application server 440 may process the extracted data stored in the memory unit 444 to correlate the biomarker data 420a, 422a, 424a extracted from biomarker data records 420, 422, 424 with the first portion of the sample data 414a. This correlation serves to cluster biomarker data with sample data so that the sample data for the biological sample is clustered with the biomarker data for the same biological sample. In some implementations, the correlation of the biomarker data and the first portion of the sample data may be based on keyed data associated with each of the biomarker data records 420, 422, 424 and the sample data records 414. For example, the keyed data may include a sample identifier or a subject identifier, e.g., a subject from which the sample is derived.

[0120] The application server 440 provides the extracted biomarker data 420a, 422a, 424a and the extracted first portion of the sample data 414a as an input to a vector generation unit 450. The vector generation unit 250 is used to generate a data structure based on the extracted biomarker data 420a, 422a, 424a and the extracted first portion of the sample data 414a. The generated data structure is a feature vector 460 that includes a plurality of values that numerical represents the extracted biomarker data 420a, 422a, 424a and the extracted first portion of the sample data 414a. The feature vector 460 may include a field for each type of biomarker and each type of sample data. For example, the feature vector 460 may include one or more fields corresponding to (i) one or more types of DNA sequencing data such as single variants / point mutations, insertions and deletions, substitutions, translocations, fusions, breaks, duplications, amplifications, loss, copy numbers, repeats, total mutational burden, microsatellite instability, (ii) one or more types of in situ hybridization data such as DNAcopy number, gene copies, gene translocations, (iii) one or more types of RNA data such as gene expression or gene fusion, (iv) one or more types of protein data such as presence, level or cellular location obtained using immunohistochemistry, and (v) one or more types of sample data such as disease or disorder, sample type, patient details, or the like.

[0121] The vector generation unit 450 is configured to assign a weight to each field of the feature vector 460 that indicates an extent to which the extracted biomarker data 420a, 422a, 424a and the extracted first portion of the sample data 414a includes the data represented by each field. In one implementation, for example, the vector generation unit 450 may assign a ‘1’ to each field of the feature vector that corresponds to a feature found in the extracted biomarker data 420a, 422a, 424a and the extracted first portion of the sample data 414a. In such implementations, the vector generation unit 450 may, for example, also assign a ‘0’ to each field of the feature vector that corresponds to a feature not found in the extracted biomarker data 420a, 422a, 424a and the extracted first portion of the sample data 414a. The output of the vector generation unit 450 may include a data structure such as a feature vector 460 that can be used to train the machine learning model 470.

[0122] The application server 440 can label the training feature vector 460. Specifically, the application server can use the extracted second portion of the sample data 414to label the generated feature vector 460. The label of the training feature vector 460 generated based on the verified sample type 414a-4 can be used to predict the tissue or organ that was the origin for a biological sample represented by the sample record 414 and having medical condition 414a-1 defined by the specific set of biomarkers 420a, 422a, 424a, each of which is described by described in the training data structure 460.

[0123] The application server 440 can train the machine learning model 470 by providing the feature vector 460 as an input to the machine learning model 470. The machine learning model 470 may process the generated feature vector 460 and generate an output 472. The application server 440 can use a loss function 480 to determine the amount of error between the output 472 of the machine learning model 470 and the value specified by the training label, which is generated based on the second portion of the extracted sample data describing the verified sample type 414a-4. The output 482 of the loss function 480 can be used to adjust the parameters of the machine learning model 470.

[0124] In some implementations, adjusting the parameters of the machine learning model 470 may include manually tuning of the machine learning model parameters modelparameters. Alternatively, in some implementations, the parameters of the machine learning model 470 may be automatically tuned by one or more algorithms of executed by the application server 442.

[0125] The application server 440 may perform multiple iterations of the process described above for each sample data record 414 stored in the sample database that correspond to a set of biomarker data 420, 422, 424 for a biological sample. This may include hundreds of iterations, thousands of iterations, tens of thousands of iterations, hundreds of thousands of iterations, millions of iterations, or more, until each of the sample data records 414 stored in the sample database 412 and having a corresponding set of biomarker data for a biological sample are exhausted, until the machine learning model 470 is trained to within a particular margin of error, or a combination thereof. A machine learning model 470 is trained within a particular margin of error when, for example, the machine learning model 470 is able to predict, based upon a set of unlabeled biomarker data along with the desired patient / sample data, an origin of a sample having the biomarker / patient / sample data. The origin may include, for example, a probability, a general indication of the confidence in the origin classification, or the like.

[0126] Embodiments may include a process for generating training data structures for training a machine learning model to predict sample type (e.g., origin). In one aspect, the process may include obtaining, from a first distributed data source, a first data structure that includes fields structuring data representing a set of one or more biomarkers associated with a biological sample, storing the first data structure in one or more memory devices, obtaining from a second distributed data source, a second data structure that includes fields structuring data representing the biological sample and origin data for the biological sample having the one or more biomarkers, storing the second data structure in the one or more memory devices, generating a labeled training data structure that structures data representing (i) the one or more biomarkers, (ii) a biological sample, (iii) an origin, and (iv) a predicted origin for the biological sample based on the first data structure and the second data structure, and training a machine learning model using the generated labeled training data.

[0127] As described herein, the sample types that the models are trained to predict can be tumor types. The tumor types used for training can be categorized into levels as described herein. See, e.g., FIGs.9A-9C and related discussion. For example, ovarian epithelial tumor can be a first level (top level) origin / type, where lower-level origins / types can include serousovarian / fallopian tube / peritoneal, clear cell ovarian cancer, endometrioid ovarian cancer, mucinous ovarian cancer, and / or ovarian carcinosarcoma / malignant mixed mesodermal tumor.

[0128] In some embodiments, the training can determine a reference vector for each tumor type of each node so as to optimize the differentiation between different tumor types of a given level. C. Production system

[0129] FIG.5 is a block diagram of a system for using a trained machine learning model 570 to predict a sample type using biomarker and sample data from a subject. Machine learning model 570 can be any model at any level as described herein, e.g., a top-level model or a second-level model. When a top-level model, machine learning model 570 can be selected first. Then, depending on which sample type is determined to have the highest probability, a corresponding second-level model can be selected for the next selection in the tree. Thus, multiple models may be called using the framework described in FIG.5.

[0130] The machine learning model 570 includes a machine learning model that has been trained using the process described with reference to the system of FIG.4 above. For example, machine learning model 570 was trained to predict sample type (e.g., top level or lower levels) using patient sample data that comprises data representing a tissue / organ where the sample was obtained 414a-2 and a sample format 414a-3. In the example of FIG.4, a medical condition 414a-1 was optionally used to train the model and there may be implementations of the present disclosure where the machine learning model 570 can be trained using additional sample data 414 (e.g., patient age and / or sex). The trained machine learning model 570 is capable of predicting, based on an input feature vector representative of biomarker data 520, 522, 524, the tissue / organ where the sample was collected 514a-1 and other relevant sample data 514a-2 such as sample format.

[0131] The application server 540 hosting the machine learning model 570 is configured to receive unlabeled biomarker data records 520, 522, 524. The biomarker data records 520, 522, 524 include one or more data structures that have fields structuring data that represents one or more biomarkers such as DNA 520a, protein 522a, RNA 524a, or any combination or subcombinations thereof. The received biomarker data records may include various characteristics of the biomarkers such as (i) sequencing data from DNA and / or RNA,including genetic variants, total mutational burden, or the like, (ii) one or more types of in situ hybridization data such as DNA copies, gene copies, gene translocations, (iii) one or more types of RNA data such as gene expression or gene fusion, or (iv) one or more types of protein data such as presence, level or location obtained using immunohistochemistry (IHC), for example, protein expression and / or localization. In some implementations, the biomarker data records 520, 522, 524 include one or more biomarkers and attributes listed in any one of tables herein. However, the present disclosure need not be so limited, and other biomarkers may be used as desired. For example, the biomarker data may be obtained by whole exome sequencing, whole transcriptome sequencing, whole genome sequencing, or a combination thereof. In some embodiments, the input biomarker data records comprise or consist of DNA 520 and RNA 524, obtained via whole exome sequencing and whole transcriptome sequencing, respectively.

[0132] The application server 540 hosting the machine learning model 570 can be configured to receive sample data 514 representing the collection site (where sample was obtained) of the tissue / organ data 514a-1 for the biological sample having biomarkers represented by the received biomarker data records 520, 522, 524. However, as discussed elsewhere herein, due to the potential for disease (e.g., cancer) to spread from, e.g., organ to organ, the tissue / organ 514a-1 where a sample was obtained may not be the actual sample origin.

[0133] In some implementations, the sample data 514 is received or provided by a terminal 506 over the network 530, and the biomarker data is obtained from a first distributed computer 510. The biomarker data may be derived from laboratory machinery used to perform various assays. The sample record 514 can include data representing a tissue / organ where the sample was obtained 514a-1 and other sample data 514a-2. In other implementations, the sample data 514, such as optional sample collection site 514a-1, and the biomarker data 520, 522, 524 may each be received from the terminal 505. For example, the terminal 505 may be user device of a treating physician, an employee or agent of the treating physician, a laboratory testing outfit, or other entity that inputs data representing a sample and a data representing patient attributes for the biological sample. In some implementations, the sample data 514 may include data structures structuring fields of data representing a proposed origin described by a tissue or organ name. In other implementations, the sample data 520 may include data structures structuring fields of data representing more complexsample data such as sample type, age and / or sex of the patient from which the sample is derived, or the like.

[0134] The application server 540 receives the biomarker data records 520, 522, 524 and the sample data 514. The application server 540 provides the biomarker data records 520, 522, 524, the sample data 514a-2 (e.g., biopsy type), and the tissue / organ where the biological sample was obtained 514a-1 to an extraction unit 542 that is configured to extract (i) particular biomarker data such as DNA data 520a, protein data 522a, and RNA expression data 524a, (ii) sample data 514a-2, and (iii) site of collection data 514a-1 from the fields of the biomarker data records 520, 522, 524 and the sample data 514. In some implementations, the extracted data is stored in the memory unit 544 as a buffer, cache or the like, and then provided as an input to the vector generation unit 550 when the vector generation unit 550 has bandwidth to receive an input for processing. In other implementations, the extracted data is provided directly to a vector generation unit 550 for processing. For example, in some implementations, multiple vector generation units 550 may be employed to enable parallel processing of inputs to reduce latency.

[0135] The vector generation unit 550 can generate a data structure such as a feature vector 560 that includes a plurality of fields and includes one or more fields for each type of biomarker data and one or more fields for each type of sample data. For example, each field of the feature vector 560 may correspond to (i) each type of extracted biomarker data that can be extracted from the biomarker data records 520, 522, 524 such as each type of next generation sequencing data, each type of in situ hybridization data, each type of RNA or DNA data, and each type of protein (e.g., immunohistochemistry) data (if any) and (ii) each type of sample data that can be extracted from the sample data records 514 such as each type of disease or disorder, each type of sample, and patient details (e.g., age and / or sex).

[0136] The vector generation unit 550 is configured to assign a weight to each field of the feature vector 560 that indicates an extent to which the extracted biomarker data 520a, 522a, 524a, the extracted sample data 514a-2, and the extracted collection site 514a-1 includes the data represented by each field. In one implementation, for example, the vector generation unit 550 may assign a ‘1’ to each field of the feature vector 560 that corresponds to a feature found in the extracted biomarker data 520a, 522a, 524a, the extracted sample data 514a-2, and the extracted collection site 514a-1. In such implementations, the vector generation unit 550 may, for example, also assign a ‘0’ to each field of the feature vector that corresponds toa feature not found. The output of the vector generation unit 550 may include a data structure such as a feature vector 560 that can be provided as an input to the trained machine learning model 570.

[0137] The trained machine learning model 570 (e.g., a top-level model or a lower-level model) process the generated feature vector 560 based on the adjusted parameters that were determining during the training stage. The output 572 of the trained machine learning model 570 provides an indication of the sample type (e.g., top level origin and any lower levels) for the biological sample, depending on what level of the particular model corresponds. In such implementations, the output 572 may be provided to the terminal 506 using the network 530. The terminal 506 may then generate output on a user interface 507 that indicates a predicted sample type for the biological sample having the biomarkers represented by the feature vector 560.

[0138] In other implementations, the output 572 may be provided to a prediction unit 580 that is configured to decipher the meaning of the output 572. For example, the prediction unit 580 can be configured to map the output 572 to one or more categories of confidence for determining the sample type. Then, the output of the prediction unit 528 can be used as part of message 590 that is provided to the terminal 506 using the network 530 for review by laboratory testing outfit, treating physician, or other appropriate party.

[0139] Accordingly, embodiments may include a process for using multiple trained machine learning models in a tree to predict a sample type for a sample from a subject. In one aspect, the process may include obtaining a data structure representing a set of one or more biomarkers associated with a biological sample, obtaining data representing sample data for the biological sample, optionally obtaining data representing a origin type for the biological sample (e.g., location of biopsy), generating a data structure for input to a machine learning model that structures data representing the biomarkers (and optionally other data), providing the generated data structure as an input to the machine learning model that has been trained to predict sample types (e.g., top level or lower level types in a tree) using labeled training data structures, and obtaining an output generated by the machine learning model based on the machine learning model processing of the provided data structure, and determining a predicted sample type (e.g., tumor type) for the biological sample having the one or more biomarkers based on the obtained output generated by the machine learning model.

[0140] In some implementations, the machine learning models output an indication whether the sample is more likely to be from one origin versus another, instead of or in addition to indicating that the sample is more or less likely to be from a certain origin. For example, the machine learning model may indicate that the sample is more or less likely to be of prostatic origin (i.e., from the prostate), or the machine learning module may indicate whether the sample is most likely derived from the prostate or from the colon. Any such types (e.g., origins) can be so compared.

[0141] As described herein, the sample types that the models are trained to predict can be tumor types. See, e.g., FIGs.9A-9C and related discussion. In such cases, the models in the production system can be used to predict a tumor type for a biological sample using biomarker data and optionally sample data obtained for a patient. If a lower-level origin / type is identified, both the lower levels and upper-level origin / types can be reported. For example, if lower-level tumor type of clear cell ovarian cancer is identified, the upper-level ovarian epithelial tumor can also be output. D. Use of multiple submodels per model

[0142] In some embodiments, a ML model can be an ensemble of multiple submodels. For example, a set of models can each contribute a vote or numerical output, from which an average or majority value can be determined. Accordingly, multiple models can be trained to perform the prediction / classification and the joint predictions can be used to make the classification. In this scenario, each model is allowed to “vote” and the classification receiving the majority of the votes is deemed the winner. As another example, the outputs can be averaged to obtain an overall value (e.g., an averaged probability), where the highest probability can be compared to a threshold to determine if a sample type for a given level can be called, as can be done in the embodiments described above.

[0143] The classification can be any useful classification, e.g., to characterize a phenotype. For example, the classification may provide a diagnosis (e.g., disease or healthy), prognosis (e.g., predict a better or worse outcome), theranosis (e.g., predict or monitor therapeutic efficacy or lack thereof), or other phenotypic characterization (e.g., origin of a CUPs tumor sample).

[0144] In some embodiments, a particular submodel can have its contribution weighted more, e.g., based on historical accuracy. A confidence score for a submodel can be used todetermine the weight. Accordingly, a confidence score indicating that a machine learning mode is historically accurate can be used to boost a value of output data generated by the machine learning model. Similarly, a confidence score indicating that a machine learning model is historically inaccurate can be used to reduce a value of output data generated by the machine learning model. Such boosting or reducing of the value of output data generated by a machine learning model can be achieved, for example, by using the confidence score as a multiplier of less than one for reduction and more than 1 for boosting. Other operations can also be used to adjust the value of output data such as subtracting a confidence score from the value of the output data to reduce the value of the output data or adding the confidence score to the value of the output data to boost the value of the output data. Use of confidence scores to boost or reduce the value of output data generated by the machine learning models is particularly useful when the machine learning models are configured to output probabilities that will be applied to one or more thresholds to determine whether a sample is or is not from an origin (or has a particular sample type) or is from one of two possible origins. This is because using the confidence score to adjust the output of a machine learning model can be used to move a generated output value above or below a class threshold, thereby altering a prediction by a machine learning model based on its historical accuracy. III. TRAINING MODELS

[0145] As described above, a set of ML models is used to navigate a hierarchal type tree. Each ML model is trained using a different set of training samples, although the different sets can overlap or be a subset of another set.

[0146] Every training sample can be used to train the top-level ML model, which can differentiate among the different top-level types (e.g., sample / tumor types). For a given sample origin (e.g., Bladder / Urinary Tract), a corresponding second-level ML model can be trained using only cases labeled with the child labels, e.g., bladder adenocarcinoma or urothelial carcinoma in this example. See FIG.9A. Such a model is designed to differentiate between these two subtypes. A. Generating multiple labels per case

[0147] A problem with identifying a tumor type is that all the unique combinations of primary tumor site and histology can be over 12,000. Instead of a single model differentiating between all 12,000, a hierarchical labeling system is used. A sample can receive multiple labels, e.g., when a type of a second level or lower is applicable. Having the multiple labelsallows independent training of the different models. The different labels can correspond to standard trees such as can be found at oncotree.mskcc.org / # / home. Certain labels / classifications can be combined, depending on what tree is used. For instance, clear cell ovarian and clear cell borderline ovarian are two OncoTree codes / labels. Some embodiments can combine those together as a new label that can be called clear cell ovarian.

[0148] The following process can be used for identifying labels and creating a tree. To create / map samples to a label, a primary tumor site and histology is obtained for a give specimen (sample). A label is determined from a tree (e.g., OncoTree) that may be the final tree or a larger tree than what is finally used. The label is mapped to the final tree that is used. For example, a code (Ovary, Clear cell adenocarcinoma, NOS) can be mapped to CCOV / Clear Cell Ovarian Cancer. With each sample labeled with an OncoTree code, the codes can be grouped into custom label sets. For example, OncoTree codes {CCOV, CCBOV} corresponding to OncoTree labels {Clear Cell Ovarian, Clear Cell Borderline Ovarian} can be mapped to a custom label called “Clear Cell Ovarian”.

[0149] With each specimen now labeled with the most granular / specific label available, the system can map “upwards” in the label hierarchy and produce multiple labels per specimen. For example, specimens labeled with label “Clear Cell Ovarian” can also be labeled as “Ovarian Epithelial Tumor” since Clear Cell Ovarian is a subtype of Ovarian Epithelial.

[0150] As another example, the primary tumor site is “Esophagogastric junction”, and the specimen histology is “Adenocarcinoma, NOS.” The code (Esophagogastric junction, Adenocarcinoma, NOS) maps to OncoTree: GEJ. OncoTree GEJ (Adenocarcinoma of the Gastroesophageal Junction) can map to label Adenocarcinoma of the Gastroesophageal Junction, which is the same label, in this case. Now that this specimen has label “Adenocarcinoma of the Gastroesophageal Junction”, it inherits all upstream labels as well, so this specimen can now get labeled with three distinct labels: Adenocarcinoma of the Gastroesophageal Junction, Esophagogastric Adenocarcinoma, and Esophagus / Stomach. This specimen can be used to train models at multiple levels of specificity.

[0151] As examples, the top level types (e.g., the plurality of primary tumor origins) consists of, comprises, or comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, or all 38 of prostate, bladder, endocervix, peritoneum, stomach, esophagus, ovary, parietal lobe, cervix, endometrium, liver, sigmoid colon, upper-outer quadrant of breast, uterus, pancreas, head ofpancreas, rectum, colon, breast, intrahepatic bile duct, cecum, gastroesophageal junction, frontal lobe, kidney, tail of pancreas, ascending colon, descending colon, gallbladder, appendix, rectosigmoid colon, fallopian tube, brain, lung, temporal lobe, lower third of esophagus, upper-inner quadrant of breast, transverse colon, and skin. B. Independent training for each type having subtypes

[0152] A top level ML model can use all of the training samples in a training set. For example, if 100,000 training samples exist with 26 top level types, then the training can use all 100,000 training samples. However, each of a set of second-level ML models (each corresponding to a different top-level type) are trained with a portion of the training samples that have a respective top-level type.

[0153] As an example for independently training lower-level models, we look at a tree designed to differentiate between thyroid subtypes. In some embodiments, thyroid can have the following hierarchy: (1) first / top level: Thyroid, (2) 2ndlevel: Anaplastic, Hurthle Cell, Medullary, and Well-Differentiated, and (3) 3rdlevel: Follicular and Papillary.

[0154] All thyroid labeled cases are used to train the thyroid second-level ML model. Samples labeled as Papillary and Follicular Thyroid are also labeled as Well Differentiated and also as Thyroid. For the second-level ML model, Follicular and Papillary samples are simply labeled as “Well Differentiated.” The thyroid second-level ML model is trained to differentiate between Anaplastic, Hurthle, Medullar, and Well-Differentiated only.

[0155] A third-level ML model is trained only using samples labeled as well-differentiated thyroid. This third-level ML model differentiates between Follicular and Papillary. Such a third-level model excludes Anaplastic, Hurthle, and Medullary. The second-level model and the third-level model are trained independently of each other and independent of the top-level model.

[0156] If a top-level type (e.g., adrenal) does not have any subtypes, then no second-level ML model would exist.

[0157] In the future if the tree were updated (e.g., to add or remove a sample type at any level), then only the ML model(s) affected by that change would need to be updated. In the thyroid example, if another type of thyroid well-differentiated was added to Follicular and Papillary, then the third-level model would need to be retrained. But if the overall set of training samples did not change (e.g., some of the Follicular and Papillary just received a newlabel for the third level), then the second-level model would not need to be retrained since the training set for the second-level would not change. And if one of the second-level types was removed without any Well-Differentiated having a label changed, then the Well- Differentiated third-level model would not need to be retrained. This can provide an advantage of reducing the computational effort to retrain models, since the entire model system does not need to be retrained. C. Example sets of features

[0158] Various parameters can be used in the input feature vector (also referred to as an input data structure). Such parameters include a set of biomarkers. Other parameters may also be used, e.g., demographic data of the subject and personal characteristics, such as sex, height, weight, and age. Parameters can be encoded in various ways. For example, patient sex can be encoded as 1 or 0.

[0159] The set of biomarkers can include information derived from DNA, including without limitation DNA mutations. Unless otherwise stated or obvious in context, a “mutation” as used herein may comprise any change in a gene or genome as compared to wild type, including without limitation a point mutation, sequence variant, polymorphism, deletion, insertion, indels (i.e., insertions or deletions), substitution, translocation, fusion, break, duplication, amplification, repeat, or copy number variation. In preferred embodiments, such DNA mutations comprise sequence variants and / or copy number variations. The DNA mutations can be limited to those pathogenic (P) or likely pathogenic (LP) mutations only. Example numbers of DNA mutations can be at least 10, 20, 50, 150, 200, 250, 300, 350, 400, 450, 500, 1,000, 2,000, 5,000, 10,000, or 20,000. The existence of a mutation (P / LP) anywhere in a gene can be encoded as 1 and 0 for a lack of a mutation, e.g., wildtype (WT), variant of unknown significance (VUS), and indeterminate variant (IND). In other embodiments, each possibly DNA mutation can be encoded (e.g., one-hot-encoded).

[0160] The models can be trained using any number of desired features, here biomarkers, to achieve the desired level of performance. As will be understood by those of skill in the art, multiple features may provide a more robust prediction, but too many may lead to overfitting. Such parameters can be optimized in the training and testing phases of model development. In embodiments, the features comprise DNA mutations obtained using sequencing technology such as next-generation sequencing (NGS). As desired, NGS can be used for whole genome sequencing (WGS), whole exome sequencing (WES), whole transcriptomesequencing (WTS) or NGS can be used to assess selected sets of biomarkers of interest. In various embodiments, the features comprise known cancer related genes, or subsets thereof such as all genes containing “actionable” mutations, such as those known to affect treatment. In embodiments, the features are selected from any one or more of Tables 1-7 herein. In embodiments, the features are selected from the following genes: ABL1, AKT1, AKT2, AKT3, ALK, AMER1, APC, AR, ARAF, ARID1A, ARID2, ASXL1, ATM, ATR, ATRX, AXIN1, BAP1, BARD1, BCL2, BCL9, BCOR, BLM, BMPR1A, BRAF, BRCA1, BRCA2, BRIP1, BTG1, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCDC6, CCND1, CCND2, CCND3, CD79B, CDC73, CDH1, CDK12, CDK4, CDKN1B, CDKN2A, CDKN2B, CEBPA, CHEK1, CHEK2, CIC, CNOT3, CREBBP, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CYLD, DDR2, DICER1, DNMT3A, EGFR, EP300, ERBB2, ERBB3, ERBB4, ERCC2, ESR1, EXT1, EZH2, FANCA, FANCC, FANCD2, FANCE, FANCF, FANCG, FANCL, FAS, FBXW7, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FLT4, FOXA1, FOXL2, FOXO3, FUBP1, GATA3, GNA11, GNA13, GNAQ, GNAS, GRIN2A, H3F3A, H3F3B, HIST1H3B, HNF1A, HRAS, IDH1, IDH2, IL7R, IRF4, JAK1, JAK2, JAK3, KDM5C, KDM6A, KDR, KEAP1, KIT, KLF4, KMT2A, KMT2C, KMT2D, KRAS, LRP1B, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAX, MED12, MEF2B, MEN1, MET, MITF, MLH1, MPL, MRE11, MSH2, MSH6, MSI, MTOR, MUTYH, MYC, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NPM1, NRAS, NSD1, NSD2, NT5C2, NTRK1, NTRK2, NTRK3, PALB2, PBRM1, PDE4DIP, PDGFRA, PDGFRB, PHOX2B, PIK3CA, PIK3R1, PIK3R2, PIM1, PMS1, PMS2, POLE, POT1, PPARG, PPP2R1A, PRDM1, PRKAR1A, PRKDC, PTCH1, PTEN, PTPN11, RAC1, RAD50, RAD51B, RAF1, RB1, RET, RNF43, ROS1, RUNX1, SDHAF2, SDHB, SDHC, SDHD, SETD2, SF3B1, SMAD2, SMAD4, SMARCA4, SMARCB1, SMARCE1, SMO, SOCS1, SPEN, SPOP, SRC, STAG2, STAT3, STAT5B, STK11, SUFU, SUZ12, TCF7L2, TERT, TET2, TGFBR2, TNFAIP3, TNFRSF14, TP53, TRAF7, TRRAP, TSC1, TSC2, U2AF1, VHL, WRN, WT1, and XPO1. Gene identifiers used herein are those commonly accepted in the scientific community at the time of filing and can be used to look up the genes at various well-known databases such as the HUGO Gene Nomenclature Committee (HNGC; genenames.org), NCBI’s Gene database (ncbi.nlm.nih.gov / gene), GeneCards (genecards.org), Ensembl (ensembl.org), UniProt (uniprot.org), and others.

[0161] The set of biomarkers can include RNA expression levels. Example genes can include those in the GSEA MsigDB C6 Oncogenic gene sets available at www.gsea-msigdb.org / gsea / msigdb / human / genesets.jsp?collection=C6. Example numbers of RNA expression can be at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 500, 1,000, 1,500, 2,000, 5,000, 10,000, 15,000, or 20,000 different RNA transcript expression levels. An expression level can be expressed in various units, e.g., transcripts per million.

[0162] Certain biomarkers may be selected to determine a desired phenotype, such as whether a treatment for a disease or disorder is of likely benefit, or a tumor type (e.g., tumor origin). Examples lists of features can be found in U.S. Patent Publication Nos. 2023 / 0113092, e.g., DNA markers in tables 2-116, 122-125, and 128-129 and RNA markers in tables 117-120, and 2022 / 0093217.

[0163] The selected features may provide optimal sample type prediction (e.g., using feature ranking), although selection may be made so long as the selections retain the ability to meet desired performance criteria, such as but not limited to accuracy of at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or at least 99%. In some embodiments, an importance value can be used for the feature ranking. In some embodiments, the set of biomarkers in a sample / training / reference vector may comprise the top 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the feature biomarkers with the highest importance value. In some embodiments, the sample / training / reference vector comprises the top 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 feature biomarkers with the highest importance value.

[0164] Nucleic acid used for analysis can be isolated from cells in the sample according to standard methodologies. See, e.g., Sambrook et al., 1989. The nucleic acid, for example, may be genomic DNA, circulating or cell-free DNA, fractionated or whole cell RNA, or miRNA acquired from exosomes or cell surfaces. Where RNA is used, it may be desired to convert the RNA to a complementary DNA. In one embodiment, the RNA is whole cell RNA; in another, it is poly-A RNA; in another, it is exosomal RNA. The nucleic acid may be amplified. Depending on the format of the assay for analyzing the one or more genes, the specific nucleic acid of interest is identified in the sample directly using amplification or with a second, known nucleic acid following amplification.

[0165] Numerous techniques for detecting nucleotide variants are known in the art and can be used for the method of this disclosure. The techniques can be protein-based or nucleic acid-based. In either case, the techniques can be sufficiently sensitive so as to accurately detect the small nucleotide or amino acid variations. In some embodiments, a probe is used which is labeled with a detectable marker. In certain applications, the detection may be performed by visual means (e.g., ethidium bromide staining of a gel). Alternatively, the detection may involve indirect identification of the product. Unless otherwise specified in a particular technique described below, any suitable marker known in the art can be used, including but not limited to, chemiluminescence, radioactive scintigraphy of radiolabel or fluorescent label, via a system using electrical or thermal impulse signals (Affymax Technology; Bellus, 1994), radioactive isotopes, fluorescent compounds, biotin which is detectable using streptavidin, enzymes (e.g., alkaline phosphatase), substrates of an enzyme, ligands and antibodies, etc. See Jablonski et al., Nucleic Acids Res., 14:6115-6128 (1986); Nguyen et al., Biotechniques, 13:116-123 (1992); Rigby et al., J. Mol. Biol., 113:237-251 (1977). In preferred embodiments, nucleic acid sequencing techniques such as dye- termination (Sanger) sequencing or next-generation sequencing (NGS) are used to determine nucleotide variants. 1. DNA Mutations

[0166] Example DNA mutations include copy number variation (e.g., deletions and amplifications such as duplications) and sequence variants (e.g., point mutations, polymorphisms, inversions, substitutions, translocations, rearrangements, and the like). For example, 1p19q is indicative of certain cancers such as oligodendriogliomas. A single chromosome loss of 17 is the most frequent early occurrence in ovarian cancer, and 3p deletion in clear cell kidney and trisomy 7 and 17 in papillary renal cancer are established predictors. Chromosome 6 loss, 8 gain is a marker of eye cancers. Her2 gene amplification is observed in breast cancer.

[0167] In some embodiments, each gene can be represented as a binary number indicating the present or absence of a pathogenic (P) mutation or a likely pathogenic mutation (LP). Non-binary encoding can be used to differentiate mutations at a given gene. In other embodiments, each pathogenic mutation can be treated as a separate marker.

[0168] For a copy number variation, the encoding can be as a numerical value, e.g., as different integers or as numbers with one or more decimal points of resolution, e.g., a valuelike 4.3. Any method capable of determining a DNA copy number profile of a particular sample can be used for molecular profiling. The skilled artisan is aware of and capable of using a number of different platforms for assessing whole genome copy number changes at a resolution sufficient to identify the copy number of the one or more biomarkers of the methods described herein.

[0169] For example, sequencing or ISH techniques can be used for determining copy number / gene amplification. In some embodiments, the copy number profile analysis involves amplification of whole genome DNA by a whole genome amplification method. The whole genome amplification method can use a strand displacing polymerase and random primers.

[0170] The copy number profile analysis can involve hybridization of whole genome amplified DNA with a high-density array. In a more specific aspect, the high-density array has 5,000 or more different probes. In another specific aspect, the high-density array has 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000 or more different probes. In another specific aspect, each of the different probes on the array is an oligonucleotide having from about 15 to 200 bases in length. In another specific aspect, each of the different probes on the array is an oligonucleotide having from about 15 to 200, 15 to 150, 15 to 100, 15 to 75, 15 to 60, or 20 to 55 bases in length.

[0171] In some embodiments, a microarray is employed. In some aspects as described herein, prior to or concurrent with genotyping (analysis of copy number profiles), the sample may be amplified any number of mechanisms. The most common amplification procedure used involves PCR. Other suitable amplification methods include the ligase chain reaction (LCR), transcription amplification (Kwoh et al., Proc. Natl. Acad. Sci. USA 86, 1173 (1989) and WO88 / 10315), self-sustained sequence replication (Guatelli et al., Proc. Nat. Acad. Sci. USA, 87, 1874 (1990) and WO90 / 06995), selective amplification of target polynucleotide sequences (U.S. Pat. No.6,410,276), consensus sequence primed polymerase chain reaction (CP-PCR) (U.S. Pat. No.4,437,975), arbitrarily primed polymerase chain reaction (AP-PCR) (U.S. Pat. Nos.5,413,909, 5,861,245) and nucleic acid based sequence amplification (NABSA). As is apparent from the above survey of the suitable detection techniques, it may or may not be necessary to amplify the target DNA, i.e., the gene, cDNA, mRNA, miRNA, ora portion thereof to increase the number of target DNA molecule, depending on the detection techniques used.

[0172] Nucleic acid variants can also be detected using standard electrophoretic techniques. Detection of small genetic variations can also be accomplished by a variety of hybridization- based approaches. Sequence-specific probe hybridization can be used to detect a particular nucleic acid in a mixture or mixed population comprising other species of nucleic acids. In embodiments, fragment analysis (referred to herein as “FA”) methods are used for molecular profiling. Fragment analysis (FA) includes techniques such as restriction fragment length polymorphism (RFLP) and / or (amplified fragment length polymorphism).

[0173] Examples of particular genes of interest for mutational analysis are provided in Tables 1-7 below. 2. Expression Levels

[0174] The methods and systems as described herein can comprise expression profiling, which includes assessing differential expression of one or more target genes disclosed herein. Differential expression can include overexpression and / or underexpression of a biological product, e.g., a gene, mRNA or protein, compared to a control (or a reference). The control can include similar cells to the sample but without the disease (e.g., expression profiles obtained from samples from healthy individuals). A control can be a previously determined level that is indicative of a drug target efficacy associated with the particular disease and the particular drug target. The control can be derived from the same patient, e.g., a normal adjacent portion of the same organ as the diseased cells, the control can be derived from healthy tissues from other patients, or previously determined thresholds that are indicative of a disease responding or not-responding to a particular drug target. The control can also be a control found in the same sample, e.g. a housekeeping gene or a product thereof (e.g., mRNA or protein). For example, a control nucleic acid can be one which is known not to differ depending on the cancerous or non-cancerous state of the cell.

[0175] The expression level of a control nucleic acid can be used to normalize signal levels in the test and reference populations. Illustrative control genes include, but are not limited to, e.g., β-actin, glyceraldehyde 3-phosphate dehydrogenase and ribosomal protein P1. Multiple controls or types of controls can be used. The source of differential expression can vary. For example, a gene copy number may be increased in a cell, thereby resulting in increasedexpression of the gene. Alternately, transcription of the gene may be modified, e.g., by chromatin remodeling, differential methylation, differential expression or activity of transcription factors, etc. Translation may also be modified, e.g., by differential expression of factors that degrade mRNA, translate mRNA, or silence translation, e.g., microRNAs or siRNAs. In some embodiments, differential expression comprises differential activity. For example, a protein may carry a mutation that increases the activity of the protein, such as constitutive activation, thereby contributing to a diseased state. Molecular profiling that reveals changes in activity can be used to guide treatment selection.

[0176] Any useful method of acquiring expression levels can be employed. For example, some embodiments can use whole transcriptome sequencing analysis (WTS; RNA-seq) using the Illumina NGS platform, which methodology queries over 22,000 transcripts (genes) in a single assay.

[0177] Methods of gene expression profiling include methods based on hybridization analysis of polynucleotides, and methods based on sequencing of polynucleotides. Commonly used methods known in the art for the quantification of mRNA expression in a sample include northern blotting and in situ hybridization (Parker & Barnes (1999) Methods in Molecular Biology 106:247-283); RNAse protection assays (Hod (1992) Biotechniques 13:852-854); and reverse transcription polymerase chain reaction (RT-PCR) (Weis et al. (1992) Trends in Genetics 8:263-264). Alternatively, antibodies may be employed that can recognize specific duplexes, including DNA duplexes, RNA duplexes, and DNA-RNA hybrid duplexes or DNA-protein duplexes. Representative methods for sequencing-based gene expression analysis include Serial Analysis of Gene Expression (SAGE), gene expression analysis by massively parallel signature sequencing (MPSS) and / or next generation sequencing.

[0178] In some embodiments, the RNA expression data is normalized using Trimmed Mean of M-values (TMM). See Robinson and Oshlack, A Scaling Normalization Method for Differential Expression Analysis of RNA-seq Data, Genome Biol.2010;11(3):R25. doi: 10.1186 / gb-2010-11-3-r25. Epub 2010 Mar 2. Another technique for quantification can be found in Patro, R. et al. (2017), “Salmon provides fast and bias-aware quantification of transcript expression” Nature Methods 14(4):417-9.

[0179] The degree of differential expression can also be taken into account. For example, a gene can be considered as differentially expressed when the fold-change in expressioncompared to control level is at least 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.2, 2.5, 2.7, 3.0, 4, 5, 6, 7, 8, 9 or 10-fold different in the sample versus the control. The differential expression takes into account both overexpression and underexpression. A gene or gene product can be considered up or down-regulated if the differential expression meets a statistical threshold, a fold-change threshold, or both. For example, the criteria for identifying differential expression can comprise both a p-value of 0.001 and fold change of at least 1.5- fold (up or down). One of skill will understand that such statistical and threshold measures can be adapted to determine differential expression by any molecular profiling technique disclosed herein.

[0180] To avoid overfitting or similar error, analysis of panels may require training data on tens of thousands of tumor samples. To further avoid issues faced relying on RNA transcript analysis, such as overfitting of data based on the high number of total mRNAs, we may train the systems using more limited sets of transcripts. Traditionally, proteins that have been used in IHC based tumor classification. See, e.g., Lin and Liu, Immunohistochemistry in Undifferentiated Neoplasm / Tumor of Uncertain Origin, Arch Pathol Lab Med. 2014;138:1583–1610, which reference is incorporated herein by reference in its entirety. In some embodiments, the panel of mRNA transcripts used to implement the system comprise the mRNA encoding such proteins, and may further include various isoforms or related family members thereof.

[0181] Example RNA transcripts usable as features in a ML model include ACVRL1, AFP, ALPP, AMACR, ANKRD30A, ANO1, ARG1, AR, BCL2, BCL6, CA9, CALB2, CALCA, CALD1, CCND1, CD1A, CD2, CD34, CD3G, CD5, CD79A, CD99L2, CDH17, CDH1, CDK4, CDKN2A, CDX2, CEACAM16, CEACAM18, CEACAM19, CEACAM1, CEACAM20, CEACAM21, CEACAM3, CEACAM4, CEACAM5, CEACAM6, CEACAM7, CEACAM8, CGA, CGB3, CNN1, COQ2, CPS1, CR1, CR2, CTNNB1, DES, DSC3, ENO2, ERBB2, ERG, ESR1, FLI1, FOXL2, FUT4, GATA3, GPC3, HAVCR1, HNF1B, IL12B, IMP3, INHA, ISL1, KIT, KLK3, KL, KRT10, KRT14, KRT15, KRT16, KRT17, KRT18, KRT19, KRT1, KRT20, KRT2, KRT3, KRT4, KRT5, KRT6A, KRT6B, KRT6C, KRT7, KRT8, LIN28A, LIN28B, MAGEA2, MDM2, MIB1, MITF, MLANA, MLH1, MME, MPO, MS4A1, MSH2, MSH6, MSLN, MTHFR, MUC1, MUC2, MUC4, MUC5AC, MYOD1, MYOG, NANOG, NAPSA, NCAM1, NCAM2, NKX2-2, NKX3-1, OSCAR, PAX2, PAX5, PAX8, PDPN, PDX1, PECAM1, PGR, PIP, PMEL, PMS2, POU5F1, PSAP, PTPRC, S100A10, S100A11, S100A12, S100A13, S100A14, S100A16,S100A1, S100A2, S100A4, S100A5, S100A6, S100A7A, S100A7L2, S100A7, S100A8, S100A9, S100B, S100PBP, S100P, S100Z, SALL4, SATB2, SDC1, SERPINA1, SERPINB5, SF1, SFTPA1, SMAD4, SMARCB1, SMN1, SOX2, SPN, SYP, TFE3, TFF1, TFF3, TG, TLE1, TMPRSS2, TNFRSF8, TP63, TPM1, TPM2, TPM3, TPM4, TPSAB1, TTF1, UPK2, UPK3A, UPK3B, VHL, VIL1, VIM, and WT1. Some embodiments can use one or more of these as an RNA biomarker.

[0182] Protein-based detection techniques can also be useful for expression profiling. In some cases, nucleotide variants cause amino acid substitutions or deletions or insertions or frame shift that affect the protein primary, secondary or tertiary structure. To detect the amino acid variations, protein sequencing techniques may be used. For example, a protein or fragment thereof corresponding to a gene can be synthesized by recombinant expression using a DNA fragment isolated from an individual to be tested. A cDNA fragment of no more than 100 to 150 base pairs encompassing the polymorphic locus to be determined is used. The amino acid sequence of the peptide can then be determined by conventional protein sequencing methods. Alternatively, the HPLC-microscopy tandem mass spectrometry technique can be used for determining the amino acid sequence variations. In this technique, proteolytic digestion is performed on a protein, and the resulting peptide mixture is separated by reversed-phase chromatographic separation. Tandem mass spectrometry is then performed and the data collected is analyzed. See Gatlin et al., Anal. Chem., 72:757-763 (2000).

[0183] Protein-based detection molecular profiling techniques include immunoaffinity assays based on antibodies selectively immunoreactive with mutant gene encoded protein according to the present methods. These techniques include without limitation immunoprecipitation, Western blot analysis, molecular binding assays, enzyme-linked immunosorbent assay (ELISA), enzyme-linked immunofiltration assay (ELIFA), fluorescence activated cell sorting (FACS), immunohistochemistry (IHC), and the like. For example, an optional method of detecting the expression of a biomarker in a sample comprises contacting the sample with an antibody against the biomarker, or an immunoreactive fragment of the antibody thereof, or a recombinant protein containing an antigen binding region of an antibody against the biomarker; and then detecting the binding of the biomarker in the sample. Methods for producing such antibodies are known in the art. Antibodies can be used to immunoprecipitate specific proteins from solution samples or to immunoblot proteins separated by, e.g., polyacrylamide gels. Immunocytochemical methods can also be used in detecting specific protein polymorphisms in tissues or cells. Other well-known antibody-based techniques can also be used including, e.g., ELISA, radioimmunoassay (RIA), immunoradiometric assays (IRMA) and immunoenzymatic assays (IEMA), including sandwich assays using monoclonal or polyclonal antibodies. See, e.g., U.S. Pat. Nos.4,376,110 and 4,486,530, both of which are incorporated herein by reference. IV. EXAMPLE TREES

[0184] In an implementation, a model was trained on over 230,000 cases using features selected from patient sex, DNA mutation data for 225 cancer-related genes, and WTS expression data for over 10,200 genes. The DNA and RNA data was obtained using NGS on tumor tissue via gene panels and / or WES for DNA and WTS for RNA. A. Example top and some lower levels

[0185] FIG.6 shows an example top level of a hierarchal tumor type tree according to embodiments of the present disclosure. Each node shows the PPV and sensitivity for that given top-level tumor type, along with the number of samples tested and number of samples for that given type used to train the top-level model. The solid circles correspond to nodes that have child nodes, and thus a second-level model exists.

[0186] FIG.7 shows an example top level and several lower levels of a hierarchal tumor type tree according to embodiments of the present disclosure. B. Example results for subject with 2nd-level type

[0187] FIG.8 shows results of a subject identified as having cholangiocarcinoma. The ML model system identifies the subject has having Pancreatobiliary (top level) -> Cholangiocarcinoma (2nd-level). Pancreatobiliary has a 97% match, with HCC matching at 3%. Thus, the total for the top level is 100% or 1.

[0188] When pancreatobiliary is expanded, the three subtypes are shown: Pancreatic Adenocarcinoma, Cholangiocarcinoma, and Gallbladder Cancer. The probabilities for the three subtypes are scaled so that their sum equals 97%. Cholangiocarcinoma matches at 95%, with only 1% pancreatic adenocarcinoma.

[0189] Thus, embodiments can provide the ability to dynamically provide predicted labels based on the underlying probabilities. If the probability is sufficiently high (i.e., above a threshold), then that particular type (and all types at higher levels) can be reported out. If the probability were much lower, e.g., 30%, then embodiments might determine there isinsufficient confidence to report this subtype. However, the higher-level type can still be reported if it has sufficient confidence. Thus, the level of resolution can be determined on a case-by-case basis. C. Example complete Tree

[0190] FIGS.9A-9C show an example complete tree according to embodiments of the present disclosure. The hierarchal structure is shown via indenting. For example, adrenal cortical carcinoma is a top-level type with no child nodes. And, renal cell carcinoma and Wilm’s tumor are second-level types, where kidney is the top-level type. Renal clear cell carcinoma is one of three third-level types under renal cell carcinoma.

[0191] As examples, the top level types (e.g., the plurality of primary tumor origins) consists of, comprises, or comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, or 38 top-level tumor types. As examples, any given lower-level ML model can differentiate between at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, or 38 lower-level tumor types. The number of types to include can be based on the desired level of granularity in defining the tumor type. As a non- limiting example, the tumor type may be specified to levels wherein there is no impact on patient treatment by further granularity. D. Model validation

[0192] Performance of the trained model was assessed using retrospective and prospective training sets.

[0193] The retrospective test set comprised over 23,100 non-CUP and 417 CUP cases which were sequenced using a combined WES / WTS assay to assess DNA mutations (WES) and RNA transcript levels (WTS). Performance metrics included hierarchical positive predictive value (hPPV), hierarchical sensitivity (hSens), hierarchical F1 score (hF1; see Kiritchenko, S., et al. Learning and evaluation in the presence of class hierarchies: Application to text categorization. in Advances in Artificial Intelligence (eds. Lamontagne, L. & Marchand, M.). 395–406. (Springer, 2006)), and call rate (i.e., the model called the tumor origin with high confidence).

[0194] Performance of the retrospective analysis is illustrated in FIG.10A. In the figure, performance for non-CUP samples is shown combined (“Global”), or separated into “Primary,” meaning the specimen that was sequenced was from the primary tumor, or “Metastatic,” meaning that the specimen that was sequenced was from a metastatic lesion. The designations do not reflect the clinical status of having known metastases but only whether the specimen was from the primary site or metastatic site. Metastatic samples are further delineated by the location of the metastatic site, indicated by “Met” followed by a location. For example, “Met Liver” indicates the metastatic site of the specimen was the liver. “Node” refers to lymph node. Under all criteria, the model predicted the tumor type in the retrospective setting with high performance, including a call rate of 87.77% and 95.27% for non-CUP and CUP samples, respectively, and 95.14% for the combined samples.

[0195] The same analysis was repeated for a prospective test set comprising 3200 cases that were sequenced using the combined WES / WTS assay to assess DNA mutations (WES) and RNA transcript levels (WTS). Results are shown in FIG.10B.36 of the 3200 cases were CUPs. Similar to the retrospective analysis, the call rate for CUP samples in the prospective settings was 88.9% and 95.1% for non-CUP and CUP samples, respectively. In the figure, TOP 1 shows performance metrics considering the most probably tissue of origin. TOP 2 means that the first and second most probable types were considered.

[0196] In a follow-up study provided in Example C below, GPSai, a deep learning tool built using example systems and methods provided herein, identified tumor tissue of origin with 95.0% accuracy and classified 84.0% of CUP into one of 90 cancer categories. During eight months of live clinical testing (80,308 cases), 704 patients underwent diagnosis change due to GPSai.57.4% of tumors with diagnosis change were originally labeled as CUP, while the remainder were misdiagnoses. Level 1 targeted-therapy recommendations were changed on the molecular profiling clinical report in 86.1% of cases with diagnosis change.53.6% of physicians surveyed indicated a change in treatment plan for their patient due to GPSai. See Example C. V. METHOD

[0197] Example methods are described below for training models that can form a hierarchal tree and use of such a hierarchal tree of models.A. Use of hierarchal tree of models

[0198] FIG.11A is a flowchart illustrating a method for determining a tumor type using a hierarchal tumor type tree according to embodiments of the present disclosure. Some or all of various embodiments of method 1100 may be performed or controlled by a computer system. While method 1100 is described for tumor types, embodiments may be implemented for other sample types as well.

[0199] At block 1110, a sample vector is generated using biological data measured from a biological sample of a subject having a cancer. The biological data can include a set of biomarkers, as well as other data, e.g., demographics and personal characteristics (e.g., sex, age, weight, height, etc.) as described herein. As examples, the biological sample comprises cells from a solid tissue (e.g., a solid tumor), a bodily fluid, or a combination thereof. For instance, the biological sample can comprise an FFPE tissue sample. In some embodiments, the tissue sample can be microdissected, e.g., tumor tissue can be micro dissected to enrich the percentage tumor.

[0200] In some embodiments, the set of biomarkers can include DNA mutations. In various implementations, the DNA mutations can comprise one or more point mutation, sequence variant, polymorphism, deletion, insertion, indels (i.e., insertions or deletions), substitution, translocation, fusion, break, duplication, amplification, repeat, or copy number variation. Such DNA mutations can be pathogenic (P) or likely pathogenic (LP) mutations. As examples, the set of one or more biomarkers comprises one or more biomarkers listed in any one of Tables 1-7. In some embodiments, the set of biomarkers can include a set of expression levels, which be measured from RNA and / or proteins. In some embodiments, values for the set of biomarkers can be determined using whole genome sequencing. In some embodiments, values for the set of biomarkers can be determined using whole exome sequencing and / or whole transcriptome sequencing. As examples, such values can be whether a mutation is present or an expression level. The set of biomarkers can include DNA sequence variants and RNA expression.

[0201] The biological data may be obtained from one or more assays, which may measure one or more different types of biomarkers (e.g., molecules or complexes thereof). More than one type of molecule may be analyzed in one assay, which can be referred to as a hybrid assay. In preferred embodiments, the assay uses NGS technology for both whole exome sequencing for genomic DNA, and whole transcriptome sequencing for mRNA. The one ormore assays can be used for determining a presence, level, or state of a protein or nucleic acid for each of the one or more biomarkers. As examples, the presence, level or state of at least one protein can be determined using a technique selected from immunohistochemistry (IHC), flow cytometry, an immunoassay, an antibody or functional fragment thereof, an aptamer, mass spectrometry, or any combination thereof. Optionally the presence, level or state of all of the proteins can be determined using the technique. As other examples, the presence, level or state of at least one nucleic acid can be determined using a technique selected from polymerase chain reaction (PCR), in situ hybridization, amplification, hybridization, microarray, nucleic acid sequencing, dye termination sequencing, pyrosequencing, next generation sequencing, whole exome sequencing, whole genome sequencing, whole transcriptome sequencing, or any combination thereof. Optionally the presence, level or state of all of the nucleic acids is determined using the technique. As examples, the state of the at least one nucleic acid can comprise a sequence, mutation, polymorphism, deletion, insertion, substitution, translocation, fusion, break, duplication, amplification, repeat, copy number, or any combination thereof. In embodiments, the state of the nucleic acids may be determined using NGS for gene panels, WES, WTS, WGS, or any useful combination thereof.

[0202] At block 1120, a top-level machine learning model is loaded into memory of a computer system. The model can be loaded previously or in response to the sample vector being generated. The top-level machine learning model can be trained using top-level training samples. Each top-level training sample can include a top-level reference vector measured from a top-level reference biological sample labeled with a particular top-level tumor type (or other particular top-level sample types) of N top-level tumor types, or other N sample origins when implemented to determine sample types. A top-level tumor type can be a tumor origin, and a top-level sample type can be a sample origin. As examples, N can be an integer greater than or equal to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 50, 60, 70, 80, 90 or 100.

[0203] At block 1130, the top-level machine learning model processes the sample vector to identify that the biological sample has a first top-level tumor type of the N top-level tumor types. Examples of such top-level tumor types are provided as the top levels in FIGS.9A-9C. The first top-level tumor type can have a first likelihood that is highest out of the N top-level tumor types. A first portion of the N top-level tumor types can have no child nodes in the hierarchal tumor type tree, whereas a second portion can have child nodes (second-level types), e.g., the first top-level tumor type. The first likelihood can be required to be greaterthan a threshold, e.g., 50%, 51%, 52%, 53%, or higher values. The threshold can depend on the second highest probability. For example, if the first top-level tumor type is 51% and a second probability for a second top-level tumor type is 49%, then both may be used in later steps or an indeterminant classification may be output.

[0204] At block 1140, it is determined that the first top-level tumor type has M second- level tumor types in the hierarchal tumor type tree, or other M second-level sample types. M can be an integer greater than one, e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, or 38.

[0205] At block 1150, a second-level machine learning model is identified as corresponding to the first top-level tumor type. The second-level machine learning model can be trained using second-level training samples, e.g., training samples having the first top- level tumor type. Each second-level training sample can include a second-level reference vector measured from a second-level reference biological sample labeled with a particular second-level tumor type (or other particular second-level sample type) of the M second-level tumor types, or other M second-level sample types when implemented to determine sample types.

[0206] The machine learning models can consider additional features besides the set of biomarkers. For example, at least one of the top-level machine learning model and the second-level machine learning model is further trained using one or more characteristics of the tumor. As examples, the one or more characteristics of the tumor can comprise one or more selected from a group consisting of: a collection site of the biological sample, a sample format of the biological sample, and a metastatic status. As another example, at least one of the top-level machine learning model and the second-level machine learning model is further trained using one or more characteristic of the subject. As examples, the one or more characteristic of the subject comprises sex and / or age. For instance, when the biomarkers include DNA sequence variants and RNA expression, the DNA sequence variants can be determined using whole exome sequencing, the RNA expression can be determined using whole transcriptome sequencing, and the one or more characteristic of the subject can comprise sex as a binary feature (i.e., male or female).

[0207] At block 1160, the second-level machine learning model is loaded in the memory. The model can be loaded previously or in response to the model being identified. The second- level machine learning model can be trained using a first set of second-level training samples.Each second-level training sample can include a second-level reference vector measured from a second-level reference biological sample labeled with a second-level tumor type of M second-level tumor types corresponding to the first top-level tumor type.

[0208] The top-level model can be trained on all of the training samples, but for a given top-level tumor type, the corresponding second-level model would only be trained on a subset of the training samples. Accordingly, the first set of second-level training samples can include a subset of the top-level training samples labeled with the first top-level tumor type. Each of the first set of second-level training samples can have a second label being one on the M second-level tumor types.

[0209] The same or different reference vectors can be used for the different models. For example, different sets of biomarkers can be used for different models at different levels or two different models at a same level can use different sets of biomarkers. Accordingly, in some embodiments, the second-level reference vector of each of the first set of second-level training samples can be a same vector used to train the top-level machine learning model. Similarly, the biomarkers input into a model for the sample vector can vary among the different models. Thus, the biomarkers input into the top-level model can be the same or different than the biomarkers input into a lower-level model. Accordingly, in some embodiments, all values of the sample vector are processed by both the top-level machine learning model and the second-level machine learning model.

[0210] At block 1170, the second-level machine learning model processes the sample vector to identify the biological sample has a first second-level tumor type of the M second- level tumor types. The first second-level tumor type can have a second likelihood that is highest out of the M second-level tumor types.

[0211] The system can provide a report comprising the first top-level tumor type and / or the first second-level tumor type. The report can further comprise a diagnosis, prognosis, or theranosis for the cancer in the subject, based at least in part on the biological data measured from the biological sample from the subject. The theranosis can comprise one or more treatment of likely benefit, lack of benefit, or indeterminate benefit in treating the subject, optionally based at least in part on the first second-level tumor type.

[0212] As described in section I.C, the second likelihood can be scaled using the first likelihood. For example, the second likelihood can be scaled using the first likelihood (e.g.,multiplied by) to obtain a modified second likelihood, and the modified second likelihood can be reported.

[0213] If the second likelihood is not sufficiently high (e.g., greater than a threshold), only the top-level tumor type may be reported. Similarly, if the likelihood for any type at any level is not sufficiently high, any higher-level tumor type having a sufficiently high likelihood can be reported. Accordingly, if first likelihood is greater than a first threshold, the method can further comprise determining the second likelihood is less than a second threshold and reporting only the first top-level tumor type and not the first second-level tumor type.

[0214] In some embodiments, the system can run every sample through all models. The system can perform post-processing to determine which tumor types to report. For examples, probabilities for each tumor type can be analyzed and compared to respective thresholds and / or to each other to determine which probabilities and which tumor types to report. Accordingly, in some embodiments, the system can determine that a second top-level tumor type of the N top-level tumor types has K second-level tumor types in the hierarchal tumor type tree, with K being an integer greater than one. An additional second-level machine learning model corresponding to the second top-level tumor type can be identified. The additional second-level machine learning model can be loaded into memory. This additional model can be trained using a second set of second-level training samples. Each training sample of the second set can include an additional second-level reference vector measured from an additional second-level reference biological sample labeled with a second second- level tumor type of the K second-level tumor types corresponding to the second top-level tumor type. The additional second-level machine learning model can process the sample vector to determine probabilities for the K second-level tumor types.

[0215] Additional levels can also be analyzed, as described herein. Accordingly, some embodiments can determine that the first second-level tumor type has J third-level tumor types in the hierarchal tumor type tree, with J being an integer greater than one; identify a third-level machine learning model corresponding to the first second-level tumor type; loading the third-level machine learning model trained using a first set of third-level training samples; and process, using the third-level machine learning model, the sample vector to determine probabilities for the J third-level tumor types. Each of the first set of third-level training samples can include an a third-level reference vector measured from a third-levelreference biological sample labeled with a third-level tumor type of the J third-level tumor types corresponding to the first second-level tumor type.

[0216] As described in section II.D, any of the models can be an ensemble model. Accordingly, the top-level machine learning model can comprise a plurality of submodels. Processing the sample vector to identify the biological sample has the first top-level tumor type can comprise: obtaining output data from each of the plurality of submodels and using the output data from each of the plurality of submodels to identify the biological sample has the first top-level tumor type. In various embodiments, the first top-level tumor type can be determined by applying a majority rule to the output data, by using the output data as input into a dynamic voting model, or a combination thereof. Determining by the majority rule can comprise determining a number of occurrences of each top-level tumor type of the N top- level tumor types and selecting the first top-level tumor type as having the highest number of occurrences.

[0217] In various embodiments, each of the top-level machine learning models and the second-level machine learning models comprises a random forest classification algorithm, support vector machine, logistic regression, k-nearest neighbor model, neural network, naïve Bayes model, quadratic discriminant analysis, Gaussian processes model, clustering, or any combination thereof. Both the top-level machine learning model and the second-level machine learning model can be neural networks.

[0218] Each ML model can be trained to optimize an accuracy for the training set designated for that model. For example, the second-level machine learning model can be trained by: obtaining predicted tumor types generated by the second-level machine learning model from the second-level reference vectors; determining differences between the predicted tumor types generated by the second-level machine learning model and labels of the first set of second-level training samples; and adjusting one or more parameters of the second-level machine learning model based on the differences.

[0219] As described in section II and section VI below, additional ML models can be used to determine a prognosis and / or theranosis, e.g., for the first second-level tumor type (or any other tumor type at any level) based on the biological data. In embodiments, such ML model(s) can be used to determine whether a treatment for the first second-level tumor type is of likely benefit, lack of benefit, or indeterminate benefit in treating the subject, based on the biological data. In embodiments, the biomarker data obtained for the patient can be used topredict such treatment efficacies using direct biomarker-treatment associations, e.g., as described in section VI below. In such cases, the same biomarker data can be processed using ML models and also used in less complex treatment-biomarker association rules in order to identify optimal treatment for the subject. Treatments can be suggested based on any useful combination of tumor type determination via the systems and methods provided herein, ML models to predict treatment efficacy, and biomarker-treatment associations. One or more treatment can be administered to the subject based on the determining. In addition, the administration of one or more treatment of unlikely benefit may be avoided.

[0220] One or more orthogonal methods can be performed to confirm the first top-level tumor type and / or the first second-level tumor type. As examples, the orthogonal method can comprise one or more of imaging, immunohistochemistry, mutational signatures, hallmark fusions, viral reads, or any useful combination thereof. B. Training models for hierarchal tree

[0221] FIG.11B shows method 1105 of training machine learning models forming a hierarchal tumor type tree according to embodiments of the present disclosure. Aspects of method 1105 can be performed as described for method 1100, including additional steps described for method 1100. While method 1105 is described for tumor types, embodiments may be implemented for other sample types as well.

[0222] At block 1101, a computer system trains a top-level machine learning model using top-level training samples that each include biological data for a set of biomarkers. Each top- level training sample can include a top-level reference vector measured from a top-level reference biological sample labeled with a particular top-level tumor type of N top-level tumor types. N can be an integer greater than 1, e.g., as recited from method 1100. A first top-level tumor type can have M second-level tumor types in the hierarchal tumor type tree. M can be an integer greater than one.

[0223] At block 1102, the computer system can train a second-level machine learning model corresponding to the first top-level tumor type using a first set of second-level training samples, which can each include a second-level reference vector measured from a second- level reference biological sample labeled with a second-level tumor type of M second-level tumor types corresponding to the first top-level tumor type.

[0224] One or more other top-level tumor types can have multiple second-level tumor types. In such a case, for each of the one or more other top-level tumor types, the computer system can train a respective second-level machine learning model corresponding to the top- level tumor type using another set of second-level training samples. C. Use of hierarchical tree for top-level with no second levels

[0225] For some subjects, the top-level tumor type might not have multiple second-level tumor types. The following method describes such an example.

[0226] FIG.11C shows a method of determining a tumor type using a hierarchal tumor type tree according to embodiments of the present disclosure. Aspects of method 1118 can be performed as described for method 1100, including additional steps described for method 1100. While method 1118 is described for tumor types, embodiments may be implemented for other sample types as well.

[0227] At block 1111, a computer system loads into memory a top-level machine learning model trained using top-level training samples. Each top-level training sample can include a top-level reference vector measured from a top-level reference biological sample labeled with a particular top-level tumor type of N top-level tumor types. N can be an integer greater than 1, e.g., as recited from method 1100. A first top-level tumor type can have M second-level tumor types in the hierarchal tumor type tree. M can be an integer greater than one.

[0228] At block 1112, the computer system loads a second-level machine learning model corresponding to the first top-level tumor type trained using a first set of second-level training samples, which can each include a second-level reference vector measured from a second- level reference biological sample labeled with a second-level tumor type of M second-level tumor types corresponding to the first top-level tumor type.

[0229] At block 1113, a sample vector is generated using biological data measured from a biological sample of a subject. The biological data includes a set of biomarkers.

[0230] At block 1114, the top-level machine learning model processes the sample vector to identify the biological sample has a second top-level tumor type of the N top-level tumor types. The second top-level tumor type having a likelihood that is highest out of the N top- level tumor types. The second top-level tumor type may or may not have one or more second-level tumor types.VI. MOLECULAR PROFILING

[0231] Molecular profiling can be used in any of the embodiments described above. In addition to determining a tumor type, some embodiments can perform a diagnosis, prognosis, and / or theranosis. The molecular profiling of one or more targets can be used to determine or identify a therapeutic for an individual, e.g., one or more candidate treatments. For example, the presence, level or state of one or more biomarkers can be used to determine or identify a therapeutic for an individual. In various embodiments, the treatment can include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, or stem cell therapy.

[0232] Molecular profiling can be performed by any known means for detecting molecules in a biological sample. Useful biological samples include those described above. In some embodiments, fixed tissue is used to molecular profiling of tumors and blood is used as the biological sample for liquid biopsy. Molecular profiling assays can include without limitation, protein and nucleic acid analysis techniques. Protein analysis techniques include, by way of non-limiting examples, immunoassays, immunohistochemistry (IHC), and mass spectrometry. Nucleic acid analysis techniques include, by way of non-limiting examples, amplification such as polymerase chain amplification (PCR) amplification (e.g., qPCR or RT- PCR), hybridization, microarrays, in situ hybridization, and DNA sequencing or RNA sequencing (e.g., dye termination sequencing, Sanger, high throughput or next generation sequencing (NGS), pyrosequencing, and restriction fragment analysis). Other examples of useful assays include in situ hybridization (ISH); fluorescent in situ hybridization (FISH); chromogenic in situ hybridization (CISH); various types of microarray (mRNA expression arrays, low density arrays, protein arrays, etc); comparative genomic hybridization (CGH); Northern blot; Southern blot; and any other appropriate technique to assay the presence or quantity of a biological molecule of interest. In various embodiments, any one or more of these methods can be used concurrently or subsequent to each other for assessing target genes disclosed herein.

[0233] The molecular profiling can be based on either the gene, e.g., DNA sequence, and / or gene product, e.g., mRNA or protein. Such nucleic acid and / or polypeptide can be profiled as applicable as to presence or absence, level or amount, activity, mutation, sequence, haplotype, rearrangement, copy number, or other measurable characteristic.

[0234] In some embodiments, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or at least about 100 genes or gene products are profiled by at least one technique, a plurality of techniques. In some embodiments, at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 11,000, 12,000, 13,000, 14,000, 15,000, 16,000, 17,000, 18,000, 19,000, 20,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, 30,000, 31,000, 32,000, 33,000, 34,000, 35,000, 36,000, 37,000, 38,000, 39,000, 40,000, 41,000, 42,000, 43,000, 44,000, 45,000, 46,000, 47,000, 48,000, 49,000, or at least 50,000 genes or gene products are profiled using various techniques.

[0235] As part of determining a molecular profile, sequencing nucleic acid molecules can provide sequence reads that include molecular information indicating point mutations, polymorphisms, deletions, insertions, substitutions, translocations, fusions, breaks, duplications, amplification, repeats, copy numbers (copy number variation; CNV; copy number alteration; CNA), transcript levels (expression levels), or any combination thereof. Such information can be determined by analyzing the sequence reads. In preferred embodiments, NGS is used to sequence nucleic acid molecules. The high-throughput nature of NGS allows for analysis of the whole exome or whole genome and / or whole transcriptome of >22,000 genes / gene products in a single assay, thereby providing a comprehensive overview of such molecular information.

[0236] The methods provided herein are useful for identifying states of biomarkers, e.g., mutations and / or expression levels, in any type of cancer. As examples, the cancer can comprise an acute lymphoblastic leukemia; acute myeloid leukemia; adrenocortical carcinoma; AIDS-related cancer; AIDS-related lymphoma; anal cancer; appendix cancer; astrocytomas; atypical teratoid / rhabdoid tumor; basal cell carcinoma; bladder cancer; brain stem glioma; brain tumor, brain stem glioma, central nervous method atypical teratoid / rhabdoid tumor, central nervous method embryonal tumors, astrocytomas, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, pineal parenchymal tumors of intermediate differentiation, supratentorial primitive neuroectodermal tumors and pineoblastoma; breast cancer; bronchial tumors; Burkitt lymphoma; cancer of unknown primary site (CUP); carcinoid tumor; carcinoma of unknown primary site; central nervous method atypical teratoid / rhabdoid tumor; central nervous method embryonal tumors; cervical cancer; childhood cancers;chordoma; chronic lymphocytic leukemia; chronic myelogenous leukemia; chronic myeloproliferative disorders; colon cancer; colorectal cancer; craniopharyngioma; cutaneous T-cell lymphoma; endocrine pancreas islet cell tumors; endometrial cancer; ependymoblastoma; ependymoma; esophageal cancer; esthesioneuroblastoma; Ewing sarcoma; extracranial germ cell tumor; extragonadal germ cell tumor; extrahepatic bile duct cancer; gallbladder cancer; gastric (stomach) cancer; gastrointestinal carcinoid tumor; gastrointestinal stromal cell tumor; gastrointestinal stromal tumor (GIST); gestational trophoblastic tumor; glioma; hairy cell leukemia; head and neck cancer; heart cancer; Hodgkin lymphoma; hypopharyngeal cancer; intraocular melanoma; islet cell tumors; Kaposi sarcoma; kidney cancer; Langerhans cell histiocytosis; laryngeal cancer; lip cancer; liver cancer; malignant fibrous histiocytoma bone cancer; medulloblastoma; medulloepithelioma; melanoma; Merkel cell carcinoma; Merkel cell skin carcinoma; mesothelioma; metastatic squamous neck cancer with occult primary; mouth cancer; multiple endocrine neoplasia syndromes; multiple myeloma; multiple myeloma / plasma cell neoplasm; mycosis fungoides; myelodysplastic syndromes; myeloproliferative neoplasms; nasal cavity cancer; nasopharyngeal cancer; neuroblastoma; Non-Hodgkin lymphoma; nonmelanoma skin cancer; non-small cell lung cancer; oral cancer; oral cavity cancer; oropharyngeal cancer; osteosarcoma; other brain and spinal cord tumors; ovarian cancer; ovarian epithelial cancer; ovarian germ cell tumor; ovarian low malignant potential tumor; pancreatic cancer; papillomatosis; paranasal sinus cancer; parathyroid cancer; pelvic cancer; penile cancer; pharyngeal cancer; pineal parenchymal tumors of intermediate differentiation; pineoblastoma; pituitary tumor; plasma cell neoplasm / multiple myeloma; pleuropulmonary blastoma; primary central nervous method (CNS) lymphoma; primary hepatocellular liver cancer; prostate cancer; rectal cancer; renal cancer; renal cell (kidney) cancer; renal cell cancer; respiratory tract cancer; retinoblastoma; rhabdomyosarcoma; salivary gland cancer; Sézary syndrome; small cell lung cancer; small intestine cancer; soft tissue sarcoma; squamous cell carcinoma; squamous neck cancer; stomach (gastric) cancer; supratentorial primitive neuroectodermal tumors; T-cell lymphoma; testicular cancer; throat cancer; thymic carcinoma; thymoma; thyroid cancer; transitional cell cancer; transitional cell cancer of the renal pelvis and ureter; trophoblastic tumor; ureter cancer; urethral cancer; uterine cancer; uterine sarcoma; vaginal cancer; vulvar cancer; Waldenström macroglobulinemia; or Wilm’s tumor

[0237] As further examples, the cancer can comprise an acute myeloid leukemia (AML), breast carcinoma, cholangiocarcinoma, colorectal adenocarcinoma, extrahepatic bile duct adenocarcinoma, female genital tract malignancy, gastric adenocarcinoma, gastroesophageal adenocarcinoma, gastrointestinal stromal tumor (GIST), glioblastoma, head and neck squamous carcinoma, leukemia, liver hepatocellular carcinoma, low grade glioma, lung bronchioloalveolar carcinoma (BAC), non-small cell lung cancer (NSCLC), lung small cell cancer (SCLC), lymphoma, male genital tract malignancy, malignant solitary fibrous tumor of the pleura (MSFT), melanoma, multiple myeloma, neuroendocrine tumor, nodal diffuse large B-cell lymphoma, non-epithelial ovarian cancer (non-EOC), ovarian surface epithelial carcinoma, pancreatic adenocarcinoma, pituitary carcinomas, oligodendroglioma, prostatic adenocarcinoma, retroperitoneal or peritoneal carcinoma, retroperitoneal or peritoneal sarcoma, small intestinal malignancy, soft tissue tumor, thymic carcinoma, thyroid carcinoma, or uveal melanoma.

[0238] Molecular profiling can be used to provide a comprehensive view of the biological state of a sample. In an embodiment, molecular profiling is used for whole tumor profiling. Accordingly, a number of molecular approaches are used to assess the state of a tumor. The whole tumor profiling can be used for selecting a candidate treatment for a tumor. Molecular profiling can be used to select candidate therapeutics on any sample for any stage of a disease. In embodiments, the methods as described herein are used to profile a newly diagnosed cancer. The candidate treatments indicated by the molecular profiling can be used to select a therapy for treating the newly diagnosed cancer. In other embodiments, the methods as described herein are used to profile a cancer that has already been treated, e.g., with one or more standard-of-care therapy. In embodiments, the cancer is refractory to the prior treatment / s. For example, the cancer may be refractory to the standard of care treatments for the cancer. The cancer can be a metastatic cancer or other recurrent cancer. The treatments can be on-compendium or off-compendium treatments.

[0239] Molecular profiling of individual samples can be used to select one or more candidate treatments for a disorder in a subject, e.g., by identifying targets for drugs that may be effective for a given cancer. For example, the candidate treatment can be a treatment known to have an effect on cells that differentially express genes as identified by molecular profiling techniques, an experimental drug, a government or regulatory approved drug or any combination of such drugs, which may have been studied and approved for a particularindication that is the same as or different from the indication of the subject from whom a biological sample is obtain and molecularly profiled.

[0240] Molecular profiling methods according to the present disclosure can also comprise measuring epigenetic change, i.e., modification in a gene caused by an epigenetic mechanism, such as a change in methylation status or histone acetylation. Frequently, the epigenetic change will result in an alteration in the levels of expression of the gene which may be detected (at the RNA or protein level as appropriate) as an indication of the epigenetic change. Often the epigenetic change results in silencing or down regulation of the gene, referred to as “epigenetic silencing.” The most frequently investigated epigenetic change in the methods as described herein involves determining the DNA methylation status of a gene, where an increased level of methylation is typically associated with the relevant cancer (since it may cause down regulation of gene expression). Aberrant methylation, which may be referred to as hypermethylation, of the gene or genes can be detected. Typically, the methylation status is determined in suitable CpG islands which are often found in the promoter region of the gene(s). The term “methylation,” “methylation state” or “methylation status” may refers to the presence or absence of 5-methylcytosine at one or a plurality of CpG dinucleotides within a DNA sequence. CpG dinucleotides are typically concentrated in the promoter regions and exons of human genes.

[0241] Various assay procedures to directly detect methylation are known in the art and can be used in conjunction with the present methods. These assays rely onto two distinct approaches: bisulphite conversion based approaches and non-bisulphite based approaches. Non-bisulphite based methods for analysis of DNA methylation rely on the inability of methylation-sensitive enzymes to cleave methylation cytosines in their restriction.

[0242] Other techniques for DNA methylation analysis include sequencing, methylation- specific PCR (MS-PCR), melting curve methylation-specific PCR (McMS-PCR), MLPA with or without bisulfite treatment, QAMA, MSRE-PCR, MethyLight, ConLight-MSP, bisulfite conversion-specific methylation-specific PCR (BS-MSP), COBRA (which relies upon use of restriction enzymes to reveal methylation dependent sequence differences in PCR products of sodium bisulfite-treated DNA), methylation-sensitive single-nucleotide primer extension conformation (MS-SNuPE), methylation-sensitive single-strand conformation analysis (MS-SSCA), Melting curve combined bisulfite restriction analysis (McCOBRA), PyroMethA, HeavyMethyl, MALDI-TOF, MassARRAY, Quantitative analysis of methylatedalleles (QAMA), enzymatic regional methylation assay (ERMA), QBSUPT, MethylQuant, Quantitative PCR sequencing and oligonucleotide-based microarray systems, Pyrosequencing, and Meth-DOP-PCR.

[0243] A general approach to molecular profiling is as follows. First, obtain a sample comprising cells from a cancer in a subject, e.g., a tumor sample or bodily fluid sample such as described herein. In some embodiments, the sample comprises metastatic cells. Next, perform molecular profiling assays on the sample to assess one or more biomarkers and thereby obtain a molecular profile for the sample. In preferred embodiments, the molecular profiling assays comprise NGS, including WES, WGS, WTS, and / or targeted sequencing of panels of genes and / or gene products. The molecular profiling data for the sample can be input into a ML model such as described herein. In some embodiments, this comprises inputting the sample vector into a ML system comprising multiple trained ML models in order to determine a tumor type. As a non-limiting example, one may input the sample vector to determine tumor type based on a hierarchical tree as described herein.

[0244] Table 1 lists numerous biomarkers we have profiled over the past several years. As relevant molecular profiling and patient outcomes are available, any or all of these biomarkers can serve as features to input into a cognitive computing environment. The table shows molecular profiling techniques and various biomarkers assessed using those techniques. The listing is non-exhaustive, and data for all of the listed biomarkers will not be available for every patient. It will further be appreciated that various biomarkers have been profiled using multiple methods. As a non-limiting example, consider the EGFR gene expressing the Epidermal Growth Factor Receptor (EGFR) protein. As shown in Table 1, expression of EGFR protein has been detected using IHC; EGFR gene amplification, gene rearrangements, mutations and alterations have been detected with ISH, Sanger sequencing, NGS, fragment analysis, and PCR such as qPCR; and EGFR RNA expression has been detected using PCR techniques, e.g., qPCR, and DNA microarray. As a further non-limiting example, molecular profiling results for the presence of the EGFR variant III (EGFRvIII) transcript has been collected using fragment analysis (e.g., RFLP) and sequencing (e.g., NGS).

[0245] Table 2 shows exemplary molecular profiles for various tumor lineages. Data from these molecular profiles may be used as the input for NGP in order to identify one or more biosignatures of interest. In the table, the cancer lineage is shown in the column “TumorType.” The remaining columns show various biomarkers that can be assessed using the indicated methodology (i.e., immunohistochemistry (IHC), in situ hybridization (ISH), or other techniques). As explained above, the biomarkers are identified using symbols known to those of skill in the art. Under the IHC column, “MMR” refers to the mismatch repair proteins MLH1, MSH2, MSH6, and PMS2, which are each individually assessed using IHC. Under the NGS column “DNA,” “CNA” refers to copy number alteration, which is also referred to herein as copy number variation (CNV). Whole transcriptome sequencing (WTS) is used to assess all RNA transcripts in the specimen. One of skill will appreciate that molecular profiling technologies may be substituted as desired and / or interchangeable. For example, other suitable protein analysis methods can be used instead of IHC (e.g., alternate immunoassay formats), other suitable nucleic acid analysis methods can be used instead of ISH (e.g., that assess copy number and / or rearrangements, translocations and the like), and other suitable nucleic acid analysis methods can be used instead of fragment analysis. Similarly, FISH and CISH are generally interchangeable and the choice may be made based upon probe availability and the like. Tables 4-6 present panels of genomic analysis and genes that have been assessed using Next Generation Sequencing (NGS) analysis of DNA such as genomic DNA. As desired, other nucleic acid analysis methods can be used instead of NGS analysis, e.g., other sequencing (e.g., Sanger), hybridization (e.g., microarray, Nanostring) and / or amplification (e.g., PCR based) methods. The biomarkers listed in Tables 6-7 can be assessed by RNA sequencing, such as WTS. Using WTS, any fusions, splice variants, or the like can be detected. Tables 6-7 list biomarkers with commonly detected transcript alterations in cancer.

[0246] Nucleic acid analysis may be performed to assess various aspects of a gene. For example, nucleic acid analysis can include, but is not limited to, mutational analysis, fusion analysis, variant analysis, splice variants, SNP analysis and gene copy number / amplification. Such analysis can be performed using any number of techniques described herein or known in the art, including without limitation sequencing (e.g., Sanger, Next Generation, pyrosequencing), PCR, variants of PCR such as RT-PCR, fragment analysis, and the like. NGS techniques may be used to detect mutations, fusions, variants and copy number of multiple genes in a single assay. Unless otherwise stated or obvious in context, a “mutation” as used herein may comprise any change in a gene or genome as compared to wild type, including without limitation a mutation, polymorphism, deletion, insertion, indels (i.e., insertions or deletions), substitution, translocation, fusion, break, duplication, amplification,repeat, or copy number variation. Different analyses may be available for different genomic alterations and / or sets of genes. For example, Table 3 lists attributes of genomic stability that can be measured with NGS, Table 4 lists various genes that may be assessed for point mutations and indels, Table 5 lists various genes that may be assessed for point mutations, indels and copy number variations, Table 6 lists various genes that may be assessed for gene fusions via RNA analysis, e.g., via WTS, and similarly Table 7 lists genes that can be assessed for transcript variants via RNA.

[0247] As noted in Table 1, NGS can be used for whole exome sequencing (WES), whole genome sequencing (WGS), and / or whole transcriptome sequencing (WTS). Such methods can allow for simultaneous analysis of all substantially all or all exons in genomic DNA, simultaneous analysis of all substantially all or all genomic DNA, and simultaneous analysis of substantially all or all mRNA transcripts. Molecular profiling can employ any of these techniques as desired. Table 1 – Molecular Profiling Biomarkers Technique Biomarkers , ), 3 2,MSH2, MSH4, MSH6, MSI, MTAP, MUC1, MUC16, NFKB1, NFKB1A, NFKB2, NGF, NOTCH1, NPM1, NRAS, NY-ESO-1, IDNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1,Table 2 – Molecular Profiles Whole Next Generation Transcri tomePD-L1, PR, CNA TMB, TRKA / B / C LOH ng) mL1, CNA TMB, TRKA / B / C LOHLOHTable 3 – Genomic Stability Testing (DNA) Microsatellite Instability (MSI) Tumor Mutational Burden (TMB)Table 4 – Point Mutations and Indels (DNA) ABI1 CRLF2 HOXC11 MUC1 RHOHAKT1 DNM2 HOXD13 NBN SEPT5 AMER1 DNMTA HRA NDR 1 EPTCEBPA HMGN2P46 MLLT11 PPP2R1A UBR5 H HD7 HNF1A MN1 PRF1 VHLTable 5 – Point Mutations, Indels and Copy Number Variations (DNA) ABL2 CREB1 FUS MYC RUNX1ARHGEF12 DDR2 HGF NIN SMAD2 ARID1A DDX1 HIP1 N T H2 MAD4BCL9 ERCC2 KDR PDCD1LG2 TCF12 (VEGFR2) (PDL2)CCDC6 FBXO11 MAF PRRX1 TRIM33 NB1IP1 FBXW7 MALT1 P IP1 TRIP11CIITA FLT4 MRE11 RNF43 ZMYM2 LP1 FNBP1 M H2 R 1 ZNF217Table 6 – Gene Fusions (RNA) AKT3 ETV4 MAST2 NUMBL RETTable 7 – Variant Transcripts AR-V7 EGFR vIII MET Exon 14 Skipping

[0248] Abbreviations used in Tables 1-7 and throughout the specification, e.g., IHC: immunohistochemistry; ISH: in situ hybridization; CISH: colorimetric in situ hybridization; FISH: fluorescent in situ hybridization; NGS: next generation sequencing; PCR: polymerase chain reaction; CNA: copy number alteration; CNV: copy number variation; MSI: microsatellite instability; TMB: tumor mutational burden; LOH: loss of heterozygosity.

[0249] The biomarkers used for molecular profiling can include gene fusions, such as cancer related gene fusions. A number of recurrent fusion genes have been catalogued in the Mittleman database (cgap.nci.nih.gov / Chromosomes / Mitelman). Gene fusions can be determined by sequencing mRNA transcripts and in some cases by sequencing DNA. Such fusions occur in various cancers. For example, TMPRSS2-ERG, TMPRSS2-ETV and SLC45A3-ELK4 fusions can be detected to characterize prostate cancer; and ETV6-NTRK3 and ODZ4-NRG1 can be used to characterize breast cancer. The EML4-ALK, RLF-MYCL1, TGF-ALK, or CD74-ROS1 fusions can be used to characterize a lung cancer. The ACSL3- ETV1, C15ORF21-ETV1, FLJ35294-ETV1, HERV-ETV1, TMPRSS2-ERG, TMPRSS2- ETV1 / 4 / 5, TMPRSS2-ETV4 / 5, SLC5A3-ERG, SLC5A3-ETV1, SLC5A3-ETV5 or KLK2- ETV4 fusions can be used to characterize a prostate cancer. The GOPC-ROS1 fusion can be used to characterize a brain cancer. The CHCHD7-PLAG1, CTNNB1-PLAG1, FHIT- HMGA2, HMGA2-NFIB, LIFR-PLAG1, or TCEA1-PLAG1 fusions can be used to characterize a head and neck cancer. The ALPHA-TFEB, NONO-TFE3, PRCC-TFE3, SFPQ- TFE3, CLTC-TFE3, or MALAT1-TFEB fusions can be used to characterize a renal cell carcinoma (RCC). The AKAP9-BRAF, CCDC6-RET, ERC1-RETM, GOLGA5-RET, HOOK3-RET, HRH4-RET, KTN1-RET, NCOA4-RET, PCM1-RET, PRKARA1A-RET, RFG-RET, RFG9-RET, Ria-RET, TGF-NTRK1, TPM3-NTRK1, TPM3-TPR, TPR-MET, TPR-NTRK1, TRIM24-RET, TRIM27-RET or TRIM33-RET fusions can be used to characterize a thyroid cancer and / or papillary thyroid carcinoma; and the PAX8-PPARy fusion can be analyzed to characterize a follicular thyroid cancer. Fusions that are associated with hematological malignancies include without limitation TTL-ETV6, CDK6-MLL, CDK6-TLX3, ETV6-FLT3, ETV6-RUNX1, ETV6-TTL, MLL-AFF1, MLL-AFF3, MLL- AFF4, MLL-GAS7, TCBA1-ETV6, TCF3-PBX1 or TCF3-TFPT, which are characteristic of acute lymphocytic leukemia (ALL); BCL11B-TLX3, IL2-TNFRFS17, NUP214-ABL1, NUP98-CCDC28A, TAL1-STIL, or ETV6-ABL2, which are characteristic of T-cell acute lymphocytic leukemia (T-ALL); ATIC-ALK, KIAA1618-ALK, MSN-ALK, MYH9-ALK, NPM1-ALK, TGF-ALK or TPM3-ALK, which are characteristic of anaplastic large celllymphoma (ALCL); BCR-ABL1, BCR-JAK2, ETV6-EVI1, ETV6-MN1 or ETV6-TCBA1, characteristic of chronic myelogenous leukemia (CML); CBFB-MYH11, CHIC2-ETV6, ETV6-ABL1, ETV6-ABL2, ETV6-ARNT, ETV6-CDX2, ETV6-HLXB9, ETV6-PER1, MEF2D-DAZAP1, AML-AFF1, MLL-ARHGAP26, MLL-ARHGEF12, MLL-CASC5, MLL-CBL,MLL-CREBBP, MLL-DAB21P, MLL-ELL, MLL-EP300, MLL-EPS15, MLL- FNBP1, MLL-FOXO3A, MLL-GMPS, MLL-GPHN, MLL-MLLT1, MLL-MLLT11, MLL- MLLT3, MLL-MLLT6, MLL-MYO1F, MLL-PICALM, MLL-SEPT2, MLL-SEPT6, MLL- SORBS2, MYST3-SORBS2, MYST-CREBBP, NPM1-MLF1, NUP98-HOXA13, PRDM16- EVI1, RABEP1-PDGFRB, RUNX1-EVI1, RUNX1-MDS1, RUNX1-RPL22, RUNX1- RUNX1T1, RUNX1-SH3D19, RUNX1-USP42, RUNX1-YTHDF2, RUNX1-ZNF687, or TAF15-ZNF-384, which are characteristic of acute myeloid leukemia (AML); CCND1- FSTL3, which is characteristic of chronic lymphocytic leukemia (CLL); BCL3-MYC, MYC- BTG1, BCL7A-MYC, BRWD3-ARHGAP20 or BTG1-MYC, which are characteristic of B- cell chronic lymphocytic leukemia (B-CLL); CITTA-BCL6, CLTC-ALK, IL21R-BCL6, PIM1-BCL6, TFCR-BCL6, IKZF1-BCL6 or SEC31A-ALK, which are characteristic of diffuse large B-cell lymphomas (DLBCL); FLIP1-PDGFRA, FLT3-ETV6, KIAA1509- PDGFRA, PDE4DIP-PDGFRB, NIN-PDGFRB, TP53BP1-PDGFRB, or TPM3-PDGFRB, which are characteristic of hyper eosinophilia / chronic eosinophilia; and IGH-MYC or LCP1-BCL6, which are characteristic of Burkitt’s lymphoma. One of skill will understand that additional fusions, including those yet to be identified to date, can be used to guide treatment once their presence is associated with a therapeutic intervention.

[0250] As described herein, molecular profiling information can be used to identify a tumor type, such as the primary tumor site of a metastatic cancer of unknown primary (CUP). In some embodiments, the predictions can be used to assist in planning treatment of cancer patients. In some embodiments, such information is used to verify the original diagnosis of a cancer at the same time molecular profiling is used to identify treatment options. If the information differs from the original diagnosis, additional inquiry may be performed (e.g., pathologist review and orthogonal testing (e.g., IHC)) to verify the diagnosis and thus benefit patient treatment. A. Diagnosis

[0251] The molecular profiling provided by the disclosure is not limited to identifying candidate treatments for patients in need thereof. Indeed, the data derived from molecularprofiling can be used to characterize various phenotypes of interest. For example, molecular profiling may be used to screen for disease, monitor disease before and / or after treatment, or characterize a disease aggressiveness. For example, molecular profiling may be used to predict the risk that a primary tumor will metastasize. Thus, molecular profiling of a primary tumor may provide both personalized treatment options for the patient and in addition provide a metastatic potential for the tumor. The treating physician may consider the predicted metastatic potential when deciding a course of treatment for the patient.

[0252] A diagnosis may be performed based on the detection of one or more biomarkers. In another embodiment, a diagnosis of a phenotype may include the determination of a quantitative parameter (e.g., an expression level of RNA or a copy number of DNA) that is compared to one or more cutoff values that are selected to differentiate between different phenotype classifications, e.g., disease is present or not.

[0253] As examples, the phenotype can comprise detecting the presence of or likelihood of developing a tumor, neoplasm, or cancer, or characterizing the tumor, neoplasm, or cancer (e.g., stage, grade, aggressiveness, likelihood of metastasis or recurrence, etc). In some embodiments, the cancer comprises an acute myeloid leukemia (AML), breast carcinoma, cholangiocarcinoma, colorectal adenocarcinoma, extrahepatic bile duct adenocarcinoma, female genital tract malignancy, gastric adenocarcinoma, gastroesophageal adenocarcinoma, gastrointestinal stromal tumors (GIST), glioblastoma, head and neck squamous carcinoma, leukemia, liver hepatocellular carcinoma, low grade glioma, lung bronchioloalveolar carcinoma (BAC), lung non-small cell lung cancer (NSCLC), lung small cell cancer (SCLC), lymphoma, male genital tract malignancy, malignant solitary fibrous tumor of the pleura (MSFT), melanoma, multiple myeloma, neuroendocrine tumor, nodal diffuse large B-cell lymphoma, non-epithelial ovarian cancer (non-EOC), ovarian surface epithelial carcinoma, pancreatic adenocarcinoma, pituitary carcinomas, oligodendroglioma, prostatic adenocarcinoma, retroperitoneal or peritoneal carcinoma, retroperitoneal or peritoneal sarcoma, small intestinal malignancy, soft tissue tumor, thymic carcinoma, thyroid carcinoma, or uveal melanoma. The systems and methods herein can be used to characterize these and other cancers. Characterizing a phenotype can be providing a diagnosis, prognosis or theranosis of the cancer.

[0254] As described herein, the diagnosis of a tumor type may be unknown (e.g., CUP) or could be mistaken. As treatment selection may differ for various types of cancer, identifyingthe proper tumor type can alter the course of treatment. See Example C for real world examples of such scenarios in clinical practice. The systems and methods provided herein can be used to provide or confirm the diagnosis as appropriate. B. Treatment selection

[0255] A molecular profiling approach can provide a method for selecting treatments for an individual that could favorably change the clinical course of a medical condition, including without limitation cancer. Molecular profiling can provide a personalized approach to selecting treatments that are more likely to benefit a cancer. The molecular profiling methods described herein can be used to guide treatment in any desired setting, including without limitation the front-line / standard of care setting, or for patients with poor prognosis, such as those with metastatic disease or those whose cancer has progressed on standard front line therapies, or whose cancer has progressed on previous chemotherapeutic or hormonal regimens. The cancer can be a metastatic cancer or other recurrent cancer. The treatments can be on-compendium or off-compendium treatments. The treatments may be standard of care for the type of cancer in the individual, or the treatments may be typically used for other types of cancer. Thus, profiling may expand the choice of treatments for the individual.

[0256] The systems and methods provided herein may be used to classify patients as more or less likely to benefit or respond to various treatments. Unless otherwise noted, the terms “response” or “non-response,” as used herein, refer to any appropriate indication that a treatment provides a benefit to a patient (a “responder” or “benefiter”) or has a lack of benefit to the patient (a “non-responder” or “non-benefiter”). Such an indication may be determined using accepted clinical response criteria such as the standard Response Evaluation Criteria in Solid Tumors (RECIST) criteria, or other useful patient response criteria such as progression free survival (PFS), time to progression (TTP), disease free survival (DFS), time-to-next treatment (TNT, TTNT), tumor shrinkage or disappearance, or the like. RECIST is a set of rules published by an international consortium that define when tumors improve (“respond”), stay the same (“stabilize”), or worsen (“progress”) during treatment of a cancer patient. As used herein and unless otherwise noted, a patient “benefit” from a treatment may refer to any appropriate measure of improvement, including without limitation a RECIST response or longer PFS / TTP / DFS / TNT / TTNT. Beneficial or desired clinical results include, but are not limited to, alleviation or amelioration of one or more symptoms, diminishment of extent of disease, stabilized (i.e., not worsening) state of disease, preventing spread of disease, delay orslowing of disease progression, amelioration or palliation of the disease state, and remission (whether partial or total), whether detectable or undetectable. Benefit also includes prolonging survival as compared to expected survival if not receiving a treatment or if receiving a different treatment. Likewise, “lack of benefit” from a treatment may refer to any appropriate measure of worsening disease during treatment. Generally, disease stabilization is considered a benefit, although in certain circumstances, if so noted herein, stabilization may be considered a lack of benefit. A predicted or indicated benefit may be described as “indeterminate” if there is not an acceptable level of prediction of benefit or lack of benefit. In some cases, benefit is considered indeterminate if it cannot be calculated, e.g., due to lack of necessary data.

[0257] The treatment selected can be any treatment for which an association can be made to data derived in whole or in part from the molecular profiling data. For example, the treatment selected can be a targeted therapy such as a monoclonal antibody to a known cancer biomarker, or can be a therapy known to have an effect on cells that differentially express one or more genes or carry any manner of mutations or anomalies. The treatment may comprise experimental drugs or biologics, governmental or regulatory approved drug or biologics or any useful cocktails or combination thereof. In some embodiments, the selected treatments have been studied and approved for a particular indication that is the same as or different from the indication of the subject from whom the biological sample is obtained. ML models can be applied to the molecular profiling data to determine a course of treatment. The treatments can be based in part of the tumor type, as predicted or confirmed by the systems and methods provided herein.

[0258] When multiple treatment options are revealed by molecular profiling, decision rules can be put in place to prioritize the selection of a treatment regimen. For example, treatments may be prioritized based on direct results of molecular profiling, anticipated efficacy of therapeutic agent, prior history with the same or other treatments, expected side effects, availability of therapeutic agent, cost of therapeutic agent, drug-drug interactions, and other factors. Based on the recommended and prioritized therapeutic agent targets, a treating physician can decide on the course of treatment for a particular individual.

[0259] The methods described herein are used to provide personalized treatment options for cancer patients. In some embodiments, the subject has been previously treated with one or more therapeutic agents to treat the cancer. The cancer may be refractory to one of theseagents, e.g., by acquiring drug resistance mutations. Such acquired mutations may be identified over a time course using liquid biopsy. In some embodiments, the cancer is metastatic. In some embodiments, the subject has not previously been treated with one or more therapeutic agents identified by the method. Using molecular profiling, candidate treatments can be selected regardless of the stage, anatomical location, or anatomical origin of the cancer cells. Accordingly, molecular profiling methods and systems as described herein can identify treatments based on individual characteristics of diseased cells, e.g., tumor cells, and other personalized factors in a subject in need of treatment, as opposed to relying on a traditional one-size fits all approach that is conventionally used to treat individuals suffering from a disease, especially cancer. In some cases, the recommended treatments are those not typically used to treat the disease or disorder inflicting the subject. In some cases, the recommended treatments are used after standard-of-care therapies are no longer providing adequate efficacy.

[0260] The treating physician can use the results of the molecular profiling methods to optimize a treatment regimen for a patient. The candidate treatment identified by the methods as described herein can be used to treat a patient; however, such treatment is not required of the methods. Indeed, the analysis of molecular profiling results and identification of candidate treatments based on those results can be automated and does not require physician involvement.

[0261] Treatments associated with one or more of the biomarkers or sample / training / reference vectors may be determined using treatment association such as in any of International Patent Publications WO / 2007 / 137187 (Int’l Appl. No. PCT / US2007 / 069286), published November 29, 2007; WO / 2010 / 045318 (Int’l Appl. No. PCT / US2009 / 060630), published April 22, 2010; WO / 2010 / 093465 (Int’l Appl. No. PCT / US2010 / 000407), published August 19, 2010; WO / 2012 / 170715 (Int’l Appl. No. PCT / US2012 / 041393), published December 13, 2012; WO / 2014 / 089241 (Int’l Appl. No. PCT / US2013 / 073184), published June 12, 2014; WO / 2011 / 056688 (Int’l Appl. No. PCT / US2010 / 054366), published May 12, 2011; WO / 2012 / 092336 (Int’l Appl. No. PCT / US2011 / 067527), published July 5, 2012; WO / 2015 / 116868 (Int’l Appl. No. PCT / US2015 / 013618), published August 6, 2015; WO / 2017 / 053915 (Int’l Appl. No. PCT / US2016 / 053614), published March 30, 2017; WO / 2016 / 141169 (Int’l Appl. No. PCT / US2016 / 020657), published September 9, 2016; WO2018175501 (Int’l Appl. No. PCT / US2018 / 023438), published September 27, 2018; WO / 2020 / 113237 (based on Int’lPatent Appl. No. PCT / US2019 / 064078, filed December 2, 2019); WO / 2020 / 146554 (based on Int’l Patent Appl. No. PCT / US2020 / 012815, filed January 8, 2020; WO / 2021 / 112918 (based on Int’l Patent Appl. No. PCT / US2020 / 035990, filed June 3, 2020); WO / 2021 / 163706 (based on Int’l Patent Appl. No. PCT / US2021 / 018263, filed February 16, 2021); WO / 2021 / 222867 (based on Int’l Patent Appl. No. PCT / US2021 / 030351, filed April 30, 2021); WO / 2022 / 056328 (based on Int’l Patent Appl. No. PCT / US2021 / 049966, filed September 10, 2021); WO / 2022 / 103809 (based on Int’l Patent Appl. No. PCT / US2021 / 058741, filed November 10, 2021); and WO / 2022 / 132964 (based on Int’l Patent Appl. No. PCT / US2021 / 063603, file December 15, 2021); each of which publications is incorporated by reference herein in its entirety. Such rules can be updated as new information becomes available regarding various biomarkers, biosignatures, treatments, and the relationships thereof. The indication whether each treatment is likely to benefit the patient, not benefit the patient, or has indeterminate benefit may be weighted. For example, a likely or potential benefit may be a strong potential benefit or a lesser potential benefit. Such weighting can be based on any appropriate criteria, e.g., the strength of the evidence of the biomarker-treatment association, or the results of the profiling, e.g., a degree or level of over- or underexpression, mutation, or any other relevant state (e.g., wild type or altered). As the treating physician is ultimately responsible for treating their patient, such physician may use the report to assist in guiding their treatment recommendations.

[0262] Further exemplary associations between genomic and transcriptomic findings and therapy options are shown in Table 8. In Table 8, the column “Biomarker” uses gene identifiers commonly accepted in the scientific community at the time of filing and that can be used to look up the genes at various well-known databases such as the HUGO Gene Nomenclature Committee (HNGC; genenames.org), NCBI’s Gene database (ncbi.nlm.nih.gov / gene), GeneCards (genecards.org), Ensembl (ensembl.org), UniProt (uniprot.org), and others. Such identifiers may be used elsewhere in the disclosure. However, the following “Biomarkers” are genomic signatures and / or groups of genes as follows: gLOH: genomic loss of heterozygosity; HLA Genotype: human leukocyte antigen genotype; HRD: homologous recombination deficiency; HRR: homologous recombination repair genes (comprises ATM, BARD1, BRCA1, BRCA2, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, RAD54L, MLH1, MRE11, FANCA, NBN); MMR Deficiency: mismatch repair deficiency (comprises MLH1, MSH2, MSH6, PMS2); MSI: microsatellite instability; MMR Proficiency: mismatch repair proficiency (comprises MLH1,MSH2, MSH6, PMS2); MSS: microsatellite stable; TMB: tumor mutational burden (also referred to as tumor mutational load (TML)). The column “Alteration” notes the class of biological features that are associated with the various therapies. CNA refers to copy number alteration. The column “Therapy” notes therapies that are linked to the corresponding alteration / s in the biomarker / s. In some cases, certain therapies are preferential for certain cancer types as noted in parenthesis. In a non-limiting example, an RNA fusion detected involving the ALK gene may indicate brigatinib or lorlatinib as of likely benefit for NSCLC. Abbreviations in the Therapy column include: NSCLC: non-small cell lung cancer; CRC: colorectal cancer; DFSP: dermatofibrosarcoma protuberans; CUP: cancer of unknown primary; GIST: gastrointestinal stromal tumor; CNS: central nervous system. Table 8. Biomarkers and therapy associations Biomarker Alteration Therapy RNA Fusion crizotinib, ceritinib, alectinib, brigatinib (NSCLC), , e),cholangiocarcinoma) gLOH DNA Mut rucaparib (ovarian) (Gnmi) ation ce s,

[0263] In some embodiments, the genomic and transcriptomic findings provided by the molecular profiling are associated with an ongoing clinical trial. As a non-limiting example, the molecular profiling data may reveal a mutation that is a requirement for enrollment in a certain trial.

[0264] In some embodiments, therapy selection is based upon multiple marker signatures, including those that use machine learning and artificial intelligence. See, e.g., International Patent publications WO / 2020 / 113237 (based on Int’l Patent Appl. No. PCT / US2019 / 064078, filed December 2, 2019); WO / 2020 / 146554 (based on Int’l Patent Appl. No. PCT / US2020 / 012815, filed January 8, 2020; WO / 2021 / 112918 (based on Int’l Patent Appl. No. PCT / US2020 / 035990, filed June 3, 2020); WO / 2021 / 163706 (based on Int’l Patent Appl. No. PCT / US2021 / 018263, filed February 16, 2021); WO / 2021 / 222867 (based on Int’l Patent Appl. No. PCT / US2021 / 030351, filed April 30, 2021); WO / 2022 / 056328 (based on Int’l Patent Appl. No. PCT / US2021 / 049966, filed September 10, 2021); WO / 2022 / 103809 (based on Int’l Patent Appl. No. PCT / US2021 / 058741, filed November 10, 2021); and WO / 2022 / 132964 (based on Int’l Patent Appl. No. PCT / US2021 / 063603, file December 15, 2021); each of which publications is incorporated by reference herein in its entirety.

[0265] As further described herein, molecular profiling generates patient specific data that is used to determine and / or confirm a tumor type. The same data can be used to determine treatments of likely benefit or likely lack of benefit for the patient, either using biomarker- therapy association rules or ML models. See Example C herein for clinical application.

[0266] Additionally or alternatively to treatments, based on any classification of tumor type or sample type, the subject can be referred for additional screening modalities, e.g. using chest X ray, ultrasound, computed tomography, magnetic resonance imaging, or positron emission tomography. C. Report

[0267] A molecular profiling report can be delivered to the caregiver for the subject, e.g., the oncologist or other treating physician. The caregiver can use the results of the report to guide a treatment regimen for the subject. For example, the caregiver may administer one or more treatments indicated as likely benefit in the report. Similarly, the caregiver may avoid treating the patient with one or more treatments indicated as likely lack of benefit in the report. In some embodiments, such as when the report includes a biosignature indicating alikely recurrence or metastasis, the treating physician may choose, for example, a more aggressive treatment regimen, more frequent monitoring, or both. Such decisions are made by the caregiver with guidance from the report.

[0268] The report can comprise multiple sections of relevant information, including but not limited to: 1) description of the patient and sample; 2) a complete or partial listing of the biomarkers (nucleic acids, proteins, or other biological matter of interest) in the molecular profile; 3) a description of the state of one or more of the biomarkers in the molecular profile as determined for the subject; 4) a description of one or more biological signatures as determined for the molecular profile, such as microsatellite stability, tumor mutational load / burden, recurrence predictors, treatment response predictors; and / or metastasis predictors; 5) an overview of the tumor type as either determined (e.g. for CUP) or confirmed using the systems and methods provided herein; 6) one or more treatment associated with one or more of the biomarkers, groups of biomarkers, and / or biological signatures determined for the molecular profile; 7) an indication whether one or more treatment is likely to benefit the patient, not benefit the patient, or has indeterminate benefit; 8) one or more clinical trials for which the patient may be eligible based on the molecular profile; 9) an indication whether the cancer is predicted to recur and / or metastasize; and / or 10) evidence relevant to the foregoing, such as literature reports and / or clinical trial results.

[0269] The description of the molecular profile within the report can include such information as the laboratory technique used to assess each biomarker, optionally including the result and any criteria used to score each technique. By way of non-limiting example, the criteria for scoring a copy number alteration (or variation, CNA or CNV) may be a presence (i.e., a copy number that is greater or lower than the “normal” copy number present in a subject who does not have cancer, or statistically identified as present in the general population, typically diploid) or absence (i.e., a copy number that is considered the same as the “normal” copy number present in a subject who does not have cancer, or statistically identified as present in the general population, typically diploid).

[0270] The report can be computer generated, and can be a printed report, a computer file or both. The report can be made accessible via a secure web portal. The report may be displayed using any desired medium. In some embodiments, the display is a printout, a computer file, including without limitation a pdf file, or may be displayed via an applicationon a computer display such as a computer monitor, laptop display, tablet, smartphone, or other mobile device.

[0271] In an aspect, the disclosure provides a system for generating a molecular profiling report such as described above, comprising: (a) at least one host server; (b) at least one user interface for accessing the at least one host server to access and input data; (c) at least one processor for processing the inputted data; (d) at least one memory coupled to the processor for storing the processed data and instructions for: i) accessing a biomarker status (e.g., a tumor derived mutation) determined by molecular profiling methodology as described herein; and ii) identifying biomarkers, biosignatures and related data and any information derived using such data (treatments, clinical trials, phenotypes, predictions, etc., as described herein); and (e) at least one display for displaying results and outcomes of the molecular profiling. In some embodiments, the system further comprises at least one memory coupled to the processor for storing the processed data and instructions for identifying, based on the generated molecular profile according to the methods above, at least one therapy with potential benefit for treatment of the cancer; and at least one display for display thereof. The system may further comprise at least one database comprising references for various biomarker states, data for drug / biomarker associations, or both. The at least one display can be a report provided by the present disclosure. VII. MEASUREMENT AND COMPUTER SYSTEMS

[0272] FIG.12 illustrates a measurement system 1200 according to an embodiment of the present disclosure. The system as shown includes a sample 1205, such as cell-free nucleic acid molecules (e.g., DNA and / or RNA) within an assay device 1210, where an assay 1208 can be performed on sample 1205. For example, sample 1205 can be contacted with reagents of assay 1208 to provide a signal of a physical characteristic 1215 (e.g., sequence information of a cell-free nucleic acid molecule). An example of an assay device can be a flow cell that includes probes and / or primers of an assay or a tube through which a droplet moves (with the droplet including the assay). Physical characteristic 1215 (e.g., a fluorescence intensity, a voltage, or a current), from the sample is detected by detector 1220. Detector 1220 can take a measurement at intervals (e.g., periodic intervals) to obtain data points that make up a data signal. In one embodiment, an analog-to-digital converter converts an analog signal from the detector into digital form at a plurality of times.

[0273] Assay device 1210 and detector 1220 can form an assay system, e.g., a sequencing system that performs sequencing according to embodiments described herein, such as NGS. A data signal 1225 is sent from detector 1220 to logic system 1230. As an example, data signal 1225 can be used to determine sequences and / or locations in a reference genome of nucleic acid molecules (e.g., DNA and / or RNA). Data signal 1225 can include various measurements made at a same time, e.g., different colors of fluorescent dyes or different electrical signals for different molecule of sample 1205, and thus data signal 1225 can correspond to multiple signals. Data signal 1225 may be stored in a local memory 1235, an external memory 1240, or a storage device 1245. The assay system can be comprised of multiple assay devices and detectors.

[0274] Logic system 1230 may be, or may include, a computer system, ASIC, microprocessor, graphics processing unit (GPU), etc. It may also include or be coupled with a display (e.g., monitor, LED display, etc.) and a user input device (e.g., mouse, keyboard, buttons, etc.). Logic system 1230 and the other components may be part of a stand-alone or network connected computer system, or they may be directly attached to or incorporated in a device (e.g., a sequencing device) that includes detector 1220 and / or assay device 1210. Logic system 1230 may also include software that executes in a processor 1250. Logic system 1230 may include a computer readable medium storing instructions for controlling measurement system 1200 to perform any of the methods described herein. For example, logic system 1230 can provide commands to a system that includes assay device 1210 such that sequencing or other physical operations are performed. Such physical operations can be performed in a particular order, e.g., with reagents being added and removed in a particular order. Such physical operations may be performed by a robotics system, e.g., including a robotic arm, as may be used to obtain a sample and perform an assay.

[0275] Measurement system 1200 may also include a treatment device 1260, which can provide a treatment to the subject. Treatment device 1260 can determine a treatment and / or be used to perform a treatment. Examples of such treatment can include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, and stem cell transplant. Logic system 1230 may be connected to treatment device 1260, e.g., to provide results of a method described herein. The treatment device may receive inputs from other devices, such as an imaging device and user inputs (e.g., to control the treatment, such as controls over a robotic system).

[0276] Measurement system 1200 may also include a reporting device 1255, which can present results of any of the methods describe herein, e.g., as determined using the measurement system. Reporting device 1255 can be in communication with a reporting module within logic system 1230 that can aggregate, format, and send a report to reporting device 1255. The reporting module can present information determined using any of the method described herein. The information can be presented by reporting device 1255 in any format that can be recognized and interpreted by a user of the measurement system 1200. For example, the information can be presented by reporting device 1255 in a displayed, printed, or transmitted format, or any combination thereof.

[0277] Any of the computer systems mentioned herein may utilize any suitable number of subsystems. Examples of such subsystems are shown in FIG.13 in computer system 10. In some embodiments, a computer system includes a single computer apparatus, where the subsystems can be the components of the computer apparatus. In other embodiments, a computer system can include multiple computer apparatuses, each being a subsystem, with internal components. A computer system can include desktop and laptop computers, tablets, mobile phones and other mobile devices.

[0278] The subsystems shown in FIG.13 are interconnected via a system bus 75. Additional subsystems such as a printer 74, keyboard 78, storage device(s) 79, monitor 76 (e.g., a display screen, such as an LED), which is coupled to display adapter 82, and others are shown. Peripherals and input / output (I / O) devices, which couple to I / O controller 71, can be connected to the computer system by any number of means known in the art such as input / output (I / O) port 77 (e.g., USB, FireWire®). For example, I / O port 77 or external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 10 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus 75 allows the central processor 73 to communicate with each subsystem and to control the execution of a plurality of instructions from system memory 72 or the storage device(s) 79 (e.g., a fixed disk, such as a hard drive, or optical disk), as well as the exchange of information between subsystems. The system memory 72 and / or the storage device(s) 79 may embody a computer readable medium. Another subsystem is a data collection device 85, such as a camera, microphone, accelerometer, and the like. Any of the data mentioned herein can be output from one component to another component and can be output to the user.

[0279] A computer system can include a plurality of the same components or subsystems, e.g., connected together by external interface 81, by an internal interface, or via removable storage devices that can be connected and removed from one component to another component. In some embodiments, computer systems, subsystem, or apparatuses can communicate over a network. In such instances, one computer can be considered a client and another computer a server, where each can be part of a same computer system. A client and a server can each include multiple systems, subsystems, or components. In various embodiments, methods may involve various numbers of clients and / or servers, including at least 10, 20, 50, 100, 200, 500, 1,000, or 10,000 devices. Methods can include various numbers of communication messages between devices, including at least 100, 200, 500, 1,000, 10,000, 50,000, 100,000, 500,00, or one million communication messages. Such communications can involve at least 1 MB, 10 MB, 100 MB, 1 GB, 10 GB, or 100 GB of data.

[0280] Aspects of embodiments can be implemented in the form of control logic using hardware circuitry (e.g., an application specific integrated circuit or field programmable gate array) and / or using computer software stored in a memory with a generally programmable processor in a modular or integrated manner, and thus a processor can include memory storing software instructions that configure hardware circuitry, as well as an FPGA with configuration instructions or an ASIC. As used herein, a processor can include a single-core processor, multi-core processor on a same integrated chip, or multiple processing units on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will know and appreciate other ways and / or methods to implement embodiments of the present disclosure using hardware and a combination of hardware and software.

[0281] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language such as, for example, Java, C, C++, C#, Objective-C, Swift, or scripting language such as R, Perl or Python using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer readable medium for storage and / or transmission. A suitable non-transitory computer readable medium can include random access memory (RAM), a read only memory (ROM), a magnetic medium such as a hard-drive or a floppy disk, or an optical medium such as a compact disk (CD) or DVD (digital versatile disk) or Blu-ray disk, flash memory, and thelike. The computer readable medium may be any combination of such devices. In addition, the order of operations may be re-arranged. A process can be terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.

[0282] Such programs may also be encoded and transmitted using carrier signals adapted for transmission via wired, optical, and / or wireless networks conforming to a variety of protocols, including the Internet. As such, a computer readable medium may be created using a data signal encoded with such programs. Computer readable media encoded with the program code may be packaged with a compatible device (e.g., as firmware) or provided separately from other devices (e.g., via Internet download). Any such computer readable medium may reside on or within a single computer product (e.g., a hard drive, a CD, or an entire computer system), and may be present on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing any of the results mentioned herein to a user.

[0283] Any of the methods described herein may be totally or partially performed with a computer system including one or more processors, which can be configured to perform the steps. Any operations performed with a processor (e.g., aligning, determining, comparing, computing, calculating) may be performed in real-time. The term “real-time” may refer to computing operations or processes that are completed within a certain time constraint. The time constraint may be 1 minute, 1 hour, 1 day, or 7 days. Thus, embodiments can be directed to computer systems configured to perform the steps of any of the methods described herein, potentially with different components performing a respective step or a respective group of steps. Although presented as numbered steps, steps of methods herein can be performed at a same time or at different times or in a different order. Additionally, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of a step may be optional. Additionally, any of the steps of any of the methods can be performed with modules, units, circuits, or other means of a system for performing these steps.

[0284] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of embodiments of the disclosure.However, other embodiments of the disclosure may be directed to specific embodiments relating to each individual aspect, or specific combinations of these individual aspects.

[0285] The above description of example embodiments of the present disclosure has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the disclosure to the precise form described, and many modifications and variations are possible in light of the teaching above.

[0286] A recitation of “a”, “an” or “the” is intended to mean “one or more" unless specifically indicated to the contrary. The use of “or” is intended to mean an “inclusive or,” and not an “exclusive or” unless specifically indicated to the contrary. Reference to a “first” component does not necessarily require that a second component be provided. Moreover, reference to a “first” or a “second” component does not limit the referenced component to a particular location unless expressly stated. The term “based on” is intended to mean “based at least in part on.”

[0287] The claims may be drafted to exclude any element which may be optional. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely”, “only”, and the like in connection with the recitation of claim elements, or the use of a “negative” limitation.

[0288] All patents, patent applications, publications, and descriptions mentioned herein are incorporated by reference in their entirety for all purposes. None is admitted to be prior art. Where a conflict exists between the instant application and a reference provided herein, the instant application shall dominate. VIII. EXAMPLES A. Case Study 1

[0289] This case study provides an illustration of the system and methods provided herein in the clinical setting. A 30-year-old woman presented to her primary care physician with abdominal distension and gastrointestinal symptoms. Imaging revealed widespread metastatic disease, including pleural effusions, pulmonary nodules, and widespread lymphadenopathy. Biopsy performed on a soft tissue, right abdominal wall mass indicated metastatic adenocarcinoma.

[0290] It was important to identify the origin of the disease to guide proper treatment. The most likely types for a young woman are either breast or gynecological cancer. Immunohistochemistry (IHC) for conventional tissue markers was performed. By IHC, the tumor was found to be positive for CK7, TRPS1 patchy+, GATA3 rare+ and negative for CK20 and PAX8. The CK7+ status indicates epithelial origin. The negative results for CK20 and PAX8 suggest that the primary tumor was not gynecological or colorectal. The weakly positive results for TRPS1 and GATA3 were not definitive for breast diagnosis. Based on the combined results, the patient was diagnosed with a metastatic malignancy of unknown primary (CUP).

[0291] A tumor specimen was sent to our laboratories (Caris Life Sciences, Phoenix, AZ) for molecular profiling and tissue of origin analysis as described herein. The molecular profiling comprised selected IHCs and a hybrid NGS assay that performs WES and WTS. The tissue origin was predicted using the systems and methods provided herein. The analysis revealed a 99% probability that the tumor was mesothelioma, a rare form of cancer with small tumors that often do not present symptoms until stage III or IV.

[0292] The molecular profiling further revealed a BAP1 deletion, which is characteristic of mesothelioma. The patient was also identified with a BRCA1 pathogenic variant that was likely to be germline, which has also been linked to mesothelioma. The diagnosis was confirmed using IHC testing including calretinin (CALB2 gen), a standard biomarker for mesothelioma, and WT1, which is also consistent with the diagnosis.

[0293] Taken together, the molecular profiling and tumor origin analyses provided herein revealed that the patient had metastatic mesothelioma, thereby avoiding treatment with breast cancer drugs that may have been ineffective. The molecular profiling also revealed that the tumor was TMB low, microsatellite stable (MSS), and LOH low, indicating lack of benefit from immunotherapy. B. Case Study 2

[0294] This case study provides another illustration of the system and methods provided herein in the clinical setting. The patient was a male diagnosed with a pleomorphic undifferentiated sarcoma, which was confirmed with a second opinion. The oncologist believed likely treatment may be palliative chemotherapy for an uncurable disease.

[0295] A tumor specimen was sent to our laboratories (Caris Life Sciences, Phoenix, AZ) for molecular profiling and tissue of origin analysis as described herein in hopes of finding a targetable biomaker present. The molecular profiling comprised selected IHCs and a hybrid NGS assay that performs WES and WTS. The tissue origin was predicted using the systems and methods provided herein. The analysis revealed a 99% probability that the tumor was of hematological origin. Further analysis by pathologists showed that the cancer was an unusual mature B-cell lymphoma. The patient’s diagnosis was changed to lymphoma and appropriate treatment was prescribed. Radiation placed the patient into a near clinical complete remission (CR), and the treating physician plans further possible curative treatment based on the final lymphoma diagnosis. C. GPSai: A Clinically Validated AI Tool for Tissue of Origin Prediction During Routine Tumor Profiling

[0296] This Example provides an implementation of a tissue of origin classifier according to the systems and methods herein.

[0297] Although a majority of metastatic tumors are readily classified based on standard pathological and clinical evaluation, a subset is more challenging to diagnose and may present with unclear or potentially incorrect primary organ and histological diagnoses. Among these are ‘cancers of unknown primary’ (CUP), which compose ~2% of cancer cases (1-4) and are associated with lack of effective treatment options and poor outcomes (1, 5-8). Metastatic tumors may also be assigned a diagnostic label but present with diagnostic discrepancies upon further evaluation. The rate of pathology diagnostic discrepancies, including tumor lineage, is estimated to be between 6% and 9% (9-11), with one study demonstrating a change in treatment plan for 38% of cancers with a discrepant diagnosis (11). As such, accurate diagnosis of pathologically ambiguous cancers is critical to effective treatment planning.

[0298] In recent years, artificial intelligence (AI) tools have been developed to help address such diagnostic ambiguities. Gene expression profiling-based models for identification of tumor tissue of origin report diagnostic accuracies in the range of 73-96%, with varying ability to classify CUP cases (12-19). More recently, next-generation sequencing-based genomic profiling has been used to identify tumor tissue of origin and reveal actionable targets simultaneously (20-30), an approach supported by the American Society of Clinical Oncology (ASCO) (31). However, factors limiting the clinical impact of tissue of originclassifiers include low numbers of samples or cancer types used in the training set, excessive cost of the technique, low quality RNA for gene expression-based assays, contamination of the tumor biopsy with surrounding tissue (32), and limited tissue availability for testing (33, 34).

[0299] More broadly, there are multiple barriers to successful clinical implementation of AI models, with most remaining in the research phase (35). In addition to an appropriate and well-curated training set and robust clinical validation, optimization of clinical workflows and interpretation of the results in the context of other clinical information is necessary (35, 36). Accordingly, AI tools can be developed to support and augment the expert opinion of pathologists and physicians—an idea embodied in the ‘physician-in-the-loop’ paradigm (37- 39).

[0300] In this Example, a Genomic Probability Score artificial intelligence (GPSai) model system was developed through deep learning training of a curated dataset of 201,612 cases and predicts histologic diagnoses with greatly enhanced granularity and performance compared to other commercially available tests. In addition to validating the accuracy and performance of GPSai, we herein demonstrate that GPSai has clinical impact by driving changes in diagnosis and altering guideline-directed targeted-therapy eligibility in a meaningful subset of patients. Together, this study shows that GPSai is an accurate, clinically validated AI tool that supports a ‘physician-in-the-loop’ approach to the diagnosis of CUP and other pathologically ambiguous tumors. 1. Methods a) Specimen selection

[0301] Historical cases from the years 2019-2023 were used to train the model (N=201,612) and for retrospective validation (N=21,549, including 443 CUP). Cases were selected contingent on having 1) both gene expression and DNA variant data; 2) an International Classification of Disease (ICD) Primary Tumor Site (PTS); and 3) a histology that mapped to one of 90 Oncotree-based labels (40) reportable by the model (see FIGs.9A- 9C). Baseline demographic information is shown in FIGs.14A-B. Data from cases profiled between March-October 2024 after clinical launch of the GPSai deep learning model were also analyzed retrospectively (N=80,308).b) RNA and DNA sequencing

[0302] Molecular profiling was performed in our laboratories (Caris Life Sciences; Phoenix, AZ, USA), a College of American Pathologists (CAP) / Clinical Laboratory Improvement Amendments (CLIA)-certified laboratory with additional ISO15189 and ISO13485 certifications. Because this study used samples spanning six years, the methods (including materials and software) used for DNA and RNA profiling evolved over time. Initially, DNA was profiled using a 592 whole-gene panel prior to adoption of whole exome sequencing (WES) with targeted enrichment of 720 clinically relevant genes in 2020; whole exome sequencing (WTS) was performed separately. In 2023, this approach was replaced by a hybrid NGS offering, referred to as “MI Tumor Seek Hybrid,” which analyzes RNA and DNA from the same total nucleic acid extraction to perform WTS and WES, respectively.

[0303] Details of these sequencing methods are as follows. Formalin-fixed, paraffin- embedded (FFPE) slides underwent review by a board-certified pathologist to measure tumor content and mark the area for microdissection. In this example, a minimum of 20% tumor content in the area for microdissection was required for next-generation sequencing (NGS). A range of tumor content levels (20-100%) were included in the training and validation cohorts. Nucleic acid was extracted using appropriate FFPE kits for RNA, DNA, or total nucleic acid. DNA sequencing was performed using a 592-whole gene panel on the NextSeq platform or by whole exome sequencing with enrichment of 720 clinically relevant genes on the NovaSeq 6000 platform (Illumina, Inc., San Diego, CA). WTS was performed using the Illumina Novaseq 6000 platform to an average of 60M reads. Raw WTS data was demultiplexed by Illumina Dragen BioIT accelerator, trimmed, counted, PCR-duplicates removed, and aligned to human reference genome (hg19 / hg38) by STAR aligner. For transcription counting, transcripts per million (TPM) molecules were generated using the Salmon expression pipeline (Patro et al, 2017). For MI Tumor Seek Hybrid, RNA was labeled during first strand cDNA synthesis by adapter sequences on the 5’ end of the cDNA primers and whole transcriptome sequencing (WTS) was performed with 720 clinically relevant genes sequenced at increased depth. Sequencing data was extracted into split FASTQ files (RNA and DNA) for processing. DNA variants detected were mapped to reference genome hg38, and bioinformatics tools such as BWA, SamTools, Pindel, and snpFF were incorporated to perform variant calling functions; germline variants were filtered with various germline databases, including dbSNP. Genetic variants identified were interpreted by board-certified molecular geneticists and categorized as ‘pathogenic,’ ‘likely pathogenic,’ ‘variant ofunknown significance,’ ‘likely benign,’ or ‘benign,’ according to the American College of Medical Genetics and Genomics (ACMG) standards. Pathogenic and likely-pathogenic variants are counted as “reportable” to the ordering physician. c) Deep learning model development

[0304] The model was trained in PyTorch (version 2.4.1) using variant data, gene expression, and binary sex (male / female). The pathological diagnosis entered by the submitting site was mapped to an Oncotree-derived code (see FIGs.9A-C) and used as the training label. Twenty distinct neural networks were trained using training sets consisting of samples from each cancer category that summed to 201,612 total samples, with each neural network corresponding to a branch of the hierarchy. Model optimization was performed using a subset of 46,708 samples. A holdout set consisting of ~10% of available tumors in each cancer category was used for the retrospective validation (21,106 total mapped tumor samples and 443 CUP). The first network differentiated between 26 major cancer categories, and subsequent networks differentiated between 64 subcategories, which are arranged in a hierarchical manner. See FIGs.9A-9C. Given the hierarchical nature of the diagnostic labels, hierarchical metrics were used in the validations performed (41), as illustrated in FIGs.15A- 15D.

[0305] A description of how the model assigns scores to each category and how the threshold score was set to define a positive call by the model is as follows. The model assigns a non-negative score to each category so that the scores of all major categories sum to unity (1.0, or 100%). Subcategory scores also sum to their major categories’ score, as shown in FIG.16. A combination of call rate and hierarchical positive predictive value (PPV) for metastatic cases was used to set the threshold score for inclusion of the GPSai result on the final molecular report (and also defines the call rates reported herein). A GPSai score of ≥0.55 on a category label was determined as the intersection of optimal call rate and hierarchical PPV (see FIG.17), but higher scores suggest additional confidence in the classification. Given that major category scores can sum to unity, only one category can achieve a score of ≥0.55. Other example thresholds can be used, e.g., any value greater than 0.50.

[0306] RNA-seq data was downloaded for open-access cases directly from TCGA (The Cancer Genome Atlas) Genomic Data Commons portal. Pathogenic variant information was obtained from the supplemental section of Sanchez-Vega et al. (42).d) Pathology procedures for review of GPSai results

[0307] A ‘Critical Value Discrepancy’ is opened when GPSai assigns a label to a CUP case or if a non-CUP case has a GPSai result that differs from the diagnosis submitted by the ordering physician along with the patient sample. “Critical Value Discrepancies” are reviewed by American Board of Pathology-certified pathologists. The clinical report is updated (diagnosis, therapeutic biomarker panel, guideline-driven drug associations, and clinical trials) only when a CUP case has a GPSai score ≥90% and the call is supported by orthogonal evidence. When the GPSai score is ≥90% and results cannot be proven via orthogonal methods, or when there is more than one GPSai result with none ≥90%, the result is included in the clinical report and the option is provided for lineage change at the behest of the ordering physician. For non-CUP cases with ‘Critical Value Discrepancies’, a similar procedure is followed. If the submitted diagnosis has a GPSai score of 0% or if a single GPSai category has a score ≥90%, confirmatory testing is ordered. Otherwise, the result is reported with discretionary lineage change. Criteria for inclusion of a GPSai result on the clinical report is summarized in Table 9. Table 9. Pathology procedures for review of GPSai results “Critical Value Discrepancy” h y a a

[0308] Regarding physician surveys, between March and October 2024, 957 physicians who received a ‘Critical Value Discrepancy’ on the clinical report for their patient (even if this did not result in a diagnosis change) were eligible for survey participation. Molecular Science Liaisons collected survey data from 96 physicians.2. Results a) Accuracy and robustness of GPSai for identifying tumor tissue of origin

[0309] We first determined the percentage of CUP and non-CUP cases that GPSai was able to classify into one of 90 categories / subcategories reportable by the model (see FIGs.9A- 9C), revealing a call rate of 84.0% (n=372 / 443) in CUP cases and 96.3% (n=20,325 / 21,106) in non-CUP cases. We then evaluated the accuracy of the GPSai calls in the non-CUP cases compared to the pathologist-submitted diagnostic labels. The global accuracy for the top major cancer category predicted by the model (TOP1 PPV) was 95.0%, which was further improved to 98.3% when considering the top two major categories predicted by the model (TOP2 PPV). Hierarchical PPV, which incorporates accuracy of the cancer subcategories, was 93.3%. Performance of the model was only modestly reduced in samples from metastatic sites compared to primary sites. See Table 10. Hierarchical PPV and hierarchical sensitivity exceeded 91% for specific metastatic sites (liver, lung, lymph node, bone, and brain. See Table 11. In the tables, TOP1 = top major category selected by GPSai model, and TOP2 = top 2 major categories selected by the model. Table 10. GPSai model performance in retrospective validation Samples, N Call Rate, hP TOP1 TOP2 % PV, % hSens, % PPV, % PPV, % Non-CUPGlobal 21106 96.3 93.3 92.7 95.0 98.3 Primary 12668 97.2 94.4 93.7 96.3 98.9 Metastatic 8438 94.8 91.8 91.1 93.2 97.4 CUP443 90.5 N / A N / A N / A N / ATable 11. GPSai model performance in metastatic sites Samples, NCall Rate, h TOP1 TOP2 % PPV, % hSens, % PPV, % PPV, % Lymph node 1578 94.0 92.1 91.6 93.3 97.6 Liver 2048 96.6 93.1 92.5 94.9 97.4 Lung 751 94.0 91.5 91.2 92.1 97.2Bone 630 96.3 94.3 94.2 94.7 97.7 Brain 324 95.1 92.5 91.9 93.2 97.4

[0310] We next evaluated how well GPSai performed for each individual cancer category. As in the global validation, model performance for individual cancer categories was only slightly reduced for samples from metastatic sites compared to primary sites. The model was able to accurately classify commonly diagnosed tumor types (e.g., non-small cell lung cancer (NSCLC), breast, bowel, ovarian epithelial, pancreatobiliary, prostate, melanoma, cervix / uterine, and stomach / esophagus) with high accuracy (PPV) and sensitivity in both primary and metastatic tumors; many of these tumor types achieved 98-99% PPV and sensitivity. Some miscalls were observed for metastatic cervical / uterine tumors (85% PPV), which were primarily classified as ovarian / fallopian / peritoneal tumors. Miscalls for stomach / esophagus (79% PPV) were more dispersed, but NSCLC and pancreatobiliary were the most common incorrectly predicted categories. b) External TCGA validation

[0311] As an external validation, we applied the model to TCGA data from 6,820 primary and metastatic tumors. We first aligned the 26 broad cancer categories predicted by our model with TCGA designated tumor types. The external validation showed 97.9% overall concordance and 98.6% call rate for GPSai. Bowel, breast, cervix / uterine, kidney, NSCLC, oral squamous cell carcinoma (OSCC), prostate, and thyroid were the most represented categories in TCGA dataset. Only NSCLC demonstrated a cluster of miscalls in neuroendocrine (n=18) and OSCC (n=8) categories, illustrating some ambiguity between these three cancer types. A majority of “no calls” were also in NSCLC (n=37) and OSCC (n=30) categories. c) Quantification of diagnosis changes prompted by GPSai

[0312] To gain a better understanding of the clinical relevance of GPSai, we next explored its implementation in the course of routine molecular testing conducted over eight months. Out of 80,308 total profiled cases, 1.2% (n=957) of cases had a ‘Critical Value Discrepancy’ from the submitted diagnosis, and the diagnosis was ultimately changed in 73.6% (n=704) of these.57.4% (n=404) of cases with diagnosis change were submitted as CUP and 42.6% (n=300) were non-CUP cases that could be considered pathologically ambiguousmisdiagnoses. See Table 12 for numbers of cases changed. As described above, in order for a ‘Critical Value Discrepancy’ to qualify for diagnosis change on the clinical report provided to the ordering physician, the GPSai score can be required to be ≥90% and supported by confirmatory data, which could include diagnostic IHC, mutational signatures (e.g., ultraviolet, tobacco (43)), hallmark fusions (e.g., TMPRSS2:ERG), viral reads (EBV, MCPYV, HPV16 / 18), imaging, or clinical history. We specifically examined viral reads, ultraviolet and tobacco mutational signatures, and fusions among cases with a discrepant GPSai call from the outside diagnosis, which revealed that orthogonal evidence more often supported GPSai calls over the outside diagnosis. See FIG.18; Table 13. This analysis underscores the utility of GPSai to act as a ‘second opinion’ for CUP and other pathologically ambiguous metastatic tumors submitted for molecular testing, thus prompting additional workup that should be considered in conjunction with the GPSai result to achieve an accurate diagnosis. Table 12: Cases with diagnosis change due to GPSai results Count of Submitted Row Labels Lineage CUP 404 Lung Non-small cell lung cancer (NSCLC) 105 Colorectal adenocarcinoma 38 Ovarian Surface Epithelial Carcinomas 28 Cholangiocarcinoma 25 Breast carcinoma 22 Pancreatic Adenocarcinoma 21 Kidney cancer 18 Bladder carcinoma - urothelial 18 Squamous Cell Skin Cancer 11 Mesothelioma 11 Thymic Carcinoma 9 Prostatic Adenocarcinoma 9 Neuroendocrine carcinoma 9 Melanoma 9 Gastric Adenocarcinoma 8 Endometrial carcinoma 8 Hepatocellular Carcinoma 8 Esophagogastric Junction Carcinoma 7 HNSCC 4 Soft tissue sarcoma 4 Salivary gland carcinoma 4Cervical carcinoma 3 Thyroid carcinoma 3 Uterine Serous Carcinoma 3 Testicular carcinoma 3 Merkel Cell Carcinoma (MCC) 3 Anal Carcinoma 2 Lung Small Cell Cancer (SCLC) 2 Glioblastoma 2 SCC of rectal region 1 Lymphoma 1 Endometrial Stromal Sarcoma 1 Ewing Sarcoma 1 Penile carcinoma 1 Vulvar carcinoma 1 Male Genital Tract Malignancy 1 Lung Non-small cell lung cancer (NSCLC) 87 Bladder carcinoma - urothelial 11 HNSCC 10 Mesothelioma 9 Squamous Cell Skin Cancer 7 Lung Small Cell Cancer (SCLC) 6 Thymic Carcinoma 5 Breast carcinoma 5 Colorectal adenocarcinoma 4 Endometrial carcinoma 4 Esophagogastric Junction Carcinoma 3 Kidney cancer 3 Soft tissue sarcoma 3 Neuroendocrine carcinoma 3 Thyroid carcinoma 3 Melanoma 3 Hepatocellular Carcinoma 2 Prostatic Adenocarcinoma 1 Salivary gland carcinoma 1 Pancreatic Adenocarcinoma 1 Anaplastic thyroid carcinoma 1 Ovarian Surface Epithelial Carcinomas 1 Lung Non-small cell lung cancer (NSCLC) (squamous to adenocarcinoma) 1 Colorectal adenocarcinoma 22 Endometrial carcinoma 6 Neuroendocrine carcinoma 5 Prostatic Adenocarcinoma 2Lung Non-small cell lung cancer (NSCLC) 2 Ovarian Surface Epithelial Carcinomas 1 Cholangiocarcinoma 1 Salivary gland carcinoma 1 Squamous Cell Skin Cancer 1 Anal Carcinoma 1 Pancreatic Adenocarcinoma 1 Hepatocellular Carcinoma 1 Breast carcinoma 21 Lung Non-small cell lung cancer (NSCLC) 9 Ovarian Surface Epithelial Carcinomas 3 Squamous Cell Skin Cancer 3 Bladder carcinoma - urothelial 2 Hepatocellular Carcinoma 1 Anal Carcinoma 1 Salivary gland carcinoma 1 Melanoma 1 Neuroendocrine carcinoma 19 Merkel Cell Carcinoma (MCC) 6 Breast carcinoma 5 Thyroid carcinoma 1 Lung Non-small cell lung cancer (NSCLC) 1 Melanoma 1 Neuroblastoma 1 Prostatic Adenocarcinoma 1 Hepatocellular Carcinoma 1 Uterine Sarcoma 1 Meningioma 1 Endometrial carcinoma 18 Ovarian Surface Epithelial Carcinomas 9 Lung Non-small cell lung cancer (NSCLC) 2 Cholangiocarcinoma 1 Uterine Sarcoma 1 Colorectal adenocarcinoma 1 Neuroendocrine carcinoma 1 Cervical carcinoma 1 Multiple Myeloma 1 Lymphoma 1 Bladder carcinoma - urothelial 16 Prostatic Adenocarcinoma 5 Neuroendocrine carcinoma 2 Lung Non-small cell lung cancer (NSCLC) 2 Lymphoma 1Cervical carcinoma 1 Colorectal adenocarcinoma 1 Soft tissue sarcoma 1 Breast carcinoma 1 HNSCC 1 Endometrial carcinoma 1 Pancreatic Adenocarcinoma 16 Colorectal adenocarcinoma 5 Neuroendocrine carcinoma 3 Hepatocellular Carcinoma 1 Bladder carcinoma - urothelial 1 Osteosarcoma 1 Cholangiocarcinoma 1 Breast carcinoma 1 Melanoma 1 Prostatic Adenocarcinoma 1 Lung Non-small cell lung cancer (NSCLC) 1 Ovarian Surface Epithelial Carcinomas 10 Mesothelioma 5 Kidney cancer 1 Non Epithelial Ovarian Cancer (non-EOC) 1 Breast carcinoma 1 Uterine Serous Carcinoma 1 Melanoma 1 Salivary gland carcinoma 9 Squamous Cell Skin Cancer 7 Lung Non-small cell lung cancer (NSCLC) 1 Neuroendocrine carcinoma 1 Cervical carcinoma 9 Bladder carcinoma - urothelial 3 Ovarian Surface Epithelial Carcinomas 2 Neuroendocrine carcinoma 2 Vulvar carcinoma 1 Endometrial carcinoma 1 Gastric Adenocarcinoma 8 Cholangiocarcinoma 2 Colorectal adenocarcinoma 2 Melanoma 1 Ovarian Surface Epithelial Carcinomas 1 Pancreatic Adenocarcinoma 1 Neuroendocrine carcinoma 1 Cholangiocarcinoma 7 Lung Non-small cell lung cancer (NSCLC) 3Colorectal adenocarcinoma 1 Hepatocellular Carcinoma 1 Neuroendocrine carcinoma 1 Pancreatic Adenocarcinoma 1 Ovarian Surface Epithelial Carcinoma 6 Esophagogastric Junction Carcinoma 1 Colorectal adenocarcinoma 1 Endometrial carcinoma 1 Melanoma 1 Thyroid carcinoma 1 Lung Non-small cell lung cancer (NSCLC) 1 Small Bowel Adenocarcinoma 5 Pancreatic Adenocarcinoma 3 Gastric Adenocarcinoma 1 Lung Non-small cell lung cancer (NSCLC) 1 HNSCC 5 Lung Non-small cell lung cancer (NSCLC) 2Bladder carcinoma - urothelial 1 Ewing Sarcoma 1 Kidney cancer 5 Lung Non-small cell lung cancer (NSCLC) 3 Breast carcinoma 1 Lymphoma 1 Hepatocellular Carcinoma 3 Salivary gland carcinoma 1 Breast carcinoma 1 Lung Non-small cell lung cancer (NSCLC) 1 Squamous Cell Skin Cancer 3 Male Genital Tract Malignancy 1 HNSCC 1 Lung Non-small cell lung cancer (NSCLC) 1 Melanoma 3 Soft tissue sarcoma 2 Osteosarcoma 1 Thyroid Carcinoma 3 Lungcell lung cancer (NSCLC) 2 Kidney cancer 1 Spindle cell sarcoma 2 Liposarcoma 2 Lung Small Cell Cancer (SCLC) 2 Lung Non-small cell lung cancer (NSCLC) 1 Merkel Cell Carcinoma (MCC) 1Prostatic Adenocarcinoma 2 Cholangiocarcinoma 1 Osteosarcoma 1 Esophagogastric Junction Carcinoma 2 Lung Non-small cell lung cancer (NSCLC) 1 Ovarian Surface Epithelial Carcinomas 1 Soft tissue sarcoma 2 Meningioma 1 Squamous Cell Skin Cancer 1 Esophageal carcinoma 2 Lung Non-small cell lung cancer (NSCLC) 1 Neuroendocrine carcinoma 1 Vulvar Squamous Cell Carcinoma 2 Uterine Serous Carcinoma 1 Vulvar carcinoma 1 Prostate Adenocarcinoma 2 Colorectal adenocarcinoma 1 Meningioma 1 Ovarian Surface Epithelial Carcinoma non-clear cell 1 Ovarian Surface Epithelial Carcinomas 1 Uterine Serous Carcinoma 1 Ovarian Surface Epithelial Carcinomas 1 Bladder carcinoma - non-urothelial 1 Bladder carcinoma - urothelial 1 Sarcoma 1 Melanoma 1 Ependymoma 1 Astroblastoma 1 Germ cell tumor 1 Hepatocellular Carcinoma 1 Glioblastoma 1 Lung Non-small cell lung cancer (NSCLC) 1 Astrocytoma 1 Glioblastoma 1 Low Grade Glioma 1 Malignant Solitary Fibrous Tumor of the Pleura (MSFT) 1 Grand Total 704 Table 13. Fusions as Orthogonal Evidence in Cases with ‘Critical Value Discrepancies’ Fusion Supports Submitted Supports GPSai Diagnosis PredictionCD74:ROS1 0 1 ESR1:TNRC6B 0 1 FGFR2:WEE1 0 1 KIF5B:RET 0 1 PGR:NR4A3 0 1 PTPRK:RSPO3 0 2 TMPRSS2:ERG 2 3 EML4:ALK 1 0 EWSR1:FLI1 1 0 Total 4 10 d) Quantification of Level 1 targeted therapy recommendation changes prompted by GPSai

[0313] Cases that had diagnosis changes were also analyzed for Level 1 drug-association changes, which were defined as associations directed by the prescribing information on the drug’s label. National Comprehensive Cancer Network (NCCN) guideline-directed therapies not represented on the drug’s label were not considered. A case was categorized as having a Level 1 drug-association change if the case became eligible for, or ineligible for, a targeted therapy due to the diagnosis change. Over eight months of testing, GPSai prompted such drug-association changes in 83.9% of CUP (n=339) and 89.0% (n=267) of non-CUP cases. Although all of these changes were rooted in tumor type changes, 34.5% (n=209) were also driven by the presence of a companion diagnostic biomarker. See FIG.19A. A list of therapeutic eligibility and ineligibility associated with these lineage changes is shown in FIG. 20. Changes to eligibility for immunotherapies such as pembrolizumab were the most common, due to the large percentage of cases with diagnosis changes from or to NSCLC. See Table 12. Of the cases with Level 1, biomarker-driven changes, we determined that biomarkers were identified by WES in 21% of cases, by WTS in 3% of cases, and by IHC in 11% of cases. See FIG.19B. Together, these results indicate the utility of combining GPSai results with comprehensive molecular profiling to provide an accurate diagnosis and identify lineage-matched, treatment-associated biomarkers. e) Physician Survey Results

[0314] Physicians who received a ‘Critical Value Discrepancy’ on the clinical report for their patient (n=957) were given the opportunity to provide feedback on how GPSai results impacted clinical care (Table 14). Of 97 physicians who responded, 89.7% (n=87 / 97)accepted the GPSai results, and 88.7% (n=86 / 97) felt GPSai assisted in making a diagnosis. 53.6% (n=52 / 97) stated that the result changed the treatment plan for their patient, and another 9.3% (n=9 / 97) were either unsure or indicated that the patient did not return to the clinic. Of the 61 respondents that indicated a change or possible change in treatment plan, 68.9% (n=42 / 61) felt that there was a reasonable expectation of clinical benefit due to the change in treatment. Moreover, 8.8% (n=8 / 90) of patients became eligible for a clinical trial based on the change in diagnosis. Table 14. Physician Survey Results Yes No Other Unsure / Pt. Did Not o3. Discussion

[0315] Although the past two decades have witnessed a surge of diagnostic AI approaches, GPSai is one of few commercially available tests for tissue of origin identification (12, 15, 19, 30, 44). This tool provided using the systems and methods provided herein demonstrates both high accuracy (95.0%) and ability to classify CUP cases (84.0% call rate) in real-world validation cohorts, and performs well in metastatic tissue. The large, curated training set used provides excellent coverage of the 90 Oncotree-derived cancer categories predicted by themodel, demonstrating superior granularity through prediction of histologic diagnoses including hematological malignancies, multiple sarcoma malignancies, neuroendocrine tumors, and soft tissue / bone tumors. Furthermore, inclusion of GPSai into the routine, comprehensive molecular testing workflow reduces overall tissue requirements—which can be a roadblock to testing, especially for CUP (33, 34)—and allows the model to flag potential misdiagnoses for further diagnostic workup.

[0316] As an AI tool, GPSai is relevant to the physician-in-the-loop paradigm that integrates data-driven medical science with physician expertise (37), as illustrated in the schematic in FIG.21. Integration of GPSai into a comprehensive diagnostic workflow can both aid interpretation of available information and guide further testing to achieve an accurate diagnosis. In cases where there remains some ambiguity as to the tissue of origin, pertinent information is provided on the clinical report for the physician to assess in order to reach a most probable diagnosis and define the best treatment plan for their patient.

[0317] Simultaneous biomarker profiling provides additional information to support appropriate treatment planning for patients with CUP or other diagnostically ambiguous tumors. The results presented in this Example show that GPSai may impact patient care due to changes in eligibility for FDA-approved therapies or clinical trials. When quantifying only Level 1 targeted-therapy drug-associations, we found that a majority of both CUP and non- CUP cases underwent a therapy recommendation change due to the GPSai test result, with a significant subset of these being biomarker-driven. See, e.g., FIGs.19-20 and related discussion. In some cases, this change included the removal of a particular therapy recommendation due to low expectations of success. For instance, we identified nine cases of BRAF / MEK inhibitor (dabrafenib + trametinib) ineligibility due to a diagnosis change to colorectal adenocarcinoma. See FIG.20; (45). Physician feedback also indicated a high rate of treatment plan change (53.6%) due to GPSai and revealed that eight patients became eligible for a clinical trial, further expanding the potential for GPSai to impact patient treatment options. See Table 14 and related discussion.

[0318] The findings in this Example support the clinical utility of GPSai by demonstrating changes in diagnosis and subsequent treatment recommendation changes. Physician feedback on implementation of GPSai results further supports its relevance as part of a physician-in- the-loop approach. For therapy selection, site-specific treatment of CUP has led to improved outcomes, particularly among patients with high-accuracy predictions and those with moreresponsive tumor types (24, 32, 46-48). Consideration of GPSai results in tandem with other diagnostic tools and biomarker profiling will better aid physicians in providing the best care to patients presenting with CUP. This study further identified an important role in aiding challenging diagnoses outside of CUP, such as the identification of misdiagnoses that could significantly alter care for a meaningful subset of patients. In addition, AI tools such as GPSai have the potential to incur significant cost savings through increasing accuracy of diagnosis and treatment and reducing inefficiencies (49). 4. References (numbered according to the numbering in this Example)

[0319] 1. Massard C, et al. Carcinomas of an unknown primary origin--diagnosis and treatment. Nat Rev Clin Oncol.2011;8(12):701-10.

[0320] 2. Rassy E, Pavlidis N. The currently declining incidence of cancer of unknown primary. Cancer Epidemiol.2019;61:139-41.

[0321] 3. Mnatsakanyan E, et al. Cancer of unknown primary: time trends in incidence, United States. Cancer Causes Control.2014;25(6):747-57.

[0322] 4. American Cancer Society. Cancer Facts and Figures 2024 Atlanta: American Cancer Society; 2024.

[0323] 5. Huebner G, et al. Paclitaxel and carboplatin vs gemcitabine and vinorelbine in patients with adeno- or undifferentiated carcinoma of unknown primary: a randomised prospective phase II trial. Br J Cancer.2009;100(1):44-9.

[0324] 6. Hess KR, et al. Classification and regression tree analysis of 1000 consecutive patients with unknown primary carcinoma. Clin Cancer Res.1999;5(11):3403-10.

[0325] 7. Kaaks R, et al. Risk factors for cancers of unknown primary site: Results from the prospective EPIC cohort. Int J Cancer.2014;135(10):2475-81.

[0326] 8. Kang S, et al. Real-world data analysis of patients with cancer of unknown primary. Sci Rep.2021;11(1):23074.

[0327] 9. Peck M, et al. Review of diagnostic error in anatomical pathology and the role and value of second opinions in error prevention. J Clin Pathol.2018;71(11):995-1000.

[0328] 10. Abt AB, et al. The effect of interinstitution anatomic pathology consultation on patient care. Arch Pathol Lab Med.1995;119(6):514-7.

[0329] 11. Tsung JS. Institutional pathology consultation. Am J Surg Pathol. 2004;28(3):399-402.

[0330] 12. Erlander MG, et al. Performance and clinical evaluation of the 92-gene real- time PCR assay for tumor classification. J Mol Diagn.2011;13(5):493-503.

[0331] 13. Tothill RW, et al. Development and validation of a gene expression tumour classifier for cancer of unknown primary. Pathology.2015;47(1):7-12.

[0332] 14. Ma W, et al. New techniques to identify the tissue of origin for cancer of unknown primary in the era of precision medicine: progress and challenges. Brief Bioinform. 2024;25(2).

[0333] 15. Michuda J, et al. Validation of a Transcriptome-Based Assay for Classifying Cancers of Unknown Primary Origin. Mol Diagn Ther.2023;27(4):499-511.

[0334] 16. Moiso E, et al. Developmental Deconvolution for Classification of Cancer Origin. Cancer Discovery.2022;12(11):2566-85.

[0335] 17. Vibert J, et al. Identification of Tissue of Origin and Guided Therapeutic Applications in Cancers of Unknown Primary Using Deep Learning and RNA Sequencing (TransCUPtomics). J Mol Diagn.2021;23(10):1380-92.

[0336] 18. Zhao Y, et al. CUP-AI-Dx: A tool for inferring cancer tissue of origin and molecular subtype using RNA gene-expression data and artificial intelligence. EBioMedicine.2020;61:103030.

[0337] 19. Pillai R, et al. Validation and reproducibility of a microarray-based gene expression test for tumor identification in formalin-fixed, paraffin-embedded specimens. J Mol Diagn.2011;13(1):48-56.

[0338] 20. Ross JS, et al. Comprehensive Genomic Profiling of Carcinoma of Unknown Primary Site: New Routes to Targeted Therapies. JAMA Oncol.2015;1(1):40-9.

[0339] 21. Penson A, et al. Development of Genome-Derived Tumor Type Prediction to Inform Clinical Cancer Care. JAMA Oncol.2020;6(1):84-91.

[0340] 22. Marquard AM, et al. TumorTracer: a method to identify the tissue of origin from the somatic mutations of a tumor specimen. BMC Med Genomics.2015;8:58.

[0341] 23. Jiao W, et al. A deep learning system accurately classifies primary and metastatic cancers using passenger mutation patterns. Nat Commun.2020;11(1):728.

[0342] 24. Moon I, et al. Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary. Nat Med.2023;29(8):2057-67.

[0343] 25. He B, et al. A machine learning framework to trace tumor tissue-of-origin of 13 types of cancer based on DNA somatic mutation. Biochim Biophys Acta Mol Basis Dis. 2020;1866(11):165916.

[0344] 26. Liu X, et al. Predicting Cancer Tissue-of-Origin by a Machine Learning Method Using DNA Somatic Mutation Data. Front Genet.2020;11:674.

[0345] 27. Nguyen L, et al. Machine learning-based tissue of origin classification for cancer of unknown primary diagnostics using genome-wide mutation features. Nat Commun. 2022;13(1):4013.

[0346] 28. Liang Y, et al. A Deep Learning Framework to Predict Tumor Tissue-of- Origin Based on Copy Number Alteration. Front Bioeng Biotechnol.2020;8:701.

[0347] 29. Zhang Y, et al. A Novel XGBoost Method to Identify Cancer Tissue-of-Origin Based on Copy Number Variations. Front Genet.2020;11:585029.

[0348] 30. Abraham J, et al. Machine learning analysis using 77,044 genomic and transcriptomic profiles to accurately predict tumor type. Transl Oncol.2021;14(3):101016.

[0349] 31. Chakravarty D, et al. Somatic Genomic Testing in Patients with Metastatic or Advanced Cancer: ASCO Provisional Clinical Opinion. Journal of Clinical Oncology. 2022;40(11):1231-58.

[0350] 32. Laprovitera N, et al. Cancer of Unknown Primary: Challenges and Progress in Clinical Management. Cancers (Basel).2021;13(3).

[0351] 33. Zaun G, et al. Comprehensive biomarker diagnostics of unfavorable cancer of unknown primary to identify patients eligible for precision medical therapies. Eur J Cancer. 2024;200:113540.

[0352] 34. Huey RW, et al. Feasibility and value of genomic profiling in cancer of unknown primary: real-world evidence from prospective profiling study. J Natl Cancer Inst. 2023;115(8):994-7.

[0353] 35. Bi WL, et al. Artificial intelligence in cancer imaging: Clinical challenges and applications. CA: A Cancer Journal for Clinicians.2019;69(2):127-57.

[0354] 36. Shao J, et al. Novel tools for early diagnosis and precision treatment based on artificial intelligence. Chin Med J Pulm Crit Care Med.2023;1(3):148-60.

[0355] 37. Stenzinger A, et al. Artificial intelligence and pathology: From principles to practice and future applications in histomorphology and molecular profiling. Seminars in Cancer Biology.2022;84:129-43.

[0356] 38. Bigorra L, et al. A Physician-in-the-Loop Approach by Means of Machine Learning for the Diagnosis of Lymphocytosis in the Clinical Laboratory. Arch Pathol Lab Med.2022;146(8):1024-31.

[0357] 39. Kieseberg P, et al. A tamper-proof audit and control system for the doctor in the loop. Brain Inform.2016;3(4):269-79.

[0358] 40. Kundra R, Zhang H, Sheridan R, Sirintrapun SJ, Wang A, Ochoa A, et al. OncoTree: A Cancer Classification System for Precision Oncology. JCO Clinical Cancer Informatics.2021(5):221-30.

[0359] 41. Rezende PM, et al. Evaluating hierarchical machine learning approaches to classify biological databases. Brief Bioinform.2022;23(4).

[0360] 42. Sanchez-Vega F, et al. Oncogenic Signaling Pathways in The Cancer Genome Atlas. Cell.2018;173(2):321-37.e10.

[0361] 43. Alexandrov LB, et al. The repertoire of mutational signatures in human cancer. Nature.2020;578(7793):94-101.

[0362] 44. Meiri E, et al. A second-generation microRNA-based assay for diagnosing tumor tissue origin. Oncologist.2012;17(6):801-12.

[0363] 45. Corcoran RB, et al. Combined BRAF and MEK Inhibition with Dabrafenib and Trametinib in BRAF V600-Mutant Colorectal Cancer. J Clin Oncol.2015;33(34):4023- 31.

[0364] 46. Hainsworth JD, et al. Molecular gene expression profiling to predict the tissue of origin and direct site-specific therapy in patients with carcinoma of unknown primary site: a prospective trial of the Sarah Cannon research institute. J Clin Oncol.2013;31(2):217-23.

[0365] 47. Ding Y, et al. Site-specific therapy in cancers of unknown primary site: a systematic review and meta-analysis. ESMO Open.2022;7(2):100407.

[0366] 48. Moran S, et al. Epigenetic profiling to classify cancer of unknown primary: a multicentre, retrospective analysis. Lancet Oncol.2016;17(10):1386-95.

[0367] 49. Khanna NN, et al. Economics of Artificial Intelligence in Healthcare: Diagnosis vs. Treatment. Healthcare (Basel).2022;10(12).

Claims

WHAT IS CLAIMED IS:

1. A method of determining a tumor type using a hierarchal tumor type tree, the method comprising: generating a sample vector using biological data measured from a biological sample of a subject having a cancer, wherein the biological data includes a set of biomarkers; loading, into memory of a computer system, a top-level machine learning model trained using top-level training samples, each top-level training sample including a top-level reference vector measured from a top-level reference biological sample labeled with a particular top-level tumor type of N top-level tumor types, N being an integer greater than 3; processing, using the top-level machine learning model, the sample vector to identify the biological sample has a first top-level tumor type of the N top-level tumor types, the first top-level tumor type having a first likelihood that is highest out of the N top-level tumor types; determining that the first top-level tumor type has M second-level tumor types in the hierarchal tumor type tree, M being an integer greater than one; identifying a second-level machine learning model corresponding to the first top- level tumor type; loading, in the memory, the second-level machine learning model trained using a first set of second-level training samples, each including a second-level reference vector measured from a second-level reference biological sample labeled with a second-level tumor type of M second-level tumor types corresponding to the first top-level tumor type; and processing, using the second-level machine learning model, the sample vector to identify the biological sample has a first second-level tumor type of the M second-level tumor types, the first second-level tumor type having a second likelihood that is highest out of the M second-level tumor types.

2. The method of claim 1, wherein the set of biomarkers includes DNA mutations.

3. The method of claim 2, wherein the DNA mutations comprise one or more point mutation, sequence variant, polymorphism, deletion, insertion, indels (i.e., insertions ordeletions), substitution, translocation, fusion, break, duplication, amplification, repeat, or copy number variation.

4. The method of claim 2 or claim 3, wherein the DNA mutations comprise sequence variants and / or copy number variations.

5. The method of any one of claims 2-4, where the DNA mutations are pathogenic (P) or likely pathogenic (LP) mutations.

6. The method of any one of claims 2-5, wherein values for the set of biomarkers are determined using whole genome sequencing.

7. The method of any one of claims 2-5, wherein values for the set of biomarkers are determined using whole exome sequencing.

8. The method of any one of claims 2-7, wherein the set of one or more biomarkers comprises one or more biomarkers listed in any one of Tables 1-7.

9. The method of any one of claims 2-7, wherein the set of one or more biomarkers comprises one or more of the following: ABL1, AKT1, AKT2, AKT3, ALK, AMER1, APC, AR, ARAF, ARID1A, ARID2, ASXL1, ATM, ATR, ATRX, AXIN1, BAP1, BARD1, BCL2, BCL9, BCOR, BLM, BMPR1A, BRAF, BRCA1, BRCA2, BRIP1, BTG1, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCDC6, CCND1, CCND2, CCND3, CD79B, CDC73, CDH1, CDK12, CDK4, CDKN1B, CDKN2A, CDKN2B, CEBPA, CHEK1, CHEK2, CIC, CNOT3, CREBBP, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CYLD, DDR2, DICER1, DNMT3A, EGFR, EP300, ERBB2, ERBB3, ERBB4, ERCC2, ESR1, EXT1, EZH2, FANCA, FANCC, FANCD2, FANCE, FANCF, FANCG, FANCL, FAS, FBXW7, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FLT4, FOXA1, FOXL2, FOXO3, FUBP1, GATA3, GNA11, GNA13, GNAQ, GNAS, GRIN2A, H3F3A, H3F3B, HIST1H3B, HNF1A, HRAS, IDH1, IDH2, IL7R, IRF4, JAK1, JAK2, JAK3, KDM5C, KDM6A, KDR, KEAP1, KIT, KLF4, KMT2A, KMT2C, KMT2D, KRAS, LRP1B, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAX, MED12, MEF2B, MEN1, MET, MITF, MLH1, MPL, MRE11, MSH2, MSH6, MSI, MTOR, MUTYH, MYC, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NPM1, NRAS, NSD1, NSD2, NT5C2, NTRK1, NTRK2, NTRK3,PALB2, PBRM1, PDE4DIP, PDGFRA, PDGFRB, PHOX2B, PIK3CA, PIK3R1, PIK3R2, PIM1, PMS1, PMS2, POLE, POT1, PPARG, PPP2R1A, PRDM1, PRKAR1A, PRKDC, PTCH1, PTEN, PTPN11, RAC1, RAD50, RAD51B, RAF1, RB1, RET, RNF43, ROS1, RUNX1, SDHAF2, SDHB, SDHC, SDHD, SETD2, SF3B1, SMAD2, SMAD4, SMARCA4, SMARCB1, SMARCE1, SMO, SOCS1, SPEN, SPOP, SRC, STAG2, STAT3, STAT5B, STK11, SUFU, SUZ12, TCF7L2, TERT, TET2, TGFBR2, TNFAIP3, TNFRSF14, TP53, TRAF7, TRRAP, TSC1, TSC2, U2AF1, VHL, WRN, WT1, and XPO1.

10. The method of any preceding claim, wherein the set of biomarkers includes a set of expression levels.

11. The method of claim 10, wherein the set of expression levels are measured using RNA.

12. The method of claim 11, wherein the set of expression levels for the RNA are determined using whole transcriptome sequencing.

13. The method of any preceding claim, wherein the set of biomarkers is obtained by performing at least one assay, wherein the at least one assay comprises determining a presence, level, or state of a protein or nucleic acid for each of the one or more biomarkers.

14. The method of claim 13, wherein: a presence, level or state of at least one protein is determined using a technique selected from immunohistochemistry (IHC), flow cytometry, an immunoassay, an antibody or functional fragment thereof, an aptamer, mass spectrometry, or any combination thereof, wherein optionally the presence, level or state of all of the proteins is determined using the technique; and / or the presence, level or state of at least one nucleic acid is determined using a technique selected from polymerase chain reaction (PCR), in situ hybridization, amplification, hybridization, microarray, nucleic acid sequencing, dye termination sequencing, pyrosequencing, next generation sequencing, whole exome sequencing, whole genome sequencing, whole transcriptome sequencing, or any combination thereof, wherein optionally the presence, level or state of all of the nucleic acids is determined using the technique.

15. The method of claim 14, wherein the state of the at least one nucleic acid comprises a sequence, mutation, polymorphism, deletion, insertion, substitution, translocation, fusion, break, duplication, amplification, repeat, copy number, or any combination thereof.

16. The method of any preceding claim, wherein at least one of the top-level machine learning model and the second-level machine learning model is further trained using one or more characteristics of the cancer.

17. The method of claim 16, wherein the one or more characteristics of the cancer comprise one or more selected from a group consisting of: a collection site of the biological sample, a sample format of the biological sample, and a metastatic status.

18. The method of any preceding claim, wherein at least one of the top-level machine learning model and the second-level machine learning model is further trained using one or more characteristic of the subject.

19. The method of claim 18, wherein the one or more characteristic of the subject comprises sex and / or age.

20. The method of any preceding claim, wherein: the set of biomarkers include DNA sequence variants and RNA expression; and the top-level machine learning model and the second-level machine learning model are further trained using one or more characteristics of the subject.

21. The method of claim 20, wherein the DNA sequence variants are determined using whole exome sequencing, the RNA expression is determined using whole transcriptome sequencing, and the one or more characteristic of the subject comprises binary sex.

22. The method of claim 21, wherein the DNA sequence variants are pathogenic or likely pathogenic variants.

23. The method of any preceding claim, wherein a portion of the N top-level tumor types do not have a child node in the hierarchal tumor type tree.

24. The method of any preceding claim, wherein the second-level reference vector of each of the first set of second-level training samples is a same vector used to train the top-level machine learning model, and wherein all values of the sample vector are processed by both the top-level machine learning model and the second-level machine learning model.

25. The method of any preceding claim, wherein the first set of second-level training samples includes a subset of the top-level training samples labeled with the first top- level tumor type, and wherein each of the first set of second-level training samples has a second label being one on the M second-level tumor types.

26. The method of any preceding claim, further comprising: scaling the second likelihood using the first likelihood to obtain a modified second likelihood; and reporting the modified second likelihood.

27. The method of any preceding claim, wherein the first likelihood is greater than a first threshold, the method further comprising: determining the second likelihood is less than a second threshold; and reporting only the first top-level tumor type and not the first second-level tumor type.

28. The method of any preceding claim, further comprising: determining that a second top-level tumor type of the N top-level tumor types has K second-level tumor types in the hierarchal tumor type tree, K being an integer greater than one; identifying an additional second-level machine learning model corresponding to the second top-level tumor type; loading, in the memory, the additional second-level machine learning model trained using a second set of second-level training samples, each including an additional second- level reference vector measured from an additional second-level reference biological sample labeled with a second second-level tumor type of the K second-level tumor types corresponding to the second top-level tumor type; andprocessing, using the additional second-level machine learning model, the sample vector to determine probabilities for the K second-level tumor types.

29. The method of any preceding claim, further comprising: determining that the first second-level tumor type has J third-level tumor types in the hierarchal tumor type tree, J being an integer greater than one; identifying a third-level machine learning model corresponding to the first second-level tumor type; loading, in the memory, the third-level machine learning model trained using a first set of third-level training samples, each including an a third-level reference vector measured from a third-level reference biological sample labeled with a third-level tumor type of the J third- level tumor types corresponding to the first second-level tumor type ; and processing, using the third-level machine learning model, the sample vector to determine probabilities for the J third-level tumor types.

30. The method of any preceding claim, wherein the top-level machine learning model comprises a plurality of submodels, and wherein processing the sample vector to identify the biological sample has the first top-level tumor type comprises: obtaining output data from each of the plurality of submodels; and using the output data from each of the plurality of submodels to identify the biological sample has the first top-level tumor type.

31. The method of claim 30, wherein the first top-level tumor type is determined by applying a majority rule to the output data, by using the output data as input into a dynamic voting model, or a combination thereof.

32. The method of claim 31, wherein determining, by the majority rule and based on the output data, the first top-level tumor type comprises: determining a number of occurrences of each top-level tumor type of the N top- level tumor types; and selecting the first top-level tumor type as having a highest number of occurrences.

33. The method of any preceding claim, wherein each of the top-level machine learning model and the second-level machine learning model comprises a random forest classification algorithm, support vector machine, logistic regression, k-nearest neighbor model, neural network, naïve Bayes model, quadratic discriminant analysis, Gaussian processes model, clustering, or any combination thereof.

34. The method of claim 33, wherein the top-level machine learning model and the second-level machine learning model are neural networks.

35. The method of any preceding claim, wherein the biological sample comprises cells from a solid tumor, a bodily fluid, or a combination thereof.

36. The method of claim 35, wherein the biological sample comprises an FFPE tissue sample, optionally wherein the biological sample is microdissected.

37. The method of any preceding claim, further comprising performing one or more orthogonal methods to confirm the first top-level tumor type and / or the first second-level tumor type.

38. The method of claim 37, wherein the one or more orthogonal methods comprise one or more of imaging, immunohistochemistry, mutational signatures, hallmark fusions, or viral reads.

39. The method of any preceding claim, further comprising: determining, using the biological data and optionally one or more machine learning models, a diagnosis, a prognosis and / or theranosis for a cancer based at least in part on the first second-level tumor type.

40. The method of any preceding claim, further comprising: determining whether a treatment for the cancer is of likely benefit, lack of benefit, or indeterminate benefit in treating the subject, based at least in part on the first second-level tumor type.

41. The method of claim 40, further comprising administering the treatment to the subject based on the determining.

42. The method of any preceding claim, wherein the second-level machine learning model is trained by: obtaining predicted tumor types generated by the second-level machine learning model from the second-level reference vectors; determining differences between the predicted tumor types generated by the second-level machine learning model and labels of the first set of second-level training samples; and adjusting one or more parameters of the second-level machine learning model based on the differences.

43. The method of any preceding claim, wherein the N top-level tumor types and the M second-level tumor types comprise at least one top-level type and at least one second- level type according to FIGS.9A, 9B and 9C.

44. The method of any preceding claim, wherein the N top-level tumor types and all the second-level tumor types of the hierarchal tumor type tree are comprised in FIGS.9A, 9B and 9C.

45. The method of any preceding claim, wherein the hierarchal tumor type tree has third-level tumor types.

46. The method of any preceding claim, further comprising providing a report comprising the first top-level tumor type and / or the first second-level tumor type.

47. The method of claim 46, wherein the report further comprises a diagnosis, prognosis, or theranosis for the cancer in the subject, based at least in part on the biological data measured from the biological sample from the subject.

48. The method of claim 47 wherein the theranosis comprises one or more treatment of likely benefit, lack of benefit, or indeterminate benefit in treating the subject, optionally based at least in part on the first second-level tumor type.

49. The method of any preceding claims, wherein N is at least 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, and 25.

50. A method of determining a sample type using a hierarchal sample type tree, the method comprising: obtaining biological data measured from a biological sample of a subject, wherein the biological data includes a set of biomarkers; generating a sample vector using the biological data; loading, into memory of a computer system, a top-level machine learning model trained using top-level training samples, each top-level training sample including a top-level reference vector measured from a top-level reference biological sample labeled with a particular top-level sample type of N top-level sample types, N being an integer greater than 10; processing, using the top-level machine learning model, the sample vector to identify the biological sample has a first top-level sample type of the N top-level sample types, the first top-level sample type having a first likelihood that is highest out of the N top-level sample types; determining that the first top-level sample type has M second-level sample types in the hierarchal sample type tree, M being an integer greater than one; identifying a second-level machine learning model corresponding to the first top- level sample type; loading, in the memory, the second-level machine learning model trained using a first set of second-level training samples, each including a second-level reference vector measured from a second-level reference biological sample labeled with a first second-level sample type of M second-level sample types corresponding to the first top-level sample type; and processing, using the second-level machine learning model, the sample vector to identify the biological sample has a first second-level sample type of the M second-level sample types, the first second-level sample type having a second likelihood that is highest out of the M second-level sample types.

51. A method of training machine learning models forming a hierarchal sample type tree, the method comprising:training, by a computer system, a top-level machine learning model using top- level training samples that each include biological data for a set of biomarkers, each top-level training sample including a top-level reference vector measured from a top-level reference biological sample labeled with a particular top-level sample type of N top-level sample types, N being an integer greater than 3, wherein a first top-level sample type has M second-level sample types in the hierarchal sample type tree, M being an integer greater than one; and training, by the computer system, a second-level machine learning model corresponding to the first top-level sample type using a first set of second-level training samples, each including a second-level reference vector measured from a second-level reference biological sample labeled with a second-level sample type of M second-level sample types corresponding to the first top-level sample type.

52. The method of claim 51, wherein one or more other top-level sample types have multiple second-level sample types, the method further comprising: for each top-level sample type of the one or more other top-level sample types: training, by the computer system, a respective second-level machine learning model corresponding to the top-level sample type using another set of second-level training samples.

53. A method of determining a sample type using a hierarchal sample type tree, the method comprising: loading, into memory of a computer system, a top-level machine learning model trained using top-level training samples, each top-level training sample including a top-level reference vector measured from a top-level reference biological sample labeled with a particular top-level sample type of N top-level sample types, N being an integer greater than 3, wherein a first top-level sample type has M second-level sample types in the hierarchal sample type tree, M being an integer greater than one; loading, by the computer system, a second-level machine learning model corresponding to the first top-level sample type trained using a first set of second-level training samples, each including a second-level reference vector measured from a second-level reference biological sample labeled with a second-level sample type of M second-level sample types corresponding to the first top-level sample type;generating a sample vector using biological data measured from a biological sample of a subject, wherein the biological data includes a set of biomarkers; and processing, using the top-level machine learning model, the sample vector to identify the biological sample has a second top-level sample type of the N top-level sample types, the second top-level sample type having a likelihood that is highest out of the N top-level sample types.

54. The method of any one of claims 50-53, wherein at least one of the top- level machine learning model and the second-level machine learning model is further trained using one or more characteristics of a medical condition.

55. The method of claim 54, wherein the one or more characteristics of the medical condition comprise one or more selected from a group consisting of: a collection site of the biological sample, a sample format of the biological sample, and a metastatic status.

56. The method of any one of claims 50-55, wherein at least one of the top- level machine learning model and the second-level machine learning model is further trained using one or more characteristic of a subject.

57. The method of claim 56, wherein the one or more characteristic of the subject comprises sex and / or age.

58. The method of any one of claims 50-57, wherein the hierarchal sample type tree is a hierarchal tumor type tree.

59. A computer product comprising a non-transitory computer readable medium storing a plurality of instructions that, when executed, cause a computer system to perform the method of any one of the preceding claims.

60. A system comprising: the computer product of claim 59; and one or more processors configured to execute instructions stored on the computer readable medium.

61. A system comprising means for performing any of the above methods.PATENT Attorney Docket No.: 110588-0886WO1-1478774 Client Reference No.: CMI 886.601 62. A system comprising one or more processors configured to perform any of the above methods.

63. A system comprising modules that respectively perform the steps of any of the above methods.

Citation Information

Patent Citations

  • Methods and systems for utilizing quantitative imaging

    US20190180153A1

  • Convolutional neural network systems and methods for data classification

    US20200005899A1

  • Clinical concept identification, extraction, and prediction system and related methods

    US20210210184A1

  • Multi-omic search engine for integrative analysis of cancer genomic and clinical data

    US20210319907A1

  • Lung cancer biomarkers and uses thereof

    WO2010030697A1

Cited By

  • Predicting methylation status

    WO2026174323A1