Multi-region sequencing
Multi-region sequencing and machine learning models improve cancer diagnosis by accurately predicting tumor heterogeneity, ensuring effective treatments and accurate prognoses by analyzing distinct cell populations within a tumor.
Patent Information
- Application Number
- PCT/US2025/043555
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-30
- Filing Date
- 2025-08-26
- Publication Date
- 2026-03-05
AI Technical Summary
Current cancer diagnosis and treatment methods often rely on a single tissue biopsy, which may not accurately represent the heterogeneity of a tumor, leading to ineffective treatments and inaccurate prognoses due to the inability to analyze distinct cell populations within the tumor.
A method involving multi-region sequencing of nucleic acid molecules from spatially distinct regions of a tumor, using machine learning models to analyze sequence read data and predict tumor heterogeneity, enabling more accurate prediction of treatment effectiveness and prognosis.
Enhances the accuracy of predicting therapy effectiveness and prognosis by identifying distinct cell populations within a tumor, reducing the likelihood of ineffective treatments and inaccurate prognoses.
Smart Images

Figure US2025043555_05032026_PF_FP_ABST
Abstract
Description
FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT MULTI-REGION SEQUENCING CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Application No. 63 / 689,456, which was filed on August 30, 2024, and is incorporated by reference herein in its entirety. BACKGROUND
[0002] Many cancers are associated with the accumulation of genetic mutations due to the continual unregulated proliferation of cancer cells. In some cases, a single tumor can include multiple cell populations, each having distinct genetic profiles. Tumor heterogeneity is linked to different prognoses, metastasis profiles, treatments, and other clinically relevant factors. In a particular example, a tumor may include a first cell population that is susceptible to a particular immunotherapy, as well as a second cell population that is resistant to the immunotherapy. Moreover, a tumor that exclusively includes a single cell population may be associated with a different prognosis than a tumor that includes multiple cell populations. Therefore, it is desirable to identify the heterogeneity of a tumor within a patient. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, of which:
[0004] FIG.1 illustrates an example environment for predicting a heterogeneity condition of a tumor of a subject using genomic information of the subject.
[0005] FIG. 2 illustrates an example environment for identifying target regions from one or more tissue samples.
[0006] FIG. 3 illustrates an example environment for training and utilizing a predictive model to identify a heterogeneity condition of a tumor of a subject.
[0007] FIG.4 illustrates an example of training data utilized to train one or more ML models.
[0008] FIG.5 illustrates an example report summarizing predicted categories of a tumor of a subject.
[0009] FIG.6 illustrates an example environment for sequencing various nucleic acid molecules.
[0010] FIG. 7 illustrates an example process for identifying a heterogeneity condition of a tumor of a subject.
[0011] FIG.8 illustrates one or more devices configured to perform various operations described herein.
[0012] FIG.9 illustrates an example post multi-region tissue extracted H&E slide.
[0013] FIGs.10A-10D illustrate an example overview of functional genomic alterations (known / likely) as observed in the multi-region sequencing pilot data.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT
[0014] FIG.11 illustrates an example distribution of cancer cell fraction (CCF) of short variants.
[0015] FIG.12 illustrates an example clonality of short variants.
[0016] FIG.13 illustrates an example wet-lab merge of samples.
[0017] FIG.14 illustrates an example dry-lab merge of samples.
[0018] Some of the drawings submitted herewith may be better understood in color. Applicant considers the color versions of the drawings as part of the original submission and reserves the right to present color images of the drawings in later proceedings. DETAILED DESCRIPTION
[0019] Various implementations of the present disclosure relate to techniques for identifying heterogeneity of a tumor of a subject, as well as conditions associated with the heterogeneity of the tumor. In particular cases, a condition of the subject can be predicted based on the heterogeneity. In various cases, nucleic acid molecules are obtained from a first sample and a second sample collected from the tumor of the subject. In some cases, the first sample and the second sample are collected from regions of the tumor that are spatially distinct. In various cases, the first sample and the second sample are collected from regions associated with distinct visual characteristics, such as distinct histological characteristics. In various cases, the nucleic acid molecules include genomic DNA obtained from tissue biopsy samples. Sequence read data is generated by sequencing the nucleic acid molecules. The heterogeneity of the tumor, for instance, can be assessed by analyzing the sequence read data.
[0020] Various types of health-related conditions can be predicted using various techniques described herein. In some cases, these techniques are used to predict a heterogeneity condition associated with the tumor of the subject. For instance, these techniques can be used to predict a tumor evolution or a tumor progression. In particular examples, these techniques are used to determine whether the tumor is susceptible to a particular treatment. Conditions related to the cancer of the subject can also be determined, such as a predicted effective therapy to treat the pathogenic condition, a predicted stage of the pathogenic condition, or a predicted grade of the pathogenic condition. Non-pathogenic conditions can also be predicted using implementations of the present disclosure. For instance, the general health of the subject, a risk of developing a disease (e.g., a second primary cancer), a genomic age of the subject, a predicted survivability of the subject, and other conditions, can be predicted based on the heterogeneity.
[0021] Implementations of the present disclosure provide significant improvements to the technical field of cancer diagnosis, management, and treatment. Using current technologies, a patient’s tumor is typically categorized by performing a tissue biopsy on a potential tumor and also performing histological staining and additional analysis on the tissue biopsy sample. In particular cases, a single portion of the tissue biopsy may be used to extract nucleic acids for sequencing. The portion of the tissue biopsy may represent a small fraction of the patient’s tumor. In various cases, metrics associated with multiple portions of the tissueFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT biopsy and / or the tumor can greatly enhance the accuracy of predicting an effectiveness of a therapy or a prognosis of the patient. Accordingly, a clinician relying on various predictions described herein is less likely to administer or prescribe a treatment that will be ineffective. Further, the clinician is less likely to present an inaccurate prognosis to the patient. Inaccurate prognoses and ineffective treatments can cause significant emotional hardship, side effects, and financial burden on the patient, and can be avoided using various techniques described herein.
[0022] Various analyses described herein cannot be performed in the human mind, or by pen and paper. For example, the sequence read data may represent numerous (e.g., thousands) of bases to be analyzed. In various cases, it would be impossible to manually or mentally identify distinct cell populations (e.g., clones) based on the sequence read data. In various cases, it would be impossible to manually or mentally identify relevant genomic features based on the sequence read data. Further, it would be impossible to manually or mentally attribute genomic features that are relevant to the classification of the tumor from which the sequence read data was generated. Particular implementations of the present disclosure are fundamentally tied to computer technology, and do not represent mere automation of processes that are performed manually. Example Definitions
[0023] As used herein, the terms “deoxyribonucleic acid,” “DNA,” “DNA molecule,” and their equivalents, may refer to a polymer of nucleotides (also referred to as “nucleobases”) containing deoxyribose. The nucleotides in DNA include cytosine (C), guanine (G), adenine (A), and thymine (T). Each DNA nucleotide includes a deoxyribose and a phosphate group. An example single-stranded DNA (ssDNA) molecule includes a chain of covalently bonded DNA nucleotides. In the example ssDNA molecule, the phosphate group of the mth nucleotide is covalently bonded to the deoxyribose of the (m-1)th nucleotide, wherein m is a positive integer greater than 2 and less than or equal to the number of DNA nucleotides in the chain. In various examples, DNA is double-stranded and includes two ssDNA molecules that are complementary to one another and coiled around each other in a double helix form. The nucleotides of one ssDNA molecule are hydrogen bonded to the nucleotides of the other ssDNA molecule. In particular, the pyrimidines (A and T) hydrogen bond to each other, and the purines (C and G) hydrogen bond to each other.
[0024] As used herein, the terms “ribonucleic acid,” “RNA,” “RNA molecule,” and their equivalents, may refer to a polymer of nucleotides containing ribose. The nucleotides in RNA include cytosine (C), guanine (G), adenine (A), and uracil (U). Each RNA nucleotide includes a ribose and a phosphate group. In an example RNA molecule, the phosphate group of the nth nucleotide is covalently bonded to the ribose of the (n-1)th nucleotide, wherein n is a positive integer greater than 2 and less than or equal to the number of RNA nucleotides in the chain. Messenger RNA (mRNA) is a type of RNA molecule that is synthesized (orFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT “transcribed”) by RNA polymerase (an enzyme) to be complementary to a gene encoded in a DNA sequence, and is also used by a ribosome to synthesize a polypeptide or protein. An mRNA is therefore an example of a “coding RNA.” In various cases, intron sequences are removed from an mRNA via a process known as “RNA splicing.” MicroRNA (“miRNA”) are single-stranded RNA molecules that perform post- transcriptional gene expression regulation. For instance, a miRNA may bind to a complementary mRNA molecule, thereby cleaving, destabilizing, or otherwise preventing the mRNA molecule from being translated into a polypeptide or protein by a ribosome. In various examples, a miRNA has a length in a range of 21 to 23 RNA nucleotides. As used herein, the terms “non-coding RNA” may refer to a type of RNA that is not translated into a protein. Examples of non-coding RNA include miRNA, transfer RNA (tRNA), and ribosomal RNA (rRNA). The term “functional RNA,” and its equivalents, may refer to any RNA molecule that impacts a biological process. For instance, functional RNA may include mRNA, miRNA, tRNA, rRNA, and the like.
[0025] As used herein, the term “base,” and its equivalents, may refer to a monomer of a polymer. For example, a base of DNA or RNA is a nucleotide.
[0026] As used herein, the term “base pair,” and its equivalents, may refer to a pair of complementary DNA nucleotides, which are hydrogen-bonded to one another in a double-stranded DNA molecule. For example, a base pair includes a first base in a first ssDNA and a second base in a second ssDNA, wherein the first and second bases are complementary and hydrogen-bonded to one another.
[0027] As used herein, the terms “nucleotide,” “nucleobase,” “nucleic acid,” “nucleic acid molecule,” and their equivalents, may refer to an organic molecule that includes a nitrogenous base, a sugar, and a phosphate group. In various cases, a nucleotide is a monomer of DNA or RNA. A nucleotide, for instance, is a chemical structure.
[0028] As used herein, the terms “3’ end,” “3-prime end,” and their equivalents, may refer to a terminus of a single-stranded nucleotide polymer that includes a base whose third carbon in its deoxyribose or ribose is bound to a hydroxyl group while being unbound to another base.
[0029] As used herein, the terms “5’ end,” “5-prime end,” and their equivalents, may refer to a terminus of a single-stranded nucleotide polymer that includes a base whose fifth carbon in its deoxyribose or ribose ring is unbound to another base. In some cases, the fifth carbon is bound to a phosphate group.
[0030] As used herein, the “length” of a polymer refers to a number of covalently bonded monomers that are included in the polymer. For instance, the length of a DNA molecule may be the number of covalently bonded nucleotides in at least one strand of the DNA molecule and / or the number of base pairs in the DNA molecule. In various examples, the length of an RNA molecule may be the number of covalently bonded nucleotides in the RNA molecule.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT
[0031] As used herein, the term “gene,” and its equivalents, refers to a sequence of DNA nucleotides that is transcribed into a functional RNA. The functional RNA, for instance, is RNA that is translated into a polypeptide or protein (e.g., mRNA) or that has some other biological function (e.g., miRNA, tRNA, etc.). A gene is “expressed” when it is used as a template to generate a functional RNA. A subject, for instance, has numerous genes contained in the subject’s genome. A gene may include both introns and exons. As used herein, the term “intron,” and its equivalents, may refer to a subset of DNA nucleotides in a gene that is not used to code for any functional RNA that is expressed by the organism. As used herein, the term “exon,” and its equivalents, may refer to a subset of DNA nucleotides in a gene that is used to code for a functional RNA. For instance, an exon may encode a polypeptide or protein that is expressed by the organism. In various examples, a gene can be represented in data (e.g., as data representative of the sequence of DNA nucleotides in the gene) or as a chemical structure (e.g., as the sequence of DNA nucleotides itself).
[0032] As used herein, the term “genome,” and its equivalents, refers to the aggregate of genes of a subject. In various cases, a genome represents the sequences of several linear DNA molecules that are present in a subject’s chromosomes. A “reference genome” refers to an aggregation of genes of one or more reference subjects. In various cases, a genome is represented in data.
[0033] As used herein, the terms “pangenome,” “pan-genome,” “supragenome,” and their equivalents, refers to an aggregate set of genes from multiple subgroups (e.g., strains) within a population (e.g., a clade) of subjects. A pangenome, for example, indicates genes that are present in all subjects within the population, as well as genes that are present in some of the subjects of the population. A pangenome is represented in data, for instance.
[0034] As used herein, the term “transcriptome,” and its equivalents, refers to the aggregate of RNA sequences of a subject. In some cases, a transcriptome is limited to mRNA sequences. In various examples, a transcriptome is represented in data.
[0035] As used herein, the term “genomic DNA,” “gDNA,” “chromosomal DNA,” and their equivalents, may refer to DNA molecules that are obtained from a chromosome and / or nucleus of a cell.
[0036] As used herein, the terms “DNA fragment,” “fragment,” and their equivalents, may refer to DNA molecules that are excised and / or broken off from a larger DNA molecule.
[0037] As used herein, the term “promoter,” and its equivalents, may refer to a portion of a DNA molecule that binds one or more proteins in order to initiate transcription of a gene. For example, the promotor is located “upstream” of the gene. For example, the promotor is located between the 5’ end of the DNA molecule and the gene. A promotor may include one or more binding sites for RNA polymerase, and / or one or more transcription factor binding sites. In some examples, a promotor includes one or more CpG islands. A promoter, for instance, includes a transcription start site.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT
[0038] As used herein, the terms “CpG island,” “CGI,” “CpG site,” and their equivalents, may refer to a continuous portion of a DNA molecule whose sequence includes greater than a threshold amount (e.g., greater than 50%) of G-C base pairs.
[0039] As used herein, the term “enhancer,” and its equivalents, may refer to a portion of a DNA molecule that binds one or more proteins in order to increase the chance that a gene will be transcribed. For instance, an enhancer includes one or more transcription factor binding sites. In various cases, an enhancer includes one or more CpG islands.
[0040] As used herein, the term “cancer,” and its equivalents, may refer to a condition of a subject in which particular cells (referred to as “cancer cells”) divide uncontrollably in the subject’s body. In some cases, a cancer is characterized by a location or tissue type from which the cancer cells originated. In some examples, a cancer is characterized by a location or tissue type in which the cancer cells are located.
[0041] As used herein, the terms “tumor,” “neoplasm,” and their equivalents, may refer to a mass of tissue including cancer cells.
[0042] As used herein, the terms “tissue of origin,” “tissue origin,” and their equivalents, refers to a differentiated type of tissue from which cancer cells in the body of a subject began dividing uncontrollably in the subject’s body.
[0043] As used herein, the terms “liquid biopsy,” “fluid biopsy,” and their equivalents, may refer to a process of obtaining a fluid sample from a subject’s body. The sample, for instance, can be referred to as a “liquid biopsy sample.” Examples of fluids that are sampled from the body include blood, plasma, cerebrospinal fluid, sputum, stool, urine, lymphatic fluid, and saliva.
[0044] As used herein, the term “tissue biopsy,” and its equivalents, may refer to a process of obtaining a sample of cells from a subject’s body. A tissue biopsy, in various cases, is performed by cutting a mass of cells from the subject’s body. For instance, a tissue biopsy is a procedure performed by a surgeon, interventional radiologist, interventional cardiologist, or other specialized clinician. The term “tissue” or “tissue biopsy sample” can be used to refer to the sample of cells obtained using a tissue biopsy.
[0045] As used herein, the term “subject,” and its equivalents, may refer to a human or non-human animal. A subject that is receiving care from at least one care provider may be referred to as a “patient.”
[0046] As used herein, the terms “machine learning,” “ML,” “computer learning,” “artificial intelligence,” and their equivalents, may refer to the use of a computing devices to learn patterns in training data. The process of learning these patterns may be referred to as “training.” In particular cases, one or more computing devices may perform machine learning by executing a machine learning model. As used herein, the terms “machine learning model,” “ML model,” and their equivalents, may refer to data encoding instructions that, when executed by at least one computing device, causes the at least one computing device to learn patterns in training data by optimizing one or more metrics, values, or other types of parameters.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT After training, an ML model, when executed by at least one computing device, causes the at least one computing device to utilize the optimized parameters in order to perform one or more tasks.
[0047] As used herein, the term “variant,” and its equivalents, may refer to a difference between a subject genetic sequence and a reference sequence. For instance, a variant may correspond to a difference between one or more nucleotides in a genome of a subject and one or more corresponding nucleotides in at least one reference genome or pangenome. A variant may be characterized by its identity (e.g., what nucleotides are different), its position (e.g., where are the nucleotides located in the genome, what chromosome contains the nucleotides, what gene contains the nucleotides, etc.), its length (e.g., how many nucleotides are different from the reference sequence), its type (e.g., substitution, insertion, deletion, copy number alternation, rearrangement of fusion, etc.), and other features that indicates its significance and / or relevance. In some cases, a variant represents any apparent alteration in a sequence that has been read from a nucleic acid molecule with respect to the reference sequence, such as reads cleaved by restriction enzymes (RE). In various examples, a variant can be represented in data (e.g., by data characterizing the variant) or as a chemical structure (e.g., the nucleotides themselves). As used herein, the term “mutation,” and its equivalents, may refer to a change in a gene.
[0048] As used herein, the term “substitution,” and its equivalents, can refer to a nucleotide in a subject sequence that is different than an equivalent nucleotide (e.g., a nucleotide at the same position) in a reference sequence.
[0049] As used herein, the term “insertion,” and its equivalents, can refer to a nucleotide in a subject sequence that is added with respect to a reference sequence.
[0050] As used herein, the term “deletion,” and its equivalents, can refer to the removal of a nucleotide from a nucleotide sequence.
[0051] As used herein, the terms “copy number alternation,” “CNA,” “copy number variation,” “CNV,” and their equivalents, can refer to a portion of a reference sequence that is repeated.
[0052] As used herein, the terms “rearrangement of fusion,” “fusion rearrangement,” “translocation,” and their equivalents, can refer to a change in the relative position of one or more portions of a reference sequence, thereby generating a gene that was not present in the reference sequence.
[0053] As used herein, the term “sequencing,” and its equivalents, may refer to a process of identifying the order and identity of monomers in a polymer chain, such as the order and identity of nucleotides in a DNA or RNA molecule. The terms “whole genome sequencing,” “WGS,” and their equivalents, may refer to the process of sequencing an entire genome of a subject, including the introns and exons of the genes of the subject. The term “whole exome sequencing,” and its equivalents, may refer to the process of sequencing all exomes of a subject. The term “targeted sequencing,” and its equivalents, may refer to the process of sequencing a portion of the genome of a subject, such as sequencing a single gene of the subject.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT Various techniques can be utilized to sequence a DNA or RNA molecule, such as massively parallel sequencing (MPS), nanopore sequencing, direct sequencing, Sanger sequencing, or next-generation sequencing (NGS). In various cases, sequencing is performed on physical molecules (e.g., RNA or DNA) and is used to generate data.
[0054] As used herein, the terms “massive parallel sequencing,” “massively parallel sequencing,” “MPS,” and their equivalents, may refer to a technique for simultaneously performing multiple reactions that can be used to identify the order and identity of monomers in multiple polymer chains. In particular cases, massive parallel sequencing can be performed using sequencing-by-synthesis on clonally amplified DNA molecules that are located in spatially separated regions, which are individually monitored by sensors.
[0055] As used herein, the term “nanopore sequencing,” and its equivalents, may refer to a technique for identifying the order and identity of monomers in a polymer chain by transporting the polymer chain from a first space to a second space, wherein the first space and the second space are separated by a substrate, by directing the polymer chain through a small hole (known as a “nanopore”) embedded in the substrate, and monitoring a relative electrical signal (e.g., a voltage or current) between the first space and the second space.
[0056] As used herein, the term “sensor,” and its equivalents, may refer to a physical device or other apparatus that is configured to detect one or more detection signals.
[0057] As used herein, the term “detection signal,” and its equivalents, may refer to a physical signal that can be identified, characterized, or otherwise perceived by a sensor.
[0058] As used herein, the term “sequence read data,” and its equivalents, may refer to data that is indicative of an order and identity of monomers in a polymer, such as the order and identity of nucleotides in a DNA or RNA sequence. In various implementations, sequence read data is generated via a sequencing operation.
[0059] As used herein, the term “image,” and its equivalents, may refer to 2D or 3D array of data indicative of an array of pixels or voxels.
[0060] As used herein, the term “ligating,” and its equivalents, may refer to a process of joining two molecules together, for example, with a chemical bond.
[0061] As used herein, the term “adapter,” and its equivalents, may refer to an oligonucleotide that can be ligated to a target nucleic acid molecule. In various cases, an adapter prepares the target nucleic acid molecule for sequencing.
[0062] As used herein, the term “bait molecule,” and its equivalents, may refer to a nucleic acid molecule having a region that is complementary to a region of a target molecule (e.g., cfDNA). A bait molecule includes, for instance, a nucleic acid molecule that can hybridize to (i.e., is complementary to) a target molecule can be used to capture the target molecule. In some instances, the bait molecule is a captureFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT oligonucleotide (or capture probe). In some instances, the bait molecule is suitable for solution phase hybridization to the target molecule. In some instances, the bait molecule is suitable for solid phase hybridization to the target molecule. In some instances, the bait molecule is suitable for both solution-phase and solid-phase hybridization to the target molecule. The design and construction of bait molecules is described in more detail in, e.g., International Patent Application Publication No. WO 2020 / 236941.
[0063] As used herein, the term “amplifying,” and its equivalents, may refer to a process of generating copies of a target molecule, such as a nucleic acid molecule.
[0064] As used herein, the term “hybridization,” and its equivalents, may refer to a process by which to complementary single-stranded nucleic acid molecules bind to one another, thereby forming a double- stranded nucleic acid molecule. In certain examples, the double-stranded nature of the nucleic acid molecule is maintained under stringent hybridization conditions. Exemplary stringent hybridization conditions include an overnight incubation at 42 °C in a solution including 50% formamide, 5XSSC (750 mM NaCl, 75 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5XDenhardt's solution, 10% dextran sulfate, and 20 µg / ml denatured, sheared salmon sperm DNA, followed by washing the filters in 0.1XSSC at 50 °C.
[0065] As used herein, the term “complementary,” and its equivalents, may refer to a state of two single- stranded nucleic acid molecules with respective sequences that cause the nucleic acid molecules to spontaneously hybridize to one another. One nucleic acid molecule, for instance, may have a sequence that causes each nucleic acid to hydrogen bond to a respective nucleic acid in the other nucleic acid molecule.
[0066] As used herein, the terms “therapy,” “therapeutic agent,” “treatment,” and their equivalents, may refer to a composition or process that can be used to remediate a health problem. Cancer therapies, for instance, include surgery, radiation therapy (e.g., radiotherapy), chemotherapy, immunotherapy, cell-based therapies, and the like. Examples of cancer therapies include abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), ado-trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), aldesleukin (Proleukin), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), asciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicabtagene ciloleucel (Yescarta), axitinib (Inlyta), belantamab mafodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), belzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexucabtagene autoleucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Ilaris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), cemiplimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), copanlisib hydrochloride (Aliqopa),FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (Darzalex Faspro), darolutamide (Nubeqa), dasatinib (Sprycel), denileukin diftitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostarlimab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enasidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glasdegib maleate (Daurismo), hyaluronidase- zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idecabtagene vicleucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infigratinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguane I131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa), ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (Somatuline Depot), lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), Lenvatinib mesylate (Lenvima), letrozole (Femara), lisocabtagene maraleucel (Breyanzi), loncastuximab tesirine-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu 177- dotatate (Lutathera), margetuximabcmkb (Margenza), midostaurin (Rydapt), mobocertinib succinate (Exkivity), mogamulizumab-kpkc (Poteligeo), moxetumomab pasudotox-tdfk (Lumoxiti), naxitamab-gqgk (Danyelza), necitumumab (Portrazza), neratinib maleate (Nerlynx), nilotinib (Tasigna), niraparib tosylate monohydrate (Zejula), nivolumab (Opdivo), obinutuzumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), olaratumab (Lartruvo), osimertinib (Tagrisso), palbociclib (Ibrance), panitumumab (Vectibix), panobinostat (Farydak), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pexidartinib hydrochloride (Turalio), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), pralsetinib (Gavreto), radium 223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), ripretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase human (Rituxan Hycela), romidepsin (Istodax), rucaparib camsylate (Rubraca), ruxolitinib phosphate (Jakafi), sacituzumab govitecanhziy (Trodelvy), seliciclib, selinexor (Xpovio), selpercatinib (Retevmo), selumetinib sulfate (Koselugo), siltuximab (Sylvant), sipuleucel-T (Provenge), sirolimus protein-bound particles (Fyarro), sonidegib (Odomzo), sorafenib (Nexavar), sotorasib (Lumakras), sunitinib (Sutent), tafasitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), tebentafusp-tebn (Kimmtrak), temsirolimus (Torisel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tositumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), toremifene (Fareston), tucatinib (Tukysa), umbralisib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), vismodegib (Erivedge), vorinostat (Zolinza), zanubrutinibFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT (Brukinsa), ziv-aflibercept (Zaltrap), and combinations thereof. Examples of cancer therapies also include targeted antibody-based therapies (antibody-drug conjugates, antibody-radioisotope conjugates, and targeted immune cell therapies (e.g., immune effector cells genetically modified to express a chimeric antigen receptor (CAR).
[0067] As used herein, the term “treatment-responsive,” and its equivalents, may refer to a type of cancer cells that can be substantially killed using a predetermined type of therapy. For example, cancer cells of a subject may be responsive to a particular treatment if, after the subject is administered the treatment, the cancer cells are diminished by a particular progression level (e.g., radiographic progression level, marker- based progression level, such as prostate-specific antigen (PSA) progression, etc.). Accordingly, the responsiveness of the cells to the type of therapy may indicate the effectiveness of that therapy.
[0068] As used herein, the term “treatment-resistant,” and its equivalents, may refer to a type of cancer that cannot be substantially killed using a predetermined type of therapy.
[0069] As used herein, the term “metastasis profile,” and its equivalents, may refer to a propensity of a type of cancer to metastasize into one or more differentiated tumor types besides the cancer’s tissue origin. In some implementations, the metastasis profile can further indicate the type of tissue in which the cancer can or is likely to metastasize.
[0070] As used herein, the term “clinical trial,” and its equivalents, may refer to a research study used to evaluate a hypothesis based on participation by one or more subjects. In various examples, a clinical trial can be used to assess the efficacy and / or safety of a proposed therapy. A clinical trial may be performed in furtherance of approval of a treatment by a regulatory authority (e.g., the United States Food & Drug Administration (FDA)). Description of Example Implementations
[0071] Various implementations of the present disclosure will now be described with reference to the accompanying Figures.
[0072] FIG. 1 illustrates an example environment 100 for predicting a heterogeneity condition of a tumor 102 of a subject 104 using genomic information of the subject 104. In some cases, the subject 104 may have one or more types of cancer and may present to a clinical environment for a disease assessment, such as an assessment of the progression of the cancer(s) and / or an assessment of the health of the subject 104.
[0073] According to various examples, the subject 104 has one or more types of cancer, such as adrenal cancer, bladder cancer, blood cancer, bone cancer, brain cancer, breast cancer, carcinoma, cervical cancer, colon cancer, colorectal cancer, corpus uterine cancer, ear, nose and throat (ENT) cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, head and neck cancer, Hodgkin's disease, intestinal cancer, kidney cancer, larynx cancer, leukemia, liver cancer, lymph node cancer, lymphoma, lung cancer, melanoma, mesothelioma, myeloma, nasopharynx cancer, a neuroblastoma, non-Hodgkin's lymphoma, oralFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT cancer, ovarian cancer, pancreatic cancer, penile cancer, pharynx cancer, prostate cancer, rectal cancer, sarcoma, seminoma, skin cancer, stomach cancer, a teratoma, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, a vascular tumor, or combinations or metastases thereof.
[0074] In some embodiments, the subject 104 has B cell cancer (multiple myeloma), a melanoma, breast cancer, lung cancer, bronchus cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, urinary bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine cancer, endometrial cancer, cancer of an oral cavity, cancer of a pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small bowel cancer, appendix cancer, salivary gland cancer, thyroid gland cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, a cancer of hematological tissue, an adenocarcinoma, an inflammatory myofibroblastic tumor, a gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphocytic leukemia (ALL), acute myelocytic leukemia (AML), chronic myelocytic leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia Vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft-tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell cancer, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familiar hypereosinophilia, chronic eosinophilic leukemia, a neuroendocrine cancer, or a carcinoid tumor
[0075] In various cases, one or more care providers 106 (also referred to as “healthcare provider(s)”) is responsible for monitoring and / or treating the subject 104. The care provider(s) 106 may include one or more of a clinician, a surgeon, an anesthesiologist, a pathologist, a laboratory technician, or the like.
[0076] According to some implementations, the tumor 102 may be initially identified using a noninvasive technique. For example, the tumor 102 may be visualized using an imaging modality, such as ultrasound, x-ray, computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), single photon emission CT (SPECT), or any combination thereof. Using the noninvasive technique, the care provider(s) 106 may identify the presence and location of the tumor 102. In various cases, the careFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT provider(s) 106 may surgically remove a sample from the tumor 102 to determine genomic information associated with the tumor 102.
[0077] For instance, a tissue sample 108 may be obtained from the subject 104. The tissue sample 108 is, in various cases, obtained by removing cells from the tumor 102. In some cases, the tissue sample 108 is obtained from a particular spatial position of the tumor 102. In some examples, the tissue sample 108 is surgically excised from the subject 104. In various cases, the tissue sample 108 may be frozen or fixed. For instance, the tissue sample 108 may be obtained from the subject 104 and fixed, by the care provider(s) 106, using a fixative (e.g., formalin, paraformaldehyde, ethanol, etc.). In some cases, freezing or fixing the tissue may prevent degradation of the tissue sample 108 between sample collection and sample analysis. In various examples, the tissue sample 108 is analyzed as fresh tissue.
[0078] The tissue sample 108 includes nucleic acid molecules. According to some examples, the nucleic acid molecules include genomic DNA (gDNA). For instance, the nucleic acid molecules include chromosomal DNA that is located in, or extracted from, cells in the tissue sample 108. According to some cases, the DNA is extracted from nuclei and the cells in the tissue sample 108 using mechanical shearing and / or the introduction of a chemical (e.g., a detergent). The DNA may be subsequently isolated from proteins and other cellular materials. In some implementations, the nucleic acid molecules indicate an entire genome of the subject 104 and / or the tumor 102. In some examples, the nucleic acid molecules indicate a full RNA transcriptome of the subject 104 and / or the tumor 102. Thus, a genome and / or an RNA transcriptome of the subject 104 and / or the tumor 102 can be determined by sequencing the DNA in the nucleic acid molecules. In various cases, the nucleic acid molecules indicate a whole exome of the subject 104 and / or the tumor 102.
[0079] In some examples, the nucleic acid molecules include RNA. In some implementations, the nucleic acid molecules include messenger RNA (mRNA), microRNA, non-coding RNA, functional RNA, or any combination thereof. Various RNA in the nucleic acid molecules may be indicative of proteins expressed in the cells of the subject 104 and / or the tumor 102.
[0080] In various cases, the tissue sample 108 is transported to a location that is remote from the subject 104 for further processing. For example, the tissue sample 108 is removed from the subject 104 in a clinical environment (e.g., a hospital) and is then transported to a remote laboratory for further testing and analysis.
[0081] In various cases, a first target sample 110 is obtained from a first target region of the tissue sample 108 to determine the genomic information associated with the tumor 102. For example, the tissue sample 108 may contain a greater number of cells and / or a greater nucleic acid yield than the sample throughput of the sample analysis described herein. The first target sample 110 includes a first subset of the nucleic acid molecules (also referred to as “first nucleic acid molecules”) 111. In some examples, the first targetFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT sample 110 represents a portion of the tissue sample 108. In various cases, a remaining portion of the tissue sample 108 may be fixed (e.g., treated with a fixation agent) or frozen for additional analysis.
[0082] A sequencer 112 is configured to generate first sequence read data 114 indicating the sequences of the first nucleic acid molecules 111. The sequencer 112, for instance, includes one or more devices that are configured to generate the first sequence read data 114 by processing at least a portion of the first target sample 110. For instance, the sequencer 112 may be configured to generate first sequence read data 114 by processing the first target sample 110. In some cases, the first nucleic acid molecules 111are extracted from the first target sample 110. The extraction can be performed by the sequencer 112, by another device, manually (e.g., by a laboratory technician), or any combination thereof. Any appropriate extraction method known to those of ordinary skill in the art can be utilized.
[0083] In various cases, the sequencer 112 is configured to perform one or more processes (e.g., chemical reactions) on the first nucleic acid molecules 111 in order to prepare the first nucleic acid molecules 111 for sequencing. For instance, the sequencer 112 may ligate adapters onto the first nucleic acid molecules 111 and / or amplify the first nucleic acid molecules 111, such that numerous copies of the ligated first nucleic acid molecules 111 are available for sequencing. Examples of the adapters include, for example, amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences. The first nucleic acid molecules 111 (e.g., the ligated nucleic acid molecules) may be amplified by generating multiple copies of the first nucleic acid molecules 111 using one or more techniques such as polymerase chain reaction (PCR), a non-PCR amplification technique, or an isothermal amplification technique.
[0084] The sequencer 112 may identify the length, position, and identity of the bases in the first nucleic acid molecules 111 by sequencing the first nucleic acid molecules 111 (e.g., the amplified and / or ligated first nucleic acid molecules 111). In various implementations, the sequencer 112 utilizes first-generation sequencing (e.g., Sanger sequencing), second-generation sequencing (e.g., massive parallel sequencing), third-generation sequencing (e.g., nanopore sequencing), or a combination thereof. In some cases, the sequencer 112 is configured to sequence substantially all of the nucleotides of all of the first nucleic acid molecules 111 fragments obtained from the first target sample 110. In some examples, the sequencer 112 is configured to perform targeted sequencing. For instance, the sequencer 112 may determine whether the first nucleic acid molecules 111 fragments contain one or more predetermined sequences at one or more genomic locations.
[0085] In various cases, the sequencer 112 includes one or more sensors that are configured to detect physical signals (also referred to as “detection signals”) that are indicative of the nucleotide sequences of the first nucleic acid molecules 111. The sequencer 112 may perform sequencing-by-synthesis. For example, the sequencer 112 may include one or more optical sensors configured to detect optical signalsFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT emitted from fluorescently tagged nucleotide triphosphates (NTPs) that are joined together in a synthesized DNA strand using the ligated first nucleic acid molecules 111 as templates. The optical signals detected by the optical sensor(s), for instance, are indicative of the sequences of the first nucleic acid molecules 111. The sequencer 112 may perform nanopore sequencing. In various cases, the sequencer 112 includes one or more electrical sensors configured to measure an electrical signal (e.g., an electrical current) across a substrate as the ligated first nucleic acid molecules 111 are directed through a nanopore extending through the substrate. The electrical signal over time, in various cases, is indicative of the sequences of the first nucleic acid molecules 111 in the tissue sample 108. The sequencer 112, in various implementations, is configured to generate the first sequence read data 114 as digital data based on the analog signals detected by the sensor(s). For instance, the sequencer 112 includes one or more analog to digital converters (ADCs). In various cases, the sequencer 112 includes at least one processor configured to generate the first sequence read data 114.
[0086] In some implementations, the sequencer 112 performs RNA sequencing (RNA-seq) on the first nucleic acid molecules 111. For example, the first nucleic acid molecules 111 include RNA that is extracted from the first target sample 110. In some examples, the RNA in the first nucleic acid molecules 111 is fragmented. In various implementations, complementary DNA (cDNA) is generated using reverse transcriptase, such that the cDNA includes sequences that are complementary to the RNA in the first nucleic acid molecules 111 from the first target sample 110. The cDNA, according to various cases, can be sequenced using the DNA sequencing techniques described above. Accordingly, in some cases, the first sequence read data 114 indicates sequences of RNA present in the first target sample 110, which may be indicative of the transcriptome of the subject 104 and / or the tumor 102.
[0087] In various cases, the sequencer 112 performs sequencing on a subset of the first nucleic acid molecules 111. For instance, the sequencer 112 may perform targeted sequencing on one or more predetermined genes. In various cases, the genes include one or more of ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (C11orf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESR1, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3,FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-1β, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI- H, mTOR, PARP, PD-1, PDGFR, PDGFRα, PDGFRβ, PD-L1, PI3Kδ, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, or VEGFB. In some examples, the genes include one or more of POLE, TP53, CTNNNB1, L1CAM, PTEN, ERBB2, PMS2, MSH2, MSH6, MLH1, an estrogen receptor (ER) gene, or a progesterone receptor (PR) gene. The sequencer 112, in some cases, may refrain from sequencing at least a portion of the first nucleic acid molecules 111 that do not correspond to the subset.
[0088] In various implementations of the present disclosure, the tumor 102 is heterogenous and includes cells with genetic variability. For instance, the first target region may include first cells that include first nucleic acid molecules 111 with first sequences, and a second target region of the tissue sample 108 may include second cells that include a second subset of the nucleic acid molecules (also referred to as “second nucleic acid molecules”) 115. The second nucleic acid molecules 115, in various examples, may have second sequences that are different that the first sequences.
[0089] If the sequencer 112 is only configured to identify the first sequences, rather than both the first sequences and the second sequences, then the care provider(s) 106 may be unable to appreciate the heterogeneity of the tumor 102 and / or the tissue sample 108. Without context regarding the heterogeneity,FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT the care provider(s) 106 may be unable to accurately characterize the cancer of the subject 104. For instance, the care provider(s) 106 may be unable to accurately identify whether a therapy is predicted to be effective at treating the tumor 102, whether a therapy is predicted to be ineffective at treating the tumor 102, a survivability of the subject 104 (e.g., a likelihood that the subject 104 will survive by a predetermined date or time), an expected quality of life of the subject 104, a predicted tumor progression (e.g., a rate of growth of the tumor 102 or a likelihood that the tumor 102 will metastasize by a predetermined date), another factor relevant to the prognosis associated with the cancer of the subject 104, or any combination thereof.
[0090] In a particular case, the first cells are susceptible to an example immunotherapy, but the second cells are resistant to the example immunotherapy. If the care provider(s) 106 only has insight into genomic features of the first cells, and not the second cells, the care provider(s) 106 may administer the example immunotherapy to the subject 104 with the expectation that cells throughout the entire tumor 102 are susceptible to the example immunotherapy. In this case, however, the second cells in the tumor 102 could remain untreated after administration of the example immunotherapy. As a result, the cancer of the subject 104 may remain despite the attempted treatment, which can harm the subject 104. For example, the second cells in the tumor 102 may metastasize during a period when the example immunotherapy is administered, which could reduce the overall survivability of the subject 104.
[0091] In various implementations of the present disclosure, a condition of the subject 104 can be determined based on the first sequence read data 114 corresponding to the first target region and second sequence read data 116 corresponding to the second target region. For instance, the care provider(s) 106 may identify, in the tissue sample 108, the first target region and the second target region. The first and second target regions are, in various cases, spatially distinct. In various cases, the first and second target regions are disposed in different areas and / or volumes of the tissue sample 108.
[0092] The first target sample 110 and a second target sample 118 may be obtained, in some examples, from the first and second target regions, respectively. The sequencer 112 may be configured to generate the first sequence read data 114 based on the first target sample 110, and to generate the second sequence read data 116 based on the second target sample 118. In various cases, characteristics of the subject 104 can be determined more accurately by analyzing the first sequence read data 114 associated with the first target sample 110 and the second sequence read data 116 associated with the second target sample 118, relative to analyzing the first sequence read data 114 alone. For instance, the first sequence read data 114 may indicate a first therapy that effectively targets the first cells, and the second sequence read data 116 may indicate that the second cells are resistant to the first therapy. The second sequence read data 116 may indicate a second therapy that effectively targets the second cells.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT
[0093] In some examples, the care provider(s) 106 may identify the first and second target regions based on a visual characteristic of the tissue sample 108. For instance, the visual characteristic may include a histopathological morphology (e.g., cellular size, cellular morphology, cell type, a subcellular structure, one or more biomarker levels, cellular architecture, tissue architecture, etc.) and / or a vascularization (e.g., a vascularization pattern, a density of blood vessels, a structure or type of blood vessels, etc.). In various examples, the care provider(s) 106 may perform staining of the tissue sample 108 before analysis. For instance, the tissue sample 108 may be stained with a histological stain or an immunohistological stain. The staining may facilitate the identification of the visual characteristic. The care provider(s) 106 may identify metrics associated with more than one region of the image. The metrics may be indicative of a visual characteristic. In some examples, the metrics associated with the regions of the image are compared to determine the first and second target regions. For example, the first target region may include cells with a distinct histopathological morphology and / or vascularization than the second target region.
[0094] In some examples, the first target region is spaced apart from the second target region at a first predetermined interval (e.g., a distance) along a first axis. For instance, the predetermined interval may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 millimeters (mm). In some examples, the predetermined interval is 20, 40, 60, 80, 100, or more than 100 mm. In various implementations of the present disclosure, a third target region is spaced apart from the first target region at a second predetermined interval along a second axis. In some examples, the first axis and the second axis intersect (e.g., the first axis and the second axis are not parallel). For example, the first axis may be perpendicular to the second axis. In some instances, the first axis and the second axis intersect at a 10, 20, 30, 40, 50, 60, 70, or 80° angle. In some examples, the first axis and the second axis correspond to points located at a first and second radius, respectively, from a reference point. The reference point may correspond to the center of the tumor 102. For instance, the first axis may correspond to a circle or a sphere with the first radius from the reference point. In various cases, the first axis corresponds to a cylinder with the first radius from a reference axis that includes the reference point.
[0095] In various implementations, the first and second target samples 110 and 118 may be obtained from the first and the second target regions, respectively. For instance, the care provider(s) 106 may collect needle punch enrichment samples from the first and second target regions. In various cases, the care provider(s) 106 may perform curl tissue extraction or straight razor blade extraction at the first and the second target regions. In some implementations, a single apparatus configured to simultaneously obtain multiple needle punch enrichment samples from distinct regions in the tissue sample 108 is utilized to obtain the first target sample 110 and the second target sample 118. For example, the apparatus includes multiple needles disposed parallel to each other and separated by one or more distances (e.g., in a range of 10 to 100FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT mm), such that the multiple needles can obtain distinct target samples from the tissue sample 108 at different spatial positions.
[0096] While FIG. 1 illustrates the first and second target samples 110 and 118 in the tissue sample 108, implementations of the present disclosure are not so limited. For instance, in some examples, the care provider(s) 106 (and / or some other entity) may identify the first target region in a first tissue sample collected from the subject 104 and the second target region in a second tissue sample collected from the subject 104. Collecting more than one tissue sample from the subject 104 may enable more comprehensive sampling of the tumor 102, for instance, by facilitating analysis of regions of the tumor that are physically located at a greater distance from each other. In some examples, the first and second tissue samples may be identified based on a visual characteristic. For instance, the care provider(s) 106 may perform a procedure to examine the body of the subject 104 (e.g., colposcopy, an endoscopy, a cystoscopy, a laparoscopy, a hysteroscopy, a colonoscopy, etc.). During the procedure, the care provider(s) 106 may identify regions of interest based on a visual characteristic (e.g., a vascularization pattern). In some examples, the first tissue sample is obtained from a first region of interest, and the second tissue sample is obtained from a second region of interest. The first target sample 110 is, in various cases, obtained from the first tissue sample, and the second target sample 118 is obtained from the second tissue sample.
[0097] In various cases, the spatial positions may be provided to the sequencer 112, such that the first and second sequence read data 114 and 116 include indications of the corresponding spatial position. The spatial positions may be indicative of the predetermined interval between the first and second target regions and the corresponding axis. In some examples, the spatial positions are indicative of the positions of the first and the second target regions relative to a particular landmark (e.g., a position within the tissue sample 108, a position along an edge of the tissue sample 108, a position relative to the tumor 102, a position relative to the body of the subject 104, a position relative to one or more reference pixels in an image of the tissue sample 108, etc.).
[0098] A genomic analyzer 120 identifies genomic features 122 of the first and second nucleic acid molecules 111 and 115 by analyzing the first and second sequence read data 114 and 116. In some implementations, the first and second sequence read data 114 and 116 are combined. For instance, the first and second nucleic acid molecules 111 and 115 may be combined and provided to the sequencer 112. Accordingly, the sequencer 112 may determine combined sequence read data 123 that is indicative of the first and second sequence read data 114 and 116. In some examples, the sequencer 112 may combine the first and second sequence read data 114 and 116 into the combined sequence read data 123. Accordingly, the genomic analyzer 120 may identify genomic features 122 of the first and second nucleic acid molecules 111 and 115 by analyzing the combined sequence read data 123. In some examples, the genomic analyzerFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT 120 identifies the genomic features 122 at least in part based on the spatial positions of the first and second target regions.
[0099] In various implementations, the genomic analyzer 120 identifies, calculates, or otherwise determines the genomic features 122 based on the sequences of the first and second nucleic acid molecules 111 and 115 indicated in the first and second sequence read data 114 and 116. One or more types of features are identified by the genomic analyzer 120. The genomic features 122 may be derived from the first and second sequence read data 114 and 116. The genomic features 122 may be indicative of the associated sequence read data. For instance, the genomic features 122 may include an indication of whether each feature is associated with the first sequence read data 114, the second sequence read data 116, or the combined sequence read data 123.
[0100] In some cases, the genomic features 122 include at least one mismatch repair deficiency (MMRD) probability score. In various cases, a MMRD probability score indicates a likelihood that one or more MMR pathways of cells in the first and / or second target samples 110 and 118 are ineffective at performing mismatch repair. In some implementations, the MMRD probability score is determined by determining genomic features 122 by analyzing the first and second sequence read data 114 and 116, inputting the genomic features 122 into at least one trained machine learning model trained to generate the MMRD probability score based on previously analyzed data from a population omitting the subject 104. The genomic features 122 relevant to the MMRD probability score include, for instance, a fraction unstable score, a composite COSMIC single-base substitution signature, a COSMIC indel signature, a copy number signature, a tumor mutational burden score, a blood-based tumor mutational burden score, a germline status for a mutation in one or more genes associated with DNA mismatch repair (MMR) (also referred to as “MMR genes”), a methylation status for the one or more MMR genes, a methylation status for one or more promoters associated with the one or more MMR genes, a methylation status of one or more enhancers associated with the one or more MMR genes, or any combination thereof. Examples of the MMR genes include, for instance, MSH2, MSH6, PMS2, or MLH1.
[0101] The genomic features 122, in some examples, include at least one copy number state of one or more genetic loci indicated by the first and / or second sequence read data 114 and 116. In various implementations, a number of copies of a predetermined sequence at a given locus in the genome of the subject 104 and / or the tumor 102 (also referred to as a “copy number” of the locus) is determined. The copy number state, in various implementations, may indicate copy numbers of one or more loci in the genome of the subject 104 and / or the tumor 102. For instance, the copy number state may indicate the presence and / or amount of copies of various sequences present in the genome of the subject 104 and / or the tumor 102, which may be due to copy number variation.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT
[0102] According to various examples, the first and second sequence read data 114 and 116 may represent a genome of the subject 104 and / or the tumor 102. Various portions of the first and second sequence read data 114 and 116 are aligned with at least one reference sequence (e.g., a reference genome). The aligned data is segmented using at least one segmentation technique (e.g., a circular binary segmentation (CBS) method, a maximum likelihood method, a hidden Markov chain method, a walking Markov method, a Bayesian methods, a long-range correlation method, a change point method, or any combination thereof), thereby generating non-overlapping segments of the first and second sequence read data 114 and 116, wherein a sequence associated with a given segment is associated with the same copy number (e.g., a number of instances in which the sequence appears in the segment). Various genetic loci are binned, or otherwise sorted, with respect to the segments of the genome of the subject 104 and / or the tumor 102. The copy number state, for instance, is representative of the respective copy numbers associated with the genetic loci.
[0103] In some implementations, the genomic features 122 include the presence or absence of a variant (e.g., a pathogenic variant) in one or more genes associated with classifying the tumor 102. The genomic features 122, in some examples, are indicative of a quantity of variants in the first and / or second sequence read data 114 and 116.
[0104] In some cases, the genomic features 122 are indicative of microsatellite instability (MSI). Microsatellites are highly polymorphic DNA-repeat regions. In certain examples, “microsatellite” refers to a repetitive nucleic acid having repeat units of less than about 10 base pairs or nucleotides in length. In certain examples, a microsatellite refers to a tract of tandemly repeated (i.e., adjacent) DNA motifs ranging from one to six or up to ten nucleotides, with each motif repeated 5 to 50 repeated times. During DNA replication, mutations (e.g., insertions or deletions) are more likely to be introduced at microsatellites than various other portions of the genome. In various cases, these mutations are corrected via MMR pathways. However, if the MMR pathways are impaired (e.g., the MMR genes of the hosting cell include variants that impede function), then the mutations at the microsatellites may be substantially retained. “Microsatellite instability” refers to genetic instability in the microsatellite regions. Cancer patients with microsatellite instability classified as being high (MSI-H or MSI-High) frequently exhibit an accumulation of somatic mutations in tumor cells that leads to a range of molecular and biological changes including high tumor mutational burden, increased expression of neoantigens and abundant tumor-infiltrating lymphocytes. Chang et al. “Microsatellite Instability: A Predictive Biomarker for Cancer Immunotherapy,” Appl Immunohistochem Mol Morphol, 26(2):e15-e21 (2018). These changes have been linked to increased sensitivity to checkpoint inhibitor drugs, such as pembrolizumab, which is used to treat advanced melanoma, head and neck squamous cell carcinoma, non-small cell lung cancer (NSCLC), and classical Hodgkin lymphoma. According to various examples, “MSI score” refers to an amount of instability in oneFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT or more microsatellites. For example, an MSI score can be represented as a fraction (i.e., an “MSI fraction”) of instability in the one or more microsatellites. Other types of portions of DNA may be associated with a high likelihood of mutations. In some cases, the genomic features 122 include a fraction unstable score, indicative of mutations in the microsatellites and other portions of the genome that are prone to mutations.
[0105] In various cases, an MSI score can be determined based on a predetermined set of repetitive loci (e.g., 2000 repetitive loci, each with a minimum of 5 repeat units of mono-, di-, and trinucleotides). By evaluating the first and / or second sequence read data 114 and 116, the genomic analyzer 120 may determine lengths of repetitive sequences corresponding to the loci. If an example locus among the loci corresponds to a predetermined repeat length, the locus is considered to be “unstable.” The MSI score, for instance, is determined by determining an amount of the unstable loci (e.g., a fraction of the unstable loci with respect to the total number of repetitive loci evaluated). In some cases, the MSI score is used to determine whether the subject 104 and / or tumor 102 is MSI-High (MSI-H). For example, MSI-H status may be applicable if the MSI score is greater than a threshold (e.g., 0.5%). Techniques for determining MSI scores are described, for instance, in Woodhouse et al., “Clinical and analytical validation of FoundationOne LiquidCDx, a novel 324-Gene cfDNA-based comprehensive genomic profiling assay for cancers of solid tumor origin,” PLoS ONE 15(9) (2020).
[0106] In some implementations, the genomic features 122 are indicative of homologous recombination deficiency (HRD). Homologous recombination includes a series of pathways that enable the repair of nucleic acids. In certain examples, homologous recombination is used by cells to repair DNA double- stranded breaks (DSBs) and collapsed replication forks. DNA DSBs are often considered the most dangerous form of DNA damage and are associated with genomic instability, which can lead to cancer development when cells are unable to repair the DNA damage. In various cases, HRD, or the inability to repair DNA damage, can be determined based on mutations in a variety of genes, including BRCA1, BRCA2, ATM, BARD1, BRIP1, CDK12, CHEK1, CHEK2, FANCL, PALB2, RAD51B, RAD51C, RAD51D, RAD54L, or a combination thereof. In some examples, the genes associated with HRD may be determined based on the type of cancer of the subject 104.
[0107] In some implementations, the genomic features 122 are indicative of alteration-level clonality. Clones are genetically identical cell populations. In various cases, tumors are associated with clonal evolution due to the high rate of cell proliferation and genomic instability. For instance, a tumor may include a founding clone that is associated with tumorigenesis. Subsequently, a subclone may originate in the tumor when a cell of the founding clone acquires a mutation. The mutated cell may, in various cases, multiply to generate a subclonal population. The genomic features 122, in some examples, are indicative of a number of clones, a proportionality of each clone, or a genetic variation between each of the clones of the tumor 102. In various instances, the first target sample 110 may include cells of a first clone, and the second targetFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT sample 118 may include cells of a second clone. For instance, the genomic features 122 may indicate that the first clone differs by one mutation from the second clone.
[0108] In various implementations, the genomic features 122 are indicative of an intra-tumor heterogeneity of the tumor 102. In some cases, the intra-tumor heterogeneity is based at least in part on the alteration- level clonality. In various examples, the intra-tumor heterogeneity is based on at least one of genetic variation, phenotypic variation, or microenvironment variation (e.g., epigenetic variation) within the tumor 102. “Cell population,” as used herein, refers to one or more cells that are genetically and phenotypically similar. In some cases, the genomic features 122 are indicative of a number of cell populations or a proportion of each of the cell populations in the tumor 102. In various examples, the genomic features 122 are indicative of a metric associated with a variation (e.g., genetic variation, phenotypic variation, microenvironment variation, or a combination thereof) between the cell populations in the tumor 102.
[0109] In some implementations, the genomic features 122 include one or more mutation signatures. In various cases, a mutational signature can represent an amount and / or identity of mutations (e.g., insertions, deletions, double-base substitutions, single-base substitutions, or any combination thereof) indicated in the first and / or second nucleic acid molecules 111 and 115 from the subject 104. In some cases, the mutational signature indicates an amount (e.g., number or percentage) of individual classes of base substitutions present in the first and / or nucleic acid molecules. For instance, the classes include single-base substitutions including C>A, C>G, C>T, T>A, T>C, and T>G. A mutational signature can be derived by comparing the sequences indicated in the first and / or second sequence read data 114 and 116 to at least one reference sequence, such as a reference genome. For example, the genomic features 122 may include a Catalogue Of Somatic Mutations In Cancer (COSMIC) mutational signature, such as a COSMIC indel signature. In some cases, the genomic features 122 include a single-base substitution signature. According to some cases, the mutation signature is derived based on a model, such as an autoencoder model.
[0110] In various examples, the genomic features 122 include at least one tumor mutational burden (TMB) score. The TMB, for instance, is a measure of the number of mutations carried by the cancer cells of the tumor 102. By comparing DNA sequences from a patient’s healthy tissues and cancer cells, the number of acquired somatic mutations present in tumors, but not in normal tissues, may be determined. In some instances, driver mutations may be excluded from a TMB calculation. In certain examples, "tumor mutational burden" or “TMB score” refers to the number of somatic mutations in a tumor's genome and / or the number of somatic mutations per area of the tumor's genome. In some embodiments, TMB, as used herein, refers to the number of somatic mutations per megabase (Mb) of DNA sequenced. In some embodiments, germline (inherited) variants are excluded when determining TMB, given that the immune system has a higher likelihood of recognizing these as self. In addition, germline variants do not reflect theFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT biology of somatic mutation for the purposes of TMB determinations. In various cases, driver mutations are excluded from a TMB calculation.
[0111] In some cases, the genomic features 122 include the presence, amount, type, or any combination thereof, of one or more hotspot mutations in the first and / or second nucleic acid molecules 111 and 115. Hotspots, for instance, can refer to loci in the genome of the subject 104 and / or the tumor 102 that are prone to mutation. Examples of hotspots include CpG islands, microsatellites, centromeric DNA, telomers, subtelomeric regions, common fragile sites, palindromic AT-rich repeats (PATRRs), G-quadruplexes, R- loops, and the like.
[0112] Hotspot mutations give rise to oncological outcomes. PhyloP, SIFT, Grantham, COSMIC and PolyPhen-2 are in silico tools that can be used to assess pathogenicity of identified variants. Exemplary hotspot genes and mutations include EGFR exon 19 activating mutation, EGFR exon 19 deletion, EGFR exon 19 insertion, EGFR exon 19 sensitizing mutation, EGFR exon 20 activation mutation, EGFR exon 20 insertion, EGFR G719 mutation, EGFR L858R mutation, EGFR L861 mutation, EGFR S768 mutation, EGFR T790M mutation, C797 mutation, KIT activating mutation, KRAS activating mutation, MET activating mutation, NRAS activating mutation, PMS2 promoter mutations, among many others. Hotspot mutations also occur in the following genes: AKT2, BRCA1, BRCA2, ERC1, NSD1, POLH, PPM1G, PTEN, RAD18, RAD51, RAD51B, RB1, TERT, TP53, TP53Bp1, ALK, ARMT1, ATAD5, ATG7, ATIC, AXL, BIRC6, BRD3, BRD4, CAPRIN1, CCAR2, CCDC6, CDK5RAP2, CHD9, CIT, CTNNB1, CUL1, EBF1, EIF3E, HIP1, HMGA2, IRF2BP2, NOTCH1, NOTCH4, NPM1, OFD1, TACC1, TACC3, TERF2, TMEM106B, UBE2L3, USP10, WRDR48, YAP1, ZEB2, and ZMYND8.
[0113] In various examples, the genomic features 122 are included in the input data provided to a predictive model 124. The predictive model 124, according to various implementations, is configured to determine one or more heterogeneity indicator 125 and one or more condition indicator 126 based at least in part on the genomic features 122. The heterogeneity indicator(s) 125, in various cases, is indicative of a heterogeneity of the tumor 102 of the subject 104. The condition indicator(s) 126, in some examples, is indicative of a condition of subject 104. The predictive model 124 may, in some examples, determine the condition of the subject 104 based at least in part on the heterogeneity indicator 125.
[0114] The predictive model 124, for example, may include one or more mathematical and / or computer- based models that are configured to predict the heterogeneity indicator(s) 125 and / or the condition indicator(s) 126 based on the genomic features 122. For instance, the predictive model 124 may include a regression model, threshold rule, confidence interval, or other type of statistical model capable of categorizing the tumor 102 based on the genomic features 122. In various cases, the predictive model 124 includes at least one classifier configured to generate the heterogeneity indicator(s) 125 and / or the condition indicator(s) 126 based on the genomic features 122. In various cases, the predictive model 124 includes oneFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT or more ML models that have been pretrained to predict whether the tumor 102 is susceptible to at least one treatment (e.g., an immunotherapy).
[0115] The predictive model 124 may predict the susceptibility of the tumor 102 to various types of treatments. For example, the predictive model 124 may be configured to determine whether the tumor 102 is susceptible to a chemotherapy, an immunotherapy, a radiotherapy, a surgical intervention, a cell-based therapy, any other type of therapy described herein, or any combination thereof. In some cases, the predictive model 124 is configured to determine whether the tumor 102 is susceptible to a combination therapy that includes two or more types of treatments to be administered to the subject 104. In some examples, the heterogeneity indicator(s) 125 include a likelihood that the tumor 102 is susceptible to a given treatment or an indication (e.g., a Boolean value) that there is greater than a threshold (e.g., 90%) likelihood that the tumor 102 is susceptible to the given treatment. In some cases, the heterogeneity indicator(s) 125 include a likelihood that the tumor 102 is not susceptible (e.g., resistant) to the given treatment or an indication that there is greater than a threshold (e.g., 90%) likelihood that the tumor 102 is not susceptible to the given treatment.
[0116] The predictive model 124 may predict an evolution of the tumor 102 of the subject 104. For instance, the predictive model 124 may generate an evolutionary tree of the tumor 102. The evolutionary tree may indicate a clonal evolution of the tumor 102. The evolution of the tumor 102 may indicate when a given clone originated and / or one or more mutations associated with the given clone. The evolution of the tumor 102 may include indications of characteristics associated with the given clone. For instance, the evolution of the tumor 102 may indicate that the given clone is resistant to a treatment (e.g., a chemotherapy). In various examples, the predictive model 124 may predict the clonal evolution of the tumor 102.
[0117] The predictive model 124 may predict a progression of the tumor 102 of the subject 104. For instance, the predictive model 124 may predict that a rate of progression of the tumor 102 has decreased due to one or more mutations indicated by the evolution of the tumor 102. In various cases, the predictive model 124 may predict a time (e.g., a date, a time range) when the tumor 102 will metastasize. In some examples, the heterogeneity indicator(s) 125 include a likelihood that the tumor 102 will metastasize (e.g., to the lymph nodes, or to a particular organ) by a given time or an indication that there is greater than a threshold (e.g., 80%) likelihood that the tumor 102 will metastasize by the given time.
[0118] The characteristics of the heterogeneity of the tumor 102 indicated by the genomic features 122 may be relevant to predictions and / or determinations provided by the predictive model 124. For example, the first cells of the first target region may be susceptible to an immunotherapy, but the second cells of the second target region may be resistant to the immunotherapy. By considering the genomic features 122 ofFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT the first and second target regions, the predictive model 124 may more accurately predict the efficacy of cancer treatments for the subject 104 than models that rely solely on features of the first target region.
[0119] In some examples, the predictive model 124 may also determine the heterogeneity indicator(s) 125 and / or the condition indicator(s) 126 based on characteristics of the first cells and the second cells. For example, the predictive model 124 may determine the condition indicator(s) 126 based on the presence of an analyte (e.g., a protein or a nucleic acid associated with heterogeneity of a tumor and / or a pathological disease) in the first target sample 110 and / or the second target sample 118. For instance, the predictive model 124 may determine a prognosis associated with a breast cancer of the subject based on the presence of estrogen receptor (ER), progesterone receptor (PR), human epidermal growth factor receptor 2 (HER2), HER3, Ki-67, CA 15-3, E-Cadherin, cyclin D1, vascular endothelial growth factor (VEGF), or a combination thereof. In some examples, the first nucleic acids are extracted from a portion of the first target sample 110. The presence of the analyte may be determined in the remaining portion of the first target sample 110. For instance, the remaining portion of the first target sample 110 may be stained an immunostain that targets an antigen. In various examples, an imaging device 128 captures at least one image 130 of the stained first target sample 110.
[0120] In various cases, the image analyzer 132 may analyze the immunostained depictions of the first target sample 110 in the image(s) 130 in order to determine whether the first cells of the first target sample 110 express the targeted antigen. For instance, the image analyzer 132 may include a CNN trained to identify whether the first cells express the antigen. In some cases, the image analyzer 132 may predict that the first cells express the antigen by determining that a signal (e.g., an amount of a signal in the image(s) representing light emitted by the immunostain) is greater than a threshold. According to some examples, the image analyzer 132 may generate a phenotypic indicator 134 that represents whether the antigen is present on, in, or otherwise expressed by, the first cells and / or the second cells. In some examples, an analyte profile may be generated based on the phenotypic indicator 134. The analyte profile may indicate the presence or a level of one or more analytes in the first target sample 110 and / or the second target sample 118.
[0121] In various implementations, the image analyzer 132 may analyze the image(s) 130 to identify the first and second target regions. The image(s) 130 may be two-dimensional (2D) or a three-dimensional (3D) image(s). For instance, the care provider(s) 106 and / or an entity (e.g., a computing device) may identify, in the image(s) 130 of the tissue sample 108, metric associated with visual characteristics (e.g., cellular morphology) of the first and second target regions.
[0122] The input data for the predictive model 124, in some cases, further includes the phenotypic indicator 134 and / or the analyte profile. For instance, the predictive model 124 may determine that there is a greater likelihood that the first cells and / or the second cells are susceptible to an immunotherapy targeting theFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT antigen if the phenotypic indicator 134 indicates that the first cells and / or the second cells express the antigen. In some cases, the predictive model 124 may determine that there is a minimal likelihood that the first cells and / or the second cells are susceptible to the immunotherapy if the phenotypic indicator 134 indicates that the first cells and / or the second cells do not express the antigen. In various cases, the predictive model 124 may determine the heterogeneity indicator(s) 125 and / or the condition indicator(s) 126 based on the phenotype of the first cells and the second cells. For instance, the image analyzer 132 may analyze the immunostained depictions of the first target sample 110 in the image(s) 130 in order to determine physical properties (e.g., cell morphology, cellular function, gene expression, protein expression, cell cycle regulation, etc.) of the first cells and the second cells. For instance, the image analyzer 132 may include a CNN trained to classify a cell morphology of the first cells and the second cells. The phenotypic indicator 134 may represent the cell morphology of the first cells and the second cells.
[0123] A report generator 136 is configured to generate a report 138 based, at least in part, on the heterogeneity indicator(s) 125 and / or the condition indicator(s) 126. The report 138, for example, includes consumable data that can inform the care provider(s) 106 about the heterogeneity indicator(s) 125 and / or the condition indicator(s) 126 of the subject 104. In various implementations, the report 138 may indicate the results of additional analyses, such as the results of a histological study, whole transcriptome sequencing, RNA sequencing, whole exome sequencing (WES), whole genome sequencing, a gene expression profiling test, a cancer (e.g., DNA) hotspot panel test, a DNA methylation test, a tumor mutational burden (TMB) test, a DNA fragmentation test, an RNA fragmentation test, a microsatellite instability (MSI) test, a tumor mutational burden (TMB) test, or a viral status test. The performance of such tests is within the ordinary skill of the art, with additional detail provided elsewhere herein. The report 138, for example, may include a genomic profile of the subject 104 based on various combinations of the above analyses and tests. The report 138 may, in some examples, include an indication of the spatial positions of the first and second target regions relative to the tissue sample 108, the tumor 102, or the body of the subject 104.
[0124] In some implementations, the report 138 indicates that a follow-up test of the subject 104 is indicated. For instance, in response to determining that the categorization of the disease is inconclusive, the report generator 136 may generate the report 138 to indicate that one or more additional tests (e.g., a histological study, genome sequencing, exome sequencing, additional DNA sequencing, RNA sequencing, transcriptome sequencing, etc.) should be performed in order to identify the cancer of the subject 104.
[0125] In various cases, the report 138 is output to a clinical device 140. For example, the report generator 136 transmits the report 138 to the clinical device 140. In various implementations, the clinical device 140 is a computing device that is operated by, owned by, or otherwise associated with the care provider(s) 106. For instance, the clinical device 140 may be a desktop computer, a laptop computer, a smart phone, or someFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT other computing device associated with the care provider(s) 106. The clinical device 140, in various cases, outputs the report 138 to the care provider(s) 106. In some cases, the clinical device 140 includes a display (e.g., a screen) that visually presents the report 138. In various cases, the clinical device 140 includes a speaker that outputs a sound indicative of the report 138. The clinical device 140, in various cases, may output the information in the report 138 using one or more output mechanisms or devices.
[0126] The care provider(s) 106 may review the report 138 by interacting with the clinical device 140. The report 138, in various cases, may enhance the clinical decision-making of the care provider(s) 106. For instance, the care provider(s) 106 may prepare and / or administer a treatment to the subject 104 based on the report 138. For instance, the care provider(s) 106 may determine a dosage of the treatment based on the report 138. According to various implementations, the care provider(s) 106 may initiate the treatment and / or refer the subject 104 to another care provider to receive the treatment. In various cases, the care provider(s) 106 may prescribe, suggest, or administer an anticancer agent for the subject 104. For example, the care provider(s) 106 may rely on the heterogeneity indicator(s) 125 and / or the condition indicator(s) 126 reflected in the report 138 to select a treatment that the tumor 102 is predicted to be susceptible to.
[0127] In various implementations, the care provider(s) 106 may develop a diagnosis and / or prognosis of the subject 104 based on the report 138. In various implementations, the care provider(s) 106 may communicate information in the report 138 to the subject 104.
[0128] FIG. 1 illustrates various elements that can be embodied in one or more computing devices. For example, at least a portion of the functions of the sequencer 112, the genomic analyzer 120, the predictive model 124, the imaging device 128, the image analyzer 132, the report generator 136, the clinical device 140, or any combination thereof, are performed by one or more processors in at least one computing device. Examples of computing devices include server computers, desktop computers, laptop computers, tablet computers, mobile phones, wearable devices, Internet of Things (IoT) devices, and the like. In various cases, instructions for performing at least a portion of the functions of these elements are stored in memory and / or in a non-transitory computer readable medium. The instructions, for instance, are executed by the processor(s).
[0129] FIG. 1 also illustrates various types of data. For example, the first sequence read data 114, the second sequence read data 116, the genomic features 122, the heterogeneity indicator(s) 125, the condition indicator(s) 126, the image(s) 130, the phenotypic indicator 134, the report 138, or any combination thereof, includes data. The various types of data illustrated in FIG. 1 may be stored, such as in memory or in non- transitory computer readable media. In various implementations, at least a portion of the data is transmitted or otherwise output by one or more computing devices. For example, a computing device may transmit one or more communication signals to another computing device, wherein the communication signal(s) encode at least a portion of the data. Examples of communication signals include electromagnetic signals,FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT optical signals, ultrasonic signals, optical signals, and electrical signals. For example, communication signals can be transmitted wirelessly and / or in a wired fashion. The communication signals, for instance, are transmitted over one or more wireless channels and / or one or more wired channels (e.g., optical cabling, electrical cabling, etc.). In various cases, the communication signal(s) are transmitted over one or more communication networks. A communication network, for instance, may be defined according to one or more physical channels, such as one or more frequency spectra. In some cases, a communication network is defined according to one or more communication protocols and / or standards. Examples of communication networks include fiber optic networks, Institute of Electrical and Electronics Engineers (IEEE) networks (e.g., WI-FI™ networks, WiMAX networks, BLUETOOTH™ networks, etc.), cellular networks (e.g., a 3rdGeneration Partnership Project (3GPP) radio network, such as a Long Term Evolution (LTE) network, a New Radio (NR) network; or a cellular core network such as a 3rdGeneration (3G) core, a 4thGeneration (4G) core, a 5thGeneration (5G) core, etc.), ultrasonic networks, and the like. In some cases, the data is broadcasted from one device to multiple other devices. In some cases, the data is unicasted from one device to another device. For instance, various forms of data described herein may be transmitted via a peer-to-peer (P2P) connection.
[0130] A particular example will now be described with reference to FIG.1. In this example, the subject 104 presents to a clinical environment for an assessment (e.g., a follow-up visit) of a colon cancer of the subject. The care provider(s) 106 may initiate and / or identify results of medical imaging (e.g., colonoscopy, computed tomography (CT) imaging) on the subject 104. In some examples, the care provider(s) 106 may determine a location of the tumor 102 on a colon of the subject 104. The care provider(s) 106 may obtain the tissue sample 108, for instance, by performing a needle biopsy procedure on the tumor 102.
[0131] In various cases, the care provider(s) 106 may determine the first target region and the second target region based on a predetermined interval and a reference point. In some examples, the care provider(s) 106 may collect, from the first target region, the first target sample 110 and, from the second target region, the second target sample 118. The care provider(s) 106 may record the spatial positions of the first target region and the second target region relative to a position along the edge of the tissue sample 108. In some examples, the first nucleic acid molecules 111 and the second nucleic acid molecules 115 are extracted from the first target sample 110 and the second target sample 118, respectively. The extracted first and second nucleic acid molecules 111 and 115 are, in various instances, sequences by the sequencer 112. In some cases, the genomic features 122 identified by the genomic analyzer 120 include a fraction unstable score and a homologous recombination deficiency.
[0132] The spatial positions of the first and second target regions and the genomic features 122 are included in input data provided to the predictive model 124. In various cases, the predictive model 124 determines whether the first cells of the first target sample 110 and / or the second cells of the second targetFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT sample 118 are resistant to an immunotherapy targeting, for instance, PD-1 (e.g., pembrolizumab, nivolumab) based on the input data. For instance, the predictive model 124 may determine that the first cells of the first target sample 110 are resistant to the immunotherapy targeting PD-1, and that the second cells of the second target sample 118 are susceptible to the immunotherapy targeting PD-1. In some examples, the predictive model 124 determines a tumor evolution of the tumor 102. For instance, the predictive model 124 may determine an evolutionary tree (e.g., a phylogenetic tree) that is indicative of clonal evolution of the tumor 102. In various cases, the predictive model 124 predicts a progression of the tumor 102. For instance, the predictive model 124 may determine that the tumor 102 will not metastasize in the next 6 months because a heterogeneity of the tumor 102 is not associated with aggressive progression. In some cases, the predictive model 124 outputs the heterogeneity indicator(s) 125 and / or the condition indicator(s) 126 based on a more sophisticated analysis of various characteristics of the spatial positions of the first and second target regions and the genomic features 122.
[0133] Accordingly, the report generator 136 may generate the report 138 to indicate a recommendation against administering the immunotherapy to the subject 104 and / or to indicate a recommendation for administering an alternative treatment (e.g., surgery, a chemotherapy, or the like) to the subject 104. Upon reviewing the report 138 on the clinical device 140, the care provider(s) 106, in some cases, administers the alternative treatment. Thus, the subject 104 may be prevented from experiencing side effects of the immunotherapy without the immunotherapy treating the tumor 102.
[0134] FIG. 2 illustrates an example environment 200 for identifying target regions from one or more tissue samples. In various cases, a first tissue sample 202 and a second tissue sample 204 are obtained from a subject. In some examples, the first and second tissue samples 202 and 204 are the tissue sample 108 described above with reference to FIG.1.
[0135] The first and second tissue samples 202 and 204 may be collected from spatially distinct positions of the body of the subject. In various examples, the first and second tissue samples 202 and 204 are collected from a tumor of the subject. In some examples, a first region and a second region of the tumor are identified. The first and second regions may be identified based on a visual characteristic (e.g., a histopathological morphology and / or a vascularization). In various cases, the first and second regions are identified based on a predetermined interval. For example, the first and second regions may be located at a first predetermined distance and a second predetermined distance from a center or an edge of the tumor. In various instances, the first and second regions are spaced apart at the predetermined interval.
[0136] The first and second tissue samples 202 and 204 may be collected, for instance, by a needle biopsy sample from the first region and the second region. In various cases, the first and second tissue samples 202 and 204 are analyzed by a care provider (e.g., a pathologist) to identify target regions. For instance, the care provider may visually inspect the first and second tissue samples 202 and 204. The care provider mayFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT visually inspect fresh, fixed, or frozen tissue. In some examples, a portion of the first and second tissue samples 202 and 204 may be treated with a fixation agent and divided into slices. A slice of the first tissue sample 202 and a slice of the second tissue sample 204 may be treated with hematoxylin and eosin (H&E). The H&E-stained slices, or images of the H&E-stained slices generated by an imaging device, may be analyzed by the care provider. Nucleic acid molecules may be extracted from a remaining portion of the first and second tissue samples 202 and 204.
[0137] Target regions may be identified in the first and second tissue samples 202 and 204. For instance, a first target region 206 and a second target region 208 may be identified in the first tissue sample 202 based on comparing first visual characteristic 210 of the first target region 206 and a second visual characteristic 212 the second target region 208. In various cases, the first target region 206 and the second target region 208 are identified based on an interval 216. For instance, the first target region may be separated, by the interval 216, from a reference point 218. In various instances, the second target region 208 is separated, by the interval 216, from the first target region 206. In some examples, a third target region 220 is identified in the second tissue sample 204 based on comparing a third visual characteristic 222 of the third target region to the first and second visual characteristics 210 and 212.
[0138] In various implementations, target samples are collected based on identifying the target regions. For instance, a first target sample may be collected from the first target region 206, a second target sample may be collected from the second target region 208, and a third target sample may be collected from the third target region 220. In some examples, the first target sample and the second target sample are the first target sample 110 and the second target sample 118, respectively, as described with reference to FIG.1. In some examples, the first target sample and the third sample are the first target sample 110 and the second target sample 118, respectively, as described with reference to FIG. 1. The target samples are collected, in some examples, be performing needle punch enrichment, curl tissue extraction, or straight razor blade extraction on each of the target regions.
[0139] FIG. 3 illustrates an example environment 300 for training and utilizing a predictive model 302 to identify a heterogeneity condition of a tumor of a subject. The predictive model 302, for instance, is the predictive model 124 described above with reference to FIG.1. In various implementations, the predictive model 302 includes a classifier 304, which may include one or more ML models. A trainer 306, for instance, is configured to optimize various parameters 308 of the classifier 304 and / or the predictive model 302 based on training data 310.
[0140] The training data 310 includes example features 312, example heterogeneities 313, example conditions 314, and / or other data. The example features 312, in various cases, are obtained based on genomic information (e.g., DNA and / or RNA) of individuals within a population 316. In some examples, the example features 312 are obtained based on phenotypic information (e.g., analyte expression, cellularFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT morphology, etc). The example features 312 may include example genomic features, example analyte profiles, example region locations (e.g., spatial positions of target samples obtained from the individuals within the population 316), or a combination thereof. In various examples, the training data 310 is pre- classified data that includes the example features 312 and labels indicating example heterogeneities and conditions associated with the example features 312. The example heterogeneities 313 may include categorization of tumor heterogeneity. For instance, the example heterogeneities 313 may include indications of whether the individuals within the population 316 have one or more heterogeneity conditions associated with a tumor. The example conditions 314 may include categorizations of pathologies experienced by the individuals within the population 316. For example, the example heterogeneities 313 and / or the example conditions 314 may be generated based on clinical evaluations of the individuals within the population 316, such as by one or more care providers.
[0141] The predictive model 302 and / or the classifier 304 may include one or more model types. For instance, the predictive model 302 and / or the classifier 304 may include an artificial neural network. An artificial neural network includes various layers that respectively process input data. For example, an artificial neural network includes an input layer, one or more hidden layers, and an output layer. The input layer performs a pre-processing operation on the input data. The hidden layer(s) may perform various processing operations on the output from the input layer. The output layer, in various cases, processes the output from the hidden layer(s). Each layer, in some cases, includes one or more nodes, which are defined by individual operations. In various cases, the hidden layer(s) include nodes that are connected to each other in parallel and / or series. Examples of artificial neural networks include feedforward neural networks, multi-layer perceptrons (MLPs), convolutional neural networks (CNNs), and backpropagation models. In various implementations, the operations performed by the layers and / or nodes within an artificial neural network included in the classifier predictive model and / or the classifier 304 is defined according to the parameters 308. For example, the parameters 308 may include weights, thresholds, filters, kernels, or other data objects that are utilized to perform operations of the classifier 304.
[0142] In some implementations, the predictive model 302 and / or the classifier 304 include a nearest- neighbor model. One example of a nearest-neighbor model includes a k-nearest neighbor model. For example, a nearest-neighbor model defines various “neighbors,” which are points within a feature space, with associated class labels. When a new data point is mapped to the feature space, the new data point is classified based on the proximity (e.g., Euclidian distance, Manhattan distance, Minkowski distance, etc.) of its “neighbors” to the new data point as well as their associated classes. In some cases, the new data point is classified as belonging to a particular class if greater than a threshold number of neighbors within a threshold distance of the new data point are members of the class. For instance, the parameters 308 may include k (e.g., the number of neighbors compared to the new data point), the threshold distance, and so on.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT
[0143] In various cases, the predictive model 302 and / or the classifier 304 include a regression analysis model. The regression analysis model, for example, is defined by a regression function that defines relationships between one or more independent variables and one or more dependent variables. The regression function may further define one or more unknown parameters that define a relationship between the independent and dependent variables. In various implementations, the unknown parameters and / or the type of regression function (e.g., linear, quadratic, etc.), is defined according to the parameters 308.
[0144] In some cases, the predictive model 302 and / or the classifier 304 include a clustering model. In various cases, a clustering model maps various data points (e.g., training data) to a feature space. Based on the proximity of groups of those data points in the features pace, one or more “clusters” are defined. An additional data point may be classified according to one or more of the clusters based on its proximity to the clusters (e.g., a center of the clusters, a boundary of the cluster, etc.). Examples of clustering models include k-means clustering, mean-shift clustering, expectation-maximization (EM) clustering, and agglomerative hierarchical clustering. The parameter(s) 308, for example, include a threshold proximity within which a new data point is classified within a cluster, a density of points used to define a cluster, and the like.
[0145] In various examples, the predictive model 302 and / or the classifier 304 include a principal component analysis model. In various implementations, a principal component analysis defines a collection principal components of unit vectors within a coordinate space based on a data set (e.g., training data). The model, for example, is an orthogonal linear transformation of the data set. Various weights of the model, for example, are included in the parameter(s) 308.
[0146] In some examples, the predictive model 302 and / or the classifier 304 may include a gradient boosting model. For example, the gradient boosting model is defined as a collection of prediction models (e.g., decision trees) that iteratively classify observed data. In various cases, the type of prediction model, weights in the prediction models, and the like, are defined by the parameter(s) 308.
[0147] In some examples, the predictive model 302 and / or the classifier 304 may include a random forest model. A random forest model, for instance, may include multiple decision trees that classify data in an ensemble fashion. In various implementations, the decision trees are defined by the parameter(s) 308.
[0148] In various implementations of the present disclosure, the trainer 306 is configured to optimize the parameters 308 based on the training data 310. For example, the trainer 306 may input first example features (corresponding to a first individual among the population 316) among the example features 312 into the predictive model 302, and may receive a predicted category. The trainer 306 may compute a loss (e.g., determine a discrepancy) between a first example category (corresponding to the first individual) among the example heterogeneities 313 and / or the example conditions 314 and the predicted category.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT Further, the trainer 306 may alter the parameters 308 in order to minimize the loss. In various cases, the trainer 306 optimizes the parameters 308 iteratively based on the entire set of the training data 310.
[0149] In various implementations, the optimization of the parameters 308 enables the predictive model 302 to identify predictive attributes of the example features 312 that are correlated to or otherwise associated with the example heterogeneities 313 and / or the example conditions 314. For instance, the predictive model 302 may determine that a particular sequence represented in the example features 312 is highly correlated with susceptibility of tumor cells to a treatment (e.g., an immunotherapy). The predictive model 302 may therefore classify heterogeneity conditions associated with a tumor based on features outside of the example features 312 by recognizing or otherwise identifying the predictive attributes.
[0150] Once the parameters 308 are optimized and / or training of the predictive model 302 is complete, the predictive model 302 may be ready to classify a new set of data. For example, the predictive model 302 may receive input data including features 318 of a subject. The features 318, for instance, may include one or more of the predictive attributes. The predictive model 302 may perform various operations on the input data based on the trained classifier 304 and the optimized parameters 308. In various cases, the predictive model 302 outputs data including one or more heterogeneity indicators 319 and / or one or more condition indicators 320 based on the features 318. The heterogeneity indicator(s) 319, for instance, include one or more predicted heterogeneity conditions associated with a tumor of the subject. The condition indicator(s) 320, for instance, may include one or more predicted pathological conditions associated with the subject.
[0151] Although FIG.3 is primarily described as referring to supervised learning, implementations are not so limited. In various cases, the training data 310 omits the example heterogeneities 313 and the example conditions 314 and the trainer 306 is configured to optimize the parameters 308 using the example features 312 and an unsupervised learning technique.
[0152] FIG. 4 illustrates an example of training data 400 utilized to train one or more ML models. For example, the training data 400 may be the training data 310 described above with reference to FIG.3.
[0153] The training data 400, in various cases, may represent m samples, wherein m is a positive integer. In some cases, the m samples are respectively obtained from m individuals within a population, although implementations are not so limited. For example, in some cases, multiple samples may be obtained from the same individual at different times.
[0154] The training data 400 includes first to mth example features 402-1 to 402-m. For example, the first to mth example features 402-1 to 402-m include features derived from genomic DNA of a tumor in the respective m samples.
[0155] The training data 400 may further include first to mth example categories 404-1 to 404-m. The first to mth example categories 404-1 to 404-m, for instance, include one or more heterogeneity conditions associated with tumors (e.g., a tumor evolution, a tumor progression, a susceptibility of the tumor cells toFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT a treatment, a resistance of the tumor cells to a treatment) represented by the m samples. In some examples, the first to mth example categories 404-1 to 404-m, for instance, include one or more conditions of subjects associated with the m samples (e.g., a type or a subtype of a pathological condition, a metastasis profile, a survivability, a symptom, a risk of developing, a stage, a grade, an ECOG performance of a cancer, a general health, a genomic age, or a risk of developing a pathological condition).
[0156] FIG.5 illustrates an example report 500 summarizing predicted categories of a tumor of a subject. In various cases, the report 500 is the report 138 described above with reference to FIG.1. The report 500, for instance, may be displayed to a patient and / or care provider. In some cases, the report 500 is generated based on features of a sample (e.g., a tissue biopsy sample) obtained from the subject.
[0157] The report 500 includes a tissue origin 502 of the tumor. The tissue origin 502, for instance, indicates a histological tissue type 504, a primary site designation 506 (e.g., whether the tumor is a primary tumor), cell subtype 507, or any combination, of the cancer.
[0158] In various cases, the report 500 includes one or more therapy indicators 508. For instance, the therapy indicator(s) 508 convey whether the tumor is predicted to be resistant to one or more predetermined therapies and / or whether the tumor is predicted to be responsive to one or more predetermined therapies.
[0159] In some examples, the report 500 includes one or more prognostic indicators 510. The prognostic indicator(s) 510, for instance, indicate a prognosis of the subject in view of the categorized tumor. For example, the prognostic indicator(s) 510 may indicate a survivability, a recoverability, a quality-of-life indicator, or other information indicative of the prognosis of the subject. The prognostic indicator(s) 510 may indicate a cancer type, a cancer subtype, a cancer stage, a cancer grade, an Eastern Cooperative Oncology Group (ECOG) performance status associated with a cancer, or other information indicative of the cancer prognosis of the subject.
[0160] The report 500 may include a trial qualification 512 of the subject. The trial qualification 512, for instance, indicates whether the subject is predicted to qualify for a predetermined clinical trial (e.g., whether the subject is eligible for the predetermined clinical trial).
[0161] The report 500 may include one or more heterogeneity indicators 513. The heterogeneity indicator(s) 513, for instance, indicate a heterogeneity of the tumor. For instance, the heterogeneity indictor(s) 513 may indicate a tumor evolution, a tumor progression. In some examples, the heterogeneity indicator(s) 513 indicate an alteration-level clonality (e.g., a number of clones within the tumor, a proportionality of each of the clones, or a genetic variation between each of the clones). In various cases, the heterogeneity indicator(s) 513 indicate an intra-tumor heterogeneity (e.g., a number of cell populations within the tumor, a proportion of each of the cell populations, or a variation between the cell populations).FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT
[0162] The report 500, in various implementations, includes a metastasis profile 514 of the subject. The metastasis profile 514, for instance, indicates a likelihood that the tumor will metastasize (e.g., at a particular point in time), one or more tissues in which the cancer is predicted to metastasize, or the like.
[0163] In various cases, the report 500 includes recommended follow-up tests 516. For example, the report 500 may include a recommendation to perform whole genome sequencing on the subject, particularly in cases if the cancer cannot be categorized above a threshold certainty.
[0164] The report 500 may include a genomic profile 518 of the subject. In various cases, the genomic profile 518 includes or is generated based on the results of DNA analysis, RNA analysis, exome analysis, or a combination thereof.
[0165] FIG.6 illustrates an example environment 600 for sequencing various nucleic acid molecules 602. In various implementations, the nucleic acid molecules 602 include gDNA. The nucleic acid molecules 602, in various cases, are extracted from a sample, such as a biological sample obtained from a subject. In some implementations, the nucleic acid molecules 602 include DNA that is complementary to RNA present in the sample.
[0166] The nucleic acid molecules 602, in various cases, are ligated with adapters 604. For examples, the adapters 604 are hybridized to the nucleic acid molecules 602. The adapters 604, for example, include additional nucleic acid molecules. In various implementations, the adapters 604 have a shorter length than the nucleic acid molecules 602 being sequenced. For instance, the adapters 604 include amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences. Although FIG. 6 illustrates adapters 604 being ligated to one end of each of the nucleic acid molecules 602, implementations are not so limited. For example, the adapters 604 may be ligated to both ends of each of the nucleic acid molecules 602.
[0167] In various examples, the nucleic acid molecules 602 ligated with the adapters 604 are amplified in order to generate amplified molecules 606. Various amplification techniques can be performed. For instance, the amplified molecules 606 are generated using PCR, a non-PCR amplification technique, an isothermal amplification technique, or any combination thereof.
[0168] Amplified molecules 606 may be captured by bait molecules 610 and sequenced. In some implementations, the amplified molecules 606 are sequenced via sequencing-by-synthesis. In various cases, fluorescently tagged deoxyribonucleotide triphosphates (dNTP) 612 are utilized to synthesize a strand that is complementary to DNA strands bound to the substrate 608. When a dNTP 612 is added to the strand (e.g., by an enzyme), the dNTP 612 emits an optical signal 614. In various implementations, the frequency of the optical signal 614 is dependent on the type of dNTP 612 from which the optical signal 614 is emitted. By detecting the optical signals 614 as the strand is being synthesized, the sequence of the original nucleic acid molecules 602 can be derived.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT
[0169] In some implementations, the amplified molecules 606 are sequenced via nanopore sequencing. For instance, the amplified molecules 606 are directed through a nanopore 616 extending through a substrate 618. In various cases, the amplified molecules 606 are negatively charged, such that they can be directed through the nanopore 616 by imposing an electrical field across the substrate 618. In various cases, the amplified molecules 606 and the nanopore 616 are in the presence of a charged solution. Thus, charged solutes traveling through the nanopore 616 can be monitored by reviewing an electrical signal (e.g., a current) sensed between electrodes 620 on either side of the substrate 618. As an amplified molecule 606 is directed through the nanopore 616, the individual bases within the amplified molecule 606 will block the nanopore 616, which may decrease the amount of charged solutes traveling through the nanopore 616 and consequently, the magnitude of the electrical signal detected by the electrodes 620. Each of the four types of bases within the amplified molecules 606, may block the nanopore 616 to a different extent. Therefore, the sequence of the nucleic acid molecules 602 can be derived by analyzing the measured electrical signal with respect to time as the amplified molecules 606 are directed through the nanopore 616.
[0170] FIG.7 illustrates an example process 700 for identifying a heterogeneity condition of a tumor of a subject. In various implementations, the process 700 is performed by an entity including at least one processor, at least one computing device, a medical device, a device configured to collect a tissue sample, the sequencer 112, the genomic analyzer 120, the predictive model 124, the image analyzer 132, the report generator 136, the clinical device 140, the predictive model 302, or any combination thereof.
[0171] At 702, the entity identifies first sequence read data associated with a first region of the tumor of the subject. For instance, the entity receives a plurality of nucleic acid molecules in a sample from the subject. The sample may include a tissue sample. The nucleic acid molecules, for instance, include genomic DNA from a tissue biopsy sample. One or more adapters are ligated onto at least some of the nucleic acid molecules. The ligated molecules are amplified and captured. In various cases, all or a subset of the captured molecules are sequenced to obtain a plurality of sequence reads that represent the sequenced amplified nucleic acid molecules, thereby generating the sequence read data. In particular examples, the sequence read data includes endpoint counts of fragments at multiple genomic positions within at least one locus of the genome of the sample.
[0172] At 704, the entity identifies second sequence read data associated with a second region of the tumor of the subject. In various examples, the first region and the second region may be located in the tissue biopsy sample. In some cases, the first region is located in a first tissue biopsy sample, and the second region is located in a second tissue biopsy sample. The first region and the second region may be determined, in some examples, based on a predetermined interval. For instance, the first region and the second region may be spaced apart at the predetermined interval along an axis. In various cases, the first region and the second region are spaced apart at the predetermined interval from a reference point. In someFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT instances, the reference point corresponds to a center of the tumor. In various implementations, the first region and the second region are determined based on a visual characteristic. For instance, the entity may compare a visual characteristic of the first region to a visual characteristic of the second region. The entity may determine that the first region includes cells from a different clone than the second region. In some examples, the entity may collect a first sample from the first region and a first sample from the second region. For instance, the entity may collect needle punch enrichment samples from the first region and the second region. The first sequence read data is, in various cases, derived from the first sample, and the second sequence read data is derived from the second sample.
[0173] At 706, the entity identifies genomic features based on the first and second sequence read data. The genomic features may be based on one or more sequences of genomic DNA indicated by the first and second sequence read data. In some examples, the first and second sequence read data may be combined into a combined sequence read data. The entity may identify genomic features based on the combined sequence read data. In various examples, the first and second sequence read data indicate spatial positions of the first region and the second region.
[0174] At 708, the entity classifies a condition of the tumor based on the genomic features. In some cases, the entity utilized an ML-based classifier to predict whether the tumor is associated with the condition. The ML-based classifier, for instance, is pre-trained based on data obtained from a population of individuals that omits the subject. In some cases, the classifier includes at least one of an ANN, a logistic regression model, a decision tree, a KNN model, a support vector machine (SVM), or a naïve Bayes classifier. In some cases, the classifier outputs a likelihood that the tumor is associated with a particular condition (or is not associated with a particular condition). In some cases, the classifier outputs an indication that the tumor is associated with the particular condition (or is not associated with the particular condition) when the likelihood exceeds a threshold likelihood.
[0175] Various types of conditions can be predicted using the process 700. In some cases, the entity predicts a condition associated with the tumor, such as a tumor evolution, a tumor progression, an effectiveness of a therapy for treatment of the tumor (e.g., a susceptibility of the tumor to a treatment, a resistance of the tumor to the treatment). In some examples, the entity predicts, based on the condition associated with the tumor, a subtype, metastasis profile, survivability, symptom, risk of developing, stage, grade, or ECOG performance of a cancer of the subject. In some cases, the entity predicts whether a particular therapy will be effective to treat the pathological condition and / or whether the pathological condition is resistant to the particular therapy. In some cases, the entity predicts a general health, genomic age, or risk of developing a pathological condition, of the subject. According to some cases, the entity generates a comprehensive metric or “score” indicating a level of health or disease of the subject.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT
[0176] FIG. 8 illustrates one or more devices 800 configured to perform various operations described herein. The device(s) 800 include one or more processor(s) 802. In some implementations, the processor(s) 802 includes a central processing unit (CPU), a graphics processing unit (GPU), both CPU and GPU, or other processing unit or component known in the art.
[0177] The processor(s) 802 is operably connected to memory 804. In various implementations, the memory 804 is volatile (such as random access memory (RAM)), non-volatile (such as read only memory (ROM), flash memory, etc.) or some combination of the two. The memory 804 stores instructions that, when executed by the processor(s) 802, causes the processor(s) 802 to perform various operations. In various examples, the memory 804 stores methods, threads, processes, applications, objects, modules, any other sort of executable instruction, or a combination thereof. In some cases, the memory 804 stores files, databases, or a combination thereof. In some examples, the memory 804 includes, but is not limited to, RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory, or any other memory technology. In some examples, the memory 804 includes one or more of CD-ROMs, digital versatile discs (DVDs), content-addressable memory (CAM), or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by the processor(s) 802. For instance, the memory 804 stores instructions that, when executed by the processor(s) 802, causes the processor(s) 802 to perform operations of the genomic analyzer 120, the predictive model 124, the image analyzer 132, and the report generator 136. For instance, the memory 804 may store instructions that cause the processor(s) 802 to determine, based on analyzing first and second sequence read data, genomic features associated with the first and second sequence read data.
[0178] The processor(s) 802 is operably connected to one or more input devices 806 and one or more output devices 808. Collectively, the input device(s) 806 and the output device(s) 808 function as an interface between at least one user and the device(s) 800. The input device(s) 806 is configured to receive an input from a user and includes at least one of a keypad, a cursor control, a touch-sensitive display, a voice input device (e.g., a microphone), a haptic feedback device (e.g., a gyroscope), or any combination thereof. The output device(s) 808 includes at least one of a display, a speaker, a haptic output device, a printer, or any combination thereof. In various examples, the processor(s) 802 causes a display among the input device(s) 806 to visually output various data described herein. In some implementations, the input device(s) 806 includes one or more touch sensors, the output device(s) 808 includes a display screen, and the touch sensor(s) are integrated with the display screen.
[0179] In various implementations, the processor(s) 802 is operably connected to one or more transceivers 810 that transmit and / or receive data over one or more communication networks 812. For example, the transceiver(s) 810 includes a network interface card (NIC), a network adapter, a local area network (LAN)FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT adapter, or a physical, virtual, or logical address to connect to the various external devices and / or systems. In various examples, the transceiver(s) 810 includes any sort of wireless transceivers capable of engaging in wireless communication (e.g., radio frequency (RF) communication). For example, the communication network(s) 812 includes one or more wireless networks that include a 3rd Generation Partnership Project (3GPP) network, such as a Long Term Evolution (LTE) radio access network (RAN) (e.g., over one or more LTE bands), a New Radio (NR) RAN (e.g., over one or more NR bands), or a combination thereof. In some cases, the transceiver(s) 810 includes other wireless modems, such as a modem for engaging in WI-FI®, WIGIG®, WIMAX®, BLUETOOTH®, or infrared communication over the communication network(s) 812.
[0180] The device(s) 800 may further include the sequencer 112. In various implementations, the sequencer 112 includes one or more fluidic circuits 814 configured to receive a sample 816 derived from a subject 818. The sequencer 112, in various cases, may be configured to generate data indicative of one or more sequences of nucleic acid molecules (e.g., DNA and / or RNA) present in the sample 816. In various cases, the sequencer 112 introduces one or more reagents 819 to the fluidic circuit(s) 814 in order to prepare for and perform sequencing of the nucleic acid molecules. Further, the sequencer 112 may include one or more sensors 820 configured to measure or otherwise detect detection signals from the fluidic circuit(s) 814, which may be indicative of the sequences of the nucleic acid molecules. According to various implementations, the sensor(s) 820 may further include one or more ADCs. The sequencer 112, in various cases, outputs sequence read data to the processor(s) 802 for additional processing.
[0181] FIG. 9 illustrates an example post multi-region tissue extracted H&E slide, showing the regions where the tissue has been extracted using the needle punch enrichment process and the spatial positions of the regions.
[0182] FIGs.10A-10D illustrate an example overview of functional genomic alterations (known / likely) as observed in the multi-region sequencing pilot data. The X-axis represents the 83 individual specimens while the Y-axis represents the functional genomic alterations identified in these 83 specimens. The plot also includes annotations of complex signatures such as tumor mutational burden (TMB), microsatellite instability (MSI), homologous repair deficiency (HRD); demographics features such as age; clinicopathological features such as computational tumor purity, specimen extraction method (curl / specimen), tumor type, and block ID.
[0183] FIG. 11 illustrates an example distribution of cancer cell fraction (CCF) of short variants (known / likely oncogenic, variants of unknown significance, noncoding short variants), grouped by the percentage occurrence of the short variant amongst multi-specimens per tumor block. A = 1 / 4 or 1 / 3, B = 2 / 4, C = 2 / 3 or 3 / 4, D = 3 / 3 or 4 / 4.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT
[0184] FIG. 12 illustrates an example clonality of short variants (known / likely oncogenic, variants of unknown significance, noncoding short variants) evaluated using the multi-region sequencing cohort as well as the cancer cell fraction of individual short variants.
[0185] FIG. 13 illustrates an example wet-lab merge of samples. The multi-region sequencing workflow illustrated generates one aliquot per specimen and each of these aliquots are sequenced individually, to generate genomics data per aliquot. Each of these spatially derived aliquots can be combined into a single aliquot (e.g., single spatially representative aliquot) and sequenced, to generate tumor-block-representative genomics data that may be used by downstream computational algorithms to deconvolute tumor heterogeneity and subclonal reconstruction on a tumor-block-level specimen.
[0186] FIG. 14 illustrates an example dry-lab merge of samples. The multi-region sequencing workflow illustrated generates one aliquot per specimen and each of these aliquots are sequenced individually, to generate genomics data per aliquot. Each of these spatially derived aliquots can be sequenced individually and the sequenced reads from each of these aliquots can be merged computationally, to create a single spatially representative cohort of sequenced reads, and to generate tumor-block-representative genomics data that may be used by downstream computational algorithms to deconvolute tumor heterogeneity and subclonal reconstruction on a tumor-block -level specimen.
[0187] Example Clauses 1. A method, including: receiving a tissue sample derived from a tumor of a subject; identifying, in the tissue sample, a plurality of target regions, the target regions being spatially distinct; collecting, from the target regions, a plurality of target samples; providing a plurality of nucleic acid molecules obtained from the target samples; ligating one or more adapters onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying the one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules; capturing amplified nucleic acid molecules from the amplified nucleic acid molecules; sequencing, by a sequencer, all or a subset of the captured amplified nucleic acid molecules to obtain a plurality of sequence reads that represent the sequenced amplified nucleic acid molecules thereby generating sequence read data; receiving, at one or more processors, the sequence read data for the plurality of sequence reads; and determining, using the one or more processors using the sequence read data, a heterogeneity condition of the tumor of the subject. 2. The method of clause 1, wherein the tissue sample includes fresh tissue, frozen tissue, or fixed tissue. 3. The method of clause 1 or 2, wherein identifying the plurality of target regions includes: identifying a heterogeneous visual characteristic of the plurality of target regions in the tissue sample, theFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT heterogeneous visual characteristic including a heterogeneous histopathological morphology of the target regions or a heterogenous vascularization of the target regions. 4. The method of any of clauses 1-3, wherein collecting the plurality of target samples includes performing needle punch enrichment on the target regions. 5. The method of any of clauses 1-4, wherein the plurality of target regions are spaced apart from each other at a predetermined interval along an axis of the tissue sample. 6. The method of any of clauses 1-5, wherein the plurality of target regions includes a first target region and a second target region, the plurality of target samples includes a first target sample obtained from the first target region and a second target sample obtained from the second target region, and wherein the sequence read data includes: first sequence read data indicating first sequence reads associated with the first target sample and a first label indicating the first target region; and second sequence read data indicating second sequence reads associated with the second target sample and a second label indicating the second target region. 7. The method of any of clauses 1-6, further including: determining, based on the sequence read data, an additional condition of the subject including a mutational profile, a fraction unstable score, a homologous recombination deficiency, an alteration-level clonality, or an intra-tumor heterogeneity, wherein determining the heterogeneity condition of the tumor of the subject is further based on the additional condition of the subject. 8. The method of any of clauses 1-7, wherein the heterogeneity condition indicates an evolution of the tumor of the subject, a predicted progression of the tumor of the subject, or a predicted effective therapy for treatment of the tumor of the subject. 9. A method, including: receiving target samples derived from distinct spatial positions in a tumor of a subject; identifying sequence read data of the target samples; and determining a heterogeneity condition of the tumor of the subject based on the sequence read data. 10. The method of clause 9, wherein one or more of the target samples are obtained from a tissue sample obtained from the tumor of the subject. 11. The method of clause 10, wherein the tissue sample includes fresh tissue, frozen tissue, or fixed tissue derived from the tumor of the subject. 12. The method of any of clauses 9-11, wherein one or more of the target samples include cells obtained from the tumor of the subject. 13. The method of any of clauses 9-12, wherein the target samples are respectively derived from spatially distinct target regions in a single sample obtained from the subject. 14. The method of clause 9, wherein the target samples are respectively derived from spatially distinct target regions in one or more samples obtained from the subject.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT 15. The method of clause 14, further including: storing an indication of locations of the target regions in the one or more samples, wherein determining the heterogeneity condition of the tumor of the subject is further based on the locations of the target regions. 16. The method of clause 14 or 15, further including: identifying the spatially distinct target regions in the one or more samples. 17. The method of clause 16, wherein identifying the spatially distinct target regions includes: determining that the spatially distinct target regions in the one or more samples have different visual characteristics. 18. The method of clause 17, wherein the different visual characteristics include different histopathological morphologies or different vascularization. 19. The method of any of clauses 16-18, wherein identifying the target regions includes: identifying an image of the one or more samples; determining, based on the image, metrics associated with different locations in the one or more samples; and selecting, among the different locations in the one or more samples, the spatially distinct target regions by comparing the metrics. 20. The method of clause 19, wherein the image includes a three-dimensional (3D) image. 21. The method of clause 19 or 20, wherein the image depicts the one or more samples stained with a histological stain or an immunohistological stain. 22. The method of any of clauses 19-21, wherein the metrics include at least one of a cell size, a cell morphology, a cell type, a subcellular structure, or a biomarker level. 23. The method of any of clauses 16-22, further including: collecting the one or more samples from the subject. 24. The method of any of clauses 14-23, wherein the target samples include needle punch enrichment samples obtained from the one or more samples. 25. The method of any of clauses 14-24, wherein the target samples are obtained by performing curl tissue extraction or straight razor blade extraction on the one or more samples. 26. The method of any of clauses 14-25, wherein the distinct spatial positions include locations that are spaced apart at a predetermined interval along an axis. 27. The method of clause 26, wherein the predetermined interval is a first predetermined interval, the axis is a first axis, and the locations are first locations, and wherein the distinct spatial positions also include second locations spaced apart at a second predetermined interval along a second axis. 28. The method of any of clauses 14-27, further including extracting DNA or RNA from one or more of the target samples. 29. The method of clause 28, wherein the DNA includes genomic DNA or complementary DNA (cDNA).FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT 30. The method of clause 28 or 29, wherein the RNA includes messenger RNA, microRNA, or non- coding RNA. 31. The method of any of clauses 14-30, further including isolating biological analytes from one or more of the target samples. 32. The method of clause 31, further including: generating an analyte profile indicative of the biological analytes in the one or more target samples. 33. The method of any of clauses 14-32, wherein the sequence read data indicates nucleic acid molecules in a combined sample that includes the target samples. 34. The method of any of clauses 9-33, wherein the sequence read data indicates a quantity and / or presence of variants present in fragments in the target samples. 35. The method of clause 34, wherein the variants include at least one of a substitution, an insertion, a deletion, a copy number alternation, or a rearrangement of fusion. 36. The method of any of clauses 9-35, wherein the sequence read data further indicates at least one of the distinct spatial positions associated with at least one of the target samples corresponding to one or more of the sequence reads. 37. The method of any of clauses 9-36, further including: receiving a plurality of nucleic acid molecules obtained from the target samples; ligating one or more adapters onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying the one or more ligated nucleic acid molecules; capturing all or a subset of the amplified nucleic acid molecules; and sequencing, by a sequencer, the captured nucleic acid molecules to obtain a plurality of sequence reads that represent the captured nucleic acid molecules, thereby generating the sequence read data for a genome of the target samples. 38. The method of clause 37, wherein the one or more adapters include amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences. 39. The method of clause 37 or 38, wherein the captured nucleic acid molecules are captured from the amplified nucleic acid molecules by hybridization to one or more bait molecules. 40. The method of clause 39, wherein the one or more bait molecules include one or more additional nucleic acid molecules, each of the one or more additional nucleic acid molecules including a region that is complementary to a region of a captured nucleic acid molecule. 41. The method of any of clauses 37-40, wherein amplifying the one or more ligated nucleic acid molecules includes performing a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT 42. The method of any of clauses 37-41, wherein sequencing the captured nucleic acid molecules includes use of a massively parallel sequencing (MPS) technique, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing. 43. The method of any of clauses 37-42, wherein sequencing the captured nucleic acid molecules includes next-generation sequencing (NGS). 44. The method of any of clauses 37-43, wherein the sequencer includes a next-generation sequencer. 45. The method of any of clauses 37-44, wherein sequencing the captured nucleic acid molecules includes sequencing-by-synthesis or nanopore sequencing. 46. The method of any of clauses 9-45, further including: generating ligated molecules by ligating adapters onto nucleic acid molecules of the target samples; generating amplified ligated molecules by amplifying the ligated molecules; generating, using the amplified ligated molecules, detection signals; detecting, by at least one sensor, the detection signals; and generating the sequence read data based on the detection signals. 47. The method of clause 46, wherein the detection signals include electrical signals and / or optical signals. 48. The method of clause 46 or 47, wherein generating, using the amplified ligated molecules, the detection signals includes: synthesizing, by a polymerase using fluorescently tagged nucleotide triphosphates (NTPs), a synthesized nucleic acid molecule that is complementary to one of the amplified ligated molecules, and wherein detecting, by the at least one sensor, the detection signals includes: detecting, by at least one optical sensor, optical signals emitted by the fluorescently tagged NTPs upon binding to the synthesized nucleic acid molecule, the optical signals being indicative of at least one sequence of the nucleic acid molecules of the target samples. 49. The method of any of clauses 46-48, wherein generating, using the amplified ligated molecules, the detection signals includes: directing the amplified ligated molecules through a nanopore extending from a first space to a second space through a substrate, and wherein detecting, by the at least one sensor, the detection signals includes: detecting, by sensors disposed in the first space and the second space, an electrical signal over time, the electrical signal being indicative of at least one sequence of the nucleic acid molecules of the target samples. 50. The method of any of clauses 46-49, wherein the sequence read data indicates a full genome or RNA transcriptome of the target samples. 51. The method of any of clauses 46-50, wherein the sequence read data indicates a whole exome of the target samples. 52. The method of clause 46, wherein the sequence read data indicates a predetermined panel of genes of the target samples.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT 53. The method of clause 52, wherein the predetermined panel includes one or more of ABL1, ACVR1B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APC, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1, AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BCOR, BCORL1, BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1, BTG1, BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1, CCND2, CCND3, CCNE1, CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1, CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1R, CSF3R, CTCF, CTNNA1, CTNNB1, CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1L, EED, EGFR, EMSY (C11orf30), EP300, EPHA3, EPHB1, EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1, ESR1, ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1, FLT3, FOXL2, FUBP1, GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1, HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1, IDH2, IGF1R, IKBKE, IKZF1, INPP4B, IRF2, IRF4, IRS2, JAK1, JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1, MERTK, MET, MITF, MKNK1, MLH1, MPL, MRE11A, MSH2, MSH3, MSH6, MST1R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1, NOTCH1, NOTCH2, NOTCH3, NPM1, NRAS, NT5C2, NTRK1, NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1, PARP2, PARP3, PAX5, PBRM1, PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1, PRKAR1A, PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51, RAD51B, RAD51C, RAD51D, RAD52, RAD54L, RAF1, RARA, RB1, RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1, SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1, SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11, SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1, VEGFA, VHL, WHSC1, WHSC1L1, WT1, XPO1, XRCC2, ZNF217, ZNF703, ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-1β, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, MSI-H, mTOR, PARP, PD-1, PDGFR, PDGFRα, PDGFRβ, PD-L1, PI3Kδ, PIGF, PTCH, RAF, RANKL, RET, ROS1, SLAMF7, VEGF, VEGFA, or VEGFB.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT 54. The method of any of clauses 9-53, further including: determining, based on the sequence read data, a mutational profile, one or more pathogenic variants, a fraction unstable score, a homologous recombination deficiency, an alteration-level clonality, or an intra-tumor heterogeneity of the target samples. 55. The method of clause 54, wherein determining the fraction unstable score further includes determining a microsatellite instability (MSI) fraction. 56. The method of clause 54 or 55, wherein the one or more pathogenic variants include variants of at least one of POLE, TP53, CTNNNB1, L1CAM, PTEN, ERBB2, PMS2, MSH2, MSH6, MLH1, an estrogen receptor (ER) gene, or a progesterone receptor (PR) gene. 57. The method of any of clauses 54-56, wherein the mutational profile, one or more pathogenic variants, a fraction unstable score, a homologous recombination deficiency, an alteration-level clonality, or an intra-tumor heterogeneity is associated with the target samples. 58. The method of any of clauses 54-57, further including: generating input features based on the sequence read data, the input features being indicative of the mutational profile, the one or more pathogenic variants, the fraction unstable score, the homologous recombination deficiency, the alteration- level clonality, the intra-tumor heterogeneity, or the heterogeneity condition of the subject. 59. The method of clause 58, wherein generating input features based on the sequence read data includes: inputting, into a machine learning (ML) model configured to detect the input features, the sequence read data. 60. The method of clause 59, wherein the ML model includes a neural network. 61. The method of clause 60, wherein the neural network includes multiple layers, an individual layer among the multiple layers including a transformation defined by one or more parameters, and wherein extracting the input features from the sequence read data includes generating an output by applying the transformation to the input, the input being based on the sequence read data. 62. The method of any of clauses 59-61, wherein the input features are determined based at least in part on pre-classified data, the pre-classified data being generated by: identifying training sequence read data associated with samples corresponding to a plurality of individuals omitting the subject; and generating the pre-classified data by labeling the training data with labels indicative of heterogeneity conditions of the plurality of subjects. 63. The method of clause 62, further including: training the ML model to identify attributes, indicated by the training data, that are predictive of the heterogeneity conditions of the plurality of subjects, wherein the input features are instances of the attributes identified via the training of the ML model. 64. The method of clause 63, wherein determining the heterogeneity condition of the subject based on the sequence read data includes:FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT classifying, using a classifier, the heterogeneity condition of the subject based on the input features. 65. The method of clause 64, wherein the classifier includes an ML-based classifier. 66. The method of clause 65, wherein training the ML-based classifier using training data includes performing supervised learning on the ML-based classifier. 67. The method of clause 65 or 66, wherein training the ML-based classifier using the training data includes performing unsupervised learning on the ML-based classifier. 68. The method of any of clauses 65-67, wherein training the ML-based classifier using the training data includes optimizing parameters of the ML-based classifier using the training data. 69. The method of any of clauses 65-68, wherein the ML-based classifier includes at least one of: an artificial neural network (ANN); a logistic regression model; a random forest model; a decision tree; a k- nearest neighbor (KNN) model; a support vector machine (SVM); or a naïve Bayes classifier. 70. The method of any of clauses 9-69, wherein determining the heterogeneity condition of the subject based on the sequence read data includes: predicting whether the subject has a pathological condition. 71. The method of any of clauses 9-70, wherein the heterogeneity condition includes at least one of: a tumor evolution of the subject; a predicted tumor progression of the subject; a predicted pathologic condition of the subject; a predicted pathologic condition subtype of the subject; a metastasis profile of the subject; a predicted survivability of the subject; a predicted symptom of the subject; a predicted effective therapy to treat the predicted pathologic condition of the subject; a predicted resistance of the subject to a treatment of the predicted pathologic condition; a general health of the subject; a genomic age of the subject; a risk of the subject developing the predicted pathologic condition; a predicted stage of the predicted pathologic condition of the subject; a predicted grade of the predicted pathologic condition of the subject; or a predicted Eastern Cooperative Oncology Group (ECOG) performance status of the subject. 72. The method of any of clauses 9-71, wherein the heterogeneity condition includes a health metric and / or a disease metric of the subject. 73. The method of clause 72, wherein the health metric and / or the disease metric include a quantitative metric. 74. The method of any of clauses 9-73, further including: generating, based on the heterogeneity condition, a genomic profile of the subject. 75. The method of clause 74, wherein the genomic profile includes results from at least one of: a comprehensive genomic profiling test; a whole genome sequencing (WGS) test; a whole exome sequencing (WES) test; a gene expression profiling test; a cancer hotspot panel test; a DNA methylation test; a DNA fragmentation test; an RNA fragmentation test; or an RNA sequencing test.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT 76. The method of clause 74 or 75, wherein the genomic profile of the subject includes: results from a nucleic acid sequencing-based test. 77. The method of any of clauses 74-76, further including: selecting, based on the genomic profile and / or the heterogeneity condition, an anticancer agent for administration to the subject. 78. The method of clause 77, further including: administering the anticancer agent to the subject. 79. The method of any of clauses 74-78, further including: applying, based on the genomic profile, an anticancer therapy to the subject. 80. The method of clause 79, wherein the anticancer therapy includes at least one of chemotherapy, radiation therapy, immunotherapy, a targeted therapy, or surgery. 81. The method of any of clauses 9-80, further including: identifying, based on the heterogeneity condition, a suggested treatment decision for the subject. 82. The method of clause 81, wherein the suggested treatment decision includes radiotherapy and / or chemotherapy. 83. The method of any of clauses 9-82, further including: generating a report indicating the heterogeneity condition; and outputting the report. 84. The method of clause 83, wherein outputting the report includes: transmitting data indicating the report to an external device. 85. The method of clause 84, wherein the external device is associated with the subject and / or a healthcare provider. 86. The method of clause 84 or 85, wherein the data is transmitted over one or more communication networks. 87. The method of any of clauses 84-86, wherein the data is transmitted over a peer-to-peer connection. 88. The method of any of clauses 83-87, wherein outputting the report includes: visually presenting, by a display, the report. 89. The method of any of clauses 83-88, further including: determining, based on the heterogeneity condition, one or more therapies to treat the tumor of the subject, wherein the report further indicates the one or more therapies. 90. The method of clause 89, wherein the heterogeneity condition is associated with at least one type or subtype of cancer. 91. The method of any of clauses 83-90, wherein the report indicates the distinct spatial positions associated with the target samples. 92. The method of any of clauses 9-91, further including: generating, based on the heterogeneity condition, a therapy for the subject.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT 93. The method of clause 92, wherein the therapy includes a dosage of one or more therapeutic agents predicted to treat the tumor of the subject. 94. The method of any of clauses 9-93, further including: determining, based on the heterogeneity condition, whether the subject is eligible for a clinical trial. 95. A system, including: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations including: receiving target samples derived from distinct spatial positions in a tumor of a subject; identifying sequence read data of the target samples; and determining a heterogeneity condition of the tumor of the subject based on the sequence read data. 96. The system of clause 95, wherein the operations further include: identifying spatially distinct target regions in one or more samples derived from the tumor of the subject, wherein the target samples are respectively derived from the spatially distinct target regions. 97. The system of clause 96, wherein identifying the spatially distinct target regions includes: determining that the spatially distinct target regions have different visual characteristics. 98. The system of clause 97, wherein the different visual characteristics include a different histopathological morphologies or different vascularization. 99. The system of any of clauses 96-98, wherein identifying the target regions includes: identifying an image of the one or more samples derived from the tumor of the subject; determining, based on the image, metrics associated with different locations in the one or more samples; and selecting, among the different locations in the one or more samples, the spatially distinct target regions by comparing the metrics. 100. The system of any of clauses 96-99, further including: causing a device to collect the target samples based on a visual characteristic or at a predetermined interval along an axis. 101. The system of clause 100, further including: a device configured to perform needle punch enrichment to collect the target samples from one or more samples derived from the tumor of the subject. 102. The system of any of clauses 95-101, further including: a sequencer configured to generate the sequence read data by sequencing a plurality of nucleic acid molecules in the target samples. 103. The system of any of clauses 95-102, further including: a transceiver configured to transmit data indicating the heterogeneity condition of the subject. 104. The system of any of clauses 95-103, further including: an output device configured to output an indication of the heterogeneity condition of the subject. 105. A non-transitory computer readable medium storing instructions for performing operations including: receiving target samples derived from distinct spatial positions in a tumor of a subject;FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT identifying sequence read data of the target samples; and determining a heterogeneity condition of the tumor of the subject based on the sequence read data. Conclusion
[0188] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.
[0189] The features disclosed in the foregoing description, or the following claims, or the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for attaining the disclosed result, as appropriate, may, separately, or in any combination of such features, be used for realizing implementations of the disclosure in diverse forms thereof.
[0190] As will be understood by one of ordinary skill in the art, each implementation disclosed herein can comprise, consist essentially of or consist of its particular stated element, step, or component. Thus, the terms “include” or “including” should be interpreted to recite: “comprise, consist of, or consist essentially of.” The transition term “comprise” or “comprises” means has, but is not limited to, and allows for the inclusion of unspecified elements, steps, ingredients, or components, even in major amounts. The transitional phrase “consisting of” excludes any element, step, ingredient or component not specified. The transition phrase “consisting essentially of” limits the scope of the implementation to the specified elements, steps, ingredients or components and to those that do not materially affect the implementation. As used herein, the term “based on” is equivalent to “based at least partly on,” unless otherwise specified.
[0191] Unless otherwise indicated, all numbers expressing quantities, properties, conditions, and so forth used in the specification and claims are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. When further clarity is required, the term “about” has the meaning reasonably ascribed to it by a person skilled in the art when used in conjunction with a stated numerical value or range, i.e., denoting somewhat more or somewhat less than the stated value or range, to within a range of ±20% of the stated value; ±19% of the stated value; ±18% of the stated value; ±17% of the stated value; ±16% of the stated value; ±15% of the stated value; ±14% of the stated value; ±13% of the stated value; ±12% of the stated value; ±11% of the stated value; ±10% of the stated value; ±9% of the stated value; ±8% of the stated value;FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT ±7% of the stated value; ±6% of the stated value; ±5% of the stated value; ±4% of the stated value; ±3% of the stated value; ±2% of the stated value; or ±1% of the stated value.
[0192] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. Any numerical value, however, inherently contains certain errors necessarily resulting from the standard deviation found in their respective testing measurements.
[0193] The terms “a,” “an,” “the,” and similar referents used in the context of describing implementations (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. Recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein is intended merely to better illuminate implementations of the disclosure and does not pose a limitation on the scope of the disclosure. No language in the specification should be construed as indicating any non-claimed element essential to the practice of implementations of the disclosure.
[0194] Groupings of alternative elements or implementations disclosed herein are not to be construed as limitations. Each group member may be referred to and claimed individually or in any combination with other members of the group or other elements found herein. It is anticipated that one or more members of a group may be included in, or deleted from, a group for reasons of convenience and / or patentability. When any such inclusion or deletion occurs, the specification is deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.
[0195] Unless otherwise indicated, the practice of the present disclosure can employ conventional techniques of immunology, molecular biology, microbiology, cell biology and recombinant DNA. These methods are described in the following publications. See, e.g., Sambrook, et al. Molecular Cloning: A Laboratory Manual, 2nd Edition (1989); F. M. Ausubel, et al. eds., Current Protocols in Molecular Biology, (1987); the series Methods IN Enzymology (Academic Press, Inc.); M. MacPherson, et al., PCR: A Practical Approach, IRL Press at Oxford University Press (1991); MacPherson et al., eds. PCR 2: Practical Approach, (1995); Harlow and Lane, eds. Antibodies, A Laboratory Manual, (1988); and R. I. Freshney, ed. Animal Cell Culture (1987).
[0196] Tumor mutational burden (TMB) is a measure of the number of mutations carried by tumor cells. By comparing DNA sequences from a patient’s healthy tissues and tumor cells, the number of acquiredFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT somatic mutations present in tumors, but not in normal tissues, may be determined. In some instances, driver mutations may be excluded from a TMB calculation.
[0197] In certain examples, "tumor mutational burden" or “TMB” refers to the number of somatic mutations in a tumor's genome and / or the number of somatic mutations per area of the tumor's genome. In some embodiments, TMB, as used herein, refers to the number of somatic mutations per megabase (Mb) of DNA sequenced. In some embodiments, germline (inherited) variants are excluded when determining TMB, given that the immune system has a higher likelihood of recognizing these as self. In various cases, driver mutations are excluded from a TMB calculation.
[0198] Microsatellites are highly polymorphic DNA-repeat regions. In certain examples, “microsatellite” refers to a repetitive nucleic acid having repeat units of less than about 10 base pairs or nucleotides in length. In certain examples, a microsatellite refers to a tract of tandemly repeated (i.e. adjacent) DNA motifs ranging from one to six or up to ten nucleotides, with each motif repeated 5 to 50 repeated times. “Microsatellite instability” refers to genetic instability in the microsatellite regions. Cancer patients with microsatellite instability classified as being high (MSI-H or MSI-High) frequently exhibit an accumulation of somatic mutations in tumor cells that leads to a range of molecular and biological changes including high tumor mutational burden, increased expression of neoantigens and abundant tumor-infiltrating lymphocytes. Chang et al. “Microsatellite Instability: A Predictive Biomarker for Cancer Immunotherapy,” Appl Immunohistochem Mol Morphol, 26(2):e15-e21 (2018). These changes have been linked to increased sensitivity to checkpoint inhibitor drugs, such as pembrolizumab, which is used to treat advanced melanoma, head and neck squamous cell carcinoma, non-small cell lung cancer (NSCLC), and classical Hodgkin lymphoma.
[0199] A viral status test refers to a test that identifies the presence of viral RNA or DNA in a subject. The test can identify viral load and / or viral identity. For example, the viral status test can identify the presence of viral RNA or DNA associated with the occurrence of certain cancers. Examples of such viruses include Hepatitis B Virus (HBV) and Hepatitis C Virus (HCV), Kaposi Sarcoma-Associated Herpesvirus (KSHV), Merkel Cell Polyomavirus (MCV), Human Papillomavirus (HPV), Human Immunodeficiency Virus Type 1 (HIV-1, or HIV), Human T-Cell Lymphotropic Virus Type 1 (HTLV-1), and Epstein-Barr Virus (EBV).
[0200] Cancer “hotspot” mutations give rise to oncological outcomes. PhyloP, SIFT, Grantham, COSMIC and PolyPhen-2 are in silico tools that can be used to assess pathogenicity of identified variants. Exemplary hotspot genes and mutations include EGFR exon 19 activating mutation, EGFR exon 19 deletion, EGFR exon 19 insertion, EGFR exon 19 sensitizing mutation, EGFR exon 20 activation mutation, EGFR exon 20 insertion, EGFR G719 mutation, EGFR L858R mutation, EGFR L861 mutation, EGFR S768 mutation, EGFR T790M mutation, C797 mutation, KIT activating mutation, KRAS activating mutation, MET activating mutation, NRAS activating mutation, PMS2 promoter mutations, among many others. HotspotFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT mutations also occur in the following genes: AKT2, BRCA1, BRCA2, ERC1, NSD1, POLH, PPM1G, PTEN, RAD18, RAD51, RAD51B, RB1, TERT, TP53, TP53Bp1, ALK, ARMT1, ATAD5, ATG7, ATIC, AXL, BIRC6, BRD3, BRD4, CAPRIN1, CCAR2, CCDC6, CDK5RAP2, CHD9, CIT, CTNNB1, CUL1, EBF1, EIF3E, HIP1, HMGA2, IRF2BP2, NOTCH1, NOTCH4, NPM1, OFD1, TACC1, TACC3, TERF2, TMEM106B, UBE2L3, USP10, WRDR48, YAP1, ZEB2, and ZMYND8.
[0201] A “DNA methylation test” refers to an assay, which can be commercially available, for distinguishing methylated versus unmethylated cytosine loci in DNA. Techniques for measuring cytosine methylation include bisulfite-based methylation assays. The addition of bisulfite to DNA results in the methylation of unmethylated cytosine and its ultimate conversion to the nucleotide uracil. Uracil has similar binding properties to thiamine in the DNA sequence. Previously methylated cytosine does not undergo similar chemical conversion on exposure to bisulfite. Bisulfite assays can thus be used to discriminate previously methylated versus unmethylated cytosine.
[0202] An exemplary quantitative methylation detection assay combines bisulfite treatment and restriction analysis COBRA, which uses methylation sensitive restriction endonucleases, gel electrophoresis, and detection based on labeled hybridization probes. (Ziong and Laird, Nucleic Acid Res. 199725; 2532-4). Another exemplary detection assay is the methylation specific polymerase chain reaction PCR (MSPCR) for amplification of DNA segments of interest. This assay can be performed after sodium bisulfite conversion of cytosine and uses methylation sensitive probes. Other detection assays include the Quantitative Methylation (QM) assay, which combines PCR amplification with fluorescent probes designed to bind to putative methylation sites; MethyLightTM(Qiagen, Redwood City, CA) a quantitative methylation detection assay that uses fluorescence-based PCR (Eads, et al., Cancer Res.1999; 59:2302-2306); and Ms- SNuPE, a quantitative technique for determining differences in methylation levels in CpG sites. As with other techniques, Ms-SNuPE also requires bisulfite treatment to be performed first, leading to the conversion of unmethylated cytosine to uracil while methyl cytosine is unaffected. PCR primers specific for bisulfite converted DNA are then used to amplify the target sequence of interest. The amplified PCR product is isolated and used to quantitate the methylation status of the CpG site of interest. (Gonzalgo and Jones Nuclei Acids Res1997; 25:252-31).
[0203] In particular embodiments, pyrosequencing can be used to detect marker methylation. Pyrosequencing is a method of DNA sequencing that relies on detection of the release of pyrophosphates as DNA is synthesized (and is therefore a “sequencing by synthesis” technique). To assess methylation by pyrosequencing, a DNA sample can be incubated with sodium bisulfite, converting unmethylated cytosine to uracil. The presence of uracil will result in thymine incorporation during PCR amplification. Therefore, sequencing results that include thymine at a nucleotide position that is known to encode cytosine can be interpreted as unmethylated sites. In contrast cytosines present in the sequencing results indicate that theFMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT site was methylated in the original DNA sample, because methylation protects cytosine from conversion to uracil upon treatment. Bisulfite treatment can also be performed on control samples with known methylation patterns, to reduce or eliminate false positive results. Commercially available pyrosequencing machines include Pyro Mark Q96 (Qiagen, Hilden, Germany). For more details on methods to use pyrosequencing for measurement of methylation, see Delaney et al. Methods Mol Biol. 20151343: 249- 264. Pyrosequencing is especially useful for detecting methylation in the CpG sites within genes.
[0204] In particular embodiments, a protein marker is detected by contacting a sample with reagents (e.g., antibodies), generating complexes of reagent and marker(s), and detecting the complexes. Particular embodiments for detecting and measuring protein levels can use methods including agglutination, chemiluminescence, electro-chemiluminescence (ECL), enzyme-linked immunoassays (ELISA), immunoassay, immunoblotting, immunodiffusion, immunoelectrophoresis, immunofluorescence, immunohistochemistry, immunoprecipitation, mass-spectrometry, and western blot. See also, e.g., E. Maggio, Enzyme-Immunoassay (1980), CRC Press, Inc., Boca Raton, Fla; and U.S. Pat. Nos. 4,727,022; 4,659,678; 4,376,110; 4,275,149; 4,233,402; and 4,230,797.
[0205] Read depth refers to the number of times that a specific genomic site is sequenced during a sequencing run.
[0206] Certain implementations are described herein, including the best mode known to the inventors for carrying out implementations of the disclosure. Of course, variations on these described implementations will become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventor expects skilled artisans to employ such variations as appropriate, and the inventors intend for implementations to be practiced otherwise than specifically described herein. Accordingly, the scope of this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by implementations of the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.
Claims
FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT CLAIMS What is claimed is:
1. A method, comprising: receiving target samples derived from distinct spatial positions in a tumor of a subject; identifying sequence read data of the target samples; and determining a heterogeneity condition of the tumor of the subject based on the sequence read data.
2. The method of claim 1, wherein the target samples are respectively derived from spatially distinct target regions in one or more samples obtained from the subject.
3. The method of claim 2, further comprising: storing an indication of locations of the target regions in the one or more samples, wherein determining the heterogeneity condition of the tumor of the subject is further based on the locations of the target regions.
4. The method of claim 2, further comprising: identifying the spatially distinct target regions in the one or more samples; identifying an image of the one or more samples; determining, based on the image, metrics associated with different locations in the one or more samples, the metrics comprising at least one of a cell size, a cell morphology, a cell type, a subcellular structure, or a biomarker level; and selecting, among the different locations in the one or more samples, the spatially distinct target regions by comparing the metrics.
5. The method of claim 2, wherein the target samples comprise needle punch enrichment samples obtained from the one or more samples.
6. The method of claim 2, wherein the distinct spatial positions comprise locations that are spaced apart at a predetermined interval along an axis.
7. The method of claim 2, wherein the sequence read data indicates nucleic acid molecules in a combined sample that comprises the target samples.
8. The method of claim 1, wherein the sequence read data further indicates at least one of the distinct spatial positions associated with at least one of the target samples.
9. The method of claim 1, further comprising: determining, based on the sequence read data, a mutational profile, one or more pathogenic variants, a fraction unstable score, a homologous recombination deficiency, an alteration-level clonality, or an intra-tumor heterogeneity of the target samples.
10. The method of claim 9, further comprising:FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT generating input features based on the sequence read data, the input features being indicative of the mutational profile, the one or more pathogenic variants, the fraction unstable score, the homologous recombination deficiency, the alteration-level clonality, the intra-tumor heterogeneity, or the heterogeneity condition of the subject.
11. The method of claim 10, wherein generating input features based on the sequence read data comprises inputting, into a machine learning (ML) model configured to detect the input features, the sequence read data, the method further comprising: training the ML model to identify attributes, indicated by training data, that are predictive of heterogeneity conditions of a plurality of individuals, the training data being indicative of training sequence read data associated with samples corresponding to the plurality of individuals omitting the subject and the heterogeneity conditions of the plurality of individuals, wherein the input features are instances of the attributes identified via the training of the ML model.
12. The method of claim 11, wherein determining the heterogeneity condition of the subject based on the sequence read data comprises: classifying, using a classifier, the heterogeneity condition of the subject based on the input features.
13. The method of claim 1, wherein the heterogeneity condition comprises at least one of: a tumor evolution of the subject; a predicted tumor progression of the subject; a predicted pathologic condition of the subject; a predicted pathologic condition subtype of the subject; a metastasis profile of the subject; a predicted survivability of the subject; a predicted symptom of the subject; a predicted effective therapy to treat the predicted pathologic condition of the subject; a predicted resistance of the subject to a treatment of the predicted pathologic condition; a general health of the subject; a genomic age of the subject; a risk of the subject developing the predicted pathologic condition; a predicted stage of the predicted pathologic condition of the subject; a predicted grade of the predicted pathologic condition of the subject; or a predicted Eastern Cooperative Oncology Group (ECOG) performance status of the subject.
14. The method of claim 1, further comprising:FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT generating a report indicating the heterogeneity condition; and outputting the report.
15. The method of claim 14, further comprising: determining, based on the heterogeneity condition, one or more therapies to treat the tumor of the subject, wherein the report further indicates the one or more therapies.
16. A system, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: receiving target samples derived from distinct spatial positions in a tumor of a subject; identifying sequence read data of the target samples; and determining a heterogeneity condition of the tumor of the subject based on the sequence read data.
17. The system of claim 16, further comprising: a transceiver configured to transmit data indicating the heterogeneity condition of the subject.
18. A method, comprising: receiving a tissue sample derived from a tumor of a subject; identifying, in the tissue sample, target regions, the target regions being spatially distinct; collecting, from the target regions, target samples; providing a plurality of nucleic acid molecules obtained from the target samples; ligating one or more adapters onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying the one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules; capturing amplified nucleic acid molecules from the amplified nucleic acid molecules; sequencing, by a sequencer, all or a subset of the captured amplified nucleic acid molecules to obtain a plurality of sequence reads that represent the sequenced amplified nucleic acid molecules thereby generating sequence read data; receiving, at one or more processors, the sequence read data for the plurality of sequence reads; and determining, using the one or more processors using the sequence read data, a heterogeneity condition of the tumor of the subject.FMI Docket No.: 0203-WO / 0154-CG L&H Docket No.: F171-6011PCT 19. The method of claim 18, wherein the tissue sample comprises fresh tissue, frozen tissue, or fixed tissue.
20. The method of claim 18, wherein identifying the plurality of target regions comprises: identifying a heterogeneous visual characteristic of the plurality of target regions in the tissue sample, the heterogeneous visual characteristic comprising a heterogeneous histopathological morphology of the target regions or a heterogenous vascularization of the target regions.