Identifying and correcting false positive variants
By distinguishing false positive variants through sequence read data analysis, the method addresses genomic testing inaccuracies caused by nucleic acid degradation, improving analysis accuracy and enabling methylation status inference.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- FOUNDATION MEDICINE INC
- Filing Date
- 2025-10-30
- Publication Date
- 2026-05-07
AI Technical Summary
Genomic testing inaccuracies arise from misclassifying degraded nucleic acid molecules as mutations, leading to false positive variants due to storage, which can erroneously identify health-related conditions.
Techniques to distinguish between true positive and false positive variants by processing sequence read data, identifying cytosine-to-thymine substitutions as likely degradation-induced errors, and inferring methylation status without formal methylation sequencing.
Enhances the accuracy and reliability of genomic analyses by correcting false positives, allowing for precise methylation status inference at genomic positions of interest.
Smart Images

Figure US2025053425_07052026_PF_FP_ABST
Abstract
Description
FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCTIDENTIFYING AND CORRECTING FALSE POSITIVE VARIANTSCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional App. No. 63 / 714,814, which was filed on October 31 , 2024 and is incorporated by reference herein in its entirety.BACKGROUND
[0002] Many individuals rely on genetic testing to identify whether they have, or are predicted to develop, various health related conditions. In some cases, single gene testing can be used to assess whether an individual has a particular genetic mutation that is relevant to whether the individual has a genetic disorder or a propensity for disease.
[0003] In general, mutations are identified by obtaining a sample containing nucleic acid molecules (e.g., DNA) from the individual, sequencing the nucleic acid molecules, and comparing the sequences to a reference genome. Differences between the reference genome and the sequences can be reported as mutations or other types of variants.
[0004] However, samples are often stored prior to sequencing. In some cases, storage causes the nucleic acid molecules to degrade. After sequencing, some of these degraded nucleic acid molecules may be misclassified as mutations, which can result in erroneously identifying health related conditions.BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of illustrative embodiments and the accompanying drawings, of which:
[0006] FIG. 1 illustrates an example environment for identifying false positive variants.
[0007] FIGS. 2A and 2B illustrate examples of various ground truth, degraded, and corrected nucleic acid molecules whose sequences can be analyzed for false positive variants using various techniques described herein. FIG. 2A illustrates an example of degradation and correction of unmethylated nucleic acid molecules. FIG. 2B illustrates an example of degradation and correction of methylated nucleic acid molecules.
[0008] FIG. 3 illustrates an example VAF distribution utilized to distinguish between false positive and true positive variants.
[0009] FIG. 4 illustrates an example report summarizing predicted categories of a cancer of a subject.
[0010] FIG. 5 illustrates an example process for identifying false positive variants due to the degradation of nucleotides in a sample.
[0011] FIG. 6 illustrates an example environment for sequencing various nucleic acid molecules.
[0012] FIG. 7 illustrates one or more devices configured to perform various operations described herein.DETAILED DESCRIPTION
[0013] Various implementations of the present disclosure relate to techniques for distinguishing between true positive and false positive variants in sequenced nucleic acid molecules. In particular cases, techniques for distinguishing between true positive and false positive substitution variants are described. In some examples described herein, particular types of substitutions are the result of degradation of nucleic acid molecules during storage. While some of these substitutions can be chemically corrected with corrective enzymes, others result in erroneously reported variants when the degraded nucleic acid molecules are sequenced.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0014] According to various examples of the present disclosure, these substitutions can be detected as false positive variants by processing sequence read data representative of the sequences of the degraded nucleic acid molecules. For example, groups of cytosine-to-thiamine (C>T) substitutions within the sequence read data may be more likely to be the result of degradation, rather than variants in the ground truth nucleic acid molecules prior to storage. In some cases, the false positive variants can be used to infer the methylation status of the ground truth nucleic acid molecules. In various cases, the false positive variants are either not reported, or not relied upon to generate a genomic report.
[0015] Implementations of the present disclosure provide significant improvements to the technical field of genomic testing. By identifying false positive variants, and refraining from relying on the false positive variants for genomic analyses, the accuracy and reliability of the genomic analyses can be enhanced. Moreover, in some examples described herein, methylation status at various genomic positions of interest can be inferred without the expense and complexity of formal methylation sequencing.
[0016] Various analyses described herein cannot be performed in the human mind, or by pen and paper. For example, nucleic acid molecules cannot be sequenced in the human mind. Moreover, analyses described herein rely on analyzing numerous genomic locations, such as hundreds or thousands of genomic locations, and are too extensive to be performed in the human mind or using pen and paper.Example Definitions
[0017] As used herein, the terms "deoxyribonucleic acid,” "DNA,” "DNA molecule,” and their equivalents, may refer to a polymer of nucleotides (also referred to as "nucleobases”) containing deoxyribose. The nucleotides in DNA include cytosine (C), guanine (G), adenine (A), and thymine (T). Each DNA nucleotide includes a deoxyribose and a phosphate group. An example single-stranded DNA (ssDNA) molecule includes a chain of covalently bonded DNA nucleotides. In the example ssDNA molecule, the phosphate group of the mth nucleotide is covalently bonded to the deoxyribose of the ( / 77-1 )th nucleotide, wherein m is a positive integer greater than 2 and less than or equal to the number of DNA nucleotides in the chain. In various examples, DNA is double-stranded and includes two ssDNA molecules that are complementary to one another and coiled around each other in a double helix form. The nucleotides of one ssDNA molecule are hydrogen bonded to the nucleotides of the other ssDNA molecule. In particular, the pyrimidines (A and T) hydrogen bond to each other, and the purines (C and G) hydrogen bond to each other.
[0018] As used herein, the terms "ribonucleic acid,” "RNA,” "RNA molecule,” and their equivalents, may refer to a polymer of nucleotides containing ribose. The nucleotides in RNA include cytosine (C), guanine (G), adenine (A), and uracil (U). Each RNA nucleotide includes a ribose and a phosphate group. In an example RNA molecule, the phosphate group of the nth nucleotide is covalently bonded to the ribose of the (n-1 )th nucleotide, wherein n is a positive integer greater than 2 and less than or equal to the number of RNA nucleotides in the chain. Messenger RNA (mRNA) is a type of RNA molecule that is synthesized (or "transcribed”) by RNA polymerase (an enzyme) to be complementary to a gene encoded in a DNA sequence, and is also used by a ribosome to synthesize a polypeptide or protein. An mRNA is therefore an example of a "coding RNA.” In various cases, intron sequences are removed from an mRNA via a process known as "RNA splicing.” MicroRNA ("miRNA”) are single-stranded RNA molecules that perform post- transcriptional gene expression regulation. For instance, a miRNA may bind to a complementary mRNA molecule, thereby cleaving, destabilizing, or otherwise preventing the mRNA molecule from being translated into a polypeptide or protein by a ribosome. In various examples, a miRNA has a length in a range of 21 to 23 RNA nucleotides. As usedFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT herein, the terms "non-coding RNA” may refer to a type of RNA that is not translated into a protein. Examples of noncoding RNA include miRNA, transfer RNA (tRNA), and ribosomal RNA (rRNA). The term "functional RNA,” and its equivalents, may refer to any RNA molecule that impacts a biological process. For instance, functional RNA may include mRNA, miRNA, tRNA, rRNA, and the like.
[0019] As used herein, the term "base,” and its equivalents, may refer to a monomer of a polymer. For example, a base of DNA or RNA is a nucleotide.
[0020] As used herein, the term "base pair,” and its equivalents, may refer to a pair of complementary DNA nucleotides, which are hydrogen-bonded to one another in a double-stranded DNA molecule. For example, a base pair includes a first base in a first ssDNA and a second base in a second ssDNA, wherein the first and second bases are complementary and hydrogen-bonded to one another.
[0021] As used herein, the terms "nucleotide,” "nucleobase,” "nucleic acid,” "nucleic acid molecule,” and their equivalents, may refer to an organic molecule that includes a nitrogenous base, a sugar, and a phosphate group. In various cases, a nucleotide is a monomer of DNA or RNA. A nucleotide, for instance, is a chemical structure.
[0022] As used herein, the terms "3' end,” "3-prime end,” and their equivalents, may refer to a terminus of a singlestranded nucleotide polymer that includes a base whose third carbon in its deoxyribose or ribose is bound to a hydroxyl group while being unbound to another base.
[0023] As used herein, the terms "5' end,” "5-prime end,” and their equivalents, may refer to a terminus of a singlestranded nucleotide polymer that includes a base whose fifth carbon in its deoxyribose or ribose ring is unbound to another base. In some cases, the fifth carbon is bound to a phosphate group.
[0024] As used herein, the "length” of a polymer refers to a number of covalently bonded monomers that are included in the polymer. For instance, the length of a DNA molecule may be the number of covalently bonded nucleotides in at least one strand of the DNA molecule and / or the number of base pairs in the DNA molecule. In various examples, the length of an RNA molecule may be the number of covalently bonded nucleotides in the RNA molecule.
[0025] As used herein, the term "gene,” and its equivalents, refers to a sequence of DNA nucleotides that is transcribed into a functional RNA. The functional RNA, for instance, is RNA that is translated into a polypeptide or protein (e.g., mRNA) or that has some other biological function (e.g., miRNA, tRNA, etc.). A gene is "expressed” when it is used as a template to generate a functional RNA. A subject, for instance, has numerous genes contained in the subject's genome. A gene may include both introns and exons. As used herein, the term "intron,” and its equivalents, may refer to a subset of DNA nucleotides in a gene that is not used to code for any functional RNA that is expressed by the organism. As used herein, the term "exon,” and its equivalents, may refer to a subset of DNA nucleotides in a gene that is used to code for a functional RNA. For instance, an exon may encode a polypeptide or protein that is expressed by the organism. In various examples, a gene can be represented in data (e.g., as data representative of the sequence of DNA nucleotides in the gene) or as a chemical structure (e.g., as the sequence of DNA nucleotides itself).
[0026] As used herein, the term "genome,” and its equivalents, refers to the aggregate of genes of a subject. In various cases, a genome represents the sequences of several linear DNA molecules that are present in a subject's chromosomes. A "reference genome” refers to an aggregation of genes of one or more reference subjects. In various cases, a genome is represented in data.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0027] As used herein, the terms "pangenome, ” "pan-genome,” "supragenome, ” and their equivalents, refers to an aggregate set of genes from multiple subgroups (e.g., strains) within a population (e.g., a clade) of subjects. A pangenome, for example, indicates genes that are present in all subjects within the population, as well as genes that are present in some of the subjects of the population. A pangenome is represented in data, for instance.
[0028] As used herein, the term "transcriptome, ” and its equivalents, refers to the aggregate of RNA sequences of a subject. In some cases, a transcriptome is limited to mRNA sequences. In various examples, a transcriptome is represented in data.
[0029] As used herein, the term "genomic DNA,” "gDNA,” "chromosomal DNA,” and their equivalents, may refer to DNA molecules that are obtained from a chromosome and / or nucleus of a cell.
[0030] As used herein, the terms "DNA fragment,” "fragment,” and their equivalents, may refer to DNA molecules that are excised and / or broken off from a larger DNA molecule.
[0031] As used herein, the terms "cell-free DNA,” "cfDNA,” and their equivalents, may refer to DNA fragments that are non-encapsulated and obtained outside of cells within a sample (e.g., a liquid biopsy sample).
[0032] As used herein, the terms "circulating tumor DNA,” "ctDNA,” and their equivalents, may refer to a cfDNA molecule that originates from a cancer cell.
[0033] As used herein, the term "promoter,” and its equivalents, may refer to a portion of a DNA molecule that binds one or more proteins in order to initiate transcription of a gene. For example, the promotor is located "upstream” of the gene. For example, the promotor is located between the 5' end of the DNA molecule and the gene. A promotor may include one or more binding sites for RNA polymerase, and / or one or more transcription factor binding sites. In some examples, a promotor includes one or more CpG islands. A promoter, for instance, includes a transcription start site.
[0034] As used herein, the terms "CpG island,” "CGI,” "CpG site,” and their equivalents, may refer to a continuous portion of a DNA molecule whose sequence includes greater than a threshold amount (e.g., greater than 50%) of G-C base pairs.
[0035] As used herein, the term "enhancer,” and its equivalents, may refer to a portion of a DNA molecule that binds one or more proteins in order to increase the chance that a gene will be transcribed. For instance, an enhancer includes one or more transcription factor binding sites. In various cases, an enhancer includes one or more CpG islands.
[0036] As used herein, the term "cancer,” and its equivalents, may refer to a condition of a subject in which particular cells (referred to as "cancer cells”) divide uncontrollably in the subject's body. In some cases, a cancer is characterized by a location or tissue type from which the cancer cells originated. In some examples, a cancer is characterized by a location or tissue type in which the cancer cells are located.
[0037] As used herein, the terms "tumor,” "neoplasm,” and their equivalents, may refer to a mass of tissue including cancer cells.
[0038] As used herein, the terms "tissue of origin,” "tissue origin,” and their equivalents, refers to a differentiated type of tissue from which cancer cells in the body of a subject began dividing uncontrollably in the subject's body.
[0039] As used herein, the terms "liquid biopsy,” "fluid biopsy,” and their equivalents, may refer to a process of obtaining a fluid sample from a subject's body. The sample, for instance, can be referred to as a "liquid biopsy sample.” Examples of fluids that are sampled from the body include blood, plasma, cerebrospinal fluid, sputum, stool, urine, lymphatic fluid, and saliva.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0040] As used herein, the term "tissue biopsy,” and its equivalents, may refer to a process of obtaining a sample of cells from a subject's body. A tissue biopsy, in various cases, is performed by cutting a mass of cells from the subject's body. For instance, a tissue biopsy is a procedure performed by a surgeon, interventional radiologist, interventional cardiologist, or other specialized clinician. The term "tissue” or "tissue biopsy sample” can be used to refer to the sample of cells obtained using a tissue biopsy.
[0041] As used herein, the term "subject,” and its equivalents, may refer to a human or non-human animal. A subject that is receiving care from at least one care provider may be referred to as a "patient.”
[0042] As used herein, the terms "machine learning,” "ML,” "computer learning,” "artificial intelligence,” and their equivalents, may refer to the use of a computing devices to learn patterns in training data. The process of learning these patterns may be referred to as "training.” In particular cases, one or more computing devices may perform machine learning by executing a machine learning model. As used herein, the terms "machine learning model,” "ML model,” and their equivalents, may refer to data encoding instructions that, when executed by at least one computing device, causes the at least one computing device to learn patterns in training data by optimizing one or more metrics, values, or other types of parameters. After training, an ML model, when executed by at least one computing device, causes the at least one computing device to utilize the optimized parameters in order to perform one or more tasks.
[0043] As used herein, the term "variant,” and its equivalents, may refer to a difference between a subject genetic sequence and a reference sequence. For instance, a variant may correspond to a difference between one or more nucleotides in a genome of a subject and one or more corresponding nucleotides in at least one reference genome or pangenome. A variant may be characterized by its identity (e.g., what nucleotides are different), its position (e.g., where are the nucleotides located in the genome, what chromosome contains the nucleotides, what gene contains the nucleotides, etc.), its length (e.g., how many nucleotides are different from the reference sequence), its type (e.g., substitution, insertion, deletion, copy number alternation, rearrangement of fusion, etc.), and other features that indicates its significance and / or relevance. In some cases, a variant represents any apparent alteration in a sequence that has been read from a nucleic acid molecule with respect to the reference sequence, such as reads cleaved by restriction enzymes (RE). In various examples, a variant can be represented in data (e.g., by data characterizing the variant) or as a chemical structure (e.g., the nucleotides themselves). As used herein, the term "mutation,” and its equivalents, may refer to a change in a gene.
[0044] As used herein, the term "substitution,” and its equivalents, can refer to a nucleotide in a subject sequence that is different than an equivalent nucleotide (e.g., a nucleotide at the same position) in a reference sequence.
[0045] As used herein, the term "insertion,” and its equivalents, can refer to a nucleotide in a subject sequence that is added with respect to a reference sequence.
[0046] As used herein, the term "deletion,” and its equivalents, can refer to the removal of a nucleotide from a nucleotide sequence.
[0047] As used herein, the terms "copy number alternation,” "CNA,” "copy number variation,” "CNV,” and their equivalents, can refer to a portion of a reference sequence that is repeated.
[0048] As used herein, the terms "rearrangement of fusion,” "fusion rearrangement,” "translocation,” and their equivalents, can refer to a change in the relative position of one or more portions of a reference sequence, thereby generating a gene that was not present in the reference sequence.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0049] As used herein, the term "sequencing,” and its equivalents, may refer to a process of identifying the order and identity of monomers in a polymer chain, such as the order and identity of nucleotides in a DNA or RNA molecule. The terms "whole genome sequencing,” "WGS,” and their equivalents, may refer to the process of sequencing an entire genome of a subject, including the introns and exons of the genes of the subject. The term "whole exome sequencing,” and its equivalents, may refer to the process of sequencing all exomes of a subject. The term "targeted sequencing,” and its equivalents, may refer to the process of sequencing a portion of the genome of a subject, such as sequencing a single gene of the subject. Various techniques can be utilized to sequence a DNA or RNA molecule, such as massively parallel sequencing (MPS), nanopore sequencing, direct sequencing, Sanger sequencing, or nextgeneration sequencing (NGS). In various cases, sequencing is performed on physical molecules (e.g., RNA or DNA) and is used to generate data.
[0050] As used herein, the terms "massive parallel sequencing,” "massively parallel sequencing,” "MPS,” and their equivalents, may refer to a technique for simultaneously performing multiple reactions that can be used to identify the order and identity of monomers in multiple polymer chains. In particular cases, massive parallel sequencing can be performed using sequencing-by-synthesis on clonally amplified DNA molecules that are located in spatially separated regions, which are individually monitored by sensors.
[0051] As used herein, the term "nanopore sequencing,” and its equivalents, may refer to a technique for identifying the order and identity of monomers in a polymer chain by transporting the polymer chain from a first space to a second space, wherein the first space and the second space are separated by a substrate, by directing the polymer chain through a small hole (known as a "nanopore”) embedded in the substrate, and monitoring a relative electrical signal (e.g., a voltage or current) between the first space and the second space.
[0052] As used herein, the term "sensor,” and its equivalents, may refer to a physical device or other apparatus that is configured to detect one or more detection signals.
[0053] As used herein, the term "detection signal,” and its equivalents, may refer to a physical signal that can be identified, characterized, or otherwise perceived by a sensor.
[0054] As used herein, the term "sequence read data,” and its equivalents, may refer to data that is indicative of an order and identity of monomers in a polymer, such as the order and identity of nucleotides in a DNA or RNA sequence. In various implementations, sequence read data is generated via a sequencing operation.
[0055] As used herein, the term "image,” and its equivalents, may refer to 2D or 3D array of data indicative of an array of pixels or voxels.
[0056] As used herein, the term "ligating,” and its equivalents, may refer to a process of joining two molecules together, for example, with a chemical bond.
[0057] As used herein, the term "adapter,” and its equivalents, may refer to an oligonucleotide that can be ligated to a target nucleic acid molecule. In various cases, an adapter prepares the target nucleic acid molecule for sequencing.
[0058] As used herein, the term "bait molecule,” and its equivalents, may refer to a nucleic acid molecule having a region that is complementary to a region of a target molecule (e.g., cfDNA). A bait molecule includes, for instance, a nucleic acid molecule that can hybridize to ( / .e., is complementary to) a target molecule can be used to capture the target molecule. In some instances, the bait molecule is a capture oligonucleotide (or capture probe). In some instances, the bait molecule is suitable for solution phase hybridization to the target molecule. In some instances, theFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT bait molecule is suitable for solid phase hybridization to the target molecule. In some instances, the bait molecule is suitable for both solution-phase and solid-phase hybridization to the target molecule. The design and construction of bait molecules is described in more detail in, e.g., International Patent Application Publication No. WO 2020 / 236941 .
[0059] As used herein, the term "amplifying,” and its equivalents, may refer to a process of generating copies of a target molecule, such as a nucleic acid molecule.
[0060] As used herein, the term "hybridization,” and its equivalents, may refer to a process by which to complementary single-stranded nucleic acid molecules bind to one another, thereby forming a double-stranded nucleic acid molecule. In certain examples, the double-stranded nature of the nucleic acid molecule is maintained under stringent hybridization conditions. Exemplary stringent hybridization conditions include an overnight incubation at 42 °C in a solution including 50% formamide, 5XSSC (750 mM NaCI, 75 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5XDenhardt's solution, 10% dextran sulfate, and 20 pig / ml denatured, sheared salmon sperm DNA, followed by washing the filters in 0.1XSSC at 50 °C.
[0061] As used herein, the term "complementary,” and its equivalents, may refer to a state of two single-stranded nucleic acid molecules with respective sequences that cause the nucleic acid molecules to spontaneously hybridize to one another. One nucleic acid molecule, for instance, may have a sequence that causes each nucleic acid to hydrogen bond to a respective nucleic acid in the other nucleic acid molecule.
[0062] As used herein, the terms "therapy,” "treatment,” and their equivalents, may refer to a composition or process that can be used to remediate a health problem. In some cases, the therapy includes administration of one or more therapeutic agents (e.g., medications). Anticancer therapies, for instance, include surgery, radiotherapy (also referred to as "radiation therapy”), chemotherapy, immunotherapy, cell-based therapies, and the like. Examples of cancer therapies include abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), ado-trastuzumab emtansine (Kadcyla), afatinib dimaleate (Gilotrif), aldesleukin (Proleukin), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), asciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicabtagene ciloleucel (Yescarta), axitinib (Inlyta), belantamab mafodotin-blmf (Blenrep), belimumab (Benlysta), belinostat (Beleodaq), belzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib (Bosulif), brentuximab vedotin (Adcetris), brexucabtagene autoleucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib (Cabometyx), cabozantinib (Cabometyx, Cometriq), canakinumab (Haris), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), cemiplimab-rwlc (Libtayo), ceritinib (LDK378 / Zykadia), cetuximab (Erbitux), cobimetinib (Cotellic), copanlisib hydrochloride (Aliqopa), crizotinib (Xalkori), dabrafenib (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (Darzalex Faspro), darolutamide (Nubeqa), dasatinib (Sprycel), denileukin diftitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostarlimab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elotuzumab (Empliciti), enasidenib mesylate (Idhifa), encorafenib (Braftovi), enfortumab vedotin-ejfv (Padcev), entrectinib (Rozlytrek), enzalutamide (Xtandi), erdafitinib (Balversa), erlotinib (Tarceva), everolimus (Afinitor), exemestane (Aromasin), fam-trastuzumab deruxtecan-nxki (Enhertu), fedratinib hydrochloride (Inrebic), fulvestrant (Faslodex), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib (Xospata), glasdegib maleate (Daurismo), hyaluronidase-zzxf (Phesgo), ibrutinib (Imbruvica), ibritumomab tiuxetan (Zevalin), idecabtageneFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT vicleucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infigratinib phosphate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguane 1131 (Azedra), ipilimumab (Yervoy), isatuximab-irfc (Sarclisa), ivosidenib (Tibsovo), ixazomib citrate (Ninlaro), lanreotide acetate (Somatuline Depot), lapatinib (Tykerb), larotrectinib sulfate (Vitrakvi), Lenvatinib mesylate (Lenvima), letrozole (Femara), lisocabtagene maraleucel (Breyanzi), loncastuximab tesirine-lpyl (Zynlonta), lorlatinib (Lorbrena), lutetium Lu 177-dotatate (Lutathera), margetuximabcmkb (Margenza), midostaurin (Rydapt), mobocertinib succinate (Exkivity), mogamulizumab-kpkc (Poteligeo), moxetumomab pasudotox- tdfk (Lumoxiti), naxitamab-gqgk (Danyelza), necitumumab (Portrazza), neratinib maleate (Nerlynx), nilotinib (Tasigna), niraparib tosylate monohydrate (Zejula), nivolumab (Opdivo), obinutuzumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), olaratumab (Lartruvo), osimertinib (Tagrisso), palbociclib (Ibrance), panitumumab (Vectibix), panobinostat (Farydak), pazopanib (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pexidartinib hydrochloride (Turalio), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), pralsetinib (Gavreto), radium 223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), ribociclib (Kisqali), ripretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase human (Rituxan Hycela), romidepsin (Istodax), rucaparib camsylate (Rubraca), ruxolitinib phosphate (Jakafi), sacituzumab govitecanhziy (Trodelvy), seliciclib, selinexor (Xpovio), selpercatinib (Retevmo), selumetinib sulfate (Koselugo), siltuximab (Sylvant), sipuleucel-T (Provenge), sirolimus protein-bound particles (Fyarro), sonidegib (Odomzo), sorafenib (Nexavar), sotorasib (Lumakras), sunitinib (Sutent), tafasitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen (Nolvadex), tazemetostat hydrobromide (Tazverik), tebentafusp-tebn (Kimmtrak), temsirolimus (Torisel), tepotinib hydrochloride (Tepmetko), tisagenlecleucel (Kymriah), tisotumab vedotin-tftv (Tivdak), tocilizumab (Actemra), tofacitinib (Xeljanz), tositumomab (Bexxar), trametinib (Mekinist), trastuzumab (Herceptin), tretinoin (Vesanoid), tivozanib hydrochloride (Fotivda), toremifene (Fareston), tucatinib (Tukysa), umbralisib tosylate (Ukoniq), vandetanib (Caprelsa), vemurafenib (Zelboraf), venetoclax (Venclexta), vismodegib (Erivedge), vorinostat (Zolinza), zanubrutinib (Brukinsa), ziv-aflibercept (Zaltrap), and combinations thereof. Examples of cancer therapies also include targeted antibody-based therapies (antibody-drug conjugates, antibody-radioisotope conjugates, and targeted immune cell therapies (e.g., immune effector cells genetically modified to express a chimeric antigen receptor (CAR).
[0063] As used herein, the term "treatment-responsive,” and its equivalents, may refer to a type of cancer cells that can be substantially killed using a predetermined type of therapy. For example, cancer cells of a subject may be responsive to a particular treatment if, after the subject is administered the treatment, the cancer cells are diminished by a particular progression level (e.g., radiographic progression level, marker-based progression level, such as prostate-specific antigen (PSA) progression, etc.). Accordingly, the responsiveness of the cells to the type of therapy may indicate the effectiveness of that therapy.
[0064] As used herein, the term "treatment-resistant,” and its equivalents, may refer to a type of cancer that cannot be substantially killed using a predetermined type of therapy.
[0065] As used herein, the term "metastasis profile,” and its equivalents, may refer to a propensity of a type of cancer to metastasize into one or more differentiated tumor types besides the cancer's tissue origin. In some implementations, the metastasis profile can further indicate the type of tissue in which the cancer can or is likely to metastasize.
[0066] As used herein, the term "clinical trial,” and its equivalents, may refer to a research study used to evaluate a hypothesis based on participation by one or more subjects. In various examples, a clinical trial can be used to assessFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT the efficacy and / or safety of a proposed therapy. A clinical trial may be performed in furtherance of approval of a treatment by a regulatory authority (e.g., the United States Food & Drug Administration (FDA)).Description of Example Implementations
[0067] Various implementations of the present disclosure will now be described with reference to the accompanying Figures.
[0068] FIG. 1 illustrates an example environment 100 for identifying false positive variants. In various implementations, a fresh sample 104 is obtained from a subject 102. In some cases, the subject 102 lacks any apparent disease or other pathological condition. For example, the subject 102 may present to a clinical environment for an assessment of a condition of the body of the subject 102, such as the general health or well-being of the subject 102.
[0069] In various implementations, the subject 102 has a disease or a suspected disease. The subject 102, for instance, may present to the clinical environment with a lesion. In various cases, the lesion may be a tumor that includes cancer cells. According to various examples, the subject 102 has one or more types of cancer, such as adrenal cancer, bladder cancer, blood cancer, bone cancer, brain cancer, breast cancer, carcinoma, cervical cancer, colon cancer, colorectal cancer, corpus uterine cancer, ear, nose and throat (ENT) cancer, endometrial cancer, esophageal cancer, gastrointestinal cancer, head and neck cancer, Hodgkin's disease, intestinal cancer, kidney cancer, larynx cancer, leukemia, liver cancer, lymph node cancer, lymphoma, lung cancer, melanoma, mesothelioma, myeloma, nasopharynx cancer, a neuroblastoma, non-Hodgkin's lymphoma, oral cancer, ovarian cancer, pancreatic cancer, penile cancer, pharynx cancer, prostate cancer, rectal cancer, sarcoma, seminoma, skin cancer, stomach cancer, a teratoma, testicular cancer, thyroid cancer, uterine cancer, vaginal cancer, a vascular tumor, or combinations or metastases thereof.
[0070] In some embodiments, the subject 102 has a B cell cancer (multiple myeloma), a melanoma, breast cancer, lung cancer, bronchus cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, urinary bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine cancer, endometrial cancer, cancer of an oral cavity, cancer of a pharynx, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small bowel cancer, appendix cancer, salivary gland cancer, thyroid gland cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, a cancer of hematological tissue, an adenocarcinoma, an inflammatory myofibroblastic tumor, a gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphocytic leukemia (ALL), acute myelocytic leukemia (AML), chronic myelocytic leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia Vera, Hodgkin lymphoma, non-Hodgkin lymphoma (NHL), soft-tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endotheliosarcoma, lymphangiosarcoma, lymphangioendotheliosarcoma, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinomas, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell cancer, essential thrombocythemia, agnogenic myeloidFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familiar hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine cancers, or a carcinoid tumor.
[0071] In some embodiments, the subject 102 has acute lymphoblastic leukemia (Philadelphia chromosome positive), acute lymphoblastic leukemia (precursor B-cell), acute myeloid leukemia (FLT3+), acute myeloid leukemia (with an IDH2 mutation), anaplastic large cell lymphoma, basal cell carcinoma, B-cell chronic lymphocytic leukemia, bladder cancer, breast cancer (HER2 overexpressed / amplified), breast cancer (HER2+), breast cancer (HR+, HER2-), cervical cancer, cholangiocarcinoma, chronic lymphocytic leukemia, chronic lymphocytic leukemia (with 17p deletion), chronic myelogenous leukemia, chronic myelogenous leukemia (Philadelphia chromosome positive), classical Hodgkin lymphoma, colorectal cancer, colorectal cancer (dMMR / MSI-H), colorectal cancer (KRAS wild type), cryopyrin- associated periodic syndrome, a cutaneous T-cell lymphoma, dermatofibrosarcoma protuberans, a diffuse large B-cell lymphoma, fallopian tube cancer, a follicular B-cell non-Hodgkin lymphoma, a follicular lymphoma, gastric cancer, gastric cancer (HER2+), gastroesophageal junction (GEJ) adenocarcinoma, a gastrointestinal stromal tumor, a gastrointestinal stromal tumor (KIT+), a giant cell tumor of the bone, a glioblastoma, granulomatosis with polyangiitis, a head and neck squamous cell carcinoma, a hepatocellular carcinoma, Hodgkin lymphoma, juvenile idiopathic arthritis, lupus erythematosus, a mantle cell lymphoma, medullary thyroid cancer, melanoma, a melanoma with a BRAF V600 mutation, a melanoma with a BRAF V600E or V600K mutation, Merkel cell carcinoma, multicentric Castleman's disease, multiple hematologic malignancies including Philadelphia chromosome-positive ALL and CML, multiple myeloma, myelofibrosis, a non-Hodgkin's lymphoma, a nonresectable subependymal giant cell astrocytoma associated with tuberous sclerosis, a non-small cell lung cancer, a non-small cell lung cancer (ALK+), a non-small cell lung cancer (PD-L1 +), a non-small cell lung cancer (with ALK fusion or ROS1 gene alteration), a non-small cell lung cancer (with BRAF V600E mutation), a non-small cell lung cancer (with an EGFR exon 19 deletion or exon 21 substitution (L858R) mutations), a non-small cell lung cancer (with an EGFR T790M mutation), a non-small cell lung cancer KRAS (+ / - G12C), a non-small cell lung cancer TMB-H, a non-small cell lung cancer MET exon 14 skipping, a non-small cell lung cancer ERBB2 inframe indel, a non-small cell lung cancer EGFR exon 20 indel, a neurotrophic tyrosine receptor kinase (NTRK)-positive cancer, ovarian cancer, ovarian cancer (with a BRCA mutation), pancreatic cancer, a pancreatic, gastrointestinal, or lung origin neuroendocrine tumor, a pediatric neuroblastoma, a peripheral T-cell lymphoma, peritoneal cancer, prostate cancer, a renal cell carcinoma, a small lymphocytic lymphoma, a soft tissue sarcoma, a solid tumor (MSI-H / dMMR), a squamous cell cancer of the head and neck, a squamous non-small cell lung cancer, thyroid cancer, a thyroid carcinoma, urothelial cancer, a urothelial carcinoma, or Waldenstrom's macroglobulinemia.
[0072] In some examples, the fresh sample 104 includes a tissue biopsy sample. For instance, the fresh sample 104 is obtained by removing cells from the lesion and / or from a portion of the subject 102. In some cases, the tissue biopsy sample is surgically excised from the subject 102. 1 n some cases, the fresh sample 104 includes a liquid biopsy sample. The fresh sample 104, for instance, includes blood, plasma, cerebrospinal fluid, sputum, stool, urine, lymphatic fluid, saliva, or some other fluid obtained from the body of the subject 102. In some cases, a blood sample is obtained intravenously from the subject 102. The fresh sample 104, according to various examples, is a plasma sample obtained from the blood of the subject 102. The fresh sample 104, for instance, can be obtained in a minimally invasive procedure, which could be performed by a medical technician rather than a surgeon.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0073] The fresh sample 104 contains various nucleic acid molecules 106. According to some examples, the nucleic acid molecules 106 include genomic DNA (gDNA). For instance, the nucleic acid molecules 106 include chromosomal DNA that is located in, or extracted from, cells in the fresh sample 104. According to some cases, the DNA is extracted from nuclei and the cells in the fresh sample 104 using mechanical shearing and / or the introduction of a chemical (e.g., a detergent). The DNA may be subsequently isolated from proteins and other cellular materials. In some implementations, the nucleic acid molecules 106 indicate an entire genome of the subject 102 and / or the lesion. Thus, a genome of the subject 102 and / or the lesion can be determined by sequencing the DNA in the nucleic acid molecules 106.
[0074] In some examples, the nucleic acid molecules 106 include RNA. In some implementations, the nucleic acid molecules 106 include messenger RNA (mRNA), microRNA, non-coding RNA, functional RNA, or any combination thereof. Various RNA in the nucleic acid molecules 106 may be indicative of proteins expressed in the cells of the subject 102 and / or the lesion.
[0075] In some cases, the fresh sample 104 includes cell-free DNA (cfDNA). In examples in which the subject 102 has cancer (e.g., the lesion is a cancerous tumor), the cfDNA, for instance, includes circulating tumor DNA (ctDNA) and / or non-ctDNA. In cases wherein the lesion is a tumor, cancer cells within the lesion will lyse and release the ctDNA into the bloodstream of the subject 102. These cancer cells, for example, include circulating tumor cells (CTCs). Further, other cells additionally release non-ctDNA into the bloodstream of the subject. In general, the cfDNA includes fragments with lengths that are in a range of 1 to 500, 3 to 500, or 100 to 500 bases long. For instance, the cfDNA includes fragments that are about 170 bases long and / or fragments that are about 340 bases long. For example, the cfDNA includes fragments that are 100 to 240 bases long and / or fragments that are 270 to 410 bases long.
[0076] In various cases, the fresh sample 104 is transported to a location that is remote from the subject 102 for further processing. For example, the fresh sample 104 is removed from the subject 102 in a clinical environment (e.g., a hospital) and is then transported to a remote location for further processing.
[0077] According to various cases, the fresh sample 104 is disposed in a storage container 108. The storage container 108, for instance, may include a vial, a flask, a tube, a cup, or the like. In some cases, the storage container 108 encloses the fresh sample 104 to prevent contamination from an external environment from entering the storage container 108. In some examples, the storage container 108 includes a polymer (e.g., polypropylene).
[0078] The fresh sample 104, for example, is stored in the storage container 108 in a controlled temperature environment prior to processing. In some cases, the fresh sample 104 is stored in a refrigerator and / or freezer. In some cases, the fresh sample 104 is in a frozen form. For example, the fresh sample 104 is refrigerated at a temperature of 40° Celsius (C), stored at an ultralow temperature freezer in a range of -50°C to -800°C, stored in liquid nitrogen at an even lower temperature, or any combination thereof. For instance, the fresh sample 104 is exposed to a temperature that is above 0°C.
[0079] In some cases, the fresh sample 104 is stored for an extended period of time. In some cases, the fresh sample 104 is stored for greater than 24 hours. The fresh sample 104 may be stored for days, weeks, months, or years, for instance. For example, the fresh sample 104 is sored for greater than a threshold amount of time.
[0080] The fresh sample 104 transitions to a degraded sample 110 after being disposed in the storage container 108. In various cases, the nucleic acid molecules 106 transform into degraded nucleic acid molecules 112 within the degraded sample 110. The degradation of the nucleic acid molecules 106 into the degraded nucleic acid moleculesFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT112 may be the result of chemical degradation. According to some cases, one or more nucleotides (e.g., bases and / or base pairs) within the degraded nucleic acid molecules 112 may be different from those originally present in the nucleic acid molecules 106. That is, the degraded nucleic acid molecules 112 may include erroneous nucleotides that are not present in the ground truth nucleic acid molecules 106, due to the degradation.
[0081] Corrected nucleic acid molecules 114 are generated by adding a corrective enzyme 116 to the degraded nucleic acid molecules 112. For example, at least some of the erroneous nucleotides may be converted back into ground truth nucleotides via the corrective enzyme 116. In particular examples, the corrective enzyme 116 includes uracil-DNA glycosylase. The erroneous nucleotides, for instance, include cytosines that have been degraded into uracils within the degraded nucleic acid molecules 112 via deamination. The uracil-DNA glycosylase, for instance, may convert the uracils back into cytosines, which match the original ground truth sequences of the nucleic acid molecules 106.
[0082] Although the corrective enzyme 116 transforms some of the bases in the degraded nucleic acid molecules 112 into their ground truth forms (as originally included in the nucleic acid molecules 106 of the fresh sample 104), some bases in the corrected nucleic acid molecules 114 remain erroneous. In some cases, while unmethylated cytosines are degraded into uracil during storage, methylated cytosines are degraded into thiamine. Uracil-DNA glycosylase, for instance, is unable to convert the thiamines into cytosines. Moreover, because thiamine is a DNA base, it would be challenging to chemically correct the thiamines present in the degraded nucleic acid molecules 112 as a result of degradation without converting ground truth thiamines originally present in the nucleic acid molecules 106.
[0083] A sequencer 118 is configured to generate sequence read data 120 based on the corrected nucleic acid molecules 114. The sequencer 118, for instance, includes one or more devices that are configured to generate the sequence read data 120 by processing at least a portion of the corrected nucleic acid molecules 114. In some cases, the corrected nucleic acid molecules 114 are extracted from the degraded sample 110 after the corrective enzyme 116 is applied. The extraction can be performed by the sequencer 118, by another device, manually (e.g., by a laboratory technician), or any combination thereof. Any appropriate extraction method known to those of ordinary skill in the art can be utilized.
[0084] In various cases, the sequencer 118 is configured to perform one or more processes (e.g., chemical reactions) on the corrected nucleic acid molecules 114 in order to prepare the corrected nucleic acid molecules 1 14 for sequencing. For instance, the sequencer 118 may ligate adapters onto the corrected nucleic acid molecules 114 and / or amplify the corrected nucleic acid molecules 114, such that numerous copies of the ligated corrected nucleic acid molecules 114 are available for sequencing. Examples of the adapters include, for example, amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences. The corrected nucleic acid molecules 114 (e.g., the ligated corrected nucleic acid molecules 114) may be amplified by generating multiple copies of the corrected nucleic acid molecules 114 using one or more techniques such as polymerase chain reaction (PGR), a non-PCR amplification technique, or an isothermal amplification technique.
[0085] The sequencer 118 may identify the length, position, and identity of the bases in the corrected nucleic acid molecules 114 by sequencing the corrected nucleic acid molecules 114 (e.g., the amplified and / or ligated corrected nucleic acid molecules 114). In various implementations, the sequencer 118 utilizes first-generation sequencing (e.g., Sanger sequencing), second-generation sequencing (e.g., massive parallel sequencing), third-generation sequencing (e.g., nanopore sequencing), or a combination thereof. In some cases, the sequencer 118 is configured to sequenceFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT substantially all of the nucleotides of all of the corrected nucleic acid molecules 114 fragments obtained from the degraded sample 110 after application of the corrective enzyme 116. 1 n some examples, the sequencer 1181s configured to perform targeted sequencing. For instance, the sequencer 118 may determine whether the corrected nucleic acid molecules 114 fragments contain one or more predetermined sequences at one or more genomic locations.
[0086] In various cases, the sequencer 118 includes one or more sensors that are configured to detect physical signals (also referred to as "detection signals”) that are indicative of the nucleotide sequences of the corrected nucleic acid molecules 114. The sequencer 118 may perform sequencing-by-synthesis. For example, the sequencer 118 may include one or more optical sensors configured to detect optical signals emitted from fluorescently tagged nucleotide triphosphates (NTPs) that are joined together in a synthesized DNA strand using the ligated corrected nucleic acid molecules 114 as templates. The optical signals detected by the optical sensor(s), for instance, are indicative of the sequences of the corrected nucleic acid molecules 114. The sequencer 118 may perform nanopore sequencing. In various cases, the sequencer 118 includes one or more electrical sensors configured to measure an electrical signal (e.g., an electrical current) across a substrate as the ligated corrected nucleic acid molecules 114 are directed through a nanopore extending through the substrate. The electrical signal over time, in various cases, is indicative of the sequences of the corrected nucleic acid molecules 1 14. The sequencer 118, in various implementations, is configured to generate the sequence read data 120 as digital data based on the analog signals detected by the sensor(s). For instance, the sequencer 118 includes one or more analog to digital converters (ADCs). In various cases, the sequencer 118 includes at least one processor configured to generate the sequence read data 120.
[0087] In some implementations, the sequencer 118 performs RNA sequencing (RNA-seq) on the corrected nucleic acid molecules 114. For example, the corrected nucleic acid molecules 114 include RNA. In some examples, the RNA in the corrected nucleic acid molecules 114 is fragmented. In various implementations, complementary DNA (cDNA) is generated using reverse transcriptase, such that the cDNA includes sequences that are complementary to the RNA in the corrected nucleic acid molecules 114. The cDNA, according to various cases, can be sequenced using the DNA sequencing techniques described above. Accordingly, in some cases, the sequence read data 120 indicates sequences of RNA present in the corrected nucleic acid molecules 114, which may be indicative of the transcriptome of the subject 102 and / or the lesion.
[0088] In various cases, the sequencer 118 performs sequencing on a subset of the corrected nucleic acid molecules 114. For instance, the sequencer 118 may perform targeted sequencing on one or more predetermined genes, such as any of the genes described herein. The sequencer 118, in some cases, may refrain from sequencing at least a portion of the corrected nucleic acid molecules 114 that do not correspond to the subset.
[0089] However, because the corrected nucleic acid molecules 114 include erroneous bases, the sequence read data 120 also includes erroneous bases. Conventionally, variants can be identified by identifying discrepancies between the sequence read data 120 and a reference genome. However, in implementations of FIG. 1 , the erroneous bases will identify false positive variants 122 in addition to true variants present in the nucleic acid molecules 106 of the fresh sample 104. For instance, methylated cytosines in the nucleic acid molecules 106 that have degraded into thiamine in the degraded nucleic acid molecules 112 may be marked as cytosine-to-thiamine variants in the sequence read data 120. The corrective enzyme 116, for instance, is unable to chemically correct this type of degradation. Accordingly, this type of degradation may continue to be present in the corrected nucleic acid molecules 114 after administration of theFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT corrective enzyme 116, and may falsely be identified as a variant in the sequence read data 120 even though the degradation is not a variant that was present in the original nucleic acid molecules 106 of the fresh sample 104.
[0090] In various implementations of the present disclosure, a variant analyzer 124 is configured to differentiate between the false positive variants 122 and additional variants 126 indicated by the sequence read data 120. Initially, the variant analyzer 124 aligns the sequence read data 120 with a reference genome. In various cases, the sequence read data 120 and the reference genome are aligned by common genomic positions. The variant analyzer 124, for instance, is further configured to compare the sequence read data 120 to a reference genome. Bases within the sequence read data 120 that differ from the reference genome are identified.
[0091] In various implementations, the variant analyzer 124 generates a variant allele frequency (VAF) distribution based on the comparison between the sequence read data 120 and the reference genome. For example, the VAF indicates, for each of a selection of genomic positions, a frequency that the sequence read data 120 indicates a different base than the reference genome. Thus, in some cases, each genomic position represented by the VAF distribution is associated with a corresponding frequency that the sequence read data indicates a different base (or base pair) than the reference genome at that genomic position.
[0092] In some cases, the VAF distribution generated by the variant analyzer 124 is limited to a subset of genomic positions within a genomic region. The subset of genomic positions may be selected based on the type of the false positive variants 122 that may be left in the corrected nucleic acid molecules 114 as a result of storage degradation and that are not corrected by the corrective enzyme 116. For example, the VAF distribution may correspond to genomic positions that are cytosines and / or guanines in the genomic region. In some cases, the VAF distribution is limited to CpG sites within the genomic region.
[0093] In some implementations, the VAF distribution reflects only one or more predetermined types of potential variants within the sequence read data 120. In some cases, the VAF distribution indicates only one or more types of substitution variants. For example, the VAF distribution may represent the frequency of bases that correspond to potential cytosine-to-thiamine substitution variants and / or potential guanine-to-adenosine (G>A) substitution variants in the sequence read data 120. Other types of variants, including other types of substitution variants, may not be represented in the VAF distribution generated by the variant analyzer 12, for instance.
[0094] Optionally, the VAF distribution is smoothed. Smoothing techniques include Gaussian smoothing, Savitzy- Golay filtering, additive smoothing, applying a moving average, kernel smoothing, performing local regression, Laplacian smoothing, and the like. Smoothing may reduce noise within the VAF distribution, in some cases.
[0095] According to various implementations, the variant analyzer 124 identifies candidate variants by analyzing the VAF distribution. The term "candidate variant,” as used herein, may represent a positive variant in the sequence read data 120. For instance, the candidate variants include both false positive and true positive variants indicated by the sequence read data 120.
[0096] In various implementations, candidate variants are identified by comparing the VAF distribution to a threshold. In some implementations, the threshold is a predetermined threshold. In some cases, the threshold is determined based on the VAF distribution itself. For example, the threshold may be a predetermined percentage of the peak (e.g., maximum) of the VAF distribution. The predetermined percentage, for instance, may be in a range of 50% to 98%, such as 50%, 60%, 70%, 80%, 90%, 95%, or 98%. In some cases, the threshold is determined based on an averageFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT(e.g., mean) of the VAF distribution. For example, the threshold may be obtained by determining a product of the average and a scalar quantity (e.g., 1.0, 1.2, 1.5, 2.0, or the like). Any VAF in the VAF distribution that exceeds the threshold may be classified as a candidate variant.
[0097] In various cases, the variant analyzer 124 classifies each candidate variant as a true positive variant or among the false positive variants 122 by comparing the candidate variants to its neighbors within the VAF distribution. In various implementations, the variant analyzer 124 defines threshold ranges of genomic positions within the VAF distribution that include each candidate variant. These threshold ranges may include a group of adjacent genomic positions within the VAF distribution. In cases in which the VAF distribution excludes one or more types of genomic positions within the genomic region, the group may include genomic positions that are not adjacent to one another in the sequence read data 120 and / or reference genome, but are nevertheless directly adjacent to one another in the VAF distribution. In various cases, the threshold ranges include a predetermined number of genomic positions, such as 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, or 15 adjacent genomic positions within the VAF distribution. For instance, an example threshold range (or group) includes 10 nearest-neighbor CpG sites. In some cases, the threshold range includes a percentage of genomic positions of the VAF distribution, such as between 5% and 20% of genomic positions in the VAF distribution.
[0098] The variant analyzer 124, for instance, determines that one or more of the false positive variants 122 are included in the threshold range (i.e., group) by determining that greater than a threshold number of genomic positions within the threshold range correspond to candidate variants. In some cases, the threshold number is 100%, 80%, 70%, 60%, 50%, or 40% of the total number of genomic positions within the threshold range. According to some cases, the variant analyzer 124 is configured to identifies any candidate variant within the threshold range as among the false positive variants 122 if more than the threshold number of genomic positions within the threshold range correspond to candidate variants. In some cases, the variant analyzer 124 identifies a center genomic position within the threshold range as corresponding to the false positive variants 122 based on determining that it is surrounded by greater than the threshold number of candidate variants.
[0099] The variant analyzer 124, in various implementations, analyzes each threshold range (including overlapping threshold ranges) of genomic positions in the VAF distribution in this manner, for instance. In various cases, the variant analyzer 124 is configured to determine whether each candidate variant represented by the VAF distribution is a false positive variant or a true positive variant.
[0100] The analysis of the threshold ranges, as well as the thresholding, addresses distinct physiological characteristics of the nucleic acid molecules 106 and the degraded nucleic acid molecules 112. In the case of uracil- DNA glycosylase as the corrective enzyme 116, false positives are expected when methylated cytosines in the nucleic acid molecules 106 are deanimated during storage. Methylation of cytosines, in general, tends to occur in continuous and extended segments throughout the genome. Thus, if the nucleic acid molecules 106 include methylated cytosines, these methylated cytosines are expected to manifest in groups. As a result, false positive thiamines transformed from the methylated cytosines would similarly occur in groups. By analyzing a range including adjacent cytosine-to-thiamine substitution candidate variants in the sequence read data 120, the variant analyzer 124 is able to leverage known characteristics of genome-wide methylation patterns to distinguish between true positive cytosine-to-thiamine substitutions and false positive cytosine-to-thiamine substitutions.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0101] Candidate variants in the sequence read data 120 that do not satisfy the conditions of the false positive variants 122, for instance, include additional variants 126. In various cases, the additional variants 126 include ground truth variants that were originally present in the nucleic acid molecules 106 of the fresh sample 104. The additional variants 126 omit the false positive variants 122, for instance.
[0102] In particular implementations, the false positive variants 122 can be used to derive the methylation states of cytosines in the nucleic acid molecules 106. As previously noted, the false positive variants 122 can indicate methylated cytosines in the nucleic acid molecules 106 that are degraded into thiamines in the degraded nucleic acid molecules 112 and that remain in the corrected nucleic acid molecules 114. Thus, the false positive variants 122 may be interpreted as methylated cytosines. In various cases, the variant analyzer 124 infers the methylation status of the nucleic acid molecules 106 without the sequencer 118 performing methylation sequencing on the corrected nucleic acid molecules 114.
[0103] In various cases, a report generator 128 is configured to generate a report 130 based, at least in part, on the additional variants 126 and / or the methylation status. In various implementations, the report generator 128 identifies one or more characteristics of the sequence read data 120 and / or the subject 102 by analyzing the additional variants 126 and / or the methylation status.
[0104] According to various cases, the additional variants 126 and / or methylation status are relevant to the condition of the subject 102. In some cases, the condition is a pathological condition. For instance, the additional variants 126 and / or methylation status indicate whether the subject 102 has cancer, or may indicate a type or subtype of the cancer of the subject 102. The additional variants 126 and / or methylation status may correspond to one or more genes relevant to the classification of the condition of the subject 102. Examples of genes with potential relevance to a determination of whether the subject 102 has a type or subtype of cancer include ABL1, ACVR1 B, AKT1 , AKT2, AKT3, ALK, ALOX12B, AMER1 , APC, AR, ARAF, ARFRP1 , ARID1A, ASXL1 , ATM, ATR, ATRX, AURKA, AURKB, AXIN1 , AXL, BAP1 , BARD1 , BCL2, BCL2L1 , BCL2L2, BCL6, BCOR, BCORL1 , BCR, BRAF, BRCA1 , BRCA2, BRD4, BRIP1 , BTG1 , BTG2, BTK, CALR, CARD11 , CASP8, CBFB, CBL, CCND1 , CCND2, CCND3, CCNE1 , CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1 , CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1 B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1 , CHEK2, CIC, CREBBP, CRKL, CSF1 R, CSF3R, CTCF, CTNNA1 , CTNNB1 , CUL3, CUL4A, CXCR4, CYP17A1 , DAXX, DDR1 , DDR2, DIS3, DNMT3A, DOT1 L, EED, EGFR, EMSY (C11orf30), EP300, EPHA3, EPHB1 , EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1 , ESR1 , ETV4, ETV5, ETV6, EWSR1 , EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1 , FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1 , FLT3, FOXL2, FUBP1 , GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11 , GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1 , HGF, HNF1A, HRAS, HSD3B1 , ID3, IDH1 , IDH2, IGF1 R, IKBKE, IKZF1 , INPP4B, IRF2, IRF4, IRS2, JAK1 , JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1 , KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1 , MAP2K2, MAP2K4, MAP3K1 , MAP3K13, MAPK1 , MCL1 , MDM2, MDM4, MED12, MEF2B, MEN1 , MERTK, MET, MITF, MKNK1 , MLH1 , MPL, MRE11A, MSH2, MSH3, MSH6, MST1 R, MTAP, MTOR, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, NBN, NF1 , NF2, NFE2L2, NFKBIA, NKX2-1 , NOTCH1 , NOTCH2, NOTCH3, NPM1 , NRAS, NT5C2, NTRK1 , NTRK2, NTRK3, NUTM1 , P2RY8, PALB2, PARK2, PARP1 , PARP2, PARP3, PAX5, PBRM1 , PDCD1 , PDCD1 LG2, PDGFRA, PDGFRB, PDK1 , PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1 , PIM1 , PMS2, POLD1 , POLE, PPARG, PPP2R1A,FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCTPPP2R2A, PRDM1 , PRKAR1A, PRKCI, PTCH1 , PTEN, PTPN11 , PTPRO, QKI, RAC1 , RAD21 , RAD51 , RAD51 B, RAD51C, RAD51 D, RAD52, RAD54L, RAF1 , RARA, RB1 , RBM10, REL, RET, RICTOR, RNF43, ROS1 , RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1 , SGK1 , SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1 , SMO, SNCAIP, SOCS1 , SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11 , SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1 , TSC2, TYRO3, U2AF1 , VEGFA, VHL, WHSC1 , WHSC1L1, WT1 , XPO1 , XRCC2, ZNF217, or ZNF703. In some cases, the genes include at least one estrogen receptor (ER) gene and / or at least one progesterone receptor (PR) gene. In some cases, the genes include one or more of ABL, ALK, ALL, B4GALNT1 , BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1 , CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1 , HER2, HR, IDH2, IL-1 |3, IL-6, IL-6R, JAK1 , JAK2, JAK3, KIT, KRAS, MEK, MET, mTOR, PARP, PD-1 , PDGFR, PDGFRo, PDGFRp, PD-L1 , PI3K5, PIGF, PTCH, RAF, RANKL, RET, ROS1 , SLAMF7, VEGF, VEGFA, or VEGFB. In some examples, the genes include one or more of TP53, CTNNNB1 , L1CAM, PTEN, POLE, MKI67, FAT3, TAF1 , ZFHX3, RPL22, SPTA1 , FAM135B, CSMD3, GIGYF2, CSDE1 , MLL4, ATR, CTNNB1 , USH2A, LIMCH1 , RRN3P2, FBXW7, CDH19, USP9X, COL11A1 , BOOR, ARID1A, ZNF770, ARID5B, SLC9A11 , KRAS, PNN, INPP4A, CTCF, CHD4, AMY2B, RBMX, PPP2R1A, TNFAIP6, PIK3R1 , SGK1 , HOXA7, METTL14, HPD, MIR1277, CCND1 , MECOM, NFE2L2, or ESR1.
[0105] In some examples, the additional variants 126 and / or methylation status is utilized to generate a copy number state of one or more genetic loci indicated by the sequence read data 120. In various implementations, a number of copies of a predetermined sequence at a given locus in the genome of the subject 102 and / or the lesion (also referred to as a "copy number” of the locus) is determined. The copy number state, in various implementations, may indicate copy numbers of one or more loci in the genome of the subject 102 and / or the lesion. For instance, the copy number state may indicate the presence and / or amount of copies of various sequences present in the genome of the subject 102 and / or the lesion, which may be due to copy number variation.
[0106] In some examples, the additional variants 126 and / or methylation status is used to generate a mutation signature. In various cases, a mutational signature can represent an amount and / or identity of mutations (e.g., insertions, deletions, double-base substitutions, single-base substitutions, or any combination thereof) indicated in the nucleic acid molecules 106 from the subject 102. In some cases, the mutational signature indicates an amount (e.g., number or percentage) of individual classes of base substitutions present in the nucleic acid molecules 106. For instance, the classes include single-base substitutions including C>A, C>G, C>T, T>A, T>C, and T>G. A mutational signature can be derived by comparing the sequences indicated in the sequence read data 120 to at least one reference sequence, such as a reference genome. For example, the report generator 128 may determine a Catalogue Of Somatic Mutations In Cancer (COSMIC) mutational signature, such as a COSMIC indel signature. In some cases, report generator 128 determines a single-base substitution signature.
[0107] In some cases, the additional variants 126 and / or methylation status is used to identify a tumor mutational burden (TMB) score of the subject 102. Tumor mutational burden (TMB) is a measure of the number of mutations carried by tumor ceiis. By comparing DNA sequences from a patient’s healthy tissues and tumor cells, the number of acquired somatic mutations present in tumors, but not in normal tissues, may be determined. In some instances, driver mutations may be excluded from a TMB calculation. In certain examples, "tumor mutational burden" or “TMB score"FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT refers to the number of somatic mutations in a tumor's genome and / or the number of somatic mutations per area of the tumor's genome. In some embodiments, TMB, as used herein, refers to the number of somatic mutations per megabase (Mb) of DNA sequenced. In some embodiments, germline (inherited) variants are excluded when determining TMB, given that the immune system has a higher likelihood of recognizing these as self. In addition, germline variants do not reflect the biology of somatic mutation for the purposes of TMB determinations. In various cases, driver mutations are excluded from a TMB calculation.
[0108] In some cases, the additional variants 126 and / or methylation status is used to determine the presence, amount, type, or any combination thereof, of one or more hotspot mutations. Hotspots, for instance, can refer to loci in the genome of the subject 102 and / or the lesion that are prone to mutation. Examples of hotspots include CpG islands, microsatellites, centromeric DNA, telomers, subtelomeric regions, common fragile sites, palindromic AT-rich repeats (PATRRs), G-quadruplexes, R-loops, and the like.
[0109] Hotspot mutations give rise to oncological outcomes. PhyloP, SIFT, Grantham, COSMIC and PolyPhen-2 are in silico tools that can be used to assess pathogenicity of identified variants. Exemplary hotspot genes and mutations include EGFR exon 19 activating mutation, EGFR exon 19 deletion, EGFR exon 19 insertion, EGFR exon 19 sensitizing mutation, EGFR exon 20 activation mutation, EGFR exon 20 insertion, EGFR G719 mutation, EGFR L858R mutation, EGFR L861 mutation, EGFR S768 mutation, EGFR T790M mutation, C797 mutation, KIT activating mutation, KRAS activating mutation, MET activating mutation, NRAS activating mutation, PMS2 promoter mutations, among many others. Hotspot mutations also occur in the following genes: AKT2, BRCA1 , BRCA2, ERC1 , NSD1 , POLH, PPM1 G, PTEN, RAD18, RAD51 , RAD51 B, RB1 , TERT, TP53, TP53Bp1 , ALK, ARMT1 , ATAD5, ATG7, ATIC, AXL, BIRC6, BRD3, BRD4, CAPRIN1 , CCAR2, CCDC6, CDK5RAP2, CHD9, CIT, CTNNB1 , CUL1 , EBF1, EIF3E, HIP1 , HMGA2, IRF2BP2, NOTCH1 , NOTCH4, NPM1 , OFD1 , TACC1 , TACC3, TERF2, TMEM106B, UBE2L3, USP10, WRDR48, YAP1 , ZEB2, and ZMYND8.
[0110] One or more of the characteristics described herein can be utilized to generate the report 130. The report 130, for example, includes consumable data that can inform a care provider (also referred to as a "healthcare provider”) about the predicted condition of the subject 102. The report 130, in various cases, includes a genomic profile obtained from one or more nucleic acid sequencing-based tests. For instance, the genomic profile includes results from a comprehensive genomic profiling test. In various implementations, the report 130 may indicate the results of additional analyses, such as the results of a histological study, whole transcriptome sequencing, cfRNA sequencing, whole exome sequencing (WES), whole genome sequencing, a cancer (e.g., DNA) hotspot panel test, a DNA methylation test, a TMB test, a DNA fragmentation test, an RNA fragmentation test, or a viral status test. The performance of such tests is within the ordinary skill of the art, with additional detail provided elsewhere herein. The report 130, for example, may include a genomic profile of the subject 102 based on various combinations of the above analyses and tests.
[0111] In some implementations, the report 130 indicates that a follow-up test of the subject 102 is indicated. For instance, in response to determining that the categorization of the condition of the subject 102 is inconclusive, the report generator 128 may generate the report 130 to indicate that one or more additional tests (e.g., a histological study, genome sequencing, exome sequencing, additional DNA sequencing, RNA sequencing, transcriptome sequencing, etc.) should be performed in order to accurately identify the condition of the subject 102.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0112] In various cases, the report 130 is output to a clinical device 132. For example, the report generator 128 transmits the report 130 to the clinical device 132. In various implementations, the clinical device 132 is a computing device that is operated by, owned by, or otherwise associated with the care provider. For instance, the clinical device 132 may be a desktop computer, a laptop computer, a smart phone, or some other computing device associated with the care provider. The clinical device 132, in various cases, outputs the report 130 to the care provider. In some cases, the clinical device 132 includes a display (e.g., a screen) that visually presents the report 130. In various cases, the clinical device 132 includes a speaker that outputs a sound indicative of the report 130. The clinical device 132, in various cases, may output the information in the report 130 using one or more output mechanisms or devices.
[0113] The care provider may review the report 130 by interacting with the clinical device 132. The report 130, in various cases, may enhance the clinical decision-making of the care provider. For instance, the care provider may prepare and / or administer a therapy to the subject 102 based on the report 130. According to various implementations, the care provider may initiate the therapy and / or refer the subject 102 to another care provider to receive the therapy. In various cases, if the predicted condition of the subject 102 is a disease (e.g., cancer), the care provider may prescribe, recommend, or administer an agent in order to treat the disease the subject 102.
[0114] In various implementations, the care provider may develop a diagnosis and / or prognosis of the subject 102 based on the report 130. In various implementations, the care provider may communicate information in the report 130 to the subject 102.
[0115] FIG. 1 illustrates various elements that can be embodied in one or more computing devices. For example, at least a portion of the functions of the sequencer, the variant analyzer 124, the report generator 128, the clinical device 132, or a combination thereof are performed by one or more processors in at least one computing device. Examples of computing devices include server computers, desktop computers, laptop computers, tablet computers, mobile phones, wearable devices, Internet of Things (loT) devices, and the like. In various cases, instructions for performing at least a portion of the functions of these elements are stored in memory and / or in a non-transitory computer-readable medium. The instructions, for instance, are executed by the processor(s).
[0116] FIG. 1 also illustrates various types of data. For example, the sequence read data 120, the false positive variants 122, the additional variants 126, the report 130, or any combination thereof, includes data. The various types of data illustrated in FIG. 1 may be stored, such as in memory or in non-transitory computer-readable media. In various implementations, at least a portion of the data is transmitted or otherwise output by one or more computing devices. For example, a computing device may transmit one or more communication signals to another computing device, wherein the communication signal(s) encode at least a portion of the data. Examples of communication signals include electromagnetic signals, optical signals, ultrasonic signals, optical signals, and electrical signals. For example, communication signals can be transmitted wirelessly and / or in a wired fashion. The communication signals, for instance, are transmitted over one or more wireless channels and / or one or more wired channels (e.g., optical cabling, electrical cabling, etc.). In various cases, the communication signal(s) are transmitted over one or more communication networks. A communication network, for instance, may be defined according to one or more physical channels, such as one or more frequency spectra. In some cases, a communication network is defined according to one or more communication protocols and / or standards. Examples of communication networks include fiber optic networks, Institute of Electrical and Electronics Engineers (IEEE) networks (e.g., WI-FI™ networks, WiMAX networks, BLUETOOTH™FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT networks, etc.), cellular networks (e.g., a 3rdGeneration Partnership Project (3GPP) radio network, such as a Long Term Evolution (LTE) network, a New Radio (NR) network; or a cellular core network such as a 3rdGeneration (3G) core, a 4thGeneration (4G) core, a 5thGeneration (5G) core, etc.), ultrasonic networks, and the like. In some cases, the data is broadcasted from one device to multiple other devices. In some cases, the data is unicasted from one device to another device. For instance, various forms of data described herein may be transmitted via a peer-to-peer (P2P) connection.
[0117] FIGS. 2A and 2B illustrate examples of various ground truth, degraded, and corrected nucleic acid molecules whose sequences can be analyzed for false positive variants using various techniques described herein. FIG. 2A illustrates an example of degradation and correction of unmethylated nucleic acid molecules. For instance, ground truth unmethylated nucleic acid molecules 202 are present in a fresh sample obtained from a subject. The ground truth unmethylated nucleic acid molecules 202, for instance, include a CpG island. The cytosines in the CpG island, for example, are unmethylated in the ground truth unmethylated nucleic acid molecules 202. The ground truth unmethylated nucleic acid molecules 202 are representative of a condition of the subject, in various cases.
[0118] After being stored for an extended period of time, the ground truth unmethylated nucleic acid molecules 202 are converted into degraded unmethylated nucleic acid molecules 204. Specifically, the unmethylated cytosines from the CpG island of the ground truth unmethylated nucleic acid molecules 202 are converted into uracils via deanimation. These uracils represent erroneous nucleotides 206. That is, if the degraded unmethylated nucleic acid molecules 204 are sequenced, the erroneous nucleotides 206 will result in sequencing errors that are not reflective of the sequence of the ground truth unmethylated nucleic acid molecules 202.
[0119] In various cases, the erroneous nucleotides 206 may be chemically corrected by an enzyme. For example, corrected nucleic acid molecules 208 are generated by applying a UNG enzyme to the degraded unmethylated nucleic acid molecules 204. The UNG enzyme may convert the erroneous nucleotides 206 of the degraded unmethylated nucleic acid molecules 204 back into the cytosines of the corrected nucleic acid molecules 208. Thus, degraded cytosines in the ground truth unmethylated nucleic acid molecules 202 can be chemically corrected prior to sequencing.
[0120] FIG. 2B illustrates an example of degradation and correction of methylated nucleic acid molecules, instance, ground truth methylated nucleic acid molecules 210 are present in a fresh sample obtained from a subject. The ground truth methylated nucleic acid molecules 210, for instance, include a CpG island. The cytosines in the CpG island, for example, are methylated in the ground truth methylated nucleic acid molecules 210. The ground truth methylated nucleic acid molecules 210 are representative of a condition of the subject, in various cases.
[0121] After being stored for an extended period of time, the ground truth methylated nucleic acid molecules 210 are converted into degraded methylated nucleic acid molecules 212. Specifically, the methylated cytosines from the CpG island of the ground truth methylated nucleic acid molecules 210 are converted into thiamines via deanimation. These thiamines represent erroneous nucleotides 214. That is, if the degraded unmethylated nucleic acid molecules 204 are sequenced, the erroneous nucleotides 206 will result in sequencing errors that are not reflective of the sequence of the ground truth methylated nucleic acid molecules 210.
[0122] In various cases, the erroneous nucleotides 214 cannot be chemically corrected by an enzyme. For example, if a UNG enzyme is applied to the degraded methylated nucleic acid molecules 212, the erroneous nucleotides 214 remain.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0123] Various implementations of the present disclosure can be used to detect the erroneous nucleotides 214 by processing sequence read data obtained from sequencing the degraded methylated nucleic acid molecules 212.
[0124] According to various cases, a variant analyzer may identify candidate variants corresponding to the erroneous nucleotides 214 by comparing sequence read data representing the degraded methylated nucleic acid molecules 212 to a reference genome. The variant analyzer may infer that the candidate cytosine-to-thiamine variants are the result of the degradation of the methylated cytosines by determining that the candidate cytosine-to-thiamine variants are located proximate to other candidate cytosine-to-thiamine variants, which are more likely to correspond to the degradation of methylated cytosines than individual substitution variants in the ground truth methylated nucleic acid molecules 210. In some cases, the identified erroneous nucleotides 214 are used to infer methylation patterns in the ground truth methylated nucleic acid molecules 210 without the performance of methylation sequencing.
[0125] FIG. 3 illustrates an example VAF distribution 300 utilized to distinguish between false positive and true positive variants. The VAF distribution 300 includes various bars corresponding to the frequency of one or more types of variants at given genomic positions.
[0126] For example, a horizontal axis of the VAF distribution 300 represents at least a subset of genomic positions in a genomic region. In some cases, the genomic positions represented by the VAF distribution 300 are only a portion of the genomic positions in the genomic region. In some cases, the VAF distribution 300 only represents one or more types of base pairs in the genomic region, such as genomic positions that contain only cytosines or guanines within a reference genome. In some cases, the genomic positions reflected in the VAF distribution 300 include CpG sites. The genomic region, in various cases, is only a portion of the reference genome.
[0127] In various cases, the vertical axis of the VAF distribution 300 represents the frequency of one or more types of variants appearing in sequence read data of a sample. For example, the VAF distribution 300 may exclusively represent the frequency of detected cytosine-to-thiamine and / or guanine-to-adenosine substitution variants detected in the sample. In some cases, the vertical axis is in units of percentage, such that the VAF distribution 300 represents the respective percentages that the types of variant(s) appear at each of the genomic positions represented by the horizontal axis. In some cases, the vertical axis is in units of number of detected variants at the given genomic positions.
[0128] According to some cases, a variant is detected at a particular genomic position if the VAF at that genomic position is above a detection threshold. This thresholding may prevent the reporting of erroneous sequence reads or noise as true variants of the sample. However, if the sample was chemically degraded (e.g., due to storage) such that one or more of the nucleotides in the sample are erroneous, these erroneous nucleotides may be erroneously reported as variant using conventional thresholding techniques.
[0129] In various implementations of the present disclosure, true and false positive variants may be distinguished by comparing the VAFs of groups of genomic positions within the sample. In various cases, a maximum 302 VAF is identified. In some cases, the maximum 302 is the maximum VAF in the VAF distribution 300. In some examples, the maximum 302 is a maximum VAF among all detected variants within all genomic positions in the region.
[0130] A threshold 304, for instance, is calculated based on the maximum 302. In some cases, the threshold 304 is a percentage of the maximum 302. For example, the threshold 304 may be 50%, 60%, 70%, or 80% of the maximum 302. In some cases, the threshold 304 is also the detection threshold utilized to detect the presence of one or more types of variants, but implementations are not so limited.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0131] According to some cases, the type(s) of variants reflected by the VAF distribution 300 are candidate variants that include both true positive variants (e.g., ground truth variants present in the nucleic acid molecules from the original sample) as well as false positive variants (e.g., degraded nucleotides that are not present in the nucleic acid molecules from the original sample). In various examples, the false positive variants are the result of chemical degradation of methylated cytosines in the original sample. Because methylated cytosines are generally clustered within samples, the false positive variants are therefore are expected to be present in clusters of genomic positions throughout the sample.
[0132] In various implementations, groups of adjacent genomic positions within the VAF distribution 300 are analyzed together. For example, these groups may include 4, 5, 6, 7, 8, 10, or 16 genomic positions reflected by the VAF distribution 300. It should be noted that in cases where the VAF distribution 300 only reflects a subset of genomic positions in a given region, the "adjacent” genomic positions in the VAF distribution 300 may not be directly adjacent within the reference genome and / or the genome of the sample being analyzed.
[0133] If greater than a threshold number of the genomic positions in a given group are above the threshold 304, one or more of the genomic positions may be associated with a false positive variant. For example, a false positive group 306 in the VAF distribution 300 illustrated in FIG. 3 includes a group of nine genomic positions, each associated with a VAF that is above the threshold 304. However, the probability that these adjacent genomic positions all have the same type of substitution variant (e.g., cytosine-to-thiamine substitution variants in CpG sites) is less than the probability that the adjacent genomic positions are methylated and have degraded into the alternative nucleotides (e.g., methylated cytosines in the CpG sites have degraded into thiamine). In various cases, each genomic position within the false positive group 306 is identified as being associated with a false positive variant. In some examples, one or more of the genomic positions within the false positive group 306 is marked as a false positive variant, such as the center genomic position of the false positive group 306.
[0134] In contrast, a true positive variant 308 is identified at a genomic position whose VAF is above the threshold 304, but whose adjacent genomic positions in the VAF distribution 300 are below the threshold 304. For example, a group of nine adjacent genomic positions including the genomic position of the true positive variant 308 may include a single genomic position whose VAF is above the threshold 304: that of the true positive variant 308. The probability that the true positive variant 308 is reflective of a true substitution variant in the ground truth sample (e.g., that the true positive variant 308 is reflective of a cytosine-to-thiamine substitution in a CpG site) is greater than the probability that the single genomic position was methylated in the original sample and has degraded into the alternative nucleotide (e.g., that the false positive variant 308 corresponds to a single methylated cytosine in a CpG site, when its neighboring cytosines are unmethylated).
[0135] FIG. 4 illustrates an example report 400 summarizing predicted categories of a cancer of a subject. In various cases, the report 400 is the report 130 described above with reference to FIG. 1. The report 400, for instance, may be displayed to a patient and / or care provider. In some cases, the report 400 is generated based on features of a sample obtained from the subject. In some cases, the report 400 is generated based on true positive variants identified in nucleic acid molecules of the sample, and is generated independently of one or more false positive variants identified in the nucleic acid molecules of the sample. In some cases, the report 400 is generated based on methylation statuses inferred from false positive variants identified in the nucleic acid molecules.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0136] The report 400 includes a tissue origin 402 of the cancer. The tissue origin 402, for instance, indicates a histological tissue type 404, a primary site 406, cell subtype 407, or any combination, of the cancer.
[0137] In various cases, the report 400 includes one or more therapy indicators 408. For instance, the therapy indicator(s) 408 convey whether the cancer is predicted to be resistant to one or more predetermined therapies and / or whether the cancer is predicted to be responsive to one or more predetermined therapies.
[0138] In some examples, the report 400 includes one or more prognostic indicators 410. The prognostic indicator(s) 410, for instance, indicate a prognosis of the subject in view of the categorized cancer. For example, the prognostic indicator(s) 410 may indicate a survivability, a recoverability, a quality of life indicator, or other information indicative of the prognosis of the subject.
[0139] The report 400 may include a trial qualification 412 of the subject. The trial qualification 412, for instance, indicates whether the subject is predicted to qualify for a predetermined clinical trial.
[0140] The report 400, in various implementations, includes a metastasis profile 414 of the subject. The metastasis profile 414, for instance, indicates a likelihood that the cancer will metastasize (e.g., at a particular point in time), one or more tissues in which the cancer is predicted to metastasize, or the like.
[0141] In various cases, the report 400 includes recommended follow-up tests 416. For example, the report 400 may include a recommendation to perform whole genome sequencing on the subject, particularly in cases if the cancer cannot be categorized above a threshold certainty.
[0142] The report 400 may include a genomic profile 418 of the subject. In various cases, the genomic profile 418 includes or is generated based on the results of one or more genomic analyses of the subject.
[0143] FIG. 5 illustrates an example process 500 for identifying false positive variants due to the degradation of nucleotides in a sample. The process 500 is, for example, performed by an entity including at least one of: one or more processors, a computing device, a sequencer (e.g., the sequencer 118), the variant analyzer 124, the report generator 128, or any combination thereof.
[0144] At 502, the entity generates a VAF distribution of at least one type of substitution variant by comparing sequence read data to a reference genome. The sequence read data, for instance, is obtained by sequencing nucleic acid molecules in a sample obtained from a subject. In some examples, the sample is stored for an extended period of time (e.g., days, weeks, months, years, etc.) after being extracted from the subject and before being sequenced to obtain the sequence read data. In some cases, the sample was stored at a temperature of greater than 0°C. Therefore, one or more types of nucleotides in the sample may have been chemically degraded before being sequenced. These degraded nucleotides, for instance, include false positive variants. In some implementations, the sample was treated with a corrective enzyme, such as uracil-DNA glycosylase.
[0145] In some cases, the VAF distribution is representative of a region of a genome of the sample. The VAF distribution may reflect only a subset of types of variants that are otherwise indicted by the sequence read data. In particular cases, the VAF distribution is limited to cytosine-to-thiamine substitution variants and / or guanine-to- adenosine substitution variants. That is, the VAF distribution may reflect the frequency that thiamines appear in the sequence read data at genomic positions that are associated with cytosine in the reference genome and / or the frequency that adenosines appear in the sequence read data at genomic positions that are associated with guanines in the reference genome. In some cases, only a subset of genomic positions in the region are represented by the VAFFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT distribution. For instance, the VAF distribution may only include the VAF at genomic positions associated with CpG sites in the reference genome.
[0146] At 504, the entity identifies candidate variants in the VAF distribution. According to various cases, the candidate variants include both false positive and true positive variants. To identify the candidate variants, the entity may compare the VAF distribution to a threshold. In some cases, the threshold is a percentage of a maximum VAF (e.g., a "peak” VAF) in the VAF distribution, or a maximum VAF (e.g., a "peak” VAF) in a broader VAF distribution that is not limited to one or more types of substitution variants or predetermined genomic positions. If a VAF at a particular genomic position in the VAF distribution is above the threshold, it may be classified as a candidate variant.
[0147] At 506, the entity determines that at least a portion of the candidate variants are false positive variants. Specifically, in some cases, the entity may analyze groups of adjacent genomic positions within the VAF distribution. These groups may be referred to as a "threshold range” of genomic positions and / or the VAF distribution. For instance, the entity may analyze genomic positions associated with the next-closest CpG sites, as defined by the reference genome. In some cases, the entity analyzes groups of 4, 5, 6, 7, 8, 9, 10, or 15 adjacent genomic positions in the VAF distribution.
[0148] The entity may determine how many candidate variants are present within each group (e.g., threshold range) of genomic positions. If greater than a threshold number of candidate variants are present within a given group (e.g., threshold range), the entity may infer that one or more of the candidate variants are false positive variants. Accordingly, the entity may refrain from reporting, or from generating a genomic report, based on the false positive variants. In various implementations, the entity reports, or generates a genomic report, based on additional variants that omit the false positive variants. For example, the entity may determine that the subject has a condition, such as a pathological condition, based on the additional variants. In some cases, the entity identifies whether the subject has cancer, or may classify a type of cancer that the subject has, based on the additional variants. In various implementations, the entity identifies a treatment for the pathological condition, a dosage of the treatment, or the like, based on the additional variants. In some examples, the report indicates the false positive variants.
[0149] In some cases, the entity may infer that the false positive variants are actually representative of methylation patterns within the ground truth sample, prior to degradation. For example, the entity may report, or generate the report, based on the inference that the false positive variants correspond to methylated cytosines in the ground truth sample.
[0150] FIG. 6 illustrates an example environment 600 for sequencing various nucleic acid molecules 602. In various implementations, the nucleic acid molecules 602 include cfDNA and / or gDNA. For instance, the nucleic acid molecules 602 may include ctDNA. The nucleic acid molecules 602, in various cases, are extracted from a sample, such as a biological sample obtained from a subject. In some implementations, the nucleic acid molecules 602 include DNA that is complementary to RNA present in the sample.
[0151] The nucleic acid molecules 602, in various cases, are ligated with adapters 604. For examples, the adapters 604 are hybridized to the nucleic acid molecules 602. The adapters 604, for example, include additional nucleic acid molecules. In various implementations, the adapters 604 have a shorter length than the nucleic acid molecules 602 being sequenced. For instance, the adapters 604 include amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences. Although FIG. 6 illustrates adapters 604 being ligated to one end ofFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT each of the nucleic acid molecules 602, implementations are not so limited. For example, the adapters 604 may be ligated to both ends of each of the nucleic acid molecules 602.
[0152] In various examples, the nucleic acid molecules 602 ligated with the adapters 604 are amplified in order to generate amplified molecules 606. Various amplification techniques can be performed. For instance, the amplified molecules 606 are generated using PGR, a non-PCR amplification technique, an isothermal amplification technique, or any combination thereof.
[0153] Amplified molecules 606 may be captured by bait molecules 610 and sequenced. In some implementations, the amplified molecules 606 are sequenced via sequencing-by-synthesis. In various cases, fluorescently tagged deoxyribonucleotide triphosphates (dNTP) 612 are utilized to synthesize a strand that is complementary to DNA strands bound to the substrate 608. When a dNTP 612 is added to the strand (e.g., by an enzyme), the dNTP 612 emits an optical signal 614. In various implementations, the frequency of the optical signal 614 is dependent on the type of dNTP 612 from which the optical signal 614 is emitted. By detecting the optical signals 614 as the strand is being synthesized, the sequence of the original nucleic acid molecules 602 can be derived.
[0154] In some implementations, the amplified molecules 606 are sequenced via nanopore sequencing. For instance, the amplified molecules 606 are directed through a nanopore 616 extending through a substrate 618. In various cases, the amplified molecules 606 are negatively charged, such that they can be directed through the nanopore 616 by imposing an electrical field across the substrate 618. In various cases, the amplified molecules 606 and the nanopore 616 are in the presence of a charged solution. Thus, charged solutes traveling through the nanopore 616 can be monitored by reviewing an electrical signal (e.g., a current) sensed between electrodes 620 on either side of the substrate 618. As an amplified molecule 606 is directed through the nanopore 616, the individual bases within the amplified molecule 606 will block the nanopore 616, which may decrease the amount of charged solutes traveling through the nanopore 616 and consequently, the magnitude of the electrical signal detected by the electrodes 620. Each of the four types of bases within the amplified molecules 606, may block the nanopore 616 to a different extent. Therefore, the sequence of the nucleic acid molecules 602 can be derived by analyzing the measured electrical signal with respect to time as the amplified molecules 606 are directed through the nanopore 616.
[0155] FIG. 7 illustrates one or more devices 700 configured to perform various operations described herein. The device(s) 700 include one or more processor(s) 702. In some implementations, the processor(s) 702 includes a central processing unit (CPU), a graphics processing unit (GPU), both CPU and GPU, or other processing unit or component known in the art.
[0156] The processor(s) 702 is operably connected to memory 704. In various implementations, the memory 704 is volatile (such as random access memory (RAM)), non-volatile (such as read only memory (ROM), flash memory, etc.) or some combination of the two. The memory 704 stores instructions that, when executed by the processor(s) 702, causes the processor(s) 702 to perform various operations. In various examples, the memory 704 stores methods, threads, processes, applications, objects, modules, any other sort of executable instruction, or a combination thereof. In some cases, the memory 704 stores files, databases, or a combination thereof. In some examples, the memory 704 includes, but is not limited to, RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory, or any other memory technology. In some examples, the memory 704 includes one or more of CD-ROMs, digital versatile discs (DVDs), content-addressable memory (CAM), or other optical storage, magnetic cassettes,FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by the processor(s) 702. For instance, the memory 704 stores instructions that, when executed by the processor(s) 702, causes the processor(s) 702 to perform operations of the variant analyzer 124 and / or the report generator 128.
[0157] The processor(s) 702 is operably connected to one or more input devices 706 and one or more output devices 708. Collectively, the input device(s) 706 and the output device(s) 708 function as an interface between at least one user and the device(s) 700. The input device(s) 706 is configured to receive an input from a user and includes at least one of a keypad, a cursor control, a touch-sensitive display, a voice input device (e.g., a microphone), a haptic feedback device (e.g., a gyroscope), or any combination thereof. The output device(s) 708 includes at least one of a display, a speaker, a haptic output device, a printer, or any combination thereof. In various examples, the processor(s) 702 causes a display among the input device(s) 706 to visually output various data described herein. In some implementations, the input device(s) 706 includes one or more touch sensors, the output device(s) 708 includes a display screen, and the touch sensor(s) are integrated with the display screen.
[0158] In various implementations, the processor(s) 702 is operably connected to one or more transceivers 710 that transmit and / or receive data over one or more communication networks 712. For example, the transceiver(s) 710 includes a network interface card (NIC), a network adapter, a local area network (LAN) adapter, or a physical, virtual, or logical address to connect to the various external devices and / or systems. In various examples, the transceiver(s) 710 includes any sort of wireless transceivers capable of engaging in wireless communication (e.g., radio frequency (RF) communication). For example, the communication network(s) 712 includes one or more wireless networks that include a 3rd Generation Partnership Project (3GPP) network, such as a Long Term Evolution (LTE) radio access network (RAN) (e.g., over one or more LTE bands), a New Radio (NR) RAN (e.g., over one or more NR bands), or a combination thereof. In some cases, the transceiver(s) 710 includes other wireless modems, such as a modem for engaging in WI-FI®, WIGIG®, WIMAX®, BLUETOOTH®, or infrared communication over the communication network(s) 712.
[0159] The device(s) 700 may further include the sequencer 118. In various implementations, the sequencer 118 includes one or more fluidic circuits 714 configured to receive a sample 716 derived from a subject 717. The sequencer 118, in various cases, may be configured to generate data indicative of one or more sequences of nucleic acid molecules (e.g., DNA and / or RNA) present in the sample 716. In various cases, the sequencer 118 introduces one or more reagents 718 to the fluidic circuit(s) 714 in order to prepare for and perform sequencing of the nucleic acid molecules. Further, the sequencer 118 may include one or more sensors 720 configured to measure or otherwise detect detection signals from the fluidic circuit(s) 714, which may be indicative of the sequences of the nucleic acid molecules. According to various implementations, the sensor(s) 720 may further include one or more ADCs. The sequencer 118, in various cases, outputs sequence read data to the processor(s) 702 for additional processing.Example Clauses
[0160] The following clauses provide various examples of the present disclosure.1 . A method, including: providing a plurality of nucleic acid molecules obtained from a sample from a subject; treating the plurality of nucleic acid molecules with uracil-DNA glycosylase; ligating one or more adapters onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying the one or more ligated nucleicFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT acid molecules from the plurality of nucleic acid molecules; capturing amplified nucleic acid molecules from the amplified nucleic acid molecules; sequencing, by a sequencer, all or a subset of the captured amplified nucleic acid molecules to obtain a plurality of sequence reads that represent the sequenced amplified nucleic acid molecules thereby generating sequence read data; receiving, at one or more processors, the sequence read data for the plurality of sequence reads; generating, by the one or more processors, a variant allele frequency (VAF) distribution of cytosine to thymine and guanine to adenosine substitutions in CpG sites of the sequence read data by comparing the sequence read data to a reference genome; identifying candidate variants in the VAF distribution; determining that at least a portion of candidate variants in the VAF distribution are false positive variants by determining that greater than a threshold number of the candidate variants are present within a threshold range of the VAF distribution; and generating a report based on one or more additional variants indicated by the sequence read data, the one or more additional variants omitting the false positive variants.2. The method of clause 1, wherein providing the plurality of nucleic acid molecules obtained from a sample from the subject is in response to storing the sample for a year or more.3. The method of clause 1 or 2, wherein treating the plurality of nucleic acid molecules with uracil-DNA glycosylase includes converting uracils in the nucleic acid molecules to cytosines.4. The method of any of clauses 1 to 3, wherein the threshold number of the candidate variants is about 10, and wherein the threshold range of the VAF distribution includes about 10% of the VAF distribution.5. The method of any of clauses 1 to 4, further including: determining that at least a portion of the false positive variants are indicative of ground truth methylated cytosines in the plurality of nucleic acid molecules, wherein generating the report is further based on the ground truth methylated cytosines.6. The method of any of clauses 1 to 5, further including: performing Gaussian smoothing on the VAF.7. A method, including: receiving sequence read data of nucleic acid molecules in a sample treated with a corrective enzyme after being obtained from a subject; generating a variant allele frequency (VAF) distribution of at least one type of substitution variant of the sequence read data by comparing the sequence read data to a reference genome; identifying candidate variants in the VAF distribution; determining that at least a portion of candidate variants in the VAF distribution are false positive variants by determining that greater than a threshold number of the candidate variants are present within a threshold range of the VAF distribution; and generating a report based on one or more additional variants indicated by the sequence read data, the one or more additional variants omitting the false positive variants.8. The method of clause 7, wherein the sample was stored after being obtained from the subject and before being sequenced to generate the sequence read data.9. The method of clause 8, wherein the sample was stored for greater than 1 day, one week, one month, or one year.10. The method of clause 8 or 9, wherein the sample was stored in a frozen form.11 . The method of any of clauses 8 to 10, wherein the sample was exposed to a temperature of greater than 0° Celsius during storage for greater than a threshold amount of time.12. The method of any of clauses 7 to 11, wherein the nucleic acid molecules in the sample were degraded after being obtained from the subject and before being sequenced to generate the sequence read data.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT13. The method of any of clauses 7 to 12, wherein a portion of ground truth nucleotides of the nucleic acid molecules were converted erroneous nucleotides after being obtained from the subject.14. The method of clause 13, wherein the ground truth nucleotides include cytosines, and wherein the erroneous nucleotides include at least one of uracils or thymines.15. The method of clause 13 or 14, wherein the ground truth nucleotides include methylated cytosines, and wherein the erroneous nucleotides include thymines.16. The method of any of clauses 13 to 15, wherein the ground truth nucleotides include unmethylated cytosines, and wherein the erroneous nucleotides include uracils.17. The method of any of clauses 7 to 16, wherein the corrective enzyme includes uracil-DNA glycosylase.18. The method of any of clauses 7 to 17, wherein the corrective enzyme is configured to convert uracils in nucleic acid molecules into cytosines.19. The method of any of clauses 7 to 18, wherein the VAF distribution includes candidate variants in CpG regions of the sample.20. The method of any of clauses 7 to 19, wherein the VAF distribution consists of candidate variants in CpG regions of the sample.21 . The method of any of clauses 7 to 20, wherein the at least one type of substitution variant includes at least one of: cytosine to thymine variants; or guanine to adenosine variants.22. The method of any of clauses 7 to 21 , wherein identifying the candidate variants in the VAF distribution includes: detecting a peak of the VAF distribution; and identifying the candidate variants by identifying locations in the VAF distribution that correspond to greater than a threshold percentage of the peak of the VAF distribution.23. The method of clause 22, wherein the threshold percentage is in a range of about 60% to about 90%.24. The method of clause 22 or 23, wherein the threshold number of the candidate variants is in a range of about 5 to about 100, and wherein the threshold range of the VAF distribution is less than 20% of the VAF distribution.25. The method of any of clauses 22 to 24, wherein the threshold number of the candidate variants includes about 10, and wherein the threshold range of the VAF distribution is about 10% of the VAF distribution.26. The method of any of clauses 7 to 25, further including: predicting, based on the sequence read data, a condition of the subject, wherein the report indicates the condition of the subject.27. The method of clause 26, wherein the condition includes a pathological condition.28. The method of clause 26 or 27, wherein the condition includes a cancer type or cancer subtype of the subject.29. The method of any of clauses 7 to 28, further including: predicting, based on the sequence read data, an effective therapy to treat a condition of the subject, wherein the report indicates the effective therapy.30. The method of clause 29, wherein the effective therapy includes a dosage of one or more therapeutic agents predicted to treat the condition of the subject.31 . The method of any of clauses 7 to 30, wherein the report includes an indication of the false positive variants.32. The method of clause 31 , wherein the report includes a genomic profile, the genomic profile including results from at least one of: a comprehensive genomic profiling test; a whole genome sequencing (WGS) test; a whole exome sequencing (WES) test; a gene expression profiling test; a cancer hotspot panel test; a DNA methylation test; a DNA fragmentation test; or an RNA fragmentation test.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT33. The method of clause 32, wherein the genomic profile of the subject includes: results from a nucleic acid sequencing-based test.34. The method of clause 32 or 33, further including: selecting, based on the genomic profile, an anticancer agent for administration to the subject.35. The method of clause 34, further including: administering the anticancer agent to the subject.36. The method of any of clauses 32 to 35, further including: applying, based on the genomic profile, an anticancer therapy to the subject.37. The method of clause 36, wherein the anticancer therapy includes at least one of chemotherapy, radiation therapy, immunotherapy, a targeted therapy, or surgery.38. The method of any of clauses 7 to 37, wherein the report indicates that the sample has been degraded.39. The method of any of clauses 7 to 38, further including: determining that the false positive variants correspond to ground truth methylated cytosines in the nucleic acid molecules in the sample, wherein the report indicates the ground truth methylated cytosines.40. The method of any of clauses 7 to 39, further including: outputting the report.41 . The method of clause 40, wherein outputting the report includes: transmitting data indicating the report to an external device.42. The method of clause 41 , wherein the external device is associated with the subject and / or a healthcare provider.43. The method of clause 41 or 42, wherein the data is transmitted over a peer-to-peer connection and / or over one or more communication networks.44. The method of any of clauses 40 to 43, wherein outputting the report includes: visually presenting, by a display, the report.45. The method of any of clauses 7 to 44, further including: receiving the nucleic acid molecules obtained from the sample; administering the corrective enzyme to the nucleic acid molecules; ligating one or more adapters onto one or more nucleic acid molecules from the nucleic acid molecules; amplifying the one or more ligated nucleic acid molecules; capturing all or a subset of the amplified nucleic acid molecules; and sequencing, by a sequencer, the captured nucleic acid molecules to obtain a plurality of sequence reads that represent the captured nucleic acid molecules, thereby generating the sequence read data.46. The method of clause 45, wherein the one or more adapters include amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences.47. The method of clause 45 or 46, wherein the captured nucleic acid molecules are captured from the amplified nucleic acid molecules by hybridization to one or more bait molecules.48. The method of clause 47, wherein the one or more bait molecules include one or more additional nucleic acid molecules, each of the one or more additional nucleic acid molecules including a region that is complementary to a region of a captured nucleic acid molecule.49. The method of any of clauses 45 to 48, wherein amplifying the one or more ligated nucleic acid molecules includes performing a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT50. The method of any of clauses 45 to 49, wherein sequencing the captured nucleic acid molecules includes use of a massively parallel sequencing (MPS) technique, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, direct sequencing, or Sanger sequencing.51 . The method of any of clauses 45 to 50, wherein sequencing the captured nucleic acid molecules includes next-generation sequencing (NGS).52. The method of any of clauses 45 to 51 , wherein the sequencer includes a next-generation sequencer.53. The method of any of clauses 45 to 52, wherein sequencing the captured nucleic acid molecules includes sequencing-by-synthesis or nanopore sequencing.54. The method of any of clauses 7 to 53, further including: administering the corrective enzyme to the nucleic acid molecules; generating ligated molecules by ligating adapters onto the nucleic acid molecules; generating amplified ligated molecules by amplifying the ligated molecules; generating, using the amplified ligated molecules, detection signals; detecting, by at least one sensor, the detection signals; and generating the sequence read data based on the detection signals.55. The method of clause 54, wherein the detection signals include electrical signals and / or optical signals.56. The method of clause 54 or 55, wherein generating, using the amplified ligated molecules, the detection signals includes: synthesizing, by a polymerase using fluorescently tagged nucleotide triphosphates (NTPs), a synthesized nucleic acid molecule that is complementary to one of the amplified ligated molecules, and wherein detecting, by the at least one sensor, the detection signals includes: detecting, by at least one optical sensor, optical signals emitted by the fluorescently tagged NTPs upon binding to the synthesized nucleic acid molecule, the optical signals being indicative of at least one sequence of the nucleic acid molecules of the sample.57. The method of any of clauses 54 to 56, wherein generating, using the amplified ligated molecules, the detection signals includes: directing the amplified ligated molecules through a nanopore extending from a first space to a second space through a substrate, and wherein detecting, by the at least one sensor, the detection signals includes: detecting, by sensors disposed in the first space and the second space, an electrical signal over time, the electrical signal being indicative of at least one sequence of the nucleic acid molecules of the sample.58. The method of any of clauses 54 to 57, wherein the sequence read data indicates a whole genome or RNA transcriptome of the sample.59. The method of any of clauses 54 to 58, wherein the sequence read data indicates a whole exome of the sample.60. The method of any of clauses 54 to 59, wherein the sequence read data indicates a predetermined panel of genes of the sample.61. The method of clause 60, wherein the predetermined panel includes one or more of ABL1, ACVR1 B, AKT1, AKT2, AKT3, ALK, ALOX12B, AMER1, APO, AR, ARAF, ARFRP1, ARID1A, ASXL1, ATM, ATR, ATRX, AURKA, AURKB, AXIN1 , AXL, BAP1, BARD1, BCL2, BCL2L1, BCL2L2, BCL6, BOOR, BCORL1 , BCR, BRAF, BRCA1, BRCA2, BRD4, BRIP1 , BTG1 , BTG2, BTK, CALR, CARD11, CASP8, CBFB, CBL, CCND1 , CCND2, CCND3, CCNE1 , CD22, CD274, CD70, CD74, CD79A, CD79B, CDC73, CDH1 , CDK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1 B, CDKN2A, CDKN2B, CDKN2C, CEBPA, CHEK1, CHEK2, CIC, CREBBP, CRKL, CSF1 R, CSF3R, CTCF, CTNNA1 , CTNNB1 , CUL3, CUL4A, CXCR4, CYP17A1, DAXX, DDR1, DDR2, DIS3, DNMT3A, DOT1 L, EED, EGFR, EMSY (C11orf30),FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCTEP300, EPHA3, EPHB1 , EPHB4, ERBB2, ERBB3, ERBB4, ERCC4, ERG, ERRFI1 , ESR1 , ETV4, ETV5, ETV6, EWSR1, EZH2, EZR, FAM46C, FANCA, FANCC, FANCG, FANCL, FAS, FBXW7, FGF10, FGF12, FGF14, FGF19, FGF23, FGF3, FGF4, FGF6, FGFR1, FGFR2, FGFR3, FGFR4, FH, FLCN, FLT1 , FLT3, F0XL2, FUBP1 , GABRA6, GATA3, GATA4, GATA6, GID4 (C17orf39), GNA11, GNA13, GNAQ, GNAS, GRM3, GSK3B, H3F3A, HDAC1 , HGF, HNF1A, HRAS, HSD3B1, ID3, IDH1 , IDH2, IGF1 R, IKBKE, IKZF1 , INPP4B, IRF2, IRF4, IRS2, JAK1 , JAK2, JAK3, JUN, KDM5A, KDM5C, KDM6A, KDR, KEAP1, KEL, KIT, KLHL6, KMT2A (MLL), KMT2D (MLL2), KRAS, LTK, LYN, MAF, MAP2K1 , MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MCL1, MDM2, MDM4, MED12, MEF2B, MEN1 , MERTK, MET, MITF, MKNK1 , MLH1 , MPL, MRE11A, MSH2, MSH3, MSH6, MST1 R, MTAP, MTOR, MUTYH, MYB, MYO, MYCL, MYCN, MYD88, NBN, NF1, NF2, NFE2L2, NFKBIA, NKX2-1 , NOTCH1 , NOTCH2, NOTCH3, NPM1 , NRAS, NT5C2, NTRK1 , NTRK2, NTRK3, NUTM1, P2RY8, PALB2, PARK2, PARP1 , PARP2, PARP3, PAX5, PBRM1 , PDCD1, PDCD1LG2, PDGFRA, PDGFRB, PDK1, PIK3C2B, PIK3C2G, PIK3CA, PIK3CB, PIK3R1, PIM1, PMS2, POLD1, POLE, PPARG, PPP2R1A, PPP2R2A, PRDM1 , PRKAR1A, PRKCI, PTCH1, PTEN, PTPN11, PTPRO, QKI, RAC1, RAD21, RAD51 , RAD51 B, RAD51C, RAD51 D, RAD52, RAD54L, RAF1, RARA, RB1 , RBM10, REL, RET, RICTOR, RNF43, ROS1, RPTOR, RSPO2, SDC4, SDHA, SDHB, SDHC, SDHD, SETD2, SF3B1, SGK1 , SLC34A2, SMAD2, SMAD4, SMARCA4, SMARCB1, SMO, SNCAIP, SOCS1 , SOX2, SOX9, SPEN, SPOP, SRC, STAG2, STAT3, STK11 , SUFU, SYK, TBX3, TEK, TERC, TERT, TET2, TGFBR2, TIPARP, TMPRSS2, TNFAIP3, TNFRSF14, TP53, TSC1, TSC2, TYRO3, U2AF1 , VEGFA, VHL, WHSC1 , WHSC1 L1 , WT1, XPO1 , XRCC2, ZNF217, ZNF703, ABL, ALK, ALL, B4GALNT1, BAFF, BCL2, BRAF, BRCA, BTK, CD19, CD20, CD3, CD30, CD319, CD38, CD52, CDK4, CDK6, CML, CRACC, CS1, CTLA-4, dMMR, EGFR, ERBB1, ERBB2, FGFR1-3, FLT3, GD2, HDAC, HER1, HER2, HR, IDH2, IL-1p, IL-6, IL-6R, JAK1, JAK2, JAK3, KIT, KRAS, MEK, MET, mTOR, PARP, PD-1 , PDGFR, PDGFRa, PDGFRp, PD- L1, PI3K6, PIGF, PTCH, RAF, RANKL, RET, ROS1 , SLAMF7, VEGF, VEGFA, or VEGFB.62. The method of any of clauses 7 to 61 , further including: receiving the sample.63. The method of any of clauses 7 to 62, wherein the sample includes a tissue biopsy sample, a liquid biopsy sample, or a normal control.64. The method of any of clauses 7 to 63, wherein the sample is a liquid biopsy sample and includes blood, plasma, cerebrospinal fluid, sputum, stool, urine, lymphatic fluid, or saliva.65. The method of any of clauses 7 to 64, wherein the sample is a liquid biopsy sample and includes circulating tumor cells (CTCs).66. The method of any of clauses 7 to 65, wherein the sample is a liquid biopsy sample and includes cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.67. The method of any of clauses 7 to 66, further including extracting the nucleic acid molecules from the sample.68. The method of clause 67, wherein the nucleic acid molecules include genomic DNA or cDNA.69. A non-transitory computer-readable medium configured to perform the method of any of clauses 7 to 68.70. A system, including: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations including the method of any of clauses 7 to 68.71 . The system of clause 70, further including: a sequencer configured to generate the sequence read data.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT72. The system of clause 70 or 71, further including: a transceiver configured to transmit, to an external device, a communication signal including the report.73. The system of any of clauses 70 to 72, further including: an output device configured to output the report.Conclusion
[0161] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term herein and a term in an incorporated reference, the term herein controls.
[0162] The features disclosed in the foregoing description, or the following claims, or the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for attaining the disclosed result, as appropriate, may, separately, or in any combination of such features, be used for realizing implementations of the disclosure in diverse forms thereof.
[0163] As will be understood by one of ordinary skill in the art, each implementation disclosed herein can comprise, consist essentially of or consist of its particular stated element, step, or component. Thus, the terms "include” or "including” should be interpreted to recite: "comprise, consist of, or consist essentially of.” The transition term "comprise” or "comprises” means has, but is not limited to, and allows for the inclusion of unspecified elements, steps, ingredients, or components, even in major amounts. The transitional phrase "consisting of' excludes any element, step, ingredient or component not specified. The transition phrase "consisting essentially of' limits the scope of the implementation to the specified elements, steps, ingredients or components and to those that do not materially affect the implementation. As used herein, the term "based on” is equivalent to "based at least partly on,” unless otherwise specified.
[0164] Unless otherwise indicated, all numbers expressing quantities, properties, conditions, and so forth used in the specification and claims are to be understood as being modified in all instances by the term "about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. When further clarity is required, the term "about” has the meaning reasonably ascribed to it by a person skilled in the art when used in conjunction with a stated numerical value or range, i.e., denoting somewhat more or somewhat less than the stated value or range, to within a range of ±20% of the stated value; ±19% of the stated value; ±18% of the stated value; ±17% of the stated value; ±16% of the stated value; ±15% of the stated value; ±14% of the stated value; ±13% of the stated value; ±12% of the stated value; ±11% of the stated value; ±10% of the stated value; ±9% of the stated value; ±8% of the stated value; ±7% of the stated value; ±6% of the stated value; ±5% of the stated value; ±4% of the stated value; ±3% of the stated value; ±2% of the stated value; or ±1 % of the stated value.
[0165] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. Any numerical value, however, inherently contains certain errors necessarily resulting from the standard deviation found in their respective testing measurements.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0166] The terms "a,” "an,” "the,” and similar referents used in the context of describing implementations (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. Recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., "such as”) provided herein is intended merely to better illuminate implementations of the disclosure and does not pose a limitation on the scope of the disclosure. No language in the specification should be construed as indicating any non-claimed element essential to the practice of implementations of the disclosure.
[0167] Groupings of alternative elements or implementations disclosed herein are not to be construed as limitations. Each group member may be referred to and claimed individually or in any combination with other members of the group or other elements found herein. It is anticipated that one or more members of a group may be included in, or deleted from, a group for reasons of convenience and / or patentability. When any such inclusion or deletion occurs, the specification is deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.
[0168] Unless otherwise indicated, the practice of the present disclosure can employ conventional techniques of immunology, molecular biology, microbiology, cell biology and recombinant DNA. These methods are described in the following publications. See, e.g., Sambrook, et al. Molecular Cloning: A Laboratory Manual, 2nd Edition (1989); F. M. Ausubel, et al. eds., Current Protocols in Molecular Biology, (1987); the series Methods IN Enzymology (Academic Press, Inc.); M. MacPherson, et al., PCR: A Practical Approach, IRL Press at Oxford University Press (1991); MacPherson et al., eds. PCR 2: Practical Approach, (1995); Harlow and Lane, eds. Antibodies, A Laboratory Manual, (1988); and R. I. Freshney, ed. Animal Cell Culture (1987).
[0169] Tumor mutational burden (TMB) is a measure of the number of mutations carried by tumor cells. By comparing DNA sequences from a patient’s healthy tissues and tumor cells, the number of acquired somatic mutations present in tumors, but not in normal tissues, may be determined. In some instances, driver mutations may be excluded from a TMB calculation.
[0170] In certain examples, "tumor mutational burden" or “TMB" refers to the number of somatic mutations in a tumor’s genome and / or the number of somatic mutations per area of the tumor's genome. In some embodiments, TMB, as used herein, refers to the number of somatic mutations per megabase (Mb) of DNA sequenced. In some embodiments, germline (inherited) variants are excluded when determining TMB, given that the immune system has a higher likelihood of recognizing these as self. In various cases, driver mutations are excluded from a TMB calculation.
[0171] A viral status test refers to a test that identifies the presence of viral RNA or DNA in a subject. The test can identify viral load and / or viral identity. For example, the viral status test can identify the presence of viral RNA or DNA associated with the occurrence of certain cancers. Examples of such viruses include Hepatitis B Virus (HBV) and Hepatitis C Virus (HCV), Kaposi Sarcoma-Associated Herpesvirus (KSHV), Merkel Cell Polyomavirus (MCV), Human Papillomavirus (HPV), Human Immunodeficiency Virus Type 1 (HIV-1 , or HIV), Human T-Cell Lymphotropic Virus Type 1 (HTLV-1), and Epstein-Barr Virus (EBV).FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT
[0172] Cancer "hotspot” mutations give rise to oncological outcomes. PhyloP, SIFT, Grantham, COSMIC and PolyPhen-2 are in silico tools that can be used to assess pathogenicity of identified variants. Exemplary hotspot genes and mutations include EGFR exon 19 activating mutation, EGFR exon 19 deletion, EGFR exon 19 insertion, EGFR exon 19 sensitizing mutation, EGFR exon 20 activation mutation, EGFR exon 20 insertion, EGFR G719 mutation, EGFR L858R mutation, EGFR L861 mutation, EGFR S768 mutation, EGFR T790M mutation, C797 mutation, KIT activating mutation, KRAS activating mutation, MET activating mutation, NRAS activating mutation, PMS2 promoter mutations, among many others. Hotspot mutations also occur in the following genes: AKT2, BRCA1 , BRCA2, ERC1 , MSD1 , POLH, PPM1 G, PTEN, RAD18, RAD51 , RAD51 B, RB1 , TERT, TP53, TP53Bp1 , ALK, ARMT1, ATAD5, ATG7, ATIC, AXL, BIRC6, BRD3, BRD4, CAPRIN1 , CCAR2, CCDC6, CDK5RAP2, CHD9, CIT, CTNNB1, CUL1 , EBF1 , EIF3E, HIP1 , HMGA2, IRF2BP2, NOTCH1 , NOTCH4, NPM1 , OFD1 , TACC1 , TACC3, TERF2, TMEM106B, UBE2L3, USP10, WRDR48, YAP1 , ZEB2, and ZMYND8.
[0173] A "DNA methylation test” refers to an assay, which can be commercially available, for distinguishing methylated versus unmethylated cytosine loci in DNA. Techniques for measuring cytosine methylation include bisulfite- based methylation assays. The addition of bisulfite to DNA results in the methylation of unmethylated cytosine and its ultimate conversion to the nucleotide uracil. Uracil has similar binding properties to thiamine in the DNA sequence. Previously methylated cytosine does not undergo similar chemical conversion on exposure to bisulfite. Bisulfite assays can thus be used to discriminate previously methylated versus unmethylated cytosine.
[0174] An exemplary quantitative methylation detection assay combines bisulfite treatment and restriction analysis COBRA, which uses methylation sensitive restriction endonucleases, gel electrophoresis, and detection based on labeled hybridization probes. (Ziong and Laird, Nucleic Acid Res. 199725; 2532-4). Another exemplary detection assay is the methylation specific polymerase chain reaction PGR (MSPCR) for amplification of DNA segments of interest. This assay can be performed after sodium bisulfite conversion of cytosine and uses methylation sensitive probes. Other detection assays include the Quantitative Methylation (QM) assay, which combines PGR amplification with fluorescent probes designed to bind to putative methylation sites; MethyLight™ (Qiagen, Redwood City, CA) a quantitative methylation detection assay that uses fluorescence-based PGR (Eads, et al., Cancer Res. 1999; 59:2302- 2306); and Ms-SNuPE, a quantitative technique for determining differences in methylation levels in CpG sites. As with other techniques, Ms-SNuPE also requires bisulfite treatment to be performed first, leading to the conversion of unmethylated cytosine to uracil while methyl cytosine is unaffected. PGR primers specific for bisulfite converted DNA are then used to amplify the target sequence of interest. The amplified PGR product is isolated and used to quantitate the methylation status of the CpG site of interest. (Gonzalgo and Jones Nuclei Acids Res1997; 25:252-31).
[0175] In particular embodiments, pyrosequencing can be used to detect marker methylation. Pyrosequencing is a method of DNA sequencing that relies on detection of the release of pyrophosphates as DNA is synthesized (and is therefore a "sequencing by synthesis” technique). To assess methylation by pyrosequencing, a DNA sample can be incubated with sodium bisulfite, converting unmethylated cytosine to uracil. The presence of uracil will result in thymine incorporation during PGR amplification. Therefore, sequencing results that include thymine at a nucleotide position that is known to encode cytosine can be interpreted as unmethylated sites. In contrast cytosines present in the sequencing results indicate that the site was methylated in the original DNA sample, because methylation protects cytosine from conversion to uracil upon treatment. Bisulfite treatment can also be performed on control samples withFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT known methylation patterns, to reduce or eliminate false positive results. Commercially available pyrosequencing machines include Pyro Mark Q96 (Qiagen, Hilden, Germany). For more details on methods to use pyrosequencing for measurement of methylation, see Delaney et al. Methods Mol Biol. 2015 1343: 249-264. Pyrosequencing is especially useful for detecting methylation in the CpG sites within genes.
[0176] In particular embodiments, a protein marker is detected by contacting a sample with reagents (e.g., antibodies), generating complexes of reagent and marker(s), and detecting the complexes. Particular embodiments for detecting and measuring protein levels can use methods including agglutination, chemiluminescence, electrochemiluminescence (ECL), enzyme-linked immunoassays (ELISA), immunoassay, immunoblotting, immunodiffusion, Immunoelectrophoresis, immunofluorescence, immunohistochemistry, immunoprecipitation, mass-spectrometry, and western blot. See also, e.g., E. Maggio, Enzyme-Immunoassay (1980), CRC Press, Inc., Boca Raton, Fla; and U.S. Pat. Nos. 4,727,022; 4,659,678; 4,376,110; 4,275,149; 4,233,402; and 4,230,797.
[0177] Read depth refers to the number of times that a specific genomic site is sequenced during a sequencing run.
[0178] Certain implementations are described herein, including the best mode known to the inventors for carrying out implementations of the disclosure. Of course, variations on these described implementations will become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventor expects skilled artisans to employ such variations as appropriate, and the inventors intend for implementations to be practiced otherwise than specifically described herein. Accordingly, the scope of this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by implementations of the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.
Claims
FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCTCLAIMSWhat is claimed is:
1. A method, comprising: receiving sequence read data of nucleic acid molecules in a sample treated with a corrective enzyme after being obtained from a subject; generating a variant allele frequency (VAF) distribution of at least one type of substitution variant of the sequence read data by comparing the sequence read data to a reference genome; identifying candidate variants in the VAF distribution; determining that at least a portion of candidate variants in the VAF distribution are false positive variants by determining that greater than a threshold number of the candidate variants are present within a threshold range of the VAF distribution; and generating a report based on one or more additional variants indicated by the sequence read data, the one or more additional variants omitting the false positive variants.
2. The method of claim 1 , wherein the sample was stored in a frozen form after being obtained from the subject and before being sequenced to generate the sequence read data.
3. The method of claim 2, wherein the sample was exposed to a temperature of greater than 0° Celsius during storage for greater than a threshold amount of time.
4. The method of claim 1 , wherein a portion of ground truth nucleotides of the nucleic acid molecules were converted erroneous nucleotides after being obtained from the subject.
5. The method of claim 4, wherein: the ground truth nucleotides comprise methylated cytosines, and the erroneous nucleotides comprise thymines; and / or wherein: the ground truth nucleotides comprise unmethylated cytosines, and the erroneous nucleotides comprise uracils.
6. The method of claim 1 , wherein the corrective enzyme comprises uracil-DNA glycosylase.
7. The method of claim 1 , wherein the VAF distribution comprises candidate variants in CpG regions of the sample.
8. The method of claim 1 , wherein the at least one type of substitution variant comprises at least one of: cytosine to thymine variants; or guanine to adenosine variants.
9. The method of claim 1 , wherein identifying the candidate variants in the VAF distribution comprises: detecting a peak of the VAF distribution; and identifying the candidate variants by identifying locations in the VAF distribution that correspond to greater than a threshold percentage of the peak of the VAF distribution.
10. The method of claim 9, wherein the threshold percentage is in a range of about 60% to about 90%.11 . The method of claim 9, wherein the threshold number of the candidate variants is in a range of about 5 to about 100, and wherein the threshold range of the VAF distribution is less than 20% of the VAF distribution.FMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT12. The method of claim 1, further comprising: determining that the false positive variants correspond to ground truth methylated cytosines in the nucleic acid molecules in the sample, wherein the report indicates the ground truth methylated cytosines.
13. A system, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: receiving sequence read data of nucleic acid molecules in a sample treated with a corrective enzyme after being obtained from a subject; generating a variant allele frequency (VAF) distribution of at least one type of substitution variant of the sequence read data by comparing the sequence read data to a reference genome; identifying candidate variants in the VAF distribution; determining that at least a portion of candidate variants in the VAF distribution are false positive variants by determining that greater than a threshold number of the candidate variants are present within a threshold range of the VAF distribution; and generating a report based on one or more additional variants indicated by the sequence read data, the one or more additional variants omitting the false positive variants.
14. The system of claim 13, further comprising: a sequencer configured to generate the sequence read data.
15. A method, comprising: providing a plurality of nucleic acid molecules obtained from a sample from a subject; treating the plurality of nucleic acid molecules with uracil-DNA glycosylase; ligating one or more adapters onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying the one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules; capturing amplified nucleic acid molecules from the amplified nucleic acid molecules; sequencing, by a sequencer, all or a subset of the captured amplified nucleic acid molecules to obtain a plurality of sequence reads that represent the sequenced amplified nucleic acid molecules thereby generating sequence read data; receiving, at one or more processors, the sequence read data for the plurality of sequence reads; generating, by the one or more processors, a variant allele frequency (VAF) distribution of cytosine to thymine and guanine to adenosine substitutions in CpG sites of the sequence read data by comparing the sequence read data to a reference genome; identifying candidate variants in the VAF distribution; determining that at least a portion of candidate variants in the VAF distribution are false positive variants by determining that greater than a threshold number of the candidate variants are present within a threshold range of the VAF distribution; andFMI Docket No.: 0189-WO 10142-CBL&H Docket No.: F171-6005PCT generating a report based on one or more additional variants indicated by the sequence read data, the one or more additional variants omitting the false positive variants.
16. The method of claim 15, wherein providing the plurality of nucleic acid molecules obtained from a sample from the subject is in response to storing the sample for a year or more.
17. The method of claim 15, wherein treating the plurality of nucleic acid molecules with uracil-DNA glycosylase comprises converting uracils in the nucleic acid molecules to cytosines.
18. The method of claim 15, wherein the threshold number of the candidate variants is about 10, and wherein the threshold range of the VAF distribution comprises about 10% of the VAF distribution.
19. The method of claim 15, further comprising: determining that at least a portion of the false positive variants are indicative of ground truth methylated cytosines in the plurality of nucleic acid molecules, wherein generating the report is further based on the ground truth methylated cytosines.
20. The method of claim 15, further comprising: performing Gaussian smoothing on the VAF.